Extending Czech WordNet Using a Bilingual Dictionary

BLAHUŠ, Marek a Karel PALA. Extending Czech WordNet Using a Bilingual Dictionary. In Christiane Fellbaum, Piek Vossen. 6th International Global Wordnet Conference Proceedings. Matsue, Japan: Toyohashi University of Technology. s. 50-55. ISBN 978-80-263-0244-5. 2012.

Další formáty: BibTeX LaTeX RIS

Základní údaje
Originální název	Extending Czech WordNet Using a Bilingual Dictionary
Autoři	BLAHUŠ, Marek (203 Česká republika, domácí) a Karel PALA (203 Česká republika, garant, domácí).
Vydání	Matsue, Japan, 6th International Global Wordnet Conference Proceedings, od s. 50-55, 6 s. 2012.
Nakladatel	Toyohashi University of Technology

Další údaje
Originální jazyk	angličtina
Typ výsledku	Stať ve sborníku
Obor	10201 Computer sciences, information science, bioinformatics
Stát vydavatele	Japonsko
Utajení	není předmětem státního či obchodního tajemství
Forma vydání	tištěná verze "print"
Kód RIV	RIV/00216224:14330/12:00059218
Organizační jednotka	Fakulta informatiky
ISBN	978-80-263-0244-5
Klíčová slova česky	Cornetto; XML database; DEB platform
Klíčová slova anglicky	Cornetto; XML databáze; platforma DEB
Příznaky	Mezinárodní význam, Recenzováno
Změnil	Změnil: RNDr. Pavel Šmerk, Ph.D., učo 3880. Změněno: 12. 4. 2013 10:11.

Anotace

In this paper we describe semi-automatical extending of the Czech WordNet lexical database (48,000 literals in 28,000 synsets) by translation of English literals from existing synsets in Princeton WordNet. We make use of a machine-readable bilingual dictionary to extract English-Czech translation pairs, search the English literals in Princeton WordNet and in case of a high-confidence match we transfer the literal into Czech WordNet. Along with literals, new synsets parallel to the English ones and identified by ILI are introduced into CzechWordNet, including information on their ILR (Internal Language Relations) such as hypernymy/hyponymy. The paper describes the parsing of the dictionary data, extraction of translation pairs and the criteria used for estimating the confidence level of a match. Results of the work are 36,228 added literals and 12,403 created synsets. An overview of previous similar attempts for other languages is also included.

Návaznosti
LC536, projekt VaV	Název: Centrum komputační lingvistiky
LC536, projekt VaV	Investor: Ministerstvo školství, mládeže a tělovýchovy ČR, Centrum komputační lingvistiky

VytisknoutZobrazeno: 19. 4. 2024 15:09

Extending Czech WordNet Using a Bilingual Dictionary

Další aplikace