BLAHUŠ, Marek. A Spell Checker for Esperanto. Brno: Masarykova univerzita, Fakulta informatiky, 2008, 40 pp. Bakalářská práce.
Other formats:   BibTeX LaTeX RIS
Basic information
Original name A Spell Checker for Esperanto
Name in Czech Kontrolor pravopisu pro Esperanto
Authors BLAHUŠ, Marek.
Edition Brno, 40 pp. Bakalářská práce, 2008.
Publisher Masarykova univerzita, Fakulta informatiky
Other information
Type of outcome Book on a specialized topic
Confidentiality degree is not subject to a state or trade secret
WWW URL
Organization unit Faculty of Informatics
Keywords in English spell checker; Esperanto; Hunspell; corpus; morphology; semantic classification
Tags corpus, esperanto, Hunspell, morphology, semantic classification, spell checker
Tags International impact
Changed by Changed by: Mgr. Marek Blahuš, učo 172464. Changed: 4/7/2008 01:23.
Abstract
This thesis provides a brief overview of spell checking software and describes the process of constructing a spell checker for the Esperanto language and its implementation as a dictionary (i.e. an affix file and a word list) for the Hunspell spell checker. The word list is an adaptation of word roots coming from the renowned Esperanto dictionary PIV. Recognition of morphologically complex words, which are common in Esperanto due to its agglutinative nature, is made possible by the affix file which has been built based on ready-made morpheme segmentation of word derivations appearing in the same source. Rules derived in the latter process have been improved by semantic classification of all involved roots, for which a system has been created based on corpus analysis and several specialized dictionaries, in combination with knowledge on the capability of each affix to accept roots from different semantic classes, acquired from the PMEG reference grammar. The resulting spell checker is a working proof of concept, to be further improved and integrated in the grammar checker project of the E@I organization.
Abstract (in Czech)
Tato práce poskytuje stručný přehled softwaru pro kontrolu pravopisu a popisuje proces konstrukce kontroloru pravopisu pro jazyk Esperanto a jeho implementaci jako slovníku (tj. afixového souboru a seznamu slov) pro kontrolor pravopisu Hunspell. Použitý seznam slov vznikl adaptací slovních kořenů pocházejících z uznávaného esperantského slovníku PIV. Rozpoznávání morfologicky složitých slov, která jsou v Esperantu díky jeho aglutinačnímu charakteru běžná, je umožněno afixovým souborem, který byl vybudován na základě předpřipravené morfologické segmentace slovních tvarů vyskytujících se v témže zdroji. Pravidla odvozená během tohoto procesu byla vylepšena prostřednictvím sémantické klasifikace všech zúčastněných slov, pro kterou byl vytvořen systém založený na analýze korpusu a několika odborných slovnících, v kombinaci se znalostmi o schopnosti každého afixu přijímat kořeny z různých sémantických tříd, získanými z gramatické příručky PMEG. Výsledný kontrolor pravopisu je funkčním prototypem, který bude dále vylepšován a integrován do projektu kontroloru gramatiky organizace E@I.
Abstract (in English)
This thesis provides a brief overview of spell checking software and describes the process of constructing a spell checker for the Esperanto language and its implementation as a dictionary (i.e. an affix file and a word list) for the Hunspell spell checker. The word list is an adaptation of word roots coming from the renowned Esperanto dictionary PIV. Recognition of morphologically complex words, which are common in Esperanto due to its agglutinative nature, is made possible by the affix file which has been built based on ready-made morpheme segmentation of word derivations appearing in the same source. Rules derived in the latter process have been improved by semantic classification of all involved roots, for which a system has been created based on corpus analysis and several specialized dictionaries, in combination with knowledge on the capability of each affix to accept roots from different semantic classes, acquired from the PMEG reference grammar. The resulting spell checker is a working proof of concept, to be further improved and integrated in the grammar checker project of the E@I organization.
PrintDisplayed: 24/7/2024 05:36