Závěrečná práce: Jakub Kuchár, učo 484954: Evaluation and Interpretation of Word Embeddings
Bakalářská práce
Evaluation and Interpretation of Word Embeddings
Anotace
V práci sa zaoberám problémom neadekvátnej evluácie verejných slovných modelov. Najprv stručne predstavím pojem „slovné embeddingy“ a prečo ich používam. Predstavím existujúce modely pre tieto embeddingy a taktiež niektoré z existujúcich verejných embeddingov. V hlavnej časti tejto práce evaluujem tieto verejné embeddingy na štyroch evaluačných úlohách a na základe výsledkov odporučím, ktoré embeddingy sú vhodné na ktoré úlohy.
Abstract
In my thesis, I deal with the problem of the inadequate evaluation of public word embedding models. First, I briefly introduce word embeddings and why I use them. Then, I introduce the existing models for word embeddings and also some of the existing public word embeddings. In the main part of my thesis, I evaluate public word embeddings on four evaluation tasks and according to results, I suggest which models are best suited for different evaluation tasks.
Zadání práce
The student will study word embedding models [1–3], public pre-trained word embeddings, and existing tasks for their quantitative evaluation [4–7]. The student will implement a selected set of evaluation tasks and use their implementation to evaluate and compare selected pre-trained word embeddings. Based on the results of the evaluation, the student will suggest optimal pre-trained word embeddings for the individual tasks.
Student nastuduje modely pro vektorové reprezentace slov [1–3], veřejně dostupné předtrénované reprezentace slov a existující úlohy pro jejich kvantitativní evaluaci [4–7]. Student implementuje baterii vybraných úloh pro evaluaci vektorových reprezentací slov a s využitím této baterie evaluuje a vzájemně porovná vybrané předtrénované reprezentace. Na základě své evaluace student doporučí vhodné předtrénované reprezentace pro jednotlivé úlohy.
25. 5. 2021 15:43, RNDr. Vít Starý Novotný, Ph.D., učo 409729
Práce na příbuzné téma
Seznam prací, které mají shodná klíčová slova.
-
Utilisation of language representations for Information Retrieval
Ing. Petr Mička -
Domain-specific English-Czech Neural Machine Translation
Mgr. Martin Wörgötter -
Machine Learning for Text Anomaly Detection
Ing. Alina Tsykynovska -
Pretraining and Evaluation of Czech ALBERT Language Model
RNDr. Petr Zelina, učo 469366 -
XXXXXXX cognitive auto-dispatching
Ing. Andrej Dravecký, učo 445432 -
Automatic text summarization
Mgr. Adam Hájek -
Machine translation in a specific domain
Mgr. Tereza Vrabcová, učo 485431 -
Předzpracování klinických poznámek pomocí standardizace slov a frází na základě podobnosti
Bc. Jan Halas




