Thesis/Dissertation: Petr Mička: Utilisation of language representations for Information Retrieval
Bachelor's thesis
Utilisation of language representations for Information Retrieval
Abstract
Práce je zaměřena na zlepšení kvality systémů na vyhledávání informací. Experimentuje s reprezentacemi jazyka za pomoci neuronových modelů, tzn. jako embeddingy slov nebo váhy pozornosti modelů z rodiny Transformerů. Experimentujeme s kombinováním těchto reprezentací se standardními, ale ortogonálními reprezentacemi založenými na početnosti slov, jako je TF-IDF. Naše experimenty ukazují, že vyhledávací …more
Abstract
Our work aims to create a well-performing information retrieval system utilising neural language representation models as word embeddings and attention maps of selected models of Transformers family. We also experiment with combining neural LM approaches with well-established, yet orthogonal method of TF-IDF. We show that our novel information retrieval systems can beat standard TF-IDF in quality of search and, in ensemble with TF-IDF, can deliver additional quality gains.
Thesis description
25/5/2021 15:46, Mgr. Michal Štefánik, Ph.D., UČO 422237
- Entered/Edited 7/7/2021 09:26, Helena Kryštofová
- Record made 29/4/2021 13:21, Jana Zemanová, UČO 9619
- Accessible from: 25/5/2021 12:42, Alena Dvořáková
- Thesis/dissertation received 25/5/2021 12:42, Alena Dvořáková
Attachments
2.3.3.TF_IDF___Transformer_segmentation_Cranfield.ipynb
2.3.2.TF_IDF___Word_Context_Similarity_Cranfield.ipynb
Notebooks.7z
2.4_Attention_based_IR_system_TREC_full_calculation.ipynb
2.2.TF_IDF_Cranfield.ipynb
2.4.Attention_based_IR_system_heatmap_Cranfield.ipynb
2.3.1.Word_Context_Similarity_TREC_precomputing_results_cheat_low_RAM_requered.ipynb
document_ids.gz
2.3.1.Word_Context_Similarity_Cranfield.ipynb
documents_ratings.gz
2.2.TF_IDF_TREC.ipynb
2.3.2.TF_IDF___Word_Context_Similarity_TREC.ipynb
New_Text_Document.txt
2.4.Attention_based_IR_system_heatmap_TREC.ipynb
2.3.1.Word_Context_Similarity_Cranfield_using_precomputing_results_cheat_low_RAM_requered.ipynb
Consultant
Theses on a related topic
List of theses with an identical keyword.
-
Evaluation and Interpretation of Word Embeddings
Mgr. Jakub Kuchár, UČO 484954 -
Automatic text summarization
Mgr. Adam Hájek -
Pretraining and Evaluation of Czech ALBERT Language Model
RNDr. Petr Zelina, UČO 469366 -
Machine Learning for Text Anomaly Detection
Ing. Alina Tsykynovska -
Transformer Neural Networks for Natural Language Processing
Jonáš Konečný -
Machine translation in a specific domain
Mgr. Tereza Vrabcová -
Preprocessing of clinical notes by similarity-based word and phrase standardisation
Bc. Jan Halas -
Mining Czech Clinical Notes Using the Language Modelling Technology
Mgr. Tomáš Houfek




