Bakalářská práce
Získaná ocenění: Cena děkana FI za vynikající závěrečnou práci

Similarity searching of proteins using machine learning techniques

Martin Gendiar
Anotace

Vyhľadávanie v databázach proteínov je komplexný problém, najmä z dôvodu nemožnosti ich zoradenia podľa objektívnych kritérií. V roku 2018 bol publikovaný článok s názvom The Case for Learned Index Structures, ktorý argumentuje za nový prístup k organizácii a vyhľadávaniu komplexných dát pomocou strojového učenia. Práca aplikuje tento prístup na problém podobnostného vyhľadávania v proteínových dátach …více

Abstract

Searching in protein databases is a complex problem, mainly because proteins cannot be sorted according to any objective criteria. In 2018, a paper called The Case for Learned Index Structures was published, arguing for a new paradigm for organising and searching within complex data using machine learning. This thesis applies such an approach to the problem of similarity searching in protein data and …více

Zadání práce
Searching protein databases is a difficult problem, partly because proteins cannot be sorted according to any objective criteria. A practical solution to this is similarity searching -- we can define a similarity function which determines the similarity between each pair of proteins. In 2018, a paper called The Case for Learned Index Structures has been published, arguing for a new paradigm for organizing and searching within complex data using machine learning. The goal of this thesis is to apply such an approach to the problem of similarity searching in protein data and evaluate the results. Firstly, the student will need to find a way to process data from the Protein Data Bank and represent it in a way that is suitable for machine learning. Secondly, the proteins will need to be indexed using an existing framework called Learned Metric Index (LMI) -- since the framework has never been used with this type of data, it will be necessary to identify the distinctive characteristics of protein data and modify the setup of LMI to appropriately represent similarity within this dataset. Finally, the searching efficiency of the resulting index will be evaluated.
Práce zkontrolována:
19. 5. 2022 12:42, Mgr. et Mgr. Jaroslav Oľha, Ph.D., učo 348646
Jazyk práce
angličtina angličtina
Termín obhajoby
29. 6. 2022
Práce byla úspěšně obhájena

Vedoucí

Mgr. et Mgr. Jaroslav Oľha, Ph.D., učo 348646
ANKO DITI ÚVT MU

Oponent

doc. RNDr. Vlastislav Dohnal, Ph.D., učo 2952
KSUZD FI MU

Konzultant

RNDr. Matej Antol, Ph.D., učo 325040
CERIT SC ÚVT MU

Literatura

  • ANTOL, Matej; Jaroslav OĽHA; Terézia SLANINÁKOVÁ a Vlastislav DOHNAL. Learned metric index - proposition of learned indexing for unstructured data. Information Systems. Elsevier, 2021, roč. 100, č. 101774, s. 1-12. ISSN 0306-4379. Dostupné z: https://doi.org/10.1016/j.is.2021.101774.

Masarykova univerzita Fakulta informatiky
Studijní program
Plán
Informatika
  • Přidání souboru

    Soubor nebo složku lze nahrát pomocí tlačítka Přidat.
  • Další operace se soubory

    Podrobnosti lze zjistit označením příslušného řádku.
  • Pohled pro experty

    Pro častou práci je možné zvolit režim Více možností.
  • Vyhledávání souborů

    Vyhledávaný výraz můžete zadat přímo do adresního řádku.
  • Rychlý přístup k souborům

    Pomocí funkce Nedávné je možné se rychle vrátit k právě prohlíženým souborům. Oblíbené soubory je také možné označit Hvězdičkou.