2025
High-Quality LLM Pre-Training Texts from Dictionary Data.
MEDVEĎ, Marek; Radoslav SABOL; Ondřej SOTOLÁŘ a Aleš HORÁKZákladní údaje
Originální název
High-Quality LLM Pre-Training Texts from Dictionary Data.
Autoři
Vydání
Brno, Czech Republic, Recent Advances in Slavonic Natural Language Processing, RASLAN 2025, od s. 69-84, 16 s. 2025
Nakladatel
Tribun EU
Další údaje
Jazyk
angličtina
Typ výsledku
Stať ve sborníku
Obor
10200 1.2 Computer and information sciences
Stát vydavatele
Česká republika
Utajení
není předmětem státního či obchodního tajemství
Forma vydání
tištěná verze "print"
Odkazy
Označené pro přenos do RIV
Ano
Kód RIV
RIV/00216224:14330/25:00142889
Organizační jednotka
Fakulta informatiky
ISBN
978-80-263-1858-3
ISSN
EID Scopus
Klíčová slova anglicky
large language models; LLMs; pre-training; high-quality data; dictionaries; dictionary entries; Slama models; Czech
Příznaky
Mezinárodní význam
Změněno: 13. 5. 2026 16:59, doc. RNDr. Aleš Horák, Ph.D.
Anotace
V originále
The quality of the pre-training texts is an important aspect in the development of a Large Language Model (LLM). High-quality data, such as collections of textbooks, academic papers, and educational forums, has been shown to improve model performance, generalization, and reduce biases. However, obtaining such data at scale can be challenging, especially for non-mainstream languages like Czech. In this paper, we introduce a method for generating high-quality Czech pre-training data from structured dictionary resources. By employing retrieval-augmented prompting and open-source LLMs, we transform XML-encoded lexicographic dictionary entries into fluent, semantically rich text. The resulting dataset demonstrates that dictionary-grounded generation can effectively enhance data quality. We present the results of experiments with several LLMs and the process of creating a new Czech pre-training dataset, SlamaHQTrain. This dataset was obtained by processing eight Czech dictionaries containing more than 500,000 entries and 18 million words.
Návaznosti
| EH23_025/0008710, projekt VaV |
| |
| 90254, velká výzkumná infrastruktura |
|