Bakalářská práce

Benchmarking Pragmatic Reasoning of Large Language Models in the Czech Language

Yevhenii Karpizenkov
Anotace

Velké jazykové modely se stále častěji používají jako konverzační systémy, a proto je důležité hodnotit nejen jejich faktické a sémantické schopnosti, ale také jejich schopnost vyvozovat v dialogu význam závislý na kontextu, který je ústředním předmětem pragmatiky. Současné benchmarky zaměřené na porozumění pragmatickému významu se stále soustředí převážně na angličtinu, takže čeština a další morfologicky …více

Abstract

Large language models are increasingly used as conversational systems, which makes it important to evaluate not only their factual and semantic abilities, but also their capacity to infer context-dependent meaning in dialogue, a central concern of pragmatics. Existing benchmarks for pragmatic reasoning remain mainly English-centred, leaving Czech and other morphologically rich, word-order-flexible …více

Zadání práce

The goal of this thesis is to design and develop a benchmark to evaluate pragmatic understanding and reasoning in the Czech language using large language models (LLMs).

The thesis will:

  1. Establish the theoretical foundations of cognitive pragmatics and provide a focused literature review of the linguistic aspects of Czech pragmatic phenomena, leading to a clearly defined scope and taxonomy of the phenomena to be evaluated.
  2. Review LLM benchmarks relevant to the thesis topic.
  3. Based on the theoretical and practical foundation, design benchmark tasks and the data format for the selected phenomena, including task definitions, input/output structure, dataset schema, labeling scheme, and difficulty controls.
  4. Create the dataset, ensuring linguistic naturalness and controlling for ambiguity. Implement quality assurance through review and consistency checks. The expected dataset size will be in the thousands of varied items and phenomena, in accordance with the analyzed design.
  5. Deliver a reproducible benchmark and evaluation pipeline. Define metrics, set a human baseline, and run evaluations across multiple LLMs.

The main deliverable is a reproducible benchmark that enables consistent comparison of multiple LLMs' Czech-language pragmatic competence.

Práce zkontrolována:
11. 5. 2026 07:40, doc. RNDr. Aleš Horák, Ph.D., učo 1648
Jazyk práce
angličtina angličtina
Termín obhajoby
4. 6. 2026
Práce byla úspěšně obhájena

Vedoucí

doc. RNDr. Aleš Horák, Ph.D., učo 1648
KSUZD FI MU

Oponent

RNDr. Vojtěch Kovář, Ph.D., učo 139915
ÚČJ FF MU

Konzultant

Radim Lacina, Ph.D., učo 237940
ÚČJ FF MU

  • Přidání souboru

    Soubor nebo složku lze nahrát pomocí tlačítka Přidat.
  • Další operace se soubory

    Podrobnosti lze zjistit označením příslušného řádku.
  • Pohled pro experty

    Pro častou práci je možné zvolit režim Více možností.
  • Vyhledávání souborů

    Vyhledávaný výraz můžete zadat přímo do adresního řádku.
  • Rychlý přístup k souborům

    Pomocí funkce Nedávné je možné se rychle vrátit k právě prohlíženým souborům. Oblíbené soubory je také možné označit Hvězdičkou.