Diplomová práce

Evaluation Testing of LLM Prompts Using the LLM-as-a-Judge Approach

Kristián Dobeš
Anotace

Hodnocení výstupů velkých jazykových modelů zůstává obtížným problémem, zejména u otevřených úloh, jejichž kvalita závisí na vlastnostech, jako jsou relevance, tón či doménově specifická správnost. Tradiční metriky, jako jsou BLEU, ROUGE a BERTScore, zpravidla vyžadují pevně danou referenci, a proto jsou pro tento typ úloh nedostatečné. Tato práce zkoumá přístup LLM-as-a-Judge, při němž jeden velký …více

Abstract

Evaluating large language model outputs remains difficult: traditional metrics such as BLEU, ROUGE, and BERTScore require a fixed reference, making them inadequate for open-ended tasks judged on qualities such as relevance, tone, or domain-specific correctness. This thesis examines the LLM-as-a-Judge approach, in which an LLM grades another model's output against natural-language criteria, and applies …více

Zadání práce
Demonstrate the use of LLM evaluations, with focus on the LLM-as-a-Judge approach, on a sufficiently complex practical use case.

The thesis must answer the following questions, including proof-of-concept implementation demonstrating the concepts:
1. Is LLM-as-a-Judge a relevant approach for software engineers to use in software systems incorporating LLMs?
2. How do LLM evaluations using LLM-as-a-Judge affect the development process?
3. How do LLM evaluations take advantage of prompt and context engineering and input data engineering?

The thesis may follow the following structure to provide relevant answers to these fundamental questions:
1. Determine particular use cases benefiting from the use of LLMs and a set of representative sample uses.
2. Research and explain the LLM-as-a-Judge approach to LLM evaluations, currently used metrics, relevant best practices and currently recognized limits of the approach.
3. Determine how to apply the LLM-as-a-Judge approach in your particular selected use case.
4. Research and compile a list of current available mainstream approaches and tooling for LLM-as-a-Judge. Pick one particular tool/library for your use case.
5. Implement evaluations using your tool/library of choice and demonstrate evaluations of your representative sample uses.
6. Collect, apply and explain best practices for LLM evaluations using the LLM-as-a-Judge paradigm as demonstrated in your implementation.
Práce zkontrolována:
20. 5. 2026 11:14, RNDr. Ondřej Krajíček, učo 39489
Jazyk práce
angličtina angličtina
Termín obhajoby
17. 6. 2026
Práce byla úspěšně obhájena

Vedoucí

RNDr. Ondřej Krajíček, učo 39489
KPSK FI MU

Oponent

RNDr. Samuel Pastva, Ph.D., učo 410286
KPSK FI MU

Konzultant

Bc. Tomáš Grbálik
abs FI MU

  • Přidání souboru

    Soubor nebo složku lze nahrát pomocí tlačítka Přidat.
  • Další operace se soubory

    Podrobnosti lze zjistit označením příslušného řádku.
  • Pohled pro experty

    Pro častou práci je možné zvolit režim Více možností.
  • Vyhledávání souborů

    Vyhledávaný výraz můžete zadat přímo do adresního řádku.
  • Rychlý přístup k souborům

    Pomocí funkce Nedávné je možné se rychle vrátit k právě prohlíženým souborům. Oblíbené soubory je také možné označit Hvězdičkou.