Závěrečná práce: Kristián Dobeš: Evaluation Testing of LLM Prompts Using the LLM-as-a-Judge Approach
Diplomová práce
Evaluation Testing of LLM Prompts Using the LLM-as-a-Judge Approach
Anotace
Hodnocení výstupů velkých jazykových modelů zůstává obtížným problémem, zejména u otevřených úloh, jejichž kvalita závisí na vlastnostech, jako jsou relevance, tón či doménově specifická správnost. Tradiční metriky, jako jsou BLEU, ROUGE a BERTScore, zpravidla vyžadují pevně danou referenci, a proto jsou pro tento typ úloh nedostatečné. Tato práce zkoumá přístup LLM-as-a-Judge, při němž jeden velký …více
Abstract
Evaluating large language model outputs remains difficult: traditional metrics such as BLEU, ROUGE, and BERTScore require a fixed reference, making them inadequate for open-ended tasks judged on qualities such as relevance, tone, or domain-specific correctness. This thesis examines the LLM-as-a-Judge approach, in which an LLM grades another model's output against natural-language criteria, and applies …více
Zadání práce
The thesis must answer the following questions, including proof-of-concept implementation demonstrating the concepts:
1. Is LLM-as-a-Judge a relevant approach for software engineers to use in software systems incorporating LLMs?
2. How do LLM evaluations using LLM-as-a-Judge affect the development process?
3. How do LLM evaluations take advantage of prompt and context engineering and input data engineering?
The thesis may follow the following structure to provide relevant answers to these fundamental questions:
1. Determine particular use cases benefiting from the use of LLMs and a set of representative sample uses.
2. Research and explain the LLM-as-a-Judge approach to LLM evaluations, currently used metrics, relevant best practices and currently recognized limits of the approach.
3. Determine how to apply the LLM-as-a-Judge approach in your particular selected use case.
4. Research and compile a list of current available mainstream approaches and tooling for LLM-as-a-Judge. Pick one particular tool/library for your use case.
5. Implement evaluations using your tool/library of choice and demonstrate evaluations of your representative sample uses.
6. Collect, apply and explain best practices for LLM evaluations using the LLM-as-a-Judge paradigm as demonstrated in your implementation.
20. 5. 2026 11:14, RNDr. Ondřej Krajíček, učo 39489
Konzultant
abs FI MU
Práce na příbuzné téma
Seznam prací, které mají shodná klíčová slova.
-
Design and Implementation of an Automated Testing Tool for Web Page Chatbots
Bc. Dominik Haspra -
Formulace statistických příkladů a projektů v éře generativní umělé inteligence
Bc. Lenka Šoltésová -
Anotace analýzy sentimentu a postoje pomocí generativního jazykového modelu
Mgr. Daniel Šurda -
AI-Based Threat Simulation and Mitigation in Practical Laboratories
Mgr. Kenan Fejzic, učo 565128 -
LLMs with Test Feedback for Program Synthesis
Bc. Marcel Nadzam -
Syntetická datová sada pro detekci propagandy
Bc. Pavel František Oujeský -
Extrakce informací ze sportovních přenosů
Ing. Jakub Dvořák -
Predicting Facebook Ad Campaign Performance Using Few-Shot Learning with Large Language Models
Mgr. Shahadat Hussain




