Testing of detection tools for AI-generated text

J 2023

Testing of detection tools for AI-generated text

WEBER-WULFF, Debora, Alla ANOHINA-NAUMECA, Sonja BJELOBABA, Tomáš FOLTÝNEK, Jean GUERRERO-DIB et. al.

Basic information

Original name

Testing of detection tools for AI-generated text

Authors

WEBER-WULFF, Debora (276 Germany), Alla ANOHINA-NAUMECA (428 Latvia), Sonja BJELOBABA (752 Sweden), Tomáš FOLTÝNEK (203 Czech Republic, guarantor, belonging to the institution), Jean GUERRERO-DIB (484 Mexico), Olumide POPOOLA, Petr ŠIGUT (703 Slovakia, belonging to the institution) and Lorna WADDINGTON

Edition

International Journal for Educational Integrity, 2023, 1833-2595

Other information

Language

English

Type of outcome

Článek v odborném periodiku

Field of Study

10200 1.2 Computer and information sciences

Country of publisher

Germany

Confidentiality degree

není předmětem státního či obchodního tajemství

References:

URL

Impact factor

Impact factor: 4.600 in 2022

RIV identification code

RIV/00216224:14330/23:00132774

Organization unit

Faculty of Informatics

DOI

http://dx.doi.org/10.1007/s40979-023-00146-z

UT WoS

001129231700001

Keywords in English

Artifcial intelligence; Generative pre-trained transformers; Machine-generated text; Detection of AI-generated text; Academic integrity; ChatGPT; AI detectors

Abstract

V originále

Recent advances in generative pre-trained transformer large language models have emphasised the potential risks of unfair use of artifcial intelligence (AI) generated content in an academic environment and intensifed eforts in searching for solutions to detect such content. The paper examines the general functionality of detection tools for AI-generated text and evaluates them based on accuracy and error type analysis. Specifcally, the study seeks to answer research questions about whether existing detection tools can reliably diferentiate between human-written text and ChatGPTgenerated text, and whether machine translation and content obfuscation techniques afect the detection of AI-generated text. The research covers 12 publicly available tools and two commercial systems (Turnitin and PlagiarismCheck) that are widely used in the academic setting. The researchers conclude that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques signifcantly worsen the performance of tools. The study makes several signifcant contributions. First, it summarises up-to-date similar scientific and non-scientifc eforts in the feld. Second, it presents the result of one of the most comprehensive tests conducted so far, based on a rigorous research methodology, an original document set, and a broad coverage of tools. Third, it discusses the implications and drawbacks of using detection tools for AI-generated text in academic settings.

Citovat

WEBER-WULFF, Debora, Alla ANOHINA-NAUMECA, Sonja BJELOBABA, Tomáš FOLTÝNEK, Jean GUERRERO-DIB, Olumide POPOOLA, Petr ŠIGUT and Lorna WADDINGTON. Testing of detection tools for AI-generated text. International Journal for Educational Integrity. 2023, vol. 19, No 26, p. 1-39. ISSN 1833-2595. Available from: https://dx.doi.org/10.1007/s40979-023-00146-z.

@article{2355979,
   author = {WeberandWulff, Debora and AnohinaandNaumeca, Alla and Bjelobaba, Sonja and Foltýnek, Tomáš and GuerreroandDib, Jean and Popoola, Olumide and Šigut, Petr and Waddington, Lorna},
   article_number = {26},
   doi = {http://dx.doi.org/10.1007/s40979-023-00146-z},
   keywords = {Artifcial intelligence; Generative pre-trained transformers; Machine-generated text; Detection of AI-generated text; Academic integrity; ChatGPT; AI detectors},
   language = {eng},
   issn = {1833-2595},
   journal = {International Journal for Educational Integrity},
   title = {Testing of detection tools for AI-generated text},
   url = {https://link.springer.com/article/10.1007/s40979-023-00146-z},
   volume = {19},
   year = {2023}
}

TY  - JOUR
ID  - 2355979
AU  - Weber-Wulff, Debora - Anohina-Naumeca, Alla - Bjelobaba, Sonja - Foltýnek, Tomáš - Guerrero-Dib, Jean - Popoola, Olumide - Šigut, Petr - Waddington, Lorna
PY  - 2023
TI  - Testing of detection tools for AI-generated text
JF  - International Journal for Educational Integrity
VL  - 19
IS  - 26
SP  - 1-39
EP  - 1-39
SN  - 18332595
KW  - Artifcial intelligence
KW  - Generative pre-trained transformers
KW  - Machine-generated text
KW  - Detection of AI-generated text
KW  - Academic integrity
KW  - ChatGPT
KW  - AI detectors
UR  - https://link.springer.com/article/10.1007/s40979-023-00146-z
N2  - Recent advances in generative pre-trained transformer large language models have emphasised the potential risks of unfair use of artifcial intelligence (AI) generated content in an academic environment and intensifed eforts in searching for solutions to detect such content. The paper examines the general functionality of detection tools for AI-generated text and evaluates them based on accuracy and error type analysis. Specifcally, the study seeks to answer research questions about whether existing detection tools can reliably diferentiate between human-written text and ChatGPTgenerated text, and whether machine translation and content obfuscation techniques afect the detection of AI-generated text. The research covers 12 publicly available tools and two commercial systems (Turnitin and PlagiarismCheck) that are widely used in the academic setting. The researchers conclude that the available detection tools are neither accurate nor reliable and have a main bias towards classifying the output as human-written rather than detecting AI-generated text. Furthermore, content obfuscation techniques signifcantly worsen the performance of tools. The study makes several signifcant contributions. First, it summarises up-to-date similar scientific and non-scientifc eforts in the feld. Second, it presents the result of one of the most comprehensive tests conducted so far, based on a rigorous research methodology, an original document set, and a broad coverage of tools. Third, it discusses the implications and drawbacks of using detection tools for AI-generated text in academic settings.
ER  -

WEBER-WULFF, Debora, Alla ANOHINA-NAUMECA, Sonja BJELOBABA, Tomáš FOLTÝNEK, Jean GUERRERO-DIB, Olumide POPOOLA, Petr ŠIGUT and Lorna WADDINGTON. Testing of detection tools for AI-generated text. \textit{International Journal for Educational Integrity}. 2023, vol.~19, No~26, p.~1-39. ISSN~1833-2595. Available from: https://dx.doi.org/10.1007/s40979-023-00146-z.

Detailed Information on Publication Record