Cheng Xu

Thesis: 
Trustworthy Evaluation of Large Language Models: Benchmark Data Contamination and Defence

While Large Language Models (LLMs) show great promise for automated fake news detection, benchmark data contamination—particularly at the semantic level—severely undermines evaluation validity by rewarding rote memorization over genuine reasoning. To resolve this crisis, this thesis establishes a contamination-aware evaluation paradigm: it introduces the Fuzzy Deep Hybrid Network (FDHN) and LIAR2 dataset to model truthfulness as a continuous spectrum, deploys the SSA and DCR frameworks to quantify semantic leakage and isolate true generalization, and pioneers LiveFact—a dynamic, time-aware benchmark designed to test epistemic humility against real-time information streams. Together, these contributions restore integrity, robustness, and objective verifiability to AI-driven fact-checking systems.

Supervisor: 
Prof. M-Tahar Kechadi
Email: 
cheng.xu1@ucdconnect.ie
Research Group: 
Insight Centre for Data Analytics