MSc thesis · Imperial · Apr–Sep 2026
0.801AUROC on held-out LLMs
vs 0.698 linear probe
- 13.7× lower latency
- 11.5× less compute
- 13 LLMs, 7B–27B
Amortized Trustworthiness
Small language models that predict when a large one is unsure, in one pass instead of eleven.
- SSituation
- Semantic entropy is one of the best signals that an LLM is hallucinating. But it needs about 11 sampled generations and entailment clustering for every answer, which is too slow and costly for real-time use.
- TTask
- Train a small proxy model that predicts a large model’s semantic uncertainty in a single pass and still works on LLMs it has never seen.
- AAction
- Built a reusable PyTorch dataset pipeline across 13 LLMs, covering stochastic generation, hidden-state extraction and DeBERTa entailment clustering. LoRA fine-tuned Llama-3.2-3B with PEFT on text and hidden-state features. Benchmarked against linear-probe, ridge and representation-alignment baselines using leave-one-LLM-out evaluation.
- RResult
- 0.801 AUROC vs 0.698 (linear probe) and 0.729 (aligned ridge), and 0.791 on unseen Qwen and Gemma models. Transfers from TriviaQA to SQuAD without retraining. Uncertainty goes from 11 LLM passes to one: 13.7× faster, 11.5× cheaper.
- PyTorch
- LoRA / PEFT
- Transformers
- Llama-3.2-3B
- DeBERTa
- Python