Google Scholar profile →

  1. Single LLM Debate, MoLaCE: Mixture of Latent Concept Experts Against Confirmation Bias

    Hazel Kim, Philip Torr

    Preprint

    Introduces MoLaCE, an inference-time method that mixes latent-concept experts so a single LLM can resist confirmation bias—emulating multi-agent debate more efficiently and reducing echo-chamber effects.

  2. Measuring what Matters: Construct Validity in Large Language Model Benchmarks

    Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Nicolaus Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr, Cozmin Ududec, Luc Rocher, Adam Mahdi

    NeurIPS 2025 Datasets & Benchmarks

    A systematic review of 445 LLM benchmarks with 29 experts, showing how weak construct validity undermines claims about safety and capability, and offering eight recommendations for more rigorous evaluation.

  3. Detecting LLM Hallucination through Layer-wise Information Deficiency

    Hazel Kim, Tom A. Lamb, Adel Bibi, Philip Torr, Yarin Gal

    EMNLP 2025

    Detects hallucinations at test time by tracking usable-information deficiencies across model layers—especially under ambiguous or unanswerable prompts—without extra training or architecture changes.

  4. How Ambiguous Are the Rationales For Natural Language Reasoning?

    Hazel Kim

    COLING 2025

    Studies how uncertain or ambiguous rationales affect reasoning performance, and proposes a simple way for models to choose among reasoning paths when rationale quality is inconsistent.

  5. ATHENA: Mathematical Reasoning with Thought Expansion

    JB. Kim, Hazel Kim, Joonghyuk Hahn, Yo-Sub Han

    EMNLP 2023

    Presents ATHENA, an attention-based architecture that expands candidate mathematical thoughts step by step, improving math-word-problem solving under limited or varied training signals.

  6. ALP: Data Augmentation Using Lexicalized PCFGs for Few-Shot Text Classification

    Hazel Kim, Daecheol Woo, Seong Joon Oh, Jeong-Won Cha, Yo-Sub Han

    AAAI 2022

    Generates syntactically diverse, label-preserving text with lexicalized PCFGs for few-shot classification, and pairs this with augmentation-aware train/validation splitting for stronger low-resource training.

  7. LST: Lexicon-Guided Self-Training for Few-Shot Text Classification

    Hazel Kim*, Jaeman Son*, Yo-Sub Han

    Arxiv

    Improves few-shot self-training by guiding pseudo-labels with a refined lexicon, reducing overconfident early errors when only a handful of labeled examples are available.