π¬ RESEARCH
A Lie Detector Test for Language Models: Reading Knowledge a Model Won't Reveal
"Large language models can hold knowledge they do not report. A model may sandbag on a capability evaluation, or answer against what it internally knows, and its outputs alone cannot tell whether it is hiding an answer or simply does not have one. We borrow the Concealed Information Test, a forensic..."