Theory Thursday Why language models hallucinate
Art by @basilonmypizza: https://lnkd.in/eF8FkWzN - https://basilhefti.ch/
A Fata Morgana can be a pretty mirage. And for the thirsty traveler, a deadly trap: the shimmer looks like salvation but turns out to be an illusion all along.
Hallucinations in LLMs carry similar risks.
If you use an LLM to summarize a research paper, it may invent citations. If you ask for regulatory details, it might confidently output rules that never existed. Just as following a faulty satnav can send blindly trusting people driving into rivers, trusting a fluent but wrong answer can lead to costly - sometimes dangerous - mistakes.
So what are LLM hallucinations? In short: the confident generation of content that is incorrect, unverifiable, or unfaithful to reality.
And why do they happen? Current research is looking into this.
Manuel Cossio has introduced a taxonomy: hallucination can arise from biased, noisy or outdated data, from statistical and metric pressures, decoding errors, or vague context.
This aligns well with the findings by Kalai & Vempala, who identify systemic pressure: wrong incentives during training and evaluation lead to LLMs that prefer “guessing” over admitting uncertainty.
A third insight by Orgad et al. shows that models carry cues about truth in their hidden states. In other words: the problem happens in the decoding step and could be reduced if the hidden-state signals were used.
This research will certainly improve future LLMs. For now, we have to deal with hallucination. How? Two practical checks help: run a quick plausibility scan (ask: can this really be?), and request references or citations you can verify.
Importantly, both steps can also be automated, for example for crritical tasks such as cancer classification.
Like a Fata Morgana, hallucinations tempt with detail and structure. The danger lies in acting on them.
How do you prevent yourself from falling into the hallucination trap?
• Adam Tauman Kalai, Santosh Vempala. Why Language Models Hallucinate. https://lnkd.in/eDVfpNXt
• Orgad, Shoham, et al. LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations. https://lnkd.in/eaQuEXxt
• Cossio, Manuel. A Comprehensive Taxonomy of Hallucinations in Large Language Models. https://lnkd.in/eTexbvrN
• Happy Birthday paul brown
• Art: @basilonmypizza: https://lnkd.in/eF8FkWzN https://basilhefti.ch/