AI Hallucinations Explained: Why Chatbots Make Things Up
AI hallucinations explained: why chatbots guess instead of saying they don't know, what tests from 2023 to 2026 found, and 4 signs an answer may be made up.

In March 2023, lawyers in a New York personal-injury case filed a brief citing court decisions that nobody could find. ChatGPT had invented them, quotations included, and when the lawyer who had used it asked whether one of the cases was real, it said the case existed. That June a federal judge sanctioned the lawyers and their firm.1
AI hallucinations are statements a chatbot presents as fact that are false or unsupported: an invented case, a wrong date, a real source cited for something it never said. Language models hallucinate because both the way they learn to predict text and the tests used to rank them reward a confident guess over “I don’t know,” a 2026 Nature paper led by OpenAI researchers argues.2 How often they happen depends mostly on the task. In practice, the riskiest parts of any answer are the specifics it supplies from memory: names, figures, dates, quotations and sources. Four warning signs close this page, and the everyday habits for working with a chatbot are in a beginner’s guide to AI assistants.
What AI hallucinations are, in NIST’s terms
The US National Institute of Standards and Technology (NIST) calls the problem confabulation: generative AI producing and confidently presenting false or erroneous content. Its 2024 profile of generative AI risks also counts output that strays from what the user supplied or contradicts what the system said earlier in the same conversation.3
Definition
An AI hallucination is a false or unsupported statement that an AI system presents as fact, such as an invented source, a wrong figure, or a claim that contradicts the material it was given.
NIST treats confabulation as a natural result of how these systems work. A language model generates text that follows the statistical patterns of its training data, predicting the next word, and the same process that produces accurate, consistent answers can produce inaccurate or internally inconsistent ones.3 Fluency and truth come out of one machine, so a polished answer tells you nothing about whether it is right.
Two illustrative cases show the range. Ask for a summary of a long report and the summary may be accurate except for one date that appears nowhere in the report. Ask for sources on a topic and you may get a neatly formatted list in which one title was never published.
NIST adds a warning that matters for busy readers: a system can also confabulate the logic or citations meant to justify its answer, which can lead people to trust it more than they should.3 That is why the specific details deserve the most suspicion: a plausible guess gets the shape of an answer right and the particulars wrong, and a made-up justification makes it look more trustworthy, not less.
Why chatbots make things up: a guess scores better than “I don’t know”
Chatbots make things up because both main stages of building them push toward a confident answer, according to a 2026 Nature paper by Adam Tauman Kalai and colleagues. Learning from text leaves unavoidable gaps for one-off facts, and the tests that rank models then reward filling those gaps with a guess.2
The first pressure comes from pretraining, when a model learns the patterns of language from a huge body of text. Regular patterns, such as grammar and spelling, can be learned well. A one-off detail, such as a private person’s birthday mentioned once in an obituary, offers no pattern to learn. The authors show that such errors arise even if every sentence in the training data were true, while facts repeated often, such as country capitals, rarely trip models up.2
The second pressure comes from grading. Most popular benchmarks score an answer as right or wrong and give nothing for admitting uncertainty, so a model that guesses when unsure outscores one that says it does not know. The authors compare it to a student facing a hard exam question: with no penalty for wrong answers, guessing is the winning strategy.2
- The guesser answers everything. Some guesses land, one is wrong, and the wrong answer costs nothing.
- The honest model says it does not know when unsure, and scores nothing for saying so.
- Right-or-wrong scoring puts the guesser on top, which, the Nature paper argues, rewards models for guessing.
The result looks like this in practice. Asked in November 2025 what the abbreviation PGGB stands for, three popular chatbots each gave a different, confident expansion, and all three were wrong; none said it did not know or asked for context.2
Two limits apply. The paper is mainly a mathematical argument backed by stylized experiments, not a trial, and three of its four authors are or were OpenAI employees. Its one hint for users comes from the authors’ own test: when each question spelled out the scoring, including partial credit for answering “I don’t know,” four leading models declined more often as the stakes rose.2 We found no test of whether a similar line in an everyday chat does the same, so treat it as a nudge, not a safeguard. The dependable lesson is simpler: expect the least reliable answers on rare, specific facts, the ones a model saw once or never.
How often AI hallucinations happen depends on the task
No single hallucination rate exists. Published tests from 2023 to 2026 range from a few percent of summaries to most answers, depending mainly on whether the model works from material it was given or from memory, and on how obscure the facts are. Each test also defines a hallucination its own way, so the figures below show a pattern, not a league table.
| Task tested | What the test found | Evidence |
|---|---|---|
| Summarizing a supplied article (108 models, leaderboard updated 22 September 2026) | 1.8% to about 24% of summaries contained a hallucination, as judged by the company’s own detection model | Observational (model test), limited; company research4 |
| Short, hard fact questions (SimpleQA, 2024; 4,326 questions written to be hard for GPT-4) | GPT-4o answered 61% incorrectly and declined only 1% | Observational (model test), moderate; OpenAI preprint, not peer reviewed5 |
| Reference lists for short literature reviews (2023; 636 references) | 55% of GPT-3.5’s references and 18% of GPT-4’s were fabricated | Observational (model test), limited; one study6 |
| Direct questions about US federal court cases (four models from 2023) | Hallucinations on 58% to 88% of questions, depending on the model | Observational (model test), moderate7 |
| SimpleQA again, two newer OpenAI models (reported 2026) | o4-mini was wrong on 77% and declined 3%; GPT-5-mini was wrong on 21% and declined 63% | Observational (model test), limited; run by OpenAI researchers2 |
Summaries probably score best because the facts sit in front of the model. Reference lists, case law and trivia sit at the other end: they depend on one-off details, the kind the Nature paper links to unavoidable errors.2
The last row holds the quieter lesson. GPT-5-mini answered slightly fewer questions correctly than o4-mini, yet made far fewer errors, because it declined most of the questions it was unsure about.2 Fewer hallucinations came from more honest refusals, not from knowing more.
- Myth
- The newest models have stopped hallucinating.
- Fact
- On a company-run leaderboard updated in September 2026, all 108 models tested had some summaries flagged for hallucination by its detection model, even with the source text in front of them.
The lesson for your own work: move the facts into the chat. Asking “What did this year’s staff survey say about workload?” with the survey pasted in is a summarizing task. Asking the same question without it is a memory task about a one-off fact, and the risk rises.
Do search and source links stop AI hallucinations?
Tools that answer from search results or a source database hallucinated less in testing, but far from never. In the first preregisteredpreregistration: Filing a study's hypotheses and analysis plan publicly before the data are collected or examined. It stops a planned test from being rewritten once the results are in, so a reader can tell a real confirmation from a hunt through the data.Full entry in the glossary test of AI legal research tools, which answer from databases of real case law, Stanford and Yale researchers found fewer hallucinations than from a general chatbot, but still a substantial share. These tools use retrieval-augmented generation: they first search a trusted collection, then write an answer from what they retrieved. Some vendors had marketed the method as avoiding or even eliminating hallucinations, without publishing evidence.8
The study
Moderate evidence
Three legal research tools with real case law behind them, put to 202 questions
Researchers put more than 200 legal questions to Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI, and to GPT-4 for comparison. Each legal tool hallucinated on 17 to 33 percent of questions; GPT-4, with no legal database behind it, hallucinated more often than any of them. An answer counted as hallucinated if it stated something false or cited a source that did not support its claim.8
The second half of that definition is the one to carry into your own work. A citation to a real case, report or web page proves only that the source exists, not that it backs the answer’s claim, and the researchers found that spotting such errors often took close reading of the cited source. The main caveats are that the test covered one field, US law, and the tools as they stood in 2024; the products may have changed since.8
The same habit applies whether the tool is a specialist database or a general chatbot with web search switched on.
When hallucinations leave the chat: courts, refunds and a growing case list
Hallucinations turn costly when someone acts on them. A database kept by the legal researcher Damien Charlotin counted 2,077 court and tribunal decisions worldwide involving hallucinated content, mostly fake citations, as of 23 September 2026.9
In the New York case from the opening, Mata v. Avianca, the judge wrote that nothing was inherently improper about using a reliable AI tool. The fault lay in submitting fake opinions and then standing by them after the other side and the court questioned whether they existed.1
But existing rules impose a gatekeeping role on attorneys to ensure the accuracy of their filings.
In Canada, a company could not pass the blame to its bot either. In November 2022, Air Canada’s website chatbot told a customer whose grandmother had just died that they could claim a bereavement fare after travelling; the airline’s own bereavement page, which the chatbot linked to, said the policy did not cover requests made after travel. In 2024 a British Columbia tribunal rejected the airline’s suggestion that the chatbot was responsible for its own actions and ordered it to pay the fare difference.10
Both cases point the same way. In each, responsibility for the AI’s answer stayed with the people behind it: the lawyers who filed it and the airline whose chatbot gave it. The check has to come from outside the chatbot: the lawyer’s attempt to verify, asking ChatGPT whether its case was real, only produced another confident yes.1 In the Air Canada case the correction sat one click away, on the page the chatbot itself linked to, yet the tribunal found the customer’s reliance on the chatbot reasonable and held the airline responsible.10 How employers are preparing staff for risks like these is covered in what AI literacy asks of employees.
Can AI hallucinations be fixed?
AI hallucinations can be reduced, the 2026 Nature paper argues, but not by adding knowledge alone. Some questions will stay unanswerable, such as an unlisted birthday, so scoring on accuracy alone still rewards a model for guessing when it is unsure. The authors’ proposed fix is to change the scoring so that models earn credit for admitting uncertainty.2
Researchers are also building ways to flag likely hallucinations. In a 2024 Nature study, a University of Oxford team had a model answer the same question several times and checked whether the answers agreed in meaning. Answers that scattered pointed to made-up content, and the method worked on tasks it had not been tuned for. It cannot catch errors a model makes consistently, the authors note, such as a misconception learned from its training data.11
The same idea works as a rough home test. Ask for a little-known founding date in three fresh chats and get three different years, and at least two are wrong. Get the same year three times and you still need a source, because a model can repeat a mistake it learned.
Warning signs that an answer may be made up
Hallucinations cluster where the research above predicts: rare specifics, citations, memory-based answers about niche topics, and answers that never express doubt. A full fact-checking routine is a subject of its own; these four signs tell you when one is needed.
- The answer names specifics you did not supply, such as exact figures, dates, quotations, case names or paper titles. These are the details a plausible guess gets wrong.
- It cites sources. Open each one: the legal tools above cited real cases for claims those cases did not support.8
- It answers from memory about something niche. Paste in the document, or ask the assistant to search, and compare the two answers.
- It never hedges, or a fresh chat gives a different answer to the same question. Scattered answers are the signal the Oxford method looks for.11
The bottom line
AI hallucinations are not a bug the next update will remove: they follow from how language models learn and how they are scored, and they are most common on rare facts, citations and specialist questions. Keep the facts in the chat where you can, treat every specific detail as a claim to check, and read a real-looking source as a lead, not as proof.
Frequently asked questions
Is hallucination the right word for AI errors?
It is the everyday word, but not the only one. NIST, the US standards agency, uses confabulation in its 2024 generative AI profile and notes that the same errors are colloquially called hallucinations or fabrications. Researchers also use it more narrowly: a 2024 Oxford study in Nature reserves confabulation for arbitrary, incorrect answers that can change from one attempt to the next.
Do AI chatbots lie on purpose?
Not in the ordinary case. The 2026 Nature paper by Kalai and colleagues describes hallucinations as guesses made under uncertainty, which training and testing reward. A 2024 Nature study from the University of Oxford treats a model that lies in pursuit of a reward as a separate mechanism from confabulation, and says lumping such different failures under one word is unhelpful.
Can a chatbot tell when it is guessing?
Only partly. In OpenAI's 2024 SimpleQA tests, answers given with higher stated confidence were more often right, but the models consistently overstated their confidence. The Oxford team's 2024 Nature study used a different signal, whether several answers to the same question agree in meaning; it flags one kind of made-up answer but misses errors a model repeats consistently.
Sources
- Mata v. Avianca, Inc., No. 22-cv-1461 (PKC), Opinion and Order on Sanctions (Document 54). US District Court for the Southern District of New York, Judge P. Kevin Castel (22 June 2023)
- Evaluating large language models for accuracy incentivizes hallucinations. Kalai, A. T., Nachum, O., Vempala, S. S. & Zhang, E. (2026). Nature, 653, 1047-1051
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). National Institute of Standards and Technology, US Department of Commerce (July 2024)
- Hallucination Leaderboard (HHEM-2.3, last updated 22 September 2026). Vectara (company-run benchmark, accessed 24 September 2026)
- Measuring short-form factuality in large language models. Wei, J., Karina, N., Chung, H. W., et al. (2024). OpenAI, arXiv preprint 2411.04368
- Fabrication and errors in the bibliographic citations generated by ChatGPT. Walters, W. H. & Wilder, E. I. (2023). Scientific Reports, 13, 14045
- Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Dahl, M., Magesh, V., Suzgun, M. & Ho, D. E. (2024). Journal of Legal Analysis, 16(1), 64-93; full text read in arXiv preprint 2401.01301, version 2
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C. D. & Ho, D. E. (2025). Journal of Empirical Legal Studies, 22(2), 216-242
- AI Hallucination Cases Database. Charlotin, D. (last updated 23 September 2026)
- Moffatt v. Air Canada, 2024 BCCRT 149. Civil Resolution Tribunal, British Columbia, Canada (14 February 2024)
- Detecting hallucinations in large language models using semantic entropy. Farquhar, S., Kossen, J., Kuhn, L. & Gal, Y. (2024). Nature, 630, 625-630
How we researched this
We read NIST's 2024 generative AI profile, a 2026 Nature paper on why language models hallucinate, published model tests and benchmarks from 2023 to 2026, and two court and tribunal decisions, fetched in September 2026. We preferred peer-reviewed studies and primary legal documents, and label company-run benchmarks as such. One legal study was read in its accepted arXiv version because the journal page was blocked. Main limitation: every test uses its own definition and models change quickly, so rates cannot be compared directly or carried forward.



