Technology

Is ChatGPT Accurate? It Depends on the Task, and the Measured Error Rates Run From Small to 69%

There is no single accuracy figure for ChatGPT. Studies that checked its answers one by one found mostly correct replies to doctors’ questions, invented references in up to half of citations, and hallucinated answers to most questions about specific court cases. The task, the model version and whether it can search decide which of those you get.

A laptop with a blank screen on a plain desk beside an open printed reference book and a pencil, in daylight

Short answer: there is no single accuracy number, and anyone quoting one is quoting a single test. Measured results depend on the task, the model version and whether the assistant can search. Physicians grading 284 medical answers in 2023 gave a median score of 5.5 out of 6. In April 2023, 55% of the references GPT-3.5 produced for literature reviews did not exist, against 18% for GPT-4. Asked specific questions about real federal court cases, a late-2023 GPT-4 model hallucinated 58% of the time and GPT-3.5 69%. In a 2025 broadcaster study of four AI assistants, one in five answers about the news had a major accuracy problem. It is usually right on common knowledge and least reliable exactly where a precise, checkable detail is needed.

Every figure in this article belongs to a named model at a named date. Models change several times a year, so the studies below describe the shape of the problem better than they describe the product available today.

How accurate is ChatGPT, by task

The honest way to answer the question is a table, not a percentage. These are results from studies in which people checked the output item by item.

Task Study Model and date What was measured Result
References for a literature review Walters & Wilder, Scientific Reports GPT-3.5 and GPT-4, generated April 2023 Share of generated citations that were fabricated (636 in total) 55% (GPT-3.5), 18% (GPT-4)
Same task, real references only Walters & Wilder GPT-3.5 and GPT-4, April 2023 Real citations containing substantive errors 43% (GPT-3.5), 24% (GPT-4)
Medical questions written by doctors Goodman et al., JAMA Network Open GPT-3.5, with 44 questions repeated on GPT-4, January–May 2023 Physician-rated accuracy, 1–6 scale, 284 questions Median 5.5, mean 4.8
Questions about specific federal court cases Dahl et al., Journal of Legal Analysis GPT-4 (gpt-4-1106-preview), GPT-3.5 (gpt-3.5-turbo-0613), PaLM 2, Llama 2; published 2024 Hallucination rate, pooled across tasks with a checkable reference answer 58% (GPT-4), 69% (GPT-3.5), 72% (PaLM 2), 88% (Llama 2)
Legal research with retrieval Magesh et al., Journal of Empirical Legal Studies Lexis+ AI, Westlaw AI-Assisted Research, Ask Practical Law AI; published 2025 Share of answers containing a hallucination 17–33%
Questions about the news EBU and BBC, News Integrity in AI Assistants Free versions of ChatGPT, Copilot, Gemini and Perplexity, answers generated May–June 2025 Journalist-reviewed answers, more than 3,000 45% with at least one significant issue; 20% with a significant accuracy issue
Facts about people (PersonQA) OpenAI system card, company-reported o3, o4-mini and o1, April 2025 Hallucination rate on the company’s internal benchmark 33% (o3), 48% (o4-mini), 16% (o1)

Two things stand out. The spread is enormous: the same family of models scores near the top of a physician’s scale on one task and fails most questions on another. And the worst results cluster around a particular kind of request: a specific, verifiable item such as a citation, a case holding or a fact about a named person, where an answer that merely sounds plausible is simply wrong.

Why there is no single accuracy number

A percentage only means something if you know what was asked. Three variables move the result more than anything else.

The task. A question whose answer appears thousands of times in the training text, such as how photosynthesis works or what a mortgage escrow account is, is a different problem from recalling the page numbers of a 2014 journal article. The first is pattern the model has seen repeatedly. The second is a single fact that it either retained exactly or did not, and when it did not, it tends to produce something of the right shape.

The model version. In the Walters and Wilder study the fabricated-citation rate fell from 55% to 18% between GPT-3.5 and GPT-4. In the Goodman study, 44 medical answers regenerated with GPT-4 rose from a mean of 5.2 to 5.7 out of 6. Newer has not always meant fewer errors, though. OpenAI’s own April 2025 system card reported that its o3 model hallucinated on 33% of the questions in one internal benchmark where the earlier o1 model did so on 16%, and the company wrote that more research was needed to understand why.

Whether it can search. A model answering from memory is reconstructing; a model reading a retrieved page is summarising. OpenAI’s help documentation says search lets ChatGPT look up current or niche information on the web and provide cited answers. Retrieval helps, but the evidence below shows it does not remove the problem.

Does ChatGPT make things up?

Yes, and the company says so. OpenAI’s help page on the subject states that ChatGPT can produce incorrect or misleading output, can sound confident while wrong, and lists the typical forms: incorrect definitions, dates or facts, and fabricated quotes, studies, citations or references to sources that do not exist. The industry term is hallucination.

The mechanism explains where it happens. A language model produces the most plausible continuation of the text in front of it; unless it uses a search tool, it is not looking anything up. For a fact repeated widely in its training text, the plausible answer and the true one coincide. For an arbitrary detail such as a page number, a docket number or a PMID, a plausible answer of the right shape is easy to produce and usually wrong, which matches the pattern in the table above.

Does ChatGPT cite real sources?

This is the best-measured failure, because a reference either exists or it does not.

Walters and Wilder had GPT-3.5 and GPT-4 write short literature reviews on 42 topics across disciplines in the first week of April 2023, which produced 84 papers and 636 citations (222 of them from GPT-3.5), and then looked every citation up. Their findings, published in Scientific Reports in 2023:

Measure GPT-3.5 GPT-4
Citations that were fabricated 55% 18%
Real citations with substantive errors 43% 24%

The authors called GPT-4 a major improvement and also said that problems remained, which is a fair summary. Nearly one in five invented, and about a quarter of the real ones wrong in some substantive detail, is not a standard anyone would accept from a research assistant.

A smaller study in Cureus in 2023 found the same pattern. Bhattacharyya and colleagues had GPT-3.5 write 30 short medical papers containing 115 references: 47% were fabricated, 46% were real but inaccurate, and only 7% were both real and accurate.

These models are now several generations old, and neither paper describes a search tool being used; no figure here should be read as the current rate.

With search turned on, the assistant can link to pages that exist. That changes the question from "is this source real" to "does this source say what the answer claims", which still has to be checked by opening it.

Is ChatGPT accurate for medical advice?

The most careful early evaluation is Goodman and colleagues in JAMA Network Open. Thirty-three physicians across 17 specialties wrote 284 questions and graded the answers on a six-point accuracy scale, between January and May 2023. The answers came from GPT-3.5; one set of 44 questions was later repeated on GPT-4.

  • Median accuracy was 5.5 out of 6, with a mean of 4.8.
  • Easy questions scored a median of 6.0, medium 5.5 and hard 5.0.
  • Yes-or-no questions scored a median of 6.0; descriptive questions scored 5.0.
  • In the largest question set, 180 multispecialty questions, 36 answers (20%, by the paper’s own count) scored 1 or 2, at the completely incorrect end of the scale. When 34 of them were asked again 8 to 17 days later, the median rose from 2.0 to 4.0.

The authors’ conclusion was that the chatbot gave largely accurate information with important limitations. Both halves matter. A median of 5.5 is a good result. But a mean well below the median indicates a tail of badly wrong answers, one in five in the multispecialty set, delivered in the same confident register as the good ones. The re-query finding adds a second problem: the same question can receive a noticeably different answer on a different day.

The questions were written by doctors, who phrase a clinical question precisely and can recognise a wrong answer. A member of the public describing symptoms vaguely is running a different experiment, which this study did not test. It also predates every current model.

For general understanding, such as what a term on a lab report means or which questions to bring to an appointment, the evidence suggests the answers are mostly sound. For a decision about a dose, an interaction or whether a symptom needs urgent care, a mostly sound average is not the relevant statistic; the tail is. Health queries also raise a separate issue of what happens to the information typed in.

Law and news: the two hardest tests so far

Case law. Dahl and colleagues asked public models specific, verifiable questions about randomly selected federal court cases and compared the answers with the record. In the paper published in the Journal of Legal Analysis in 2024, pooled across the tasks that had a checkable reference answer, GPT-4 (the gpt-4-1106-preview snapshot) hallucinated 58% of the time, GPT-3.5 (gpt-3.5-turbo-0613) 69%, PaLM 2 72% and Llama 2 88%. The paper does not give its testing dates. The models also often accepted a false premise built into a question instead of correcting it, and systematically overestimated their confidence relative to how often they hallucinated.

These were general chatbots without access to a legal database. Magesh and colleagues ran a preregistered evaluation of three commercial legal research tools that retrieve real documents before answering (Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI), first posted in 2024 and published in 2025. Hallucinations were reduced compared with GPT-4 used as a general-purpose chatbot, but each tool still hallucinated between 17% and 33% of the time, despite vendor claims of eliminating the problem. This is the clearest evidence available that giving a model sources lowers the error rate without bringing it near zero. These products are not ChatGPT, and the figures should not be transferred to it.

News. In October 2025 the European Broadcasting Union and the BBC published News Integrity in AI Assistants. Journalists at 22 public service media organisations in 18 countries, working in 14 languages, evaluated more than 3,000 answers to news questions from the free versions of ChatGPT, Copilot, Gemini and Perplexity, generated in late May and early June 2025. 45% of the answers had at least one significant issue, 31% had significant sourcing problems, and 20% had significant accuracy issues. These are pooled figures for four assistants, and this article does not break them down by product. The answers were assessed on their sourcing as well as their accuracy, and the report names sourcing as the single biggest cause of significant issues.

5.5 / 6Median physician rating, 284 medical answers, 2023
18%Fabricated citations from GPT-4, April 2023
58%GPT-4 hallucination rate on federal case-law questions, 2024 paper
1 in 5News answers with a major accuracy issue, 2025

Can ChatGPT do math?

This article has no peer-reviewed error rate for arithmetic that it could verify, so it does not give one. What can be said is structural. A language model predicts text; it does not calculate unless it hands the problem to a tool. OpenAI’s documentation describes a code interpreter, also labelled data analysis, as the feature that supports accurate calculation, which is an indirect acknowledgement that the model on its own is not the calculator.

The sensible practice follows from that. For any number that matters, ask the assistant to run the calculation as code and show it, or redo the arithmetic independently.

Low-risk and high-risk questions

The studies sort questions into two groups fairly cleanly.

Lower risk Higher risk
Explaining a well-established concept Citations, DOIs, page numbers, quotations
Summarising or rewording text you supply Details of a specific court case, statute or regulation
Brainstorming, outlining, drafting Facts about a named, not very famous person
Translating everyday text Recent events and anything after the training cut-off
Explaining what code does when you can run it Exact figures, dates, prices and statistics
Suggesting questions to ask a professional Doses, drug interactions, legal deadlines, tax rules

The pattern on the left is that the answer is either common knowledge, grounded in text the user provided, or easily tested. The pattern on the right is that the answer hinges on one exact fact, the user cannot tell by reading whether it is right, and being wrong has a cost.

What is not worth doing

Asking it whether it is sure. The Dahl study found that models systematically overstated their confidence relative to their actual hallucination rate. A model asked "are you certain?" may apologise and change a correct answer, or restate a wrong one. Its confidence is text, not a measurement.

Asking it to verify its own references. A model that invented a citation can equally invent a confirmation of it. Verification has to happen outside the chat.

Trusting a quoted accuracy percentage. Figures such as "ChatGPT is 90% accurate" circulate widely without a task, a model version or a date attached. Without those three, the number has no meaning, and none of the studies reviewed here supports a single headline figure.

Assuming the latest model has fixed it. The trend from GPT-3.5 to GPT-4 was a clear improvement on the tasks above, but the April 2025 system card shows a newer model hallucinating more than its predecessor on one benchmark. Improvement is the general direction, not a guarantee for each release.

Assuming search makes it safe. The legal-tool and news studies both tested systems with access to real documents and still found substantial error rates.

What to actually do

OpenAI’s own guidance is to treat output as a first draft rather than a final source and to verify quotes, data, technical details and references. The steps below make that concrete.

What the answer contains How to check it Time
A citation or study Search the exact title or paste the DOI into a resolver; confirm authors, year and journal 1–2 minutes
A claim attributed to a source Open the linked page and find the sentence 1–3 minutes
A number or statistic Trace it to a primary source; if none can be found, do not use it A few minutes
A calculation Have it run as code, or redo it on a calculator Under a minute
A legal, medical or tax point Confirm against the official source or a qualified professional Varies
A recent event Use search mode and read the original report A few minutes

Ask for sources, then open them. Requesting sources is only half the step. The EBU and BBC study found sourcing was the single biggest cause of significant issues, which means a link being present does not show that the claim is in it.

Use search-enabled mode for anything factual and current. The model has a knowledge cut-off, and without search it will answer from memory about a world that has moved on.

Ask the same question twice. The Goodman study showed answers shifting between sessions. If two fresh attempts disagree on a fact, neither should be trusted without checking.

Count the checking time. The time saved by a fast answer is partly spent confirming it, and for high-risk questions it can be quicker to go to the primary source first. That trade-off sits inside a wider one about how much of a working day to hand to tools and what constant switching does to attention. The resource side of each query is covered separately in how much water AI uses.

Questions people ask

Is ChatGPT accurate? Often, but not uniformly. Physician-graded medical answers in 2023 had a median score of 5.5 out of 6, while GPT-4 fabricated 18% of citations in 2023 and hallucinated on 58% of specific case-law questions in a 2024 paper. Accuracy depends on the task, the model version and whether search is used.

How accurate is ChatGPT? No single percentage exists. Published studies report results by task, ranging from mostly correct on general medical questions to mostly wrong on specific legal case details. Any headline figure that does not state the task, the model version and the date of testing should be treated as unsupported.

How often is ChatGPT wrong? It depends on the question. In one 2025 study of four AI assistants, including ChatGPT, 20% of news answers had a significant accuracy issue. On OpenAI’s own PersonQA benchmark, reported in April 2025, hallucination rates ranged from 16% to 48% across three models.

Is ChatGPT always right? No. OpenAI’s documentation states that ChatGPT can produce incorrect or misleading output and can sound confident while doing so. Every independent evaluation reviewed here found errors, including in the newest models tested at the time, and the errors are not flagged in the text.

Is ChatGPT reliable? Reliable enough for explaining established concepts, drafting, and working with text the user supplies. Not reliable as the sole source for citations, legal specifics, exact figures or recent events. One study also found that the same question could receive a differently graded answer when asked again days later.

Does ChatGPT make things up? Yes. The term is hallucination, and OpenAI lists fabricated quotes, studies and citations among the known forms. In an April 2023 test, 55% of the references GPT-3.5 generated and 18% of those from GPT-4 did not exist. The invented material looks the same as the real material.

Why does ChatGPT give wrong answers? A language model predicts plausible text; it does not look facts up unless it uses a search tool. For widely repeated facts the plausible answer is usually the true one. For exact details such as citations, case numbers or dates, a plausible-looking answer is easy to produce and often wrong.

Is ChatGPT accurate for medical advice? In a 2023 study, physicians rated answers to 284 medical questions at a median of 5.5 out of 6, but 36 of the 180 answers in the largest question set scored 1 or 2 out of 6. That supports use for general understanding, not for decisions about doses, interactions or urgent symptoms, which belong with a clinician.

Does ChatGPT cite real sources? Not dependably when answering from memory: a 2023 study (models tested April 2023) found 18% of GPT-4 citations fabricated and 24% of the real ones containing substantive errors. With search enabled it links to existing pages, but a 2025 broadcaster study found significant sourcing problems in 31% of assistant answers.

How to check if ChatGPT is right? Verify outside the chat. Open every cited source and find the claim in it, look up references by title or DOI, redo calculations independently, and confirm legal, medical or financial points against official sources. Asking the model whether it is sure is not a check.

Vincent Brooks

Builds digital products for a living and writes about what that work reveals: how attention is engineered, what our devices can actually measure, and which of it survives a closer look.

This article summarises published evaluations of AI assistants for general information. The results describe specific model versions at specific dates and may not reflect current products. It is not medical, legal or financial advice, and nothing here replaces checking important information against a primary source or with a qualified professional.

References

  1. Walters, W.H., & Wilder, E.I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, 14045. doi:10.1038/s41598-023-41032-5
  2. Goodman, R.S., Patrinely, J.R., Stone, C.A., et al. (2023). Accuracy and Reliability of Chatbot Responses to Physician Questions. JAMA Network Open, 6(10), e2336483. doi:10.1001/jamanetworkopen.2023.36483
  3. Dahl, M., Magesh, V., Suzgun, M., & Ho, D.E. (2024). Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models. Journal of Legal Analysis, 16(1), 64–93. doi:10.1093/jla/laae003
  4. Magesh, V., Surani, F., Dahl, M., Suzgun, M., Manning, C.D., & Ho, D.E. (2025). Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools. Journal of Empirical Legal Studies, 22(2), 216–242. doi:10.1111/jels.12413
  5. European Broadcasting Union and BBC (2025). News Integrity in AI Assistants. Published 21 October 2025. ebu.ch
  6. OpenAI (2025). OpenAI o3 and o4-mini System Card, 16 April 2025 (company-reported evaluations). openai.com
  7. OpenAI Help Center. Does ChatGPT tell the truth? (company documentation). help.openai.com
  8. Bhattacharyya, M., Miller, V.M., Bhattacharyya, D., & Miller, L.E. (2023). High Rates of Fabricated and Inaccurate References in ChatGPT-Generated Medical Content. Cureus, 15(5), e39238. doi:10.7759/cureus.39238

The weekly readout

One email each Thursday: what we tested, which claim collapsed under a closer look, and the one number worth paying attention to.

NO TRACKING PIXELS · UNSUBSCRIBE IN ONE CLICK