Can You Trust the Machine? Evaluating AI Output
A practical framework for evaluating AI output: check claims, spot fake citations, ask for reasoning, and know the model's limits.
Critical Thinking · Lesson 2
A practical framework for evaluating AI output: check claims, spot fake citations, ask for reasoning, and know the model's limits.
The last lesson made the case that AI output needs checking. This one makes that concrete: given an answer on your screen, what do you actually do? Without a method, "be sceptical" collapses into either trusting everything or trusting nothing. Neither serves you well.
The stakes are practical. People increasingly use AI answers to make decisions about health, money, law, and work. The difference between a useful tool and a costly mistake is often just a few minutes of verification applied to the claims that matter. This lesson gives you a repeatable framework so that checking becomes a habit rather than an afterthought.
The foundation of trust is not the AI's tone but independent confirmation. For any claim that matters, find it in a source that exists outside the model — a reference work, a primary document, a reputable site. If the claim is important and you cannot confirm it anywhere, treat that as a red flag, not a minor gap.
Because a model predicts plausible text, it can generate references that look perfectly real — correct-sounding authors, journals, page numbers, and dates — for papers or cases that do not exist. This is not an occasional glitch but a documented, recurring pattern. Never treat a citation as proof; treat it as a lead to check. Search the title. Open the source. Confirm it says what was claimed.
You can ask a model to show its reasoning or list its assumptions. This often surfaces weak steps you can then test — but note that the explanation is itself generated text, not a transparent log of how the answer was produced, so it too must be judged. Finally, know the model's limits: it has a training cutoff after which it knows nothing directly, and unless it is explicitly retrieving live sources, it cannot report current events, prices, or your private facts.
An AI tells you a specific vitamin cures a common illness, citing "a 2019 study in a major medical journal." Apply the framework. First, the claim is high-stakes (health), so it must be verified. Second, you check the citation: you search the journal and year and find no such study, or you find a study that says something far weaker. Third, you ask the model for its reasoning and notice it cannot point to a real, locatable source. The confident answer collapses under three minutes of checking — and you have avoided acting on a fabrication.
Not every use needs this scrutiny. If you ask an AI to rephrase your email, brainstorm names for a project, or explain a well-known concept you can immediately sanity-check, exhaustive verification is wasted effort. The framework scales with stakes: the higher the cost of being wrong, the more checking is warranted. Trusting AI for low-stakes drafting while verifying high-stakes facts is not inconsistency — it is proportion.
The 2023 Mata v. Avianca case, where lawyers submitted a brief citing court decisions that ChatGPT had entirely invented, was the first widely reported instance — but it was not the last. Legal researchers and journalists have since documented a growing list of court filings, in multiple countries, containing AI-generated citations to cases that do not exist. A number of judges have issued warnings or sanctions over such filings, and some courts have introduced standing orders requiring lawyers to disclose or verify AI use.
The pattern is consistent and instructive: the models produce citations that are formatted flawlessly and sound entirely credible, precisely because generating plausible text is what they do. The lesson generalises well beyond law. Any field that relies on specific sources — medicine, journalism, academia — faces the same risk, and the same remedy: a fabricated citation is caught not by reading it more carefully but by going and checking whether the source actually exists.
Take any substantive AI answer and run it through four questions. One: does this claim matter enough to verify? Two: can I confirm the key facts in a source outside the model? Three: are any citations real, and do they say what is claimed? Four: could this depend on information after the model's training cutoff? Write your answers. If any question fails, you know exactly where the answer is weak.
Think Like a Maester: A citation you have not checked is a claim, not evidence — no matter how perfectly it is formatted.
Trusting the machine is not all-or-nothing; it is a case-by-case judgment guided by a simple framework. Check claims that matter against sources outside the model. Treat every citation as a lead to verify, not proof, because models generate plausible references for things that do not exist — a pattern documented repeatedly in real court filings since Mata v. Avianca. Ask for reasoning, but judge it, since it too is generated text. Know the limits: a training cutoff and the absence of live retrieval mean a model cannot reliably speak to current or private facts. Scale your scrutiny to the stakes, and the machine becomes a powerful assistant whose work you can actually rely on.
Mark this lesson complete to track your progress.