Compliance 12 min read

Human Judgment Is Scarce. AI Is Making It Scarcer.

J

Jared Clark

September 05, 2026

Two more AI-in-the-workplace stories broke this year, and both followed a shape that's become familiar to anyone in a quality or compliance function: a person had a clear chance to check an AI-generated claim against a primary source, and didn't. The tool wasn't the failure. The missing verification step was.

That pattern is why "human judgment" has stopped being a phrase people use in passing and started being something quality managers, compliance leads, and engineers are asked to defend in front of an auditor. Every AI vendor pitch this year comes with some version of "keep a human in the loop." Almost none of them define what that human is actually supposed to do differently than approve the output on sight.

I want to answer that directly, for the people who have to answer it on the job: what is judgment doing that AI output can't do, why is it in short supply right now, and what does applying it actually look like inside a workflow that has a deadline attached to it.

Why Judgment Is Scarce, Not Just Skippable

Judgment is scarce for a reason that has nothing to do with AI: it can't be manufactured on demand. A reviewer builds the capacity to catch a bad root-cause claim or a mismatched batch number by doing it wrong a few times early in a career and being corrected. That capacity doesn't transfer between people, it doesn't compress into a document, and it doesn't scale the way a language model scales. You can spin up a thousand more instances of an AI system this afternoon. You cannot spin up a thousand more reviewers with ten years of pattern recognition behind their eyes.

AI output, by contrast, is abundant and getting cheaper every quarter. That asymmetry is the actual scarcity problem: the thing that has to catch the machine's mistakes is fixed in supply, while the volume of machine output it has to catch mistakes in keeps climbing. A compliance team that used to review twenty deviation summaries a week might now be handed eighty AI-drafted ones, written in the same confident house style regardless of whether the underlying claim is true. The reviewers didn't get faster. The queue did.

That's the scarcity the title of this piece is pointing at. It isn't that judgment is rare in some abstract, philosophical sense. It's that the supply of trained, experienced reviewers is fixed in the short run, AI can't be trained to replace that supply, and the demand on that supply is rising faster than most quality functions have adjusted their staffing or their review procedures to match.

What Judgment Is, Precisely

Judgment isn't the same thing as knowledge, and it isn't intelligence either. It's the capacity to weigh incomplete or conflicting information against a specific context and a set of consequences that matter to real people, and to decide, knowing you might be wrong, and to answer for that decision afterward.

A language model doesn't do the last part. It produces a statistically likely continuation of a prompt. It has no memory of having been wrong before in a way that cost it something, and nothing at stake if the sentence it generates is false. Accountability isn't a feature that can be added to these systems with better prompting. It's the thing that makes judgment mean anything in the first place, and it lives entirely on the human side of the transaction.

Two Cases Worth Knowing in Detail

Two rulings from the past few years get cited constantly, and it's worth being precise about what actually happened in each, because the precision is the point.

In Mata v. Avianca, Inc. (S.D.N.Y., No. 22-cv-1461), two attorneys submitted a federal court brief citing more than half a dozen decisions that turned out not to exist, generated by a chatbot they had used for research and never checked against a case reporter. In June 2023, U.S. District Judge Kevin Castel sanctioned the attorneys and imposed a $5,000 fine. The verification step that would have caught the error, pulling the actual case and reading it, took minutes and was skipped by people whose entire profession is built on doing exactly that.

In Moffatt v. Air Canada (2024 BCCRT 149), decided by Canada's Civil Resolution Tribunal in February 2024, an airline chatbot told a grieving customer he could apply for a bereavement fare discount after his flight, which contradicted the airline's actual policy. Air Canada argued the chatbot was a separate source of information the airline wasn't responsible for. The tribunal rejected that outright and held the airline accountable for anything its agents tell a customer, human or automated.

Neither failure involved unusual technology or a particularly devious misuse of it. Both involved an obvious point where a person could have checked a specific, falsifiable claim and didn't. That pattern now has its own public record: the AI Hallucination Cases database, a tracker maintained by French legal academic Damien Charlotin, logs court filings around the world that cite fabricated or misstated AI-generated content, and it is growing steadily as more of these filings surface. It exists because the pattern in Mata wasn't a one-time event. It's a recurring failure mode with a name and a paper trail.

A Checked Path and an Unchecked Path

Here's what the difference between "using AI" and "skipping judgment" actually looks like inside a quality or compliance workflow, using a routine task most reviewers will recognize: an AI-drafted investigation summary moving toward disposition.

Step Unchecked path Checked path
Draft generated AI writes the root-cause narrative from the ticket and attached data Same draft, treated as a starting hypothesis, not a conclusion
Root-cause claim Accepted because it reads coherently and cites the right equipment ID Traced back to the actual sensor log or batch record it claims to summarize
Supporting citations Assumed accurate because the format matches prior real investigations Spot-checked against the source document, not just the summary
Disposition decision Signed off on the strength of a clean-looking draft Signed off after the specific numeric claims have been independently confirmed
Accountability if wrong Diffuse. No one recalls approving the specific claim, only the document Clear. The reviewer who verified it can say what was checked and how

Nothing in the checked path requires rejecting the tool. It requires adding roughly five minutes of source-checking before a signature, at the specific point where a wrong answer is expensive to have missed. That's the whole discipline. It isn't more complicated than that, and it isn't less important than that either.

A workable decision rule follows directly from this: if an AI-generated claim can be verified against its primary source in under five minutes, verify it before it moves forward. If it can't be verified that fast, that slowness is itself the signal to escalate rather than a reason to wave it through.

What Judgment Looks Like in Practice

Strip away the abstraction and judgment turns into a short list of habits, not a personality trait. The ones that actually show up in a working review process:

  • Trace the specific claim, not the general impression. A confident paragraph and a correct paragraph read identically. Only checking the underlying number, citation, or log entry tells you which one you have.
  • Ask what would have to be true for this to be wrong. If you can't name a way the claim could fail, you probably haven't examined it closely enough to sign off on it.
  • Treat AI output as a first draft, not a finding. The draft can save real time on structure and phrasing. It hasn't earned the status of a verified conclusion just because it's fluent.
  • Slow down exactly where the stakes are highest, not where the deadline is loosest. A rushed review on a low-consequence item costs little. The same shortcut on a disposition decision or a customer-facing claim costs a great deal.
  • Say "I don't know yet" out loud. An AI system will not hedge convincingly even when it should. A reviewer who can say a claim is unconfirmed is doing work the tool structurally cannot do.
  • Write down what you actually checked. Not "reviewed AI draft," but which specific claim, against which specific source. That's the difference between a defensible record and a rubber stamp with an audit trail attached.

None of that is glamorous, and none of it is unique to AI. It's the same discipline that verification-based work always required. The difference now is that a fluent draft makes it easier than ever to feel like the checking already happened.

Human Judgment vs. Algorithmic Output

Dimension Human judgment Algorithmic output
Basis for a decision Context, consequence, and lived stakes Statistical pattern in training data
Accountability when wrong Falls on a named, identifiable person Falls on whoever deployed it, if anyone claims it
Behavior under uncertainty Can say "I don't know" and slow down Produces a confident answer regardless
How errors present Often comes with hedges or flagged doubt Reads identically whether correct or fabricated
How it improves Reflection, correction, lived consequence Retraining on new data, at the vendor's schedule
Cost of a mistake Personal: reputational, financial, sometimes legal Diffuse, absorbed by whoever trusted it

The tribunal in Moffatt put its finger on the real issue without mentioning algorithms at all: an organization answers for what it tells people, regardless of which mechanism did the telling. Judgment is the thing that sits with that responsibility. Output doesn't sit with anything. It just gets generated and moves on to the next prompt.

The Supply Keeps Shrinking

There's a quieter mechanism behind the scarcity, and it's the one that should worry a quality function more than any single bad filing. Judgment is a skill built through repetition, and repetition is exactly what gets removed when a tool drafts every investigation, weighs every tradeoff, and summarizes every complicated record before a person looks at it.

The attorneys in Mata weren't unusually careless. They worked in a profession where checking a citation is the most basic unit of the job, and the habit had gone soft anyway. That's the mechanism to watch for inside a compliance function: not one dramatic failure everyone learns from, but a slow erosion in the reps that build the next reviewer's ability to catch a bad claim at all. A junior investigator who has reviewed two hundred AI-assisted summaries and personally traced ten of them back to source has a different skill than one who has approved two hundred and traced none. The organization chart won't show the difference. The next serious miss will.

That's why the scarcity compounds instead of holding steady. Fewer reps produce weaker judgment, weaker judgment approves more unchecked output, and the volume of unchecked output keeps rising because nothing in the system is pushing back on it. Left alone, that's a one-way slide.

Building the Habit Back Into the Workflow

None of this argues against using these tools for what they're actually good at: fast drafts, structured summaries, and surfacing information that would have taken a person much longer to assemble by hand. It argues for being honest about where the tool's contribution ends and a person's responsibility begins.

The practitioners who come out ahead over the next few years won't be the ones who refuse the tools, and they won't be the ones who defer to them either. They'll be the ones who kept a specific, checkable verification step attached to a specific point in the workflow, even after the drafts got good enough to make skipping that step feel harmless. That's a discipline worth building into a procedure rather than leaving to individual willpower, because willpower is exactly what a clean-looking draft under deadline pressure is designed to wear down.

I've written more on the mechanics of that gap between a plausible answer and a true one, and separately on how to keep AI positioned as a tool rather than an authority inside a team's actual decision-making. Both are worth reading if this is the part of the job you're responsible for protecting.

Frequently Asked Questions

Does a decision rule for AI verification need to slow down every task? No. Most AI-assisted drafting can move at full speed. The rule only needs to bind at the specific claims that would be expensive if wrong: a root-cause statement, a numeric result, a citation, a customer-facing commitment. Everything else in the draft can be trusted as phrasing and structure.

Who is accountable when an AI-assisted document turns out to contain an error? The reviewer who signed off, not the tool. That was the entire finding in Moffatt v. Air Canada: the tribunal held the airline responsible for its chatbot's statement in the same way it would for a human agent's, because the organization, not the software, is the party that made the commitment to the customer.

What's a practical first step for a team that reviews a high volume of AI-drafted content? Pick the one claim type in each document that's most expensive to get wrong, and require it to be traced to its primary source before sign-off, every time, regardless of how confident the draft reads. That single checkpoint catches most of what matters without slowing the rest of the workflow down.

Is it realistic to verify every AI-generated claim in a high-volume environment? No, and that's not the goal. The goal is verifying the specific claims where being wrong is costly, using a fast rule of thumb: if it can be checked against a source in under five minutes, check it before moving on. If it can't be checked that quickly, that's the signal to escalate rather than approve.

How do you know if a team's judgment is atrophying rather than just adapting to new tools? Watch the ratio of claims approved to claims actually traced back to a source. If that ratio has been quietly dropping while output volume rises, the team isn't adapting. It's losing the habit that verification depends on, one unchecked approval at a time.

Last updated: 2026-09-05

J

Jared Clark

Founder, Prepare for AI

Jared Clark is the founder of Prepare for AI, a thought leadership platform exploring how AI transforms institutions, work, and society.