The AI Agent Factory

6.2 A good output and a plausible one

Status
stable
6 min read
Owner
Panaversity
Approved
Panaversity ·

In everyday life. A restaurant bill is neatly printed and added up, but it still charges for a dish you never ordered.

A plausible output has the shape of a right answer. It reads smoothly, and it is organized and confident, with headings, totals and citations, the sources it names. A good output matches the outcome you asked for in the brief and the sources it worked from. AI makes plausible output easily, whether or not it is good. Brightline's AP Worker handles the bills the company owes. The memo it sent with Friday's payment run was plausible in every line.

People trust text that is easy to read, even though being easy to read has nothing to do with being true. Researchers call this processing fluency.1 So asking "does this feel right?" as you read is the weakest check you have. Run two stronger checks instead.

  1. Completeness. Match each item by its ID, not only by the count. Every invoice in the source must appear exactly once in the output: none missing, none repeated, none extra. Friday's run had 15 open invoices. A file with one row missing and another row repeated would still have 15 rows. Check that every part the brief asked for is there: the CSV (a spreadsheet file with one row per invoice), the memo and the note. Check that every decision Dave, the controller, must make is in the memo. A reader cannot see a missing row. A match shows it at once.
  2. Accuracy. Trace claims back to the source. Trace every claim a decision rests on, and a sample of the rest. A row is accurate when its amount, date, action and reason all match the files and the policy.

"All 15 open invoices were reviewed" is a claim like any other. The CSV had 14 rows. The worker's own task record, its log of the run, said "14 rows." Completeness checks find what accuracy checks cannot, because there is no row to trace.

The title reads "What you see and what you check," and the line below it reads "A polished output still needs independent verification." Two panels are joined by an arrow labeled verify. On the left, what you see, the signs of a plausible output. Fluent writing: "Ready to approve." Neat structure: headings, a table, three decisions. Confident totals: "$24,607.75 across 8 invoices." Citations: "Policy 3.4." Below them: presentation alone is not proof. On the right, what you check, four tests against the evidence. One, match every source ID exactly once: 15 source invoices, only 14 output rows. Two, trace claims to their sources: "Bank change verified" has no supporting record, and "Policy 3.4" does not exist. Three, recompute from source files: $17,452.75 after invoice 4519 is approved, still subject to approval of the whole run. Four, compare facts across files: the memo holds Tri-County, and the CSV pays it. A footer reads: fluency makes an answer convincing, evidence makes it trustworthy.

Figure 6.2. What you see and what you check. A plausible output shows fluency, structure, totals and citations. None of those is evidence. The evidence comes from four checks: match every source ID once, trace each claim, recompute each number a decision rests on, and compare the files with each other.

Check yourself

Question 1 / 8 · blueprint

0 answered

What makes an output plausible but not good?

  1. How to Think in the AI Era, Panaversity, first edition.↑

On this page