The AI Agent Factory

6.8 The same review on both AI vendors

Status
stable
10 min read
Owner
Panaversity
Approved
Panaversity ·

In everyday life. Two shops both give receipts, but neither tells you whether the price matched the shelf label. You check that yourself.

Anthropic makes Claude, and OpenAI makes ChatGPT. The Review Contract describes the work, not the tool, so it works on either AI vendor without changes. The two boxes below follow the same six lines in the same order. A number such as (6.5) points to the part of this chapter that the line matches.

Anthropic, as verified 6 October 2026.

  • Where the checks go (6.1). In the brief. Anthropic's help center lists the outcome, format and inputs to include.1 Its guide to Cowork, where Claude carries out longer tasks for you, says to review Claude's approach before you let it run.2 Neither page names review criteria, so this book's method puts them in the brief.
  • Watching the work (6.7). Progress indicators show what Claude is doing at each step. You can open the same session in another app, such as the phone app, to watch it, answer Claude's questions or redirect the work.2
  • The record (6.7). The session keeps the task's progress and its finished output. You can start a task in one app, guide it from another and get the output wherever you are.2
  • Citations (6.5). Responses that use web search include citations.3 Anthropic says to review the cited sources, because the originals may have context the summary left out.4
  • High-risk work (6.1). Some work has real consequences, such as money, messages sent in your name or important files. For that work, Anthropic says to stay close and review what Claude does, or switch from automatic approval back to manual approval, where Claude asks before each action.2
  • The AI vendor's own warning (6.3). Claude can produce responses that are incorrect or misleading, including quotes that look trustworthy but are not based on fact. Do not rely on it as your only source of truth for high-risk advice.4

OpenAI, as verified 6 October 2026.

  • Where the checks go (6.1). In the request. OpenAI's guide to Work, ChatGPT's mode for longer, multi-step work, says to add any files, constraints and review criteria ChatGPT should use.5
  • Watching the work (6.7). In Work, you can review progress, answer questions, change direction and approve important actions.5
  • The record (6.7). You review the result in the same Work chat, and you ask for changes or more work there.5
  • Citations (6.5). Responses that use web search may include citations, and a Sources view when it is available. OpenAI says citations can be incomplete, outdated or incorrect. Open a cited source to check that it supports the answer, and check its date.6
  • High-risk work (6.1). Approve important actions as the work runs.5 OpenAI's prompting guide adds: require your approval before ChatGPT sends, publishes or changes information others rely on.7
  • The AI vendor's own warning (6.3). ChatGPT can produce incorrect or misleading output, and it can sound confident when it is wrong. Check important information against reliable sources, and use ChatGPT as a first draft, not a final source.8

The comparison. The six lines match closely. Both AI vendors take a detailed request, but only OpenAI names review criteria for it. Both show progress, let you approve important actions, cite web sources when they search, and warn that output can be wrong. One gap, on both AI vendors, changes how you write the contract. Neither AI vendor's help pages promise a citation to a row in your own file. So ask for row-level evidence in the brief, and verify it: a reason, a policy section and a source on every row of the output. The records are similar: the help pages describe progress and results kept with the session or conversation, and no separate audit log. Either way, the record shows what the worker did. It does not show whether the work was right.

The title reads "The same review on both vendors," and the line below it reads "One Review Contract. The same human responsibility." A label marks the table as a chapter snapshot of 6 October 2026. It has three columns: the review task, Anthropic Claude Cowork, and OpenAI ChatGPT Work. Set the checks. Claude: put the criteria in the brief, and review Claude's approach before it runs. ChatGPT: add files, constraints and review criteria to the request. Monitor the work. Claude: follow progress, answer questions and redirect the task. ChatGPT: review progress, answer questions, change direction and approve actions. Inspect the record. Claude: review progress and finished output in the session. ChatGPT: review progress and results in the Work conversation. Verify citations. Claude: web-search responses include citations, so open and check the sources. ChatGPT: web responses may include citations and a Sources view, so check support and date. Control consequential actions. Claude: stay close for money, messages and important files, and consider manual approval. ChatGPT: require approval before sending, publishing or changing shared information. Expect possible errors. Claude: output can be incorrect or misleading, so verify consequential claims. ChatGPT: output can sound confident and still be wrong, so verify important information. A gold box reads: ask explicitly for file-and-row evidence. For this payment run, require a reason, policy section and source on every row. Do not assume automatic row-level citations. Request them and verify them. A footer reads: independent verification, the decision and the signature stay with you. A source line at the bottom reads: Chapter 6 vendor notes, Anthropic Help Center, OpenAI Help Center and ChatGPT Learn.

Figure 6.7. The same review on both AI vendors, as verified 6 October 2026. The gold box is the line to write into every contract.

What stays the same. The checks, the decision and the signature stay with you on either AI vendor. The product supplies the record, and citations when it searches the web. It never supplies the final decision.

Check yourself

Question 1 / 8 · current

0 answered

You brief ChatGPT Work to propose a payment run. Where do your review criteria go?

  1. Claude Cowork and chat are one Claude, Claude Help Center.↑
  2. Get started with Claude Cowork, Claude Help Center.↑↑2↑3↑4
  3. Claude is providing incorrect or misleading responses. What's going on?, Claude Help Center.↑↑2
  4. ChatGPT Work and Codex, OpenAI Help Center.↑↑2↑3↑4
  5. Prompting, ChatGPT Learn, OpenAI.↑
  6. Does ChatGPT tell the truth?, OpenAI Help Center.↑

On this page