SMASHTHATJOB
LATEST THURSDAY, 6 AUGUST 2026
AI Skills reality check

The Decision Layer Your AI Agent Can't Read

ChatGPT Work turns a goal into finished work. It cannot see why your numbers moved, and it guesses rather than leave a blank. Brief that gap first.

OpenAI describes ChatGPT Work, which began rolling out on 9 July, as an agent that can “turn a goal into finished work” and stay with a project for hours, breaking it into steps and completing them on its own. You connect Drive, the CRM, Slack, the ad accounts. You give it a goal. You come back after lunch to a finished deck. Most of what it produces will be correct. That is what makes the rest hard to catch.

The failure everyone braces for is a wrong number, and there are decent habits for wrong numbers. The one that reaches the client is a right number in a false frame. Picture slide three of a July report: paid social revenue down 22 percent, flagged in warning red. Every digit traces cleanly to the ad account. It is also the most misleading sentence in the deck, because the drop was deliberate. Half that budget moved mid-month into an email launch that lives in a different tool, and the creator collaboration it paid for delivers in August. The agent saw a channel falling and called it a decline.

Underneath every set of records a business keeps sits a layer that nothing stores. Call it the decision layer: what you moved and why, what is paused on purpose, what is waiting on a person, what you gave up to buy something better. The ad account logs the spend. The CRM logs the outcome. No system logs the reasoning that produced either, because the reasoning happened in a meeting, a hallway, or your own head at eleven on a Tuesday. An agent can read every tool you connect it to and remain blind to the only layer that makes the numbers mean anything.

What makes that blindness expensive is that the agent will not mention it. OpenAI’s own researchers argued in a 2025 paper that models hallucinate largely because training and evaluation reward guessing over admitting uncertainty: on a test, a wild guess sometimes scores, while a blank guarantees zero. That is a claim about incentives, not a measured error rate for any particular agent. The behaviour it predicts is the one you keep meeting. The agent reaches the decision layer, finds nothing there, and fills the space in the same steady register it used for the lines that were true.

Verification is aimed at the wrong mistake

Recomputing the load-bearing figures and tracing each to its source is real discipline, and it is the right response to a fabricated number. It cannot touch this one. The 22 percent passes every check you can run, because it is accurate. The invented part is the word “fell” doing the work of “was cut,” and no amount of re-deriving the figure will surface that.

There is a second reason after-the-fact fixing fails here. The moment you start editing the output you are proofreading its guesses one at a time, inside a document that already reads as finished. People stop reading closely at that point. Slide three was the slide you happened to catch. The rest of that discipline still matters, and it is the authorship debt you pay down before you forward anything an agent made. This particular error just sits outside its reach.

So the work moves upstream. Every decision the agent could not see is a decision you already made. You either write those decisions down before it runs or dig them out of its output afterwards, and the second version is slower, less reliable, and done under time pressure with something plausible already on the screen.

Brief the blind spot

Four lines, before you press go. It takes about five minutes and it is the same five minutes whether the task runs for ten seconds or an hour.

WHAT IT MAY READ
  Name the files and systems. Name at least one that matters and
  is NOT connected, so it knows the record is incomplete.

WHAT MOVED, AND WHY
  The decision layer, in your words. Two or three lines is
  usually the whole of it.

DESCRIBE, DON'T EXPLAIN
  It reports what changed. It does not assign a cause, and it
  does not call any result good or bad.

WHAT STAYS MINE
  The judgments it must leave blank: causes, recommendations,
  anything that reads as advice.

The third line does most of the work. An agent that is forbidden to explain cannot invent a reason, and the gaps it leaves are legible to you as gaps rather than disguised as findings. Splitting description from judgment is the same seam that runs through writing the story around your numbers, moved earlier in the process so you never have to unpick it afterwards.

The shape holds across roles because the blindness does. A finance analyst closing the month knows that the variance the agent will file under seasonality was a reclass agreed with the auditors in week two. An operations manager knows the workstream that looks stalled in the tracker is sitting with legal by design. A solo consultant knows which client stopped emailing because the work is finished and which stopped because they are unhappy. In every case the record is accurate and silent, and the agent treats silence as absence.

The fair objection is that writing a brief is the labour you were trying to delegate. If the brief takes five minutes and the task takes an hour, the trade is obviously fine. But you were always going to decide what the numbers meant. The brief only changes where in the process you do it, and it lasts in a way the output never does: next month’s close, next month’s report, next quarter’s roll-up all reuse the same four lines with the second one rewritten.

Treat the agent as the fastest researcher you have ever worked with and the worst-briefed. It will read everything you point it at and nothing you did not. Give it the production. Keep the decision layer, write it down first, and what comes back is something you can send.

Do this today

Before your next real handoff to an AI agent, spend five minutes on four lines: what it may read and what it can't see, what moved and why, a standing instruction to describe rather than explain, and the judgments you're keeping. Paste it above the task. Then read what comes back for causes it wasn't allowed to assign, and notice how few there are.

Sources

  1. https://openai.com/index/chatgpt-for-your-most-ambitious-work/
  2. https://arxiv.org/abs/2509.04664
  3. https://openai.com/index/why-language-models-hallucinate/
  4. https://www.bnnbloomberg.ca/business/artificial-intelligence/2026/07/09/openai-launches-chatgpt-work/