Yesterday my agent wrote the sentence "I cannot verify this" and meant it.
Forty lines later, in the same reply, it put the unverified thing in a markdown table with a column header and a row per item.
I read both. I acted on the table.
Neither of us did anything wrong at the level of intent. The disclaimer was honest. The table was confident. Confidence won.
You have shipped this
Before I make this about the machine, I have done the identical thing myself, more than once.
You caveat something in a Slack thread, then paste a clean summary underneath it.
You write "rough numbers" above a spreadsheet where every cell is right-aligned to two decimal places.
You tell your lead the estimate is soft, then put it in a Gantt chart with a start date and an end date.
The hedge goes in the sentence. The answer goes in the structure. And structure wins, every single time, because structure is what people act on.
What the verification did
The task was ordinary. Recommend some films, note whether they were on a specific streaming service in a specific country.
It knew, and said, that regional streaming catalogs rotate and that its recall of one was not trustworthy. Correct instinct, stated plainly, in writing.
Then it produced a recommendation table anyway.
When I finally made it check against a live availability source, per title, one request each, five of the eight titles in that table were not on the platform at all. Not close. Not recently removed. Simply not there, in that country, that day.
The strongest pick in the list, the one described to me as the best match for what I had asked, was among them.
Here is the part that should bother you more than the miss.
Nothing in that output looked uncertain. The film knowledge was solid. The quality judgments held up fine. Every rating was close enough. Only one narrow class of fact was rotten, and it was the class that changes monthly while feeling permanent.
Confident recall and correct recall produce identical-looking output. There is no tell, and that is as true of me writing an estimate as it is of a model writing a table.
The finding, which is not "verify things"
Everyone already knows to verify things. That advice has never once changed anyone's behavior, mine included, as demonstrated above.
The useful finding is narrower.
A disclaimer in prose cannot cancel an assertion in structure.
They are not weighted equally by the reader and they never were. A table is an authority format. So is a numbered list, a comparison matrix, a confidence percentage, a chart with axes, a schema with types. These carry an implicit claim that the values inside them were sourced rather than generated.
When you put unverified content into one of those containers, the container upgrades it. Your caveat sits above, in soft prose, doing nothing.
This matters more with agent output than with human output, for a boring mechanical reason. Agents produce structured output constantly, because structured output is easier to parse, easier to render, and easier to feed into the next step. The format that makes output machine-usable is the same format that makes it look verified.
So the failure scales with how well-formatted your pipeline is.
What I would change
Stop auditing the hedging language. It is decorative and everyone skims it.
Audit the shape.
Ask which claims in the output are sitting inside an authority container. Tables, matrices, typed fields, percentages, anything with a header row. For each one, ask whether the values came from a source or from recall.
If they came from recall, they do not get the container. They get a sentence, in the same soft register as the doubt, so the confidence signal and the epistemic status finally match.
That is the whole fix. Downgrade the format to match the sourcing.
The corollary is uncomfortable and I think correct. If a claim is not worth the cost of verifying, it is also not worth a table row. Either check it or say it loosely. The middle option, checking nothing and formatting it beautifully, is the one that produces the confident wrongness people get burned by.
The thing verification gave back
Worth saying, because verification usually gets sold as pure insurance and it is not.
Running the real check surfaced something that recall could not have produced at all. One of the titles missing from the platform we were checking turned out to be free, with ads, on a service neither of us had thought to mention.
That is a better answer than the one I nearly shipped. Not a corrected answer. A better one.
Verification is usually framed as the tax you pay to avoid being wrong. In practice it is also where the non-obvious answers live, because the live source knows things recall cannot.
Your turn
What is the last thing you formatted more confidently than you knew it?
Table, chart, estimate, schema, does not matter. I want the format, not the topic.
If this was useful
I work through this in public, the wins and the freezes both, mostly on LinkedIn and YouTube. If the real version of building in the open is useful to you, that is where it lives. Find me on X, GitHub, and the work at next8n.com.
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.