superpwrSign inGet started

We caught our ownAI lying.

In our own market research, before it went anywhere. That is not an embarrassing anecdote. It is the product, working, on the company that builds it.


§01

What it found

Three fabrications in one research pack.

We ran our own AI-generated competitor research through the same scorer we sell, before showing it to anyone. It scores every claim against its sources. Here is what came back.

No.What it flaggedWhat was actually trueThe tell
01A funding roundA competitor profile claimed an $8M seed from top-tier VCs. Reality: around $500K, different city, different product. Every material fact invented.Confident, specific, and attached to no source at all.
02One statistic, two companiesThe same $200 per week cap was attributed to Tesla in one place and Meta in another. Tesla is real, sourced to an internal memo. Meta was a merge.The duplication was the tell. A real figure has one owner.
03A real number, wrong jobA 1.8 hours per day figure anchoring an ROI model turned out to be a pre-AI statistic about document search. Replaced with sourced 2026 data.The hardest class: true, checkable, and measuring something else.

Number three is the one worth sitting with. It was not a hallucination. It was a real number, correctly quoted, doing a job it was never measured for. No fact-checker catches that. You catch it by asking what the source was actually measuring.


§02

And then we did it again

The same trap, four weeks later.

A page about catching fabrications that hid its own would be self-refuting, so here is the one we made in August 2026.

Working through a stack of competitor decks, we recorded their figures as facts, including a seed round of $8M. That is almost certainly the same invented number our own tooling had already caught, banked a second time by someone who did not ask where the deck came from.

The provenance answer was in the same folder: a transcript saying plainly that the decks were generated from prompts. We read the numbers first and the provenance never.

A deck is a claim like any other. The question is not whether it looks researched. It is where each number came from, and that question comes first rather than last.


§03

The platform that built itself

Every claim here was made by the thing.

01

It provisioned its own infrastructure

Domains registered and delegated, databases created and connected, keys placed, services deployed. Through chat, with no terminal open.

02

It built another product

A whole chat and build platform, with its own database, its own sign-in and its own connector. Not a demo. A running product with users.

03

It moved its own stack

A full migration between hosting providers, database and DNS, planned and executed by the platform on itself.

04

It audits its own writing

Including this site. The claims on these pages are checked by the same machinery we are describing, and several were removed for failing.

We are not the biggest customer of this product. We are the only one so far, and we would rather say that than imply otherwise.

Two longer write-ups go through this in detail: a night where six pull requests shipped without a human in the loop, and why four rival models have to agree before a contract is ratified. Both name the specific merges and checks. The overnight run and the Board of Architects.


§04

What it cannot do

A score that never says I do not know is decoration.

The grounding signal asks a narrow question: is this claim connected to something identifiable, or is it floating. That is genuinely useful and it is not the same as being right.

It will not tell you a sourced number is measuring the wrong thing. That was fabrication three above, and a human caught it. It will not tell you an argument is bad. It will not tell you the source itself is wrong. And on work with no external sources at all, judgement calls, plans, opinions, it has very little to say.

What it does reliably is separate the claims that are anchored from the ones that only sound anchored, and put the second pile in front of a person before it goes out the door. That is a smaller promise than most of this category makes, and we would rather be held to it.


Built for the people AI forgot to take care of.

What went out unchecked last week?