Build log · Part 1 of 3

Two AI Critics Played My Puzzle Game 80 Rounds and Signed Off. I Still Said No.

How an AI-written pitch survived its own fact-check, became a working puzzle prototype, and passed two rounds of automated critique — before I opened it myself and wasn’t so sure.

Two independent AI critics played the finished prototype of ChronoShift eighty rounds combined, across two rounds of testing, and both signed off on it. Not with an unqualified thumbs-up — with hedges, caveats, and specific language about exactly what they had and hadn’t verified. Critic A’s final verdict, in its own words: “a credible first curriculum with four distinct encounters, not yet unbounded generative puzzle novelty.” That is a pass. A careful, specifically-scoped pass, but a pass.

On paper, that should have been the end of this stage of the story: two rigorous, independent, adversarial testers ran the build and found it held up. This is the build log of how it got to that point — the pitch, the self-audit, the actual build, and two full rounds of AI critique — and it ends on the same contradiction it opened with. Because I opened the build myself, after both sign-offs, and I wasn’t convinced. That part of the story is Part 2. This part is about how a genuinely rigorous process produced a build that two careful critics were comfortable passing.

The pitch an AI wrote about itself

ChronoShift didn’t start as my idea in the usual sense. An AI proposed the whole app: an “AI-driven daily cognitive game,” pitched with the kind of confidence a business plan is supposed to have. The numbers were specific and aggressive — a path to roughly $10k in monthly recurring revenue from around 1,000 subscribers — and the margin claim, quoted directly from the proposal, was blunt: “Operating margins exceed 95%.”

It read like a pitch deck, because in every meaningful way it was one. Confident market framing, a clean unit-economics story, a product loop described in three crisp verbs. What it hadn’t done yet was survive contact with a second, skeptical pass.

The AI that fact-checked the AI

That second pass came from a separate research pipeline, and its job was narrow: go back through the pitch and check whether the numbers actually held up. They mostly didn’t. The proposal’s own market-size estimate was off by roughly seven times — $2.35B against a real figure closer to $16.26B for the same category and the same year. Not a rounding error; a figure that had clearly never been checked against a second source.

The margin claim didn’t fare much better. “Operating margins exceed 95%” is the kind of line that sounds great in a pitch and quietly ignores VAT and payment-processor fees. Once those were actually modeled, the real breakeven point moved from the pitch’s round 1,000 subscribers to something closer to 1,650–1,900 paying users — a meaningfully harder bar to clear, and one nobody would have known about from the original document.

The audit didn’t stop at correcting numbers. It reframed the pitch itself, landing on language that put the product ahead of the underlying technology: “a five-minute adaptive story puzzle… AI is infrastructure, not the product promise.” That distinction — AI as the thing running underneath, not the thing being sold — ends up mattering for everything that follows in this series. It’s also the first appearance of a pattern that recurs through the whole project: a claim gets made confidently, and then a separate, more skeptical process goes back and checks it. That discipline is the actual subject of this build log, more than any single number is.

What actually got built

Out of that corrected pitch came an actual prototype, and the core mechanic is easy to state in one sentence: draw a glass ramp, and redirect a falling particle through portals, accelerators, and relays into a collector. One verb — draw a line — and one object being routed through an increasingly elaborate chamber of physics toys. It plays fast, it’s legible, and it’s the kind of loop that’s easy to prototype quickly and hard to evaluate casually, because “does this feel repetitive” isn’t a question a first playtest always answers.

Which is exactly why the next step wasn’t a single playtest. It was two independent AI critics, given the finished build and told to try to break it.

Round one: two critics, twenty rounds each

Critic A and Critic B worked independently, with no visibility into each other’s sessions, twenty rounds apiece. Both converged on the same underlying problem from different angles, which is usually a sign the problem is real rather than an artifact of one tester’s approach.

Critic A named it structurally: “The route grammar repeats the same geometry, mirrored/inverted with small translation.” In plainer terms: later chambers weren’t asking for new solutions, they were serving the same solution shape back with a flip or a nudge. Critic B put a number on how often that was happening: roughly three in four completions reused or transformed a route from an earlier chamber rather than solving the chamber independently. A player — or an AI critic playing like one — could get most of the way through the build by recognizing a shape and repeating it, not by engaging with each new chamber on its own terms.

That is a legitimate, specific, falsifiable finding. It’s also exactly the kind of thing a good adversarial critic is supposed to catch before a human ever sees the build.

The fix: a fourth mechanic

The response wasn’t a difficulty-slider tweak. It was a new mechanic: vault, an upward-launch encounter built around a curved canopy, added specifically to break the geometry the critics had flagged. Where the existing chambers were variations on redirecting a falling particle downward and outward, vault forces an upward launch through a genuinely different shape — not a reskin of the same route grammar, a different grammar entirely.

Round two: the retest

Both critics came back and re-ran twenty rounds each against the updated build — the other forty rounds in that combined total of eighty. The measurable result was a real drop in route reuse, not a cosmetic one:

Route reuse across the twenty-round retest, before and after the vault mechanic shipped.
Build Routes reusing a prior solution Reuse rate
Before vault (round one) 10 of 19 ~53%
After vault (retest) 5 of 19 ~26%

Roughly half the reuse rate, with a new mechanic in place to explain the drop rather than a looser scoring pass. And the verdicts that came back were, again, specific rather than triumphant:

“A credible first curriculum with four distinct encounters, not yet unbounded generative puzzle novelty.”

Critic A Retest verdict, round two

“This is evidence for greater depth, not an exhaustive proof that no alternative straight solution exists.”

Critic B Retest verdict, round two

Read those two verdicts closely, because the hedges are doing real work. Critic A is explicit that “four distinct encounters” is not the same claim as “unbounded generative puzzle novelty” — it verified what exists, not what the system could theoretically produce forever. Critic B is explicit that a lower reuse rate is part of “evidence for greater depth,” not proof that no alternative straight solution exists anywhere in the build. Neither critic is saying the game is infinite, or perfect, or even finished. Both are saying: within the specific thing we tested, for the specific claims we can support, this passes.

Where that leaves things

So here is where the pipeline actually stood: an exploit that let players skate through on pattern-recognition instead of problem-solving, closed, with reuse cut roughly in half and a genuinely new mechanic added specifically to close it. Two independent AI critics, working separately, running a combined eighty adversarial rounds, both landed on the same hedged-but-real conclusion — ship it.

And yet. I opened the finished build myself before writing any of this up, the way anyone eventually does with something built in their name, and what five minutes of actually playing it told me didn’t fully square with the scorecard above. Not because the critics were wrong about what they measured — they weren’t. Because “is this solvable, is the exploit closed, is there route variety” and “does this feel like one idea wearing four costumes” turn out to be different questions, and nothing in that first process ever asked the second one.

What I actually sent back, and why a rigorous scorecard and a five-minute playthrough can both be telling the truth at the same time, is Part 2.

This is Part 1 of a 3-part build log

Part 2 covers the actual feedback I sent back after opening the build myself. Part 3 covers a much deeper forensic audit that went looking for what neither the critics’ scorecard nor my own playthrough had fully caught. Both are coming soon.

See the whole series →