Case study · Agent orchestration · In build

Three Claude sessions built a Roblox game in 45 hours. I reviewed the diffs.

A co-op snowball game for Roblox: roll a ball that grows with every step, ride it down a forked bobsled track, flatten a toy town for coins, merge with a friend for a bigger smash. Built Sep 27 to 29 2026 by three Claude Code sessions with me as orchestrator, reviewer, and the only human who ever pressed Play.

Build readoutIN BUILD
Commits in 45 wall-clock hours
121
Lines of strict Luau, 68 files
15k
Live-playtest screenshots as proof
101
Review findings confirmed / rejected
14 / 4

The setup

One orchestrator, two builders, one human on the critical path.

Three desktop sessions with three non-overlapping jobs. The orchestrator writes milestone briefs, reviews every commit, runs review workflows, and relays the few steps only a human can do. The builder session owns the Luau, verifies each task live through the Roblox Studio MCP, and commits once per task. The art session generates 3D assets in Tripo, reviews and colour-fixes them headlessly in Blender, and renders icons, thumbnails and fracture sets. None of the three edits another one's files.

My job was the part agents cannot do: decide what the game should feel like, reopen Studio when it closed, approve every upload and every credit spent, and play each phase before the next one started. When a friend who had never seen the game played it, her seven complaints rewrote the next brief. That feedback loop, not the code volume, is the point of the case study.

The loop

What moves when a milestone runs.

  1. 01
    Brief
    A MILESTONE-N.md with the design, phased tasks, a hard STOP for a human feel test, a parked list, and the rules repeated
  2. 02
    Build
    The builder session works the brief task by task: strict Luau, tunables in one Balance file, server authority, checks clean
  3. 03
    Verify
    Each task is played live through the Studio MCP and lands as one commit plus a screenshot in screens/
  4. 04
    Review
    The orchestrator diffs server and profile code against the rules; money-touching milestones get a multi-agent adversarial review
  5. 05
    Feel test
    Everything stops until a human plays it; feedback goes verbatim into the next phase of the brief
  6. 06
    Ship
    Summary in PROGRESS.md, push to origin, gotchas captured into a reusable skill the same day

Every task ends in a commit and a screenshot from a live Studio playtest, so progress is auditable in git, not narrated in chat.

Design decisions

What was built, and why it is built that way.

The game is the deliverable. The orchestration pattern is the engineering. Each choice below is visible in the commit log.

01

The orchestrator never writes game code.

A session that writes the code cannot review it with fresh eyes. Keeping the orchestrator to briefs, diffs and relays meant every commit got read by a context that had not just written it, and the builder could be corrected in one message instead of re-prompted from scratch.

02

Briefs are files, not chat.

Each milestone lives in a committed markdown file with phased tasks, a parked list and the standing rules. The builder reads it cold at wake-up, the file survives context compaction, and a playtest becomes a new dated section rather than a lost conversation. Five briefs, five milestones, zero drift on scope inside a phase.

03

Server authority is a rule, not a preference.

Clients send argument-free intents; the server owns coins, size, unlocks and purchases. Purchase receipts are idempotent on the purchase id and confirmed only after the save that contains that id lands. The review pass still found a gap here, which is why the rule is written down and checked, not assumed.

04

One commit per task, one screenshot per claim.

The builder is not allowed to report a task done without a live playtest and a PNG in the repo. 101 screenshots later, every feature claim on this page traces to a commit and an image. When the MCP could not verify something, the log says so, and that item went on the human test list.

05

Money code gets an adversarial review, not a second read.

The monetization milestone ran a Workflow of 58 agents: four finder dimensions (economy, compliance, security, placement) over the diff, then every finding sent to three independent refute-first verifiers, surviving on two of three. 14 findings confirmed, 4 rejected. Two were real money bugs: a repeatable double-offline purchase that compounded without limit, and coin packs sized from an income estimate that ignored rebirth and index multipliers, up to 37x too small for the players most likely to pay.

06

Art spends credits once. Colour is fixed locally.

Tripo's texture pass ignores colour words, so the art session judges variants on shape only and recolours textures headlessly in Blender, with gutter dilation so mipmaps do not bleed. 30 assets on 3,150 credits, zero paid retries, and 76 pre-fractured debris chunks cut from the same models for the destruction system.

07

Design calls go to a decision agent. Feel calls go to a human.

To stay unblocked while I was away, routine design forks were settled by a separate Opus decision agent with the brief and research as context. Anything about how the game feels stayed with me and one first-time player. The split held: the agents chose lane widths and camera math well, and missed that a new player did not know where to walk.

08

Gotchas become skills the same day.

Every stall was written into a reusable skill file: the app reaping the Rojo server after idle, Studio closing on its own, virtual keys never reaching the game, the browser pane unable to reach a localhost bridge. The asset review loop and the render pipeline became skills too. The next game starts from the protocol, not from memory.

What broke

Where the agents were wrong, and how it was caught.

Verified at 150 ms of injected lag. Glitchy in a human's hands.

The MCP cannot press real keys, so steering was verified with injected inputs and passed. A human felt snaps on every steer change. Root cause: inputs reached the server 13 to 21 ticks late and snapshots from before the input rewound the client. Fix: tick-stamped sequence-numbered intents, a client that skips snapshots predating an unacked input, and a correction fade capped at 0.4 studs per frame. Lesson: an MCP-only check is labelled unverified until a person plays it.

Two money bugs passed the builder's own tests.

The builder verified every product grant live and still shipped a double-offline consumable that could be bought again and again, doubling the same pile each time, and coin packs that undersold rebirthed players by up to 37x. Both were caught by the refute-first review, not by a re-read. Verification and review are different jobs.

The onboarding the agents understood, a first-time player did not.

A line of +1 props from spawn to the release gate, a goal chip and a glowing ROLL! arch fixed the first complaint. Then the human said the line removed any reason to wander. The next revision replaced it with value rings and a soft size gate. Two rounds of human feedback did more for the first minute than any amount of agent reasoning.

Every stall had a human-only step nobody had planned for.

The desktop app reaped the Rojo server after idle. Studio closed on its own. The browser pane could not reach the Tripo bridge on localhost, so 3D imports moved to a real Chrome. Each cost an hour before it became a rule in the brief: list the human-only steps up front and batch them.

Where it stands

Not launched. That is the honest status.

The game is in its fifth milestone, with the roll, the town destruction, a monetization catalog behind zero product ids, store art, and a launch checklist in place. It has not been published, has earned nothing, and will not launch until three people who have never seen it play for an hour on phones. The repository is private until launch, so this page quotes the commit log rather than linking it. What the build already demonstrates is the part that transfers to any AI engineering job: decomposing work into briefs agents can execute, verifying claims instead of trusting them, running adversarial review on code that touches money, and knowing which decisions must stay with a human.

Get in touch

Want the orchestration walkthrough?