LOG 002 // 02 OCT 2026

I had the flu, so I made a city builder with an agent team

I had the flu and could not give DrivEngine the attention it needed. My brain wanted a smaller problem. Then I had a question: what would happen if I made a game prototype with agents, but treated them like a real team of people?

I already had a design for an ancient Sumerian city builder. It borrowed the parts of Anno that I like: people with growing needs, production chains, trade, and a big thing to build at the end. Its environmental problem would be salt. Irrigation helps your barley fields, but the fields slowly become salty and less productive. You have to let them rest.

That became Cradle. You start with a storehouse by the river. You build a town, feed it, dig canals, send a barge to trade for metal, and raise the Ziggurat of Enki. A full game takes about 50 minutes. It runs in the browser and uses my Rust engine, quaso.

A small Cradle town with reed houses, a canal, workshops and roads

The game was one experiment. The way we made it was another.

I kept the direction, the lead ran the work

I began with one agent, Ninshubur, as team lead. I gave him the design I had written before. We talked through what could fit into a small playable game. Four population tiers, several regions, and rival city-states belonged in the larger idea. For this prototype, we kept one map, two tiers, one trade partner, and the ziggurat as the goal.

Then I asked Ninshubur to "hire" a team. He gave each member a name, a job, a model, and an effort level. Effort was the amount of thinking time assigned to a task.

  • Enki handled the simulation and balance with Opus 5.5 at high effort.
  • Gilgamesh worked on architecture and difficult design choices with Opus 5.5 at medium effort.
  • Nisaba built code tasks with a clear design, using Sonnet 5.5 at medium effort.
  • Kulla handled art and worked with Codex for images, using Sonnet 5.5 at medium effort.

Each member had a separate Claude session that I could open. That part mattered to me. I could see the work, judge it, and step in if needed. In the normal flow I spoke only to Ninshubur. I chose the direction and playtested. Ninshubur split the work, sent tasks to the members through Claude's session messaging, reviewed what came back, and brought problems to me.

Once we had the task list, Ninshubur asked which tasks we should do next. I picked them, and he ran that work with his team on his own. He chose who did what, sent the messages, handled follow-ups, and checked the results. I got a report when a task was done or when he needed a decision from me. Over time he learned my taste from those decisions and could make more of them himself.

This was closer to working with a small team than giving one agent a giant prompt and hoping it could do everything.

The task file was the hand-off

We kept the project memory in an Obsidian vault, a folder of notes, under Git. Git gave us a history of changes. The vault gave each session the same small set of facts. Its MEMORY.md file was the index. It pointed to the rules, the current design, and notes about things we had learned.

The task list held open work only. Every task had its own file. It said what we wanted, how we planned to do it, what had to happen first, and what "done" meant. The lead would write the task before sending a member to work on it. When the task was finished, the file went away. Any fact we still needed moved into a lasting note.

If you want to try this, start with a task file about this small:

Goal: What should change?
Why: What problem does it solve?
How: What approach have we agreed on?
Depends on: What must exist first?
Done when: What can we check in the running project?

"Done when" was the important line. "Make farming work" is too vague to review. "Fields gain salt while working, fallow fields recover, the farm panel shows the loss, and save/load keeps the salt" gives the member and the lead something to check.

For a new task, a member started with a fresh conversation history and read the vault. For a follow-up on unfinished work, we kept that history. I stopped the lead once when he tried to clear a member before the second half of a task. The member still knew why the first half looked the way it did.

One code writer at a time

At first it sounds useful to have several agents code at once. They all edited the same copy of the game files, so their changes and commits could get mixed up. We settled on one code task at a time. Art could run alongside it because it touched different files.

We also learned where each model helped. For a code task that needed design choices, Gilgamesh would first read the engine and write a concrete plan into the task file. Nisaba could then build from that plan. That worked better than asking it to invent the design and write the code in one go. Enki took the heavier simulation and balance work.

Kulla ran the art side. Kulla used a small script to send requests to a visible Codex Sol 6.1 session for sprites and textures, then checked the files that came back. An image looking good in chat was not enough. Its size, transparent background, and tile shape had to work in the game. I still chose the look. For the game icon, Kulla brought me options. I picked one, asked to compare a white and blue version, and kept the first pick.

Cradle's trade quay and river town

The lead still had to check

When a member finished, it sent Ninshubur one short report: the commit, the choices it made, and what it had not tested. That last part told the lead where to look. Ninshubur ran checks, read changes, looked at screenshots, and launched the desktop game. Then I played it myself.

We needed every layer of that review. One desktop-only shader crash slipped past web tests. The first itch.io zip did not load its folders because PowerShell had put backslashes in the paths inside the zip. The lead once told me itch.io had no favicon setting. It does. Agents can sound certain and still be wrong about a simple thing.

The best fixes often came from playing. On the last evening I played on itch.io, wrote down several map problems, and left the team with the list. Code and texture work ran side by side. The lead checked the result and built a new zip. I tested that build the same evening. No amount of green tests would have told us whether the shore looked wrong or a sparkle was too sharp.

There were limits too. We had to watch the model usage budget before starting a long task. The lead's own reviews used budget. His session also filled up and was summarized several times, which made the vault more important. This took about a day to set up properly. Most of the rules above came from something going wrong, not from a perfect plan on day one.

If you want to try it

Start with a project small enough to finish. Write down what you want to decide yourself. Give one agent the lead role, and give each member a narrow job and a session you can inspect. Put the shared facts in short files under Git. Make each task small enough to review and give it a clear "done when".

Ask members to say what they did not test. Let the lead check their work, but check the parts only you can judge. Play the game. Look at the image. Open the zip. Keep one code writer in a shared copy of the files. If you use separate copies for parallel work, decide how you will bring the changes together.

From the first commit to the release build, Cradle took about five days and 126 commits. It is a small game, not the whole design I started with. But there is a town to build and a ziggurat to finish. More useful to me, I learned what kind of work I can hand to an agent team, and where I still need to sit down and play.

The completed Ziggurat of Enki and Cradle's win screen

If you want to see the result, play Cradle on itch.io.

The process blueprint

Here is the setup I would use to run this experiment again. The prompts below are examples you can adapt to your own project.

I choose the next tasks. The lead schedules and dispatches them, members do the work, and the lead reviews it. The lead handles fixes and the next selected task on its own. I get completion reports or requests for decisions.

1. Start with the lead and agree on a small game

Open one session for the team lead. Give it your design, the project location, and the tools it can use. Decide what the first playable version must contain, what can wait, and which decisions belong to you.

An opening message could be:

You are the team lead for this game prototype.
Read the attached design and the existing project.
Help me choose a small playable version.
Write down the scope and split it into tasks.
Ask me which tasks we should do next.
Once I choose, schedule and dispatch them yourself.
Manage follow-ups and review the results.
Report when a task is done or needs my decision.
Use my earlier decisions to guide similar choices.
I handle publishing.

2. Give every session the same memory

Create a shared folder of notes and keep it under Git. I used Obsidian, but the important part is that every session can read the same files. A small starting layout is enough:

memory/
  MEMORY.md          Rules and links to the other notes
  design.md          Agreed scope and game rules
  team.md            Names, roles, models and session IDs
  tasks.md           Open tasks and their owners
  tasks/             One file for each open task

Ask the lead to keep this memory current. When a decision changes, change the note too. A new session should be able to start from these files without reading your whole chat history.

3. Ask the lead to pick a team, then open its sessions

Ask for a name, role, model, and effort level for each member. Start with the jobs your project needs. For Cradle, that meant simulation, architecture, code, and art.

Create a separate session for each member yourself. Give each one its role, the project and memory paths, and the lead's session ID. Record the member's session ID in the team note so the lead can send messages to it. Check that one message and reply work before handing out real work.

For generated art, I also opened a Codex session that the tech artist could send requests to. The artist checked the saved images before adding them to the game.

4. Let the lead prepare the task files

The lead creates a task file, picks one owner, and adds the task to the open list. For example:

Goal: Add salt to farmland.
Owner: Simulation member
Why: Irrigation needs a cost the player can manage.
How: Working fields gain salt. Fallow fields recover.
     Use the rates agreed in the design note.
Depends on: Farming and save/load already work.
Done when:
- Salt lowers the harvest of a working field.
- Leaving the field fallow lets it recover.
- The farm panel shows the loss.
- Saving and loading keeps the salt value.

If the design is unclear, the lead sends the task to the architect first. The lead puts the resulting choices in the task before a code member starts.

The lead's message to a member can stay short:

Read MEMORY.md and follow the rules of work.
Your task is in memory/tasks/farmland-salt.md.
Complete it and check each "done when" item.
Report back here with:
- The commit
- Choices you made
- What you did not test
- Any blocker

5. Choose the next tasks, then let the lead run them

Once the task list exists, the lead asks which tasks you want next. Pick the work you care about. The lead then chooses the members, works out the order, and sends the tasks through session messaging. It handles that work on its own, including follow-ups.

The lead also checks the remaining usage budget, leaving room for review. In one shared copy of the project, it keeps one member writing code at a time. Art can run alongside code when the files do not overlap. The lead starts members with fresh conversation history for new tasks and keeps their history for follow-ups on unfinished work.

You speak to the lead in the normal flow. You can still open a member's session and step in when needed, but dispatching the work is the lead's job.

6. Review, then try it yourself

The member's report starts the review. The lead checks the changes against the task, reads the test results, runs the relevant builds, and inspects anything visible. If something fails, the lead sends a specific follow-up to the member and checks the fix.

The lead reports to you when a task is done or when it needs your decision. You can try the results from those reports while the lead continues with the other selected tasks. For a game, play it. For art, look at it inside the game. Before release, open the actual package. Give the lead concrete feedback, such as "this shore has a hard edge", so it can turn the problem into a task.

Keep your decisions clear. In Cradle, the lead learned my taste and could later make similar choices without asking me. Have the lead record useful decisions in memory so that understanding survives a fresh session.

7. Close the task and repeat

After review, the lead moves lasting decisions and useful lessons into the shared notes. It removes the finished task from the open list, deletes its file, and commits the memory changes so Git keeps the history.

The lead continues through the tasks you selected. When that work is finished, it asks what you want to do next. You choose the direction and priorities. The lead runs the team.

← All research logs