Since the start of 2026, there has been a flood of articles explaining “how to build AI employees.” There are more than enough already.

So I won’t write about how to build them. I’ll write about what actually worked, and what was a complete waste, after running it for 2 months. Incidents included.

Here’s the conclusion up front. What worked best wasn’t splitting up roles. It was a boring rule: “write decisions to files, don’t bury them in conversations.” When I didn’t follow it, 3 apps broke at the same time.

The current setup

I split sessions by role. Per app, the structure looks roughly like this.

                   Human (makes decisions only)
                                 │
                                PM
       ┌──────────────┬──────────┼──────────┬──────────────┐
Implementation      Legal     Design     Billing        Growth
                                 │
      Secretary (tracks progress, identifies tasks for the human)

Each app has its own set of these, so there are around 30 “employees” running at any given time. The number is high not because it’s impressive, but because I never created roles that span multiple apps (as I’ll explain later, this also caused an incident).

What worked

1. Splitting roles surfaces issues you didn’t ask about

This had the biggest effect.

Back when one session did everything, it could write code, but it never once raised something like “Doesn’t this data collection contradict your Data safety declaration in the store?” Because I never asked.

The moment I made legal its own role, these kinds of issues started coming up. In fact, the legal agent raised 4 launch blockers.

One of them was that an API key was embedded in the app itself. Had I released it, anyone could have decompiled the APK and hit the AI API on my billing account. I hadn’t noticed. The implementation agent didn’t point it out either (naturally, since that’s not its role).

Give something a role, and it starts looking at things from that role’s perspective. Exactly like a human organization.

2. “Write decisions to files”: this decided whether the setup lived or died

When multiple sessions run in parallel, you will inevitably end up with session B not knowing what session A decided.

I underestimated this, and it actually blew up.

For one app’s sake, a session changed my GitHub account name. My real email address was left in the commit history, and this was to get rid of it. As a decision, it was right.

But GitHub Pages domains are tied to the account name. That change turned the privacy policies and terms of service for the other 2 apps into 404s. The session that made the change didn’t know the other apps referenced the same domain.

As a result, I couldn’t submit to Google Play for review. I wrote up the details in a separate post.

The fix is almost anticlimactically boring. Never leave decisions in the conversation; always write them to a file. The next session can catch up by reading that file.

Whether you do this or not completely changes how long the setup survives. Honestly, it felt more important than the role split.

3. The moment I delegated decision authority, the bottleneck stopped being me

At first I had them check everything with me. “Is this wording OK?” “Is this structure OK?” Work stopped every time they asked.

Before I knew it, the bottleneck was entirely me. 30 agents working, and all of them stalled waiting for my reply.

Now I use this rule.

Unless an operation is fatal or irreversible, proceed on your own judgment without checking with the human.

There are only 3 things they’re allowed to check with me about: deletion, operations that actually incur charges, and publishing externally. For everything else, they just report the result.

It felt several times faster. It also means that me deciding on phrasing or implementation details was never valuable in the first place.

What was a complete waste

1. Starting from an org chart

This is the one people tend to do first. Draw a clean org chart up front: “There’s a PM, and under it there’s implementation…” I did it too.

It barely worked.

The reason is clear: I hadn’t done that work myself yet, so I didn’t know what roles were needed, and I was just lining up job titles. You get a plausible-looking org chart, but it doesn’t match the work that actually comes up.

What worked was the reverse order: do it yourself first, then create a role where you get stuck.

I created the legal agent because I got stuck on the store review policy requirements. I created the billing agent because I couldn’t make decisions about pricing. Getting stuck comes first, the role comes after. No other order worked.

2. Using sessions for too long

The longer the conversation, the slower the responses, and the worse the judgment.

I kept stretching sessions out because ending them felt wasteful, but it was a loss. You tell yourself “but I built up all this context,” yet most of that context is leftovers from finished work.

Now, when a session gets bloated, I write out the key points and hand off to a new session. My conclusion from actually measuring it: the cost of continuing with a slow session is bigger than the effort of writing a handoff doc.

3. Running everything on the heaviest model

I was throwing routine tasks and mechanical implementation at the smartest model too.

Just separating the steps that need deep judgment from the ones that don’t improves both cost and speed. I should have noticed this sooner.

So what did the human’s job become?

Only making decisions and taking responsibility.

AI gets through a surprising amount of work, but it won’t decide “is it OK to release this app?” In fact, the more you split roles and run them, the more decisions come up to you.

The human’s work doesn’t shrink; it shifts from tasks to decisions. That’s my honest takeaway after 2 months. If you go in expecting “handing it to AI will make things easier,” you’ll probably be let down. The density of decisions clearly went up.

And one more thing. AI will tell you things you don’t know, but it won’t tell you things you didn’t ask.

The GitHub incident was exactly that. Ask it to “change the account name,” and it will. Unless you ask “what breaks if I do that?”, nobody counts what’s going to break.

For now, knowing what to check is still the human’s job. Adding more roles didn’t take that off my plate.


The 3 apps built with this setup are in Works. The actual cloud costs of running them are all published in this post.

Next, I plan to write about how many times the apps built with this setup failed Google Play review.