I build apps by assigning roles to AI agents: implementation, QA, design, legal, secretary. Each one runs in its own session, and because they work in parallel, things move fast.

On August 15, that parallelism turned on me.

While the implementation agent was running E2E verification on the picture book app, the secretary agent started taking screenshots for the store listing. On the same device. Neither was aware of the other, and for several minutes they fought over the same screen.

What happened

I only have 1 Android device for testing.

moto g24 / ZT322LT4K9 / Android 14 / 720×1612

To take screenshots, the secretary needed to get the screen into a particular state, and it was driving the device directly with adb.

adb -s ZT322LT4K9 shell input tap 540 1200
adb -s ZT322LT4K9 shell am force-stop com.example.app

That force-stop killed the test session the implementation agent was running, partway through. A half-filled form vanished, and that day’s 1 free story was used up on something that had nothing to do with the verification. (In this app, free users get 1 story per day, so once it’s used up, you can’t test the same path until the next day.)

No error appears. All the implementation side sees is “the test that was running a moment ago suddenly jumped back to the start screen.” You start by suspecting your own code, so the road to the real cause is long.

The cause wasn’t scheduling. It was resource visibility

At first I thought, “just stagger the work times.” That was wrong.

The problem was that nowhere was it written, for anyone to see, that the device was a shared resource. Both the implementation agent’s tasks and the secretary’s tasks said “use the real device,” but nothing made explicit that both referred to the same single device. A human would know just by looking at the desk. An agent in a separate session can’t see that.

There were other shared resources with the same structure.

  • Play Console: if two agents edit a draft at the same time, whoever saves last overwrites the other
  • Free quotas and API quotas: if verification consumes them, the other side’s test fails “as designed”
  • The git working tree: one switches branches while the other is in the middle of a build
  • Simultaneous builds: they fight over the Gradle daemon and locks, and both slow down

None of these show up as “something broke.” The worst part is that the other agent’s work looks, from your side, like a failure that’s working as designed.

The rules I set

I didn’t build an elaborate locking mechanism. 3 lines are enough.

  1. Before touching the real device, check what other sessions are doing. Whoever is going to use it announces that first
  2. Fix the priority order. The implementation agent’s verification always comes first. Screenshots only after an “it’s free now” message
  3. If you change state, always put it back

The 3rd one comes from a specific incident. To get clean screenshots, I sometimes use SystemUI demo mode.

adb shell settings put global sysui_demo_allowed 1
adb shell am broadcast -a com.android.systemui.demo -e command clock -e hhmm 0900

This pins the status bar clock at 9:00 and shows full signal. Great for screenshots, but if you forget to exit, the clock stays frozen. The next agent reading logs gets confused because the device’s clock doesn’t line up with the log timestamps. I actually did this once.

adb shell am broadcast -a com.android.systemui.demo -e command exit

Never run an operation that temporarily changes the environment unless it’s paired with the operation that reverts it. I made that an operating rule.

Destructive operations need both an “announcement” and a “restore report”

The same month, I caused another incident with the same root cause. The QA agent, for verification, sideloaded a debug-signed build onto a device with production data on it, and forgot to switch it back to the Play Store version. In the course of recovering, we ended up digging up yet another problem.

All we had at the time was the advance announcement: “I’m about to sideload.” There was an announcement beforehand, but no restore report afterward. So nobody could notice it hadn’t been put back.

I changed the rule to this.

Destructive operations on a device with production data:
  announcement before  +  restore-complete report after   ← done only when both exist

With only one of them, you can’t tell from the outside whether it’s finished or not. It’s the same for automation and for human operations: a process with a “start” log but no “end” log gets left unattended.

The first thing to do in a multi-agent setup

I parallelized sessions to go faster, but if you burn half a day recovering from an accident, there’s no point. Before parallelizing, the first step should have been to list everything that’s shared and holds state.

The test is simple: “If one side touches it, does the world the other side sees change?” If so, it’s a shared resource. The real device, the store admin console, quotas, the working tree, and the production database.

Conversely, anything read-only, or anything that can be duplicated per agent, never causes accidents no matter how much you parallelize. Separate what can be separated, and protect only what can’t with rules. If I’d drawn that line at the start, that afternoon wouldn’t have been wasted.


What worked and what was a waste in this setup is in my write-up on 30 AI employees, and the setup itself is covered in how I built 3 apps in 2 months while holding a day job. The time I got stuck on the screenshot spec itself is here.