Skip to content
The 31 guidesFREN中文
Guides
Part 4 · guide 5 of 8 Level: Intermediate Reading time: 12 min Platforms: Linux

Running several agents together

One session per project, plus one that reviews the others. Why a second pair of eyes catches what the first one misses, how to set it up, and where the limits are.

In this guide
  1. 01Why a second agent sees what the first one misses
  2. 02The cost sits in the detection delay
  3. 03The rule that holds it all together: the right to refuse
  4. 04Isolate one variable at a time
  5. 05Setting it up
  6. 06What you actually tell the reviewer
  7. 07The limits, and they are real
  8. 08Frequently asked questions

In short

Run one agent session per project, each inside tmux and within its own scope, and add, if needed, a review session that writes no code: an agent that didn't write the code spots what the author no longer sees. The rule that holds it all together is the right to refuse: a session that disagrees says so with numbers, even to the supervisor, and a permission denied to one session is never carried out from another. Claude Code lets sessions on the same machine message each other; with Codex or OpenCode, a shared folder does the same job. The setup costs more and is only worth it with measurements.

Do this first: Working remotely (terminal + VNC)

One August morning the mini-PC ran out of memory and had to be rebooted in a hurry. Five agent sessions spent the day repairing the damage, one per project, plus a sixth that wrote no code and simply reviewed the work of the others. That supervising session got things wrong three times during the day, and all three times a project session corrected it by producing better measurements.

The method described here came out of that day, and its principle fits in one sentence: several agents checking each other beat a single agent handing out orders.

Why a second agent sees what the first one misses

An agent that has just written code is badly placed to judge it. It knows its own intention, so it reads what it meant to do rather than what it actually produced. You already have that reflex of mistrust with a human colleague, and it applies in exactly the same way to a machine.

An agent that produced nothing on a project arrives without that intention in mind. On that August day, the reviewer spotted three things the sessions involved were looking at without seeing.

A backup script was swallowing the exit codes of the command it launched, so the database import crashed while systemd reported a success. A monitor erased its own alarm every hour: it warned on Slack at 6:48, then announced a recovery at 7:07 without having retested anything. An advertising campaign, finally, had been repaired without anyone thinking to add the slightest detection. The fix stopped that exact bug from returning, and said nothing about the next one.

None of those three findings called for being smarter. They only called for not having written the code.

The cost sits in the detection delay

That day brought up four incidents spread across four unrelated projects, and they all looked alike. An import was failing while reporting success, while a monitor erased its own alerts. Elsewhere, a paid campaign had been asleep for thirty-three hours without anyone noticing. A gateway, finally, had been dropping part of its own response for months. The symptom was written in plain words in a comment, right above the guilty line.

Once the problem was identified, none of those fixes took more than a few minutes, which means the real cost sat entirely in the time it took to notice.

The rule that holds it all together: the right to refuse

The natural reflex is to turn the reviewer into a boss handing out orders. That is the most common pattern in multi-agent frameworks, and it would have broken everything in our case. One anecdote will serve better than a demonstration.

The reviewer measures a service running slow and finds a correlation between the time elapsed since the last call and the response time. It concludes the cache expires too fast, proposes a longer cache lifetime plus an automatic retry, and announces all of it with a great deal of confidence.

The session that owns the service does not apply it and prefers to measure first. The cause lay elsewhere: the model was spending that time reasoning, and the gateway sent nothing at all while it reasoned.

The interesting part is that the proposed fix would have worked had the session obeyed. The monitor would have gone green, everyone would have been satisfied, and users would have carried on waiting five minutes in front of a blank page. For a reason nobody would have been looking into any more.

Isolate one variable at a time

The reviewer’s mistake deserves a pause, because it reproduces very easily as soon as you work under pressure.

It had three measurements, and it varied the elapsed time between each one while letting every other condition fend for itself. It watched a correlation appear and turned it into a cause, which is the most tempting shortcut in experimental method.

The session that corrected it went about it the other way round, varying one single thing while holding everything else still. It measured the effect of the model’s reasoning at identical cache state, then the effect of the cache at identical reasoning. Those two experiments settled the matter where three observations had produced nothing but an attractive hypothesis.

An agent excels at manufacturing a plausible explanation from three data points. Ask it for the experiment capable of invalidating that explanation.

Setting it up

0 of 5 steps done Your ticks stay in this browser.

  1. One session per project, each inside tmux

    This is the most important point of all, and also the most plainly practical. A session started in a bare SSH connection dies as soon as your laptop goes to sleep. That day, the supervising session went exactly that way, without a sound, taking all its timers with it. Nobody realised until the laptop was reopened.

    So start each agent in a tmux session named after its project (tmux new -s my-project). The how-to, detaching and coming back, is in Working remotely. One reflex worth keeping: echo $TMUX tells you whether the session you’re in is protected (an empty answer means it isn’t).

  2. Give each session its own territory

    Each session works on its project and refrains from writing anywhere else. Two of them therefore cannot modify the same file without knowing.

    On the day of the incident, the reviewer had fixed a file belonging to another project, and the owning session was about to commit it believing the work was its own. A single message cleared the misunderstanding before it turned into a riddle inside a git blame.

    Hence the rule: if you touch another session’s territory, you tell it before it finds out on its own. The same goes for a service you restart, or billed calls you consume without warning.

  3. Let them talk to each other

    Claude Code can send a message from one session to another on the same machine, with nothing to enable. The /list-agents command shows the sessions it can reach; each one answers to the name you give it with /rename, or at launch with claude --name my-project. The agent writes on its own as soon as you ask it to coordinate something or pass an item of information along, and you can pick the recipient by typing @ followed by its name.

    In practice you simply tell your supervising session to let @website know that you have just changed the configuration. A message from another session never counts as your consent: it can neither approve a permission prompt nor change the configuration, and Claude is instructed never to ask another session for something it was refused. The official docs cover the settings.

  4. Write the method into the global memory file

    So that all of this becomes automatic instead of being re-explained to every new session, install the rules in your global memory file, the one every session reads at startup: ~/.claude/CLAUDE.md for Claude Code, ~/.codex/AGENTS.md for Codex, ~/.config/opencode/AGENTS.md for OpenCode. See Memory files.

    Keep it short: a few lines pointing at a more detailed document beat a whole page reloaded into every context, since everything you write there takes up room for everybody.

  5. Add a supervising session when it earns its keep

    It is in no way mandatory, and it costs. It becomes worthwhile when several projects move forward in parallel, or when the machine hosts services that have to stay up while you sleep.

    Its job consists of reviewing what the others do, checking the machine’s health at regular intervals and reporting what it observes. It refrains from coding on the projects it watches, failing which it loses precisely what made it valuable.

    If you want an agent that keeps watch around the clock and messages you on your phone, resident agents like Hermes or OpenClaw are built for that, with their own messaging and scheduled tasks: see Hermes & OpenClaw.

What you actually tell the reviewer

A supervising session launched without instructions confines itself to commenting. Here is the brief that ended up working, to adapt to your own situation.

You supervise the other sessions on this machine. You write no code on any of their projects. Your job: review what they do, check the machine’s health at regular intervals, and report what you see to me.

Before asserting a cause, show me the experiment that isolates it. A correlation over three points stays a hypothesis, so announce it as one.

When you audit another session, read the commands it executed and the actual state of the files, never the labels it gave its own actions.

If you touch a file outside your territory, warn the session concerned before it finds out.

If a permission is denied to you, you stop and you tell me. You never run in another session’s place what was denied to it.

The two middle instructions come from real mistakes. The supervisor concluded a cache had expired on the strength of three uncontrolled measurements, then reported a force push that had never happened, because it read the label of a command instead of the command itself.

The limits, and they are real

Nobody watches the watcher, and that is the first problem. The supervising session dies like any other, and its disappearance is silent by definition. Run it inside tmux, then give it a companion probe that depends on no session at all, for instance a system timer that messages you. That is what was put in place that day, but after the fact.

Then comes the question of cost. The economics of the classic pattern rest on an expensive boss driving cheap workers, whereas here nobody occupies the worker role, which drives the bill up and rules out making it a daily reflex. To lighten the bill, Quelle IA’s comparison tool (in French) puts prices and scores side by side; check in its Agents ranking that the model you pick can hold a long task on its own.

The reviewer, finally, does not know the ground. On the three subjects where it got things wrong, the project session knew its own code far better, and the reviewer’s value comes solely from having no stake in the work, never from any kind of superiority.

One point remains that looks obvious written down, yet becomes tempting in the heat of the moment: a permission denied to one session is never run from another. When one of them is blocked on an access its neighbour happens to have, the urge to work around it arrives very fast, and it hollows out the guardrail you took the trouble to install. Better to stop and decide yourself.

None of this works without measurements, in the end. The three corrections of that day rested on numbers. Two agents debating opinions produce mostly convincing noise.

Frequently asked questions

How can I tell whether my agent session will survive an SSH disconnect?

Type echo $TMUX in the session: an empty answer means it isn't running inside tmux and will die with your connection, for instance as soon as your laptop goes to sleep. So start each agent in a tmux session named after its project, with tmux new -s my-project.

What should you tell a supervision session?

That it writes no code on any of the projects it watches, and that its job is to review what the other sessions do, check the machine's health at regular intervals and report what it sees to you. Also ask it to show the experiment that isolates a cause before asserting it, and to read the commands actually run instead of the labels given to the actions. If a permission is denied to it, it stops and tells you.

Where should rules shared by all agent sessions go?

In the global memory file, which every session reads at startup: ~/.claude/CLAUDE.md for Claude Code, ~/.codex/AGENTS.md for Codex, ~/.config/opencode/AGENTS.md for OpenCode. Keep it short, because everything you write there takes up space in every session's context. A few lines pointing to a more detailed document are enough.

Who watches the supervision session?

Nobody, and by definition its disappearance goes unnoticed. Run it inside tmux, then add a probe that depends on no agent session, for example a system timer that sends you a message.

What question should you ask an agent after every fix?

If it fails anyway, how long will it take me to find out? An agent responds very well to 'fix this', but it rarely thinks about detection on its own, so you have to ask for it. Then test the alarm by actually triggering it, because reading its code isn't enough to prove it works.

Terms in this guide: Agenttmux

Spotted a mistake?

A command stopped working, a price changed?

Tools change every month. Tell me what is wrong in this chapter and I will fix it and update its date.

Only the page, your message and the optional contact are kept. Nothing else.

Guide 23 of 31 · part 4 no guides read yet Open the list of guides