staffetta — the agent ends and hands the baton to a fresh session, instead of compacting (5.5k tokens per round, not 142k) #585
RaffaeleSpezia
started this conversation in
Projects
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Sharing a project for people who leave opencode running on long jobs.
The problem. Auto-compaction fires when there is no room left, which is almost
never the moment a piece of work is actually finished. It rewrites the beginning of
the conversation, which invalidates the prefix cache — so from then on every
question pays to re-read everything. In one real session an agent hit the ceiling
four times in two hours and ended up resuming from ~142,000 tokens on every request.
What staffetta does instead. The agent decides. At a milestone it writes what
the next session needs to a file, and ends — the session really does end, and a
fresh one picks up the baton reading ~5,500 tokens. Round seven starts from the same
size as round one, because the memory file is overwritten instead of accumulated.
Memory is three files on disk:
MEMORY.md(where I am, what comes next,overwritten),
JOURNAL.md(what was done, with measurements, and what must not betried again — append-only),
AGENTS.md(the rules). Because it is on disk, you canread it, diff it, correct it by hand, and steer the next round by editing it — and
unlike a chat message it persists.
The opencode part.
adapters/opencode/loop.shruns opencode headless and pipesthe output through
tee, so the run reaches your screen and a log file. In thatstate there is no interface at all: you watch everything live without being part of
the loop, and nothing waits for you to press a key. When the agent calls
ask, theloop stops and opens the real TUI with the question already written into the memory
file — you talk, it writes the decision down, runs
resume, and the stream picksup. The client stops being the system and becomes a tool the agent operates.
The loop itself decides nothing: it ends when the model runs
--doneor--blocked, never on an inferred threshold. The round cap and the per-round timeoutare spending limits, not judgements about the work.
Measured on a rented RTX PRO 6000 with KAT-Coder V2.5 served locally by
llama.cpp: 7 milestones closed autonomously, 7 of 7 checkpoints invoked unprompted,
587 tests green at the end from a clean shell.
https://github.com/RaffaeleSpezia/staffetta — Apache-2.0, pure bash, no
dependencies. Everything under
core/works with any agent that can run a shellcommand; the opencode adapter is one file, meant to be copied and changed.
Two things I would like to hear from people here. Has anyone tried steering a run
purely by editing a file between rounds instead of typing in the TUI — does it hold
up on longer jobs? And has anyone found a way to make an agent reliably notice that
a milestone is actually closed, beyond "the tests are green"?
All reactions