AI-written. No human has read this. Every requirement below is an agent's inference.
session: e0b89429 | 2026-09-14
When Claude Code runs out of context it writes a summary of the conversation and throws the conversation away. This plugin writes the summary instead, and staples three things to it that no summary can be trusted to remember: what was actually run, what the repository actually looks like right now, and what was promised and never finished.
It also keeps the conversation it replaced, so anything the handoff left out can still be searched for afterwards.
By the same author: changelogs.core-directive.com, a changelog for every Claude Code release, written from what changed in the build.
Six parts. Only the first is a model summarising a conversation.
- The summary, written by a fork of the session itself. The fork is asked to enumerate before it narrates: a tagged one-line-per-fact inventory of the whole conversation first, the prose reading second. That is the whole trick. Prose makes the writer choose what is interesting, and the fact nobody finds interesting is the one the next session needed. Through 0.4.x the fork was asked with Claude Code's own nine-section summariser instruction plus a Work ledger section; against the same 82-fact answer key, graded blind, the inventory prompt carries 70.0% where that one carried 63.0%, and it is the only arm whose worst run beat the old prompt's best. Through 0.10.0 the fork wrote that inventory twice, first in
<analysis>tags that were dropped before the handoff was assembled, then again as the summary's first part. Now it writes it once, straight into the summary. Over the same answer key that carried 69.4% against 67.7%, which is inside the grader's noise, and it took about a third fewer output tokens (16,777 against 24,756) and 182 s per fork against 257 s. The fork is also told to answer in one reply, call no tools and leave the session's task alone, because it inherits the session's tools and a denied tool call doesn't end it, it turns into another request that gets billed too. One fork that kept working on the session's task billed 144,745 output tokens over 24 minutes. An<analysis>block a model writes anyway is still dropped, along with the<summary>wrapper tags, andanalysisCharson the row says how much was dropped. A block the model never closed is left alone unless a summary follows it, because a reply cut off inside the scratchpad has nothing else in it. Since 0.11.4 the fork does not copy the user's messages into the inventory: every typed turn already comes back verbatim (item 5), and quoting each one on its[ask]line put every prompt in the window twice. An[ask]line now names the ask in a few words and gives its status;[constraint]and[rejected]lines still quote, but only the sentence that sets the rule, and an ask that survives only inside an earlier handoff keeps its line as it stands. - The tool ledger, read off the messages, not recalled: files written (through Write and Edit, and through the shell where the command's text says so: redirects,
tee,sed -i, the destination ofmvandcp,open(..., "w")in an inline script, unless it sits inside a string the script only holds or the heredoc it sits in feeds something other than an interpreter, such as acat <<'EOF'PR body (0.11.5); each file once, with every tool that wrote it, and counted as files rather than calls; since 0.11.5 a relative path the shell wrote folds into the one absolute path Write or Edit named that it is the tail of, and stays as written when none or several match), the newest 20 calls whose output reads as a failure, and the shell commands in order. Nothing here is an exit code, because the transcript does not store one, so "no error flag" does not mean "succeeded". Only the rows since the last compaction. Until 0.2.0 every earlier compaction's rows were merged in too, and by the tenth compaction of one session that was 65k of an 82k-character handoff, a quarter of a 200k window spent re-reading history every turn. Since 0.11.3 the shell list is trimmed as well: the newest 10 commands always, then older commands that changed something, clipped to 150 characters, while the ledger stays under 6,000 characters; read-only probes (ls,grep,git status,gh pr viewand the like) go first. A count line says how many were left out and names thehandoff_lookup n=N section=ledgercall that lists every one at full length, rebuilt from the rows the compaction's JSON kept. Over the 78 stored compactions with rows, the largest list (87 commands) went from 24,833 characters to 5,894. - Session state, read live from
gitandghas the handoff is written: branch, working tree, uncommitted files, the repository's open PRs by number and title (the whole repo's, not only this session's), the last five commits, the session's model, and any agents still running. A submodule holding nothing but untracked files is not listed as a change. It describes the moment of compaction and nothing else. A row that cannot be read is left out rather than guessed at. - Commitments and open questions, from one model pass over the assistant's own turns, on
claude-opus-5-5since 0.11.1 (Sonnet 5 before). The fact it looks for is an absence: "let me check X" is easy to find, but what matters is that nothing after it ever checked X, and no pattern sees that. Two rounds of prompt wording could not move that class of fact; one model call moved it from 35.7% to 73.8%. A step handed to the user is not counted as the assistant's, and one sentence is reported once, under its most specific kind. - Every user turn, verbatim, selected by the engine's own
handleand handed back as its words alone, so nothing is rebuilt from a paraphrase. Until 0.7.0 the turn went back with its handle, which hands the engine's own copy up whole, and that copy carries every attachment the turn arrived with: the instruction bundle (CLAUDE.md, every rule file, AGENTS.md, MEMORY.md), the hook outputs, the skill and agent listings. Measured on 2026-09-17 (session 7c6495a3, depth 8): a 22.7k-character handoff came back as a 124k-token first turn, and 218k characters of it were four copies of the instruction bundle riding on 19 pinned turns, one of them a 44k-character AGENTS.md. The engine re-emits that bundle on its own after a compaction, so every copy said the same thing twice. A turn without its handle is the person's words and nothing else; thehook_additional_contextlines those turns carried (the intent ledger'sop:ids here) go with the attachments. Measured live on 2026-09-17 (engine 2.1.274, the bench'stoolscheck): the transcript after the boundary held the handoff and two pinned turns of 2,155 and 577 bytes with nothing attached to them, the engine re-emitted its instruction bundle once, and the first real turn cost 58k tokens where 0.6.0 had cost 124k. (withHandlesis a field of the forced-run recordcompact_forcewrites toruns.jsonl, not of the index row.) - The files the next window should have open, chosen by the summariser (0.3.0). Claude Code's own compaction re-attaches up to five of the most recently read files, and that path never runs when a hook answers the event, so until 0.3.0 a handoff came back with none. Now the fork ends its summary with a
<restore-files>block naming up to five files, each by absolute path withallor a line range and a reason. The code only refuses: a path the session never read or wrote (since 0.11.2 a file a shell command named counts, so one read withcator rewritten by a script can come back; the command's word only has to end the requested path, and shell-named files never feed the recency fallback), aCLAUDE.mdorAGENTS.md(the engine re-emits those itself), a duplicate, anything past the cap, and since 0.11.1 a file no longer on disk (no longer exists on disk). When nothing usable was named, the most recently touched files that still exist go instead and the row sayssource: "recency". Each file is re-read through the real Read tool, so a file edited mid-session comes back current; a read that is denied, errors or takes over 20 seconds falls back to the text of the transcript's last Read of it (source: "stored"). Each one is handed up as a realReadtool_use and its tool_result, not a narration of one, so the next window treats it exactly as a file it read. Per file 20,000 characters, 100,000 in all, both clipped rather than dropped past the file cap and dropped past the total, and never past the size guard's ceiling. Every cap is a setting below, andrecord.restorecarries what would be needed to move one: source, requested, restored, every rejection and why, and per file the lines, chars, approximate tokens, clipped chars and milliseconds.handoff_statussums them underrestores.
Parts 2 to 4 are gathered concurrently and every one of them may fail. A part that throws, times out or comes back empty is left out and named in the row; the summary alone is still the arm that scored 67.3%.
Measured paired against byte-identical model text, 82 atoms from one real
session, three grading passes each, on claude-sonnet-5:
| arm | appended | recall | net | spread |
|---|---|---|---|---|
| armA | nothing | 67.3% | 66.1% | 3.0% |
| armB | tool ledger | 68.7% | 65.0% | 4.9% |
| armD | ledger + commitments | 71.1% | 67.5% | 1.2% |
It rehearses by default. Without COMPACT_HANDOFF_LIVE it does the whole
thing, writes down what it would have handed up, and then calls next(e)
anyway, so the engine compacts exactly as it does today. Every failure path does
the same. The worst case of installing it is the behaviour you already have,
plus a log.
~/.claude/compact-handoff/, outside any repository, because the plugin's own
root is a worktree somebody may delete. COMPACT_HANDOFF_DATA_DIR moves it.
~/.claude/compact-handoff/
index.jsonl every compaction on this box, one line each
window.jsonl one occupancy reading per turn, every session
sessions/<sessionId>/
NNN-<iso>.md the handoff that was handed up
NNN-<iso>.summary.md just the model's part of it
NNN-<iso>.transcript.md the conversation as it stood before
NNN-<iso>.json the row, the ledger, feedback, lineage
NNN-<iso>.post.json what the session did in its next ten person turns
runs.jsonl this session's compactions, append-only
lookups.jsonl every history read this session made, with its size
diagnostics.jsonl probes, forced compactions, A/B arms
rehearsals/<sessionId>/ the same, for runs that did not go live
Nothing is ever rewritten. Every file is written once and every log is appended
to, because two sessions compacting in the same minute must not be able to lose
each other's rows. index.jsonl is appended through sh -c 'cat >> ...',
because the host filesystem API has no append.
Each handoff's first line is a machine-readable lineage marker:
<!-- compact-handoff: session=<id> n=003 prev=002 -->
A later compaction reads it, records depth, and prepends an "Earlier
compactions" note naming every earlier pass and how to read it back. Only a message that opens with the marker and is not a tool result counts as a handoff: since 0.11.5 a marker quoted by handoff_lookup, a Read of a stored handoff or this README's example no longer sets the depth or ends a watch. Only the
newest handoff travels in the window. Everything an earlier compaction wrote
stays on disk behind handoff_lookup and handoff_search, so history costs
context only when a session asks for it, and every such read is logged:
~/.claude/compact-handoff/lookups.jsonl every history read on this box
~/.claude/compact-handoff/sessions/<id>/lookups.jsonl this session's reads
One row per handoff_lookup, handoff_search or handoff_list call: at,
sessionId, tool, args, chars, lines and approxTokens. The token
figure is characters over four, the same estimate the size guard uses, not a
tokenizer; the name says so. handoff_status sums them under lookups.
A compaction row says what a handoff cost to write. window.jsonl says what
it costs to carry, and it is the only file here written by sessions that never
compact at all:
~/.claude/compact-handoff/window.jsonl
One row per completed turn, whatever the session: at, session, turn,
first, phase (fresh, post-compact or stock-compact), compaction, tokens, window,
percent, messages, handoffChars, handoffTokens, handoffPercent.
Nothing is sampled and nothing is conditional, because the comparison only works
if both arms are there.
turn counts this session's turns in its current window: 1 is the first turn after the session started or after a compaction. phase and compaction come from the last opener in the transcript, a handoff this plugin wrote (post-compact, compaction its lineage n) or the engine's own continuation summary (stock-compact, compaction its position, handoffChars the summary's size), because $.session.messages() answers the whole session, every window since the first. The count lives in module memory, keyed by session. Through 0.11.4 it lived in $.store, one file every session on the box shares, so every session's turns ran up one counter and first: true fired once in 1,374 rows; nothing before 0.11.5 is usable for the floor comparison below. A reload the module never saw start (/reload-plugins mid-session) writes turn: null until the next compaction, rather than a number that means something else.
The reading that matters is first: true. On a fresh session that is the
floor every window pays before any work happens: system prompt, CLAUDE.md and
AGENTS.md, the tool declarations, the first message. On a post-compact
session it is that same floor plus the handoff and its restored files. The
difference between the two is the handoff's real price, and handoffPercent is
the handoff's own share of the window, so the two together separate what this
plugin costs from what the rules files cost. Added in 0.4.2 at the operator's
direction: "This tells us how much context is our compaction summary vs claude
rules and similar."
percent is computed to one decimal from tokens / window; the engine's own
whole-number figure is the fallback when a reading carries no window.
It wrote nothing at all until 0.4.3. The reading was taken off the
turn.complete event's own messages, and that event carries no transcript:
it is answer, durationMs, aborted, turnId and reason, whatever the
declarations imply. So every turn on engine 2.1.273 threw a TypeError, the
engine printed turn.complete hook skipped: threw and no row was ever
appended. The transcript comes from $.session.messages() now, which costs a
host round trip per turn, and a turn whose transcript cannot be read still
writes its row with messages and the handoff fields null.
For ten person turns after a compaction, turn.complete rewrites NNN-<iso>.post.json beside the compaction's row. No model call: kind (handoff, or stock for a main-session compaction the engine did itself, a rehearsal, a fallback or an aborted dispatch), anchored, turnsObserved (turns the person typed; tool results and harness messages are not turns), firstUserMessage, turnsToFirstToolCall (0 is a tool call before the person said anything), reRunCommands and reReadFiles (byte-identical commands and whole-file reads the pre-compaction window already did), handoffToolsCalled, compactedAgain and done. A compaction that arrives before any turn of the previous window ends writes that window's file first, with compactedAgain: true, so the quickest re-compactions are counted rather than overwritten. The stock arm is the baseline: the same watch over the engine's own compactions, which bench/summarise_runs.py prints beside the handoff arm. It is not a random sample, since a fallback happens for a reason, so a gap between the arms is a lead and not a result. A subagent's compaction is not watched.
The watch finds where the new conversation starts by the opener the compaction put in, counted in the whole transcript, and stops anchored: false rather than guess when that opener is not there. Every record written before 0.11.5 is unusable and the summariser sets them aside: the watch sliced the whole session from the replacement's length, so a record said 510 to 894 turns, 32 to 63 re-run commands and an empty first message. The monitor lives in module memory like the turn count, so a reload mid-watch ends it, and the file keeps what it last had.
Set abStockShare in /plugin (or COMPACT_HANDOFF_AB_STOCK_SHARE) to a fraction, with live on, and that share of sessions is left to the engine's own compaction while the rest get the handoff. 0.5 gives the most data per week; 0.2 keeps most sessions on the handoff. Clear it to end the split.
- Assigned per session, never per compaction. A session that alternated would hand its next handoff the engine's summary to build on, so the arm would stop being the only difference. The arm is a hash of the session id, so it survives a reload or restart with nothing stored, and raising the share only moves sessions from handoff to stock.
- The stock arm is logged like a handoff. Its row (
disposition: abStock,outcome: stock) is filed undersessions/<id>/with the pre-compaction transcript, the engine's summary as.summary.md,tokensBefore,tokensAfter,stockSummaryChars, the engine'susageand a pricedcost. The fork never runs in that arm; the seam is still raised, so a subscriber behaves the same in both. - Every record says its arm. Compaction rows carry
ab: {arm, share, bucket}(null when the split is off),post.jsoncarriesarm, and everywindow.jsonlreading carriesarm, before the first compaction as well as after, so the arms can be checked for balance before anything differs. - Read it with
python3 bench/summarise_runs.py --ab: per arm, sessions, compactions, how many got their assigned arm, elapsed, cost, summary size, context before and at turn 1 after, turns until the next compaction, and the ten-turn watch (turns to first tool, re-ran commands, re-read files, handoff tools used, compacted again). It groups by assigned arm, so a handoff that fell back still counts against the handoff.--ab-export ab.jsonlwrites one joined line per compaction, with the file paths, for digging past the table.
A rehearsal (live off) has no split: every compaction there is already the engine's.
A row says what the compaction did and, when it did not do it, why. The fields that matter for that second case:
| field | when | what it says |
|---|---|---|
disposition |
always | replaced, rehearsed, abStock (the A/B split left this session to the engine), fellBack, passedThrough, abortedFallback, skipped (the engine tried to compact this plugin's own fork loop and was declined). Historical rows only: overBudget, written by 0.10.0 and earlier when a session had spent past its USD ceiling. That ceiling is gone, and those rows still read back through every tool |
engine |
always | the Claude Code version, from $.session.version() since 0.11.5, with CLAUDE_CODE_VERSION the fallback. Null on nearly every earlier row, because that variable is unset in an ordinary session |
fallbackReason |
every disposition but replaced |
one line naming why, e.g. noHandoff: no handoff has been written yet (fork nothing-to-fork: the fork found no warm main-thread transcript) |
forkOutcome, forkDetail |
every row whose fork was not handed up | why the fork's own answer was not used, and the fallback went to the handoff on disk. On engine 2.1.280 and later an unanswered fork records the engine's own reason: nothing-to-fork (no warm transcript yet), api-error (the detail carries the HTTP status and error kind), empty-reply or aborted, and the spend of any that made a request stays in usage. timeout is a fork that had not answered after five minutes: the compaction stops waiting and falls back, but the engine gives a plugin no way to cancel a fork, so the request runs on and still spends, and its answer is dropped when it arrives (that spend is on no row). mismatch is an answer refused by the forkInput check below, threw anything unexpected. Rows from 0.10.0 and earlier say cold where they now say nothing-to-fork (or threw, on engine 2.1.280 and later) |
cost |
always | {forkUsd, commitmentsUsd, commitmentsBasis, totalUsd, cacheReadWaivedUsd, basis, forkUsage, model, priced, pricesTaken}, or null. Since 0.4.1 cache reads are stored in forkUsage.cacheRead and priced into cacheReadWaivedUsd at list, and never added to forkUsd or totalUsd, because on a subscription they cost nothing; basis says so. Rows from 0.4.0 and earlier charged them. commitmentsBasis is measured: the usage the result reported after 0.10.0, estimate: chars/4, no cache from 0.4.0 (and since, on a result that arrives with no usage), none when no commitments pass ran, measured on rows from 0.3.0 and earlier |
maxHandoff |
every replaced and rehearsed row |
the handoff ceiling the size guard used and how it was arrived at: {window, fraction, capTokens, chars, decider}, where decider is fraction, cap, default: window unknown or override; an override past the safe cap adds overSafeCapChars and overSafeCapTokens saying by how much |
forkInput |
every row, null on the rows that never forked |
what the fork was charged to read against what the session holds: {sent, cacheRead, contextTokens, matchesContext}. sent is input plus cache read plus cache write, which is the whole conversation for a warm fork and a prefix for a cold one; matchesContext is false when sent falls more than a fifth short of contextTokens, and that is the reading that refuses the fork's answer (forkOutcome: "mismatch"). Added in 0.6.0 |
forkContext |
every row that forked | what the session looked like the instant before $.model.fork: {context, model, messages, msSinceLastFork, subagentRanThisTurn}, where context is the whole $.session.usage().context object (tokens, window, percent) and msSinceLastFork is null on a session's first fork |
parts |
every replaced and rehearsed row |
per-part sizes and timings; for the commitments pass, commitmentsVia, commitmentsModel, commitmentsPromptChars, commitmentsReplyChars, commitmentsTokensEstimated (null when the cost is measured), commitmentsCostUsd, commitmentsCostBasis, after 0.10.0 commitmentsUnanswered (the engine's reason when the call came back with no reply) and commitmentsStatus (the HTTP status on an api-error, null when no response arrived), and since 0.4.0 commitmentsRows (how many rows came back) with commitmentsHitCap and commitmentsHitCapReason (length or truncated row) saying whether the reply stopped at the 8192-token output cap |
seam |
every row this plugin handled itself | {subscribers, results}, where each result is {name, outcome, elapsedMs} and outcome is ok, threw or timedOut; {subscribers: 0} alone when nothing subscribed. Absent on passedThrough and skipped rows, which never reach the seam, and on the historical overBudget rows of 0.10.0 and earlier |
costUnknownReason |
when cost is null |
why it could not be priced. A run is never priced at 0 because its usage was missing |
costNote |
when nothing was spent | no model call was made. A pass-through costs a real zero, which is not the same as unknown |
overCeiling |
always | the size guard could not fit the replacement, because the handoff itself is larger than the ceiling and is never trimmed |
restore |
every replaced and rehearsed row |
{source, requested, restored, rejected, files, chars, approxTokens, caps, ms}; source is model, recency or none, and each file's source is fresh, stored, failed or dropped |
bench/summarise_runs.py counts fallbacks by reason and unpriced rows by
reason, so a week of silent declines is a list rather than a number.
Six, served as loopback MCP:
| tool | what it does |
|---|---|
handoff_status |
What would happen if this conversation compacted right now: live or rehearsing, where the data is, how many compactions this session has had, what the last one did, what this session has spent, how much history it has read back (lookups), and what the restores have put back (restores). |
handoff_list |
Every compaction this session has been through, oldest first, with cost, size and files. |
handoff_lookup |
Read a stored compaction back. section takes summary, ledger, state, commitments or transcript; long sections page. |
handoff_search |
Grep every stored handoff and pre-compaction transcript of this session. This is how to find what a handoff did not carry up. |
handoff_feedback |
Record that a handoff was missing or wrong, against the compaction it came from. Every handoff names it, under the lineage line; through 0.11.4 none did and it went unused. The bench reads these. |
force_compact |
Compact now instead of waiting for the window to fill. |
Four more (probe_spawn, probe_budget, ab_fork, refresh_handoff) are
registered only when COMPACT_HANDOFF_DEV=1. They exist to measure the engine,
not to be used.
claude plugin marketplace add AnExiledDev/compact-handoff-plugin
claude plugin install compact-handoff@compact-handoffSince 0.4.3 this repository is its own marketplace. .claude-plugin/marketplace.json
lists one plugin whose source is the repository root, so those two commands are
the whole install on a machine that has never seen it, and neither of them needs
a clone. The marketplace is called compact-handoff after the only plugin it
carries, which is why the install id reads compact-handoff@compact-handoff.
claude plugin install writes user scope unless you pass --scope project or
--scope local. Claude Code picks the install up on its next launch, or on
/reload-plugins in a session that is already open.
Two environment variables belong in the env block of ~/.claude/settings.json,
and the install does not write them for you:
{
"env": {
"CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1",
"COMPACT_HANDOFF_LIVE": "1"
}
}Without CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 the runtime this is built on does
not exist and the module never loads. Without COMPACT_HANDOFF_LIVE=1 it
rehearses: it does the whole job, writes down the handoff it would have handed
up, and then lets the engine compact anyway. Rehearsing stays the default after
an install on purpose. The worst case of installing something that replaces your
compaction should be the compaction you already had, so turning it live is a
second, deliberate act.
To work on the plugin rather than use it, --plugin-dir still loads a folder for
one session and takes precedence over the installed copy of the same name:
git clone https://github.com/AnExiledDev/compact-handoff-plugin.git
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 COMPACT_HANDOFF_LIVE=1 claude --plugin-dir ./compact-handoff-pluginclaude plugin marketplace update compact-handoff
claude plugin update compact-handoff@compact-handoffThe first refreshes the catalog from GitHub and the second moves the install to
the version it now names; restart, or run /reload-plugins, to load it. Pulling
a clone upgrades nothing any more, because what runs is the copy the install put
under ~/.claude/plugins/cache/, and a clone is only what --plugin-dir reads.
Claude Code will load a plugin placed at ~/.claude/skills/<name>/, and until
0.4.3 this file said to install it that way, as a symlink to a clone. Do not. A
plugin loaded out of the skills directory was measured sitting inner on the hook
chain, the skills directory is not a documented way to install anything, and what
orders two plugins loaded that way was never identified. Renaming did not move
one: aa-probe, which sorts before compact-handoff, and zz-probe, which
sorts after, behaved identically. The marketplace install at the top of this
section is the documented path, it is the same on every machine, and it carries a
version claude plugin update can move.
Where a plugin sits in the chain matters more for this plugin than for most,
because it answers session.compact with the compacted conversation and never
calls next. Whatever sits beneath it on that event is never dispatched at all.
Within the user tier, chain position is the key order under enabledPlugins
in ~/.claude/settings.json, first key outermost. Measured on 2026-09-16
against engine 2.1.273, over two real compactions, with this plugin and a probe
plugin both installed from marketplaces at user scope. With the probe's key
placed above compact-handoff@compact-handoff, the probe's session.compact
hook was dispatched, its next.trace held exactly one link (compact-handoff,
tier user, outcome returned, 20473.8 ms), its own fork of the pre-compaction
transcript answered in 1957 ms, and this plugin still wrote
disposition: "replaced". Moving the probe's entry to the front of
~/.claude/plugins/installed_plugins.json while leaving the settings key where
it was did nothing at all: the hook was never dispatched, while that same probe's
session.start, tool.call and turn.complete hooks all ran in that session.
Since 0.4.3 that is the answer for a second plugin that has to see a
compaction. Its enabledPlugins key has to sit above
compact-handoff@compact-handoff, and the only way found to put it there is to
edit ~/.claude/settings.json by hand. claude plugin install appends a new
plugin last in both files, which is why a plugin installed after this one never
receives session.compact at all. Three things are worth knowing before you lean
on it. No CLI flag for placing a key first was found. Whether a later claude plugin install, enable or disable, or a settings write by the TUI, moves a
hand-placed key back to the end has not been tested, so read the order back after
any of those. And only a manual /compact was exercised, so the auto, plugin
and precompute triggers are unmeasured, as is a chain of three. Nothing read
the engine's code for any of this: it is a measured relation between input and
output on 2.1.273 and it could change in any release.
Both write-ups are private and neither is in this repository.
Since 0.6.0 a second plugin has another way in that does not depend on the key order at all. It subscribes to the seam this plugin exposes, and the next section is how.
This plugin is written against the type declarations Claude Code prints about
itself (/plugin-types, build 2.1.269), published alongside
cc-changelog-plugin.
The ledger, grading and A/B tooling under bench/ was measured on one private
transcript; the checklist and result tables for it are not in this repository,
only the numbers the README quotes.
Alone, this plugin is the whole compaction. It answers session.compact with
the conversation it wants the session to carry and it never calls next, so the
engine's summariser does not run and neither does anything sitting beneath it on
that event. That is what it is for, and it is also the problem: a second plugin
that wants to read a conversation the moment before it is compacted has nowhere
to stand. Within the user tier the chain order is the key order under
enabledPlugins and claude plugin install appends last, so a plugin installed
after this one is never dispatched session.compact at all. The measurement is
in the section above.
Since 0.6.0 there is a seam for exactly that plugin. This one adds a noun to $
through an engine.create fold, which is the fold that builds $, and the
declarations say every noun a plugin's step adds is on every plugin's $:
$.compactHandoff = {
beforeCompact({ tool, name }), // resolves { subscribed: true, tool }
version(), // resolves "0.9.0"
};The seam carries strings and nothing else. tool is the full name of a tool
your plugin answers with a tool.call hook of its own. When a compaction starts,
this plugin raises that tool:
$.tool.call({ tool, trigger, messageCount });Your hook then runs in your plugin's own environment, which is the whole reason
the seam is shaped this way. 0.5.0 took a callback and a callback cannot cross
this boundary at all: each plugin runs in its own environment, an interface
call's arguments go through cloneInto, and cloneInto throws DataCloneError
on a function. The engine's static scan refused both modules outright in the
same run, before any of that could even be reached. Strings clone. Functions do
not.
The transcript does not travel either, and that is deliberate. messageCount is
a number, the messages stay where they are, and your hook reads the conversation
the way this plugin does, with $.model.fork over the live session. It is the
same pre-compaction transcript and the same warm cache. Measured during the spike
that led to this, on a 57k transcript, a second fork beside this plugin's own
cost $0.0012 waived and took about 2 seconds, against $0.4995 for the same fork
taken from turn.complete after the compaction had already happened. Cloning a
megabyte of messages across the boundary would buy nothing a fork does not
already give you.
Here is a whole subscriber, written the way the memory plugin writes it, so that it still works with this plugin absent:
const SEAM_TOOL = "mcp__memory-handoff__before_compact";
export const register = (on) => {
on("session.start", async ($, e, next) => {
try {
await $.compactHandoff.beforeCompact({ tool: SEAM_TOOL, name: "memory-handoff" });
await $.store.set("seam", { present: true, at: Date.now() });
} catch (error) {
// With compact-handoff absent there is no such noun, and reading it
// throws a TypeError. That throw is the detection.
await $.store.set("seam", { present: false, detail: String(error), at: Date.now() });
}
return next(e);
});
// The raise lands here, in this plugin's own environment, where a fork and
// a database and anything else this plugin owns all work normally.
on("tool.call", { tool: SEAM_TOOL }, async ($, e) => {
const outcome = await remember($, e.trigger, e.messageCount);
// Never `next(e)`. The call exists for this hook and for nothing else.
return { result: { outcome } };
});
on("session.compact", async ($, e, next) => {
const seam = await $.store.get("seam");
// Subscribed, so the raise is bringing this same compaction along in a
// moment and working here as well would be two passes for one
// compaction.
if (seam?.present !== true) {
await remember($, e.trigger, e.messages.length);
}
return next(e);
});
};Write the calls out longhand, exactly as they are above. The engine scans a
hooks module before it loads it and refuses $.compactHandoff?.beforeCompact,
const seam = $.compactHandoff, and anything else that reads a noun of $
rather than calling an event on it. Both of those spellings are why there is no
typeof check here: the try/catch is the check.
Since 0.8.0 the manifest names a type contract, types/compact-handoff.d.ts,
and it is the same two signatures the snippet above shows, written where a
compiler can read them:
"types": "./types/compact-handoff.d.ts"Run /plugin-types in the subscriber's own checkout. It copies this file
verbatim to .claude/types/claude-code-plugins/compact-handoff.d.ts under a
banner naming this plugin, its version and its tier, and references it from
.claude/types/claude-code-plugins.d.ts beside it. Point the subscriber's
tsconfig.json or jsconfig.json at that folder:
"include": [".claude/types", "hooks"]and $.compactHandoff.beforeCompact({ tool }) is typed in the subscriber with
nothing copied and nothing to keep in step by hand. A contract that could not
be read, or did not itself typecheck, is listed at the end of the index with
the reason rather than silently skipped, and claude plugin validate checks it
before any of that: it answers types ./types/compact-handoff.d.ts declares on $: $.compactHandoff.
The generated folder is not committed here. It is a per-project index of
whatever plugins that project has enabled, so it belongs in .gitignore, which
is where this repo puts it.
name is what the row calls the subscriber and it defaults to the tool name.
Name it anyway. Subscribing twice for one tool is subscribing once, so a plugin
reload costs nothing, and there is no unsubscribe: a subscription lives for as
long as the session does.
What the seam will and will not do to you:
- A raise that throws is a field on the row and changes nothing else. This plugin's own outcome is the same either way, and the test that pins that compares both rows against a run with nobody subscribed.
- A hook that answers
{ deny }is recorded asdeniedwith the refusal on the row. Denying is a legitimate answer and it costs the compaction nothing. - A raise still pending after
COMPACT_HANDOFF_SEAM_TIMEOUT_MS(default90000) is abandoned. Nothing here can cancel one, so it is left running, its answer is ignored and the compaction goes on. - Nothing is raised on a compaction this plugin only passes through. A subagent's compaction is handed back to the engine before the seam is reached, and its row carries no
seamfield at all. - With nobody subscribed the seam is one array copy and the row says
seam: { subscribers: 0 }. No environment read, no timer, no cost.
The row carries what happened:
"seam": {
"subscribers": 1,
"results": [
{
"name": "memory-handoff",
"tool": "mcp__memory-handoff__before_compact",
"outcome": "ok",
"elapsedMs": 2104
}
]
}outcome is ok, denied, threw or timedOut. A denied carries the
refusal as detail and a threw carries the first 300 characters of the error,
and elapsedMs is measured from the moment the tool was raised, so a timedOut
reads a little over the cap.
memory-handoff is the
first subscriber. Either plugin runs with the other absent, which is the point of
building it this way: with this plugin uninstalled, the memory plugin takes its
own fork from its own session.compact hook and hands the compaction back to the
engine, and with the memory plugin uninstalled the seam here has zero subscribers
and costs nothing.
Both halves of that are measured on engine 2.1.273 (2026-09-16). The noun this
plugin adds at engine.create does reach the memory plugin's $: its
session.start subscribed through $.compactHandoff.beforeCompact in about
2 ms and read 0.6.0 back from version(). And a raised tool has to be
registered: raising mcp__memory-handoff__before_compact before the memory
plugin registered it came back refused with $.tool.call: no tool named "mcp__memory-handoff__before_compact" in this session, and once registered the
raise reached the hook and answered ok in 6030 ms, inside a compaction whose
own fork took 19293 ms. Both forks read the same warm cache, 50630 of 50961
input tokens on each row. The seam never shows up in the session transcript as
a tool use, though the registered tool is listed to the model, and the memory
plugin denies any call that does not carry the seam's own fields. Should the
noun fail to reach another plugin's $ on your build, the enabledPlugins
order documented above is measured and it works.
Since 0.9.0 the fork that writes a handoff is an agent type this plugin declares
rather than an anonymous general-purpose spawn carrying its instructions in the
turn. session.start calls $.agent.register and the engine hands back
compact-handoff:handoff, which $.agent.spawn then names.
Three things move from the prompt into the spec, where the engine enforces them instead of the fork choosing to comply:
tools: ["Read", "Write", "Grep", "Glob"]. The writer reads one transcript and writes one file. It could never edit, run a command or spawn anything, and now it cannot be asked to.omitClaudeMd: true. The writer needs the transcript, never the project's instructions, and this repo'sCLAUDE.mdpulls in anAGENTS.mdlarge enough to matter against a cold fork's window (see check 2 above). Dropping it is the single largest cut to what the fork is charged to read.background: true. The handoff is written between turns; it was already running out of band and the spec says so.
The standing instructions live in the spec's prompt, so the turn the fork
receives is two lines: the transcript path and where to write the answer.
A registration that does not take loses the agent type, never the handoff.
session.start records the outcome in the plugin's store, handoff_status
reports it under agent, and a false reading routes the spawn back to
general-purpose with the standing instructions folded into the turn exactly as
before 0.9.0. Nothing about a refused registration is silent and nothing about it
stops a compaction.
The type is registered and then hidden: on("agent.offer", { agent: "compact-handoff:handoff" }, () => ({ isOffered: false })) keeps it out of the
Agent tool's list, because a session that delegates its own work to the handoff
writer gets a handoff, not the work. $.agent.spawn still reaches it by name.
Both halves are verified live on engine 2.1.274 (2026-09-17), in a headless
session with nothing but this plugin loaded. handoff_status answered
{"registered": true, "agent": "compact-handoff:handoff"}, and asked to list
every subagent_type the Agent tool offers, the same session named five and
compact-handoff:handoff was not among them.
What the test suite cannot reach: claude plugin test supplies no engine
implementation for agent.register at all, so the engine's own schema never sees
the spec there. engine-test/agent.test.ts covers the two shapes that are
testable — the spec this plugin sends, and the fallback when the registration is
refused — and the live check above covers the third.
Every one of these is also a row in /plugin, under this plugin's configuration: live, dataDir, model, maxChars, maxFraction, maxTokens, restoreFiles, restoreFileChars, restoreTotalChars, seamTimeoutMs, abStockShare, subagents, dev, refresh and refreshMs. A row set there wins over the matching variable; left empty (or 0, for a number), the variable is read as it always was, which is what a cron line or a one-off shell invocation already sets. The four on/off rows (live, subagents, dev, refresh) declare no default, so a row nobody touched stays unset and the variable decides, while a row switched off is a setting and stays off whatever the variable says. Through 0.10.0 they declared false, which the engine filled in as though it had been set, and COMPACT_HANDOFF_LIVE=1 was never read.
There is no spend ceiling. Every main-thread compaction goes to the fork whatever the session has already spent. Up to 0.10.0 a session that crossed $10 fell back to the engine for every later compaction; maxUsdPerSession and COMPACT_HANDOFF_MAX_USD_PER_SESSION were how that ceiling was set, and both are now ignored, so an old config naming either one does nothing.
| Variable | Default | Effect |
|---|---|---|
COMPACT_HANDOFF_LIVE |
off | Off, it rehearses and the engine still compacts. On, it answers the event and the engine's summariser never runs. |
COMPACT_HANDOFF_DATA_DIR |
~/.claude/compact-handoff |
Where handoffs, rows and transcripts are kept. |
COMPACT_HANDOFF_MODEL |
the session's | The model the handoff-writing fork runs on. An alias (haiku) or a full id. The commitments pass does not read it; it is fixed at claude-opus-5-5. |
COMPACT_HANDOFF_MAX_CHARS |
unset | An absolute ceiling in characters that overrides the two settings below. Never clamped, because setting it is a deliberate act; past 800000 characters (the 200k-token safe cap) the row records overSafeCapChars and overSafeCapTokens rather than refusing. The summary is never trimmed, at any size. |
COMPACT_HANDOFF_MAX_FRACTION |
0.25 |
The share of the context window a handoff may take, at four characters a token. Values outside (0, 1] fall back to the default. |
COMPACT_HANDOFF_MAX_TOKENS |
150000 |
The most a handoff may be in tokens whatever the window, so a million-token window does not hand a quarter of a million up. Clamped to 200000. The ceiling is min(fraction × window, this) × 4 characters, and 400000 characters when the window cannot be read. |
COMPACT_HANDOFF_RESTORE_FILES |
5 |
How many files the summariser may have restored after the handoff. |
COMPACT_HANDOFF_RESTORE_FILE_CHARS |
20000 |
The most of one restored file that comes back; the rest is clipped with a note saying how much. |
COMPACT_HANDOFF_RESTORE_TOTAL_CHARS |
100000 |
The most all restored files may add together, and never more than the ceiling above leaves free after the handoff. |
COMPACT_HANDOFF_SEAM_TIMEOUT_MS |
90000 |
How long one raised seam tool may run before the compaction goes on without it. Read only when something is subscribed. |
COMPACT_HANDOFF_AB_STOCK_SHARE |
unset | The live A/B split: the share of sessions left to the engine's own compaction, 0.5 for an even split. Unset, 0 or anything outside (0, 1] is no split. Read only while live is on. See "The live A/B split" above. |
COMPACT_HANDOFF_SUBAGENTS |
off | Off, a subagent's compaction is passed through with a row saying so. On, it is answered like any other. |
COMPACT_HANDOFF_DEV |
off | Registers the four measurement tools. |
COMPACT_HANDOFF_REFRESH |
off | Keeps a fallback handoff on disk for sessions a fork cannot serve (headless). |
COMPACT_HANDOFF_REFRESH_MS |
600000 |
Shortest gap between two refreshes, and never sooner than 20 new messages. |
Every row carries a cost, priced from the table in hooks/lib.js. All four
numbers in all seven rows of that table were read on 2026-09-14 off the "Model
pricing" table at https://platform.claude.com/docs/en/about-claude/pricing,
which the code records as PRICES_SOURCE beside PRICES_TAKEN; none of them is
from memory. Usage the host does not report records cost: null and a
costUnknownReason rather than a zero, so a missing measurement can never be
averaged in as a free one.
The baseline is measured, not estimated. bench/baseline.py runs both arms over
one conversation inside one session: first Claude Code's own stock compaction
prompt (lifted from the 2.1.270 binary, committed as bench/prompts/baseline.txt)
sent through the very same $.model.fork, then a real /compact that is this
plugin doing the whole job. Same conversation, same model, same cache state, both
priced off the same table, so neither the model nor the prices can be what the
difference is made of. Baseline first is forced rather than chosen: a compaction
replaces the conversation, so the second arm can only ever see the first arm's
leftovers plus the /compact turn itself.
Both rows below are Sonnet 5, 2026-09-14:
| conversation | fixture bytes | cache read | engine's own | this plugin | of which fork | of which commitments |
|---|---|---|---|---|---|---|
| small | 23,532 | 71k tokens | $0.1268 | $0.1220 | $0.0861 | $0.0359 |
| large | 584,947 | 133k tokens | $0.0728 | $0.1567 | $0.0982 | $0.0585 |
Read those two rows together or not at all. The obvious reading, that the plugin is free on a small conversation and costs double on a large one, is not what happened. At these sizes the input is nearly all cache read and cheap, so the bill is mostly output, and the engine's own prompt simply wrote a much shorter summary on the large conversation than on the small one: 10,893 output tokens against 4,254. That is one run per arm per size and summary length is the noisiest thing in this whole system, so treat these as the order of magnitude (both arms are a tenth of a dollar on Sonnet, single digit cents on Haiku) and not as a stable ratio. What is solid is the shape: the fork is most of the money and the commitments pass is the rest, and neither depends much on how long the conversation is.
The nine live checks below cost $2.32 of plugin spend across 37 compactions, on Haiku 4.5, which is the number to have in mind for everyday use. That total is the whole bench ledger, including the runs that failed and were fixed, not only the nine that pass.
$.model.fork is documented to run over this session's own cache-safe
transcript snapshot, and when it does, the fork is warm: it is charged for
the whole conversation at the cache-read rate, so forkInput.sent lands at or
just above forkInput.contextTokens and almost all of it is cacheRead. That
is the cheap case the cost table above was measured in.
Some forks come back cold instead. The fork is charged for a prefix and for
tens of thousands of uncached input tokens, and it answers over a transcript
that is not this conversation. Across the 24 rows in one box's index.jsonl
carrying both readings, the split is clean and there is nothing in between: the
ten cold forks were charged for 0.40 to 0.47 of their session's context (65k to
78k sent against 164k to 167k held), the fourteen warm ones for 1.01 to 1.04 of
theirs. Cold forks cost more than warm ones, because a cache miss is charged
at the full input rate: each of those ten paid for 40k to 52k uncached input
tokens, against 15k to 20k of cache read.
The reply reads like any other summary, so nothing but forkInput can tell.
Since 0.6.0 a fork whose input is more than a fifth short of the context is
refused: the row records forkOutcome: "mismatch" with the two numbers, the
compaction falls back to the handoff on disk, and the tokens the fork already
spent stay on the row and are priced, because they were spent either way.
Every cold row measured so far was a trigger: auto compaction of a session
holding 164k tokens or more, and all ten are paired with a second
session.compact dispatch carrying an agentId, 0.07 to 0.17 seconds later,
over a transcript one message shorter and with one more pinnable turn in it.
None of the fourteen warm rows has such a partner.
SessionCompactInput.agentId is declared as "the id of the loop compacting,
for a subagent's or a fork's own transcript". notes/design/compact-handoff-cold-forks.md in the
claude-investigations repo holds the measurement and what it could not show.
The partner is the fork itself. On engine 2.1.283 $.model.fork runs the
ordinary query loop over the main thread's transcript plus the fork prompt,
and that loop is checked for auto-compaction like any other. A session that
crossed its threshold by less than the size of the fork prompt has a fork loop
over the threshold too, so the engine compacts the fork loop, and a plugin
that passed it through as a subagent had its fork answer over the engine's
summary. This plugin now answers that compaction with { skip } (the row says
outcome: "ownFork", disposition: "skipped"), and the fork reads the whole
conversation. It knows the loop by the fork prompt in the loop's transcript,
which a genuine subagent's never holds, so a subagent compacting at the same
moment still passes through. python3 bench/coldfork.py --unfixed reproduces a
cold fork in a sandboxed session on Haiku for about $0.50; without --unfixed
the same run comes back warm and replaced.
Nine checks, each a real throwaway claude session driven through a pty with
COMPACT_HANDOFF_LIVE=1, against an isolated CLAUDE_CONFIG_DIR and an isolated
data dir. python3 bench/verify.py run runs them; table prints what they last
measured. All nine pass as of 2026-09-14 on engine 2.1.270.
| # | check | what it proves | measured |
|---|---|---|---|
| 1 | manual |
a /compact is answered and replaced |
replaced, 29,236 chars, $0.051735, 83,834 ms |
| 2 | auto |
an automatic compaction dispatches a real event, not a precompute | trigger=auto, disposition=replaced, 2 rows, no precompute row, fork warm, source handoff |
| 3 | depth |
two compactions in one session merge the earlier ledger | depth=2, prev=1, 1 earlier ledger row merged |
| 4 | instructions |
/compact <instructions> reaches the fork |
replaced, summary 10,034 chars, ledger 6,781 chars |
| 5 | blob |
the size guard trims turns and never the summary | 46,646 -> 25,005 chars, 1 turn trimmed, ceiling 30,000, untrimmable handoff 23,178, summary intact |
| 6 | concurrent |
two sessions compacting in the same minute do not collide | 2 rows from 2 sessions in 1 minute |
| 7 | tools |
handoff_lookup and handoff_search work after a compaction |
both answered, and the monitor recorded both calls |
| 8 | resume |
claude --resume still carries the handoff |
lineage marker present in the resumed transcript, and the session echoed it |
| 9 | subagent |
a subagent's compaction is passed through with a row | agentId=a3ed584c4b8e44c8e, passedThrough, reason recorded |
Two things are worth saying about how check 2 was made to work, because both
were wrong for several runs and both look like plugin bugs when they are not.
autoCompactWindow is validated to 100k-1M and a value outside that range is
dropped in silence, so a bench asking for a 30,000-token window ran on the
1,000,000 default and never compacted at all. And this repo's CLAUDE.md pulls
in a very large AGENTS.md, so a session here opens around 80k tokens against
that 100k floor: two 10k-token reads crossed the line at ten messages, on the
turn that crossed it, which is the cold-fork case in the limits below rather
than a failure of the automatic path. Given files small enough that turns finish
before the line, the fork is warm on an automatic compaction exactly as it is on
a manual one.
What could not be run: nothing in the list. There is no check for a headless
(-p / SDK) compaction because the engine refuses to dispatch one at all, which
is the first known limit below rather than an untested path.
Since 0.4.0 the commitments appendix is one $.model.complete call: an API request the engine makes in process, with no session, no settings layers and no hooks around it, so nothing below can happen to it. What it spends never reaches $.session.usage().cost, so the row prices it from the usage on the call's own result, labelled commitmentsCostBasis: "measured: the usage the result reported". Before engine 2.1.280 the call returned the reply text alone and the row carried an estimate at four characters a token, labelled "estimate: chars/4, no cache"; a result that arrives with no usage still gets that. It is never recorded as 0. The history that follows is kept because the trap is still real for anyone who spawns the CLI.
Until 0.3.0 the commitments appendix was a claude -p call, which is a full session and
loads settings like any other, so every UserPromptSubmit hook on the machine
fires on a prompt nobody typed. On this box that meant the operator's intent
ledger recorded ten commitments prompts as sentences they had typed at a
keyboard.
--setting-sources "" fixes it by dropping the user, project and local settings
layers outright, which is the whole surface hooks and plugins are declared on.
The first attempt did not: it gave the call a CLAUDE_CONFIG_DIR of its own
with an empty hooks block, and CLAUDE_CONFIG_DIR moves the user settings
layer only. The call runs with cwd at $HOME, so $HOME/.claude/settings.json
was still loaded as the project layer out of $cwd/.claude/ - the same file,
through a door the config dir does not close.
Measured on 2026-09-14 against the real ledger, two probes: with the flag, zero
captures; without it, one capture of a prompt nobody typed. Then end to end, one
live /compact whose commitments pass really ran ($0.1400, a 12,569-character
prompt, 75.9 s): the ledger held ten before it and ten after.
Dropping the credentials copy that the first attempt needed is worth as much as
fixing the leak. It wrote to one path shared by every session on the box, so two
concurrent compactions raced - one release() deleting the file while the
other's claude -p still needed it - and a token rotation during the run would
have put the new refresh token in the copy and left the operator's real one
revoked, logging their own sessions out. Credentials are not a settings source,
so the ambient config dir authenticates the call with nothing copied at all.
- A headless session (
-p/ SDK) gets nothing.$.model.forkanswersnothing-to-forkwith no warm transcript, and$.session.compactrefuses outright: "not available in a headless (-p / SDK) session yet". The disk fallback behindCOMPACT_HANDOFF_REFRESHexists for this and is off by default. - A compaction that lands before any turn has finished forks cold, and that is a real thing that happens. The fork reads the main thread's cache-safe snapshot, and there is none until a turn has completed in this process. Two live shapes hit it. A session reopened with
--resumeand pushed over the window before it has answered anything: measured at 144 ms,fellBack/noHandoff/forkOutcome: nothing-to-fork(coldon rows from 0.10.0 and earlier). And a session whose very first substantial turn crosses the window while that turn is still in flight: measured four times on fresh sessions that compacted ten to fifteen messages in, same three fields. Nothing is lost and nothing is silent - the engine's own summary is what the session carries, and the row says exactly why. One completed turn is enough to fix it, which is why the automatic path is not itself the problem: check 2 above forks warm on atrigger: autocompaction at 61 messages. - A restored file is text or nothing. Images, PDFs and notebooks the
session read cannot come back through this path; the row says
failedwith the Read tool's reason. The restored pairs are real tool blocks, so the transcript viewer shows them as Reads the model made right after the handoff, which is what they are. - A subagent's compaction is passed through unless
COMPACT_HANDOFF_SUBAGENTS=1. It is a different conversation with a different owner and nothing here has been graded on one. - About half of what a summary carries is chance. Every number above is a mean over replicates for that reason, and a difference between two single runs is not a result.
- The unsourced-claim marker is not in this plugin.
bench/invent.pytried to mark the parts of a handoff nothing in the transcript supports, and failed as a measurement: it could not separate a claim the conversation never made from a claim it made in different words, so its markers were noise on text that was fine. It is left out rather than shipped as a warning nobody can act on. CLAUDE_CONFIG_DIRdoes not isolate a subprocess from your settings, and the way it fails is quiet. It moves the user settings layer. The project layer is$cwd/.claude/settings.json, so aclaude -pspawned withcwdat$HOMEloads$HOME/.claude/settings.jsonanyway - the same file, as a different layer - and every hook in it fires on a prompt nobody typed. An emptyhooksblock in the config dir you pointed at does not help, and nothing warns you. Anyone copying this pattern should use--setting-sources ""rather than a config dir, or at minimum spawn with acwdthat has no.claude/above it.- The commitments cost was an estimate from 0.4.0 to 0.10.0. Its spend is invisible to
$.session.usage(), and before engine 2.1.280$.model.completereturned text only, socost.commitmentsUsdon those rows is characters over four at the model's list price with no cache terms. After 0.10.0 it is priced from the usage on the result. Rows from 0.3.0 and earlier carry theclaude -pfigure the CLI reported. Summing a day's rows mixes the bases;cost.commitmentsBasisis how you tell them apart. - Rows written on engine 2.1.280 to plugin 0.10.0 have no commitments. The engine began resolving
$.model.completeas a result object and the plugin still read a string, so every reply read as 0 characters (parts.commitmentsReplyChars: 0) and the section was dropped. The same change made an unanswered fork throw instead of falling back by name, so those rows sayforkOutcome: threw. - The commitments reply is capped at 8192 output tokens, and that is a
behaviour change from
claude -p.$.model.completetakes amaxTokensof at most 8192 and stops there without saying so, where the CLI would run on. A session with more unkept commitments than fit loses the tail. It does not lose it silently:parts.commitmentsHitCapistrueon that row withcommitmentsHitCapReasonnaming why (length, ortruncated rowwhen the last line opens a row and never closes it) andcommitmentsRowssays how many came back. - The declarations are a version behind the engine.
types/claude-code.d.tsout of 2.1.269 declares a tool use as{id, name, input}; 2.1.270 hands asession.compacthook{tool_use_id, tool, input, result, text}. Readinguse.nameyieldsundefinedfor every call and nothing throws — it shipped that way once and rendered "2 tool calls: 0 file writes, 0 shell commands" over two real Bash calls. The ledger readsuse.name ?? use.tool, and the tests carry fixtures for both shapes. - Every timing on a row written before 0.4.3 is
null, and cannot be recovered.$.clock.now()is declared() => numberand returns a Promise at 2.1.273, sonow() - startedAtwasNaNandJSON.stringifywrote it asnull:elapsedMs,restore.msand everyparts.*Ms, on 71 of the first 72 rows on the machine this was written on. The same arithmetic drove three deadlines, so an abandoned pending handoff was never cleared, a refresh was never due on elapsed time, andwaitForFilenever reached its budget. Every duration isDate.now()since 0.4.3, which is right whichever the engine returns, and a stamp stored by an older version is read as no reading rather than compared against.
Three suites, and all three have to pass:
bun test # the module's own functions, against a stub $
claude plugin test . # the plugin loaded into a real engine
python3 -m unittest bench/test_summarise_runs.py # the A/B report's join
The hooks module is also typechecked against the engine's own declarations:
bunx tsc -p jsconfig.json
jsconfig.json includes .claude/types, which is where /plugin-types writes claude-code.d.ts. That folder is gitignored because the declarations are Anthropic's and are rewritten by every engine release, so run /plugin-types in this checkout first, from the Claude Code version the plugin is meant to run on. Without it import('claude-code') resolves to nothing and $ is untyped, which is how a change to what $.model.complete and $.model.fork resolve once went unnoticed. The functions that call the model carry @param {import('claude-code').EngineInterface} $ so that the compiler reads those results as the engine declares them.
bunfig.toml pins [test] root = "test", and that is load-bearing rather than
tidy. Bun's positional argument is a substring filter and not a directory, so
bun test test/ still matches engine-test/, and the run fails with Cannot find module 'claude-code/testing' on a suite that was never meant for it.
claude plugin test runs every *.test.ts under the plugin root, each file in
a child of the binary, in an environment like the one the hooks run in. That
buys the one thing a stub $ cannot: the seam resolved through a live engine,
with this plugin's engine.create fold actually folded and a subscriber
actually beneath it. engine-test/seam.test.ts is that test, and breaking the
fold (dropping version from the noun) turns all three red with
$.compactHandoff.version is not a function.
Two shapes in there are forced by the engine rather than chosen:
- The subscriber is an inline plugin (
{ plugins: [probe] }), not a hook registered in the test body. The static scan reads$.<noun>.<event>off a hooks module'sregister; a hook closed over by a test body is never scanned, and every call it makes on another plugin's noun is refused at the call site withits hooks module does not call it. - That
registercloses over nothing, not even aconstat the top of the file. It is loaded the way a module is, and a name from the test file's scope fails withPROBE_TOOL is not defined. What the probe learns comes back out through the tool result.
evals/ is a third runner, and unlike the two above it costs money and answers
a different question: not whether the code is right, but whether an agent that
has this plugin loaded actually reaches for it.
claude plugin eval . --ablation with-without --no-publish \
--allow-tools 'mcp__compact-handoff__*'
Three cases, two runs each, both arms: about $0.55 a suite on 2.1.274, and
no LLM graders at all, which is why it is that cheap. Every grader is regex or
tool_used, so the whole score is free and the only spend is the agent runs
themselves. The headline number is Δ, the with-plugin score minus the
no-plugin baseline.
01-handoff-statusasks what would happen if the conversation compacted right now. Δ +0.50: the baseline can still say the word "compaction", it just cannot answer.02-search-historyasks it to search everything stored from before a compaction. Δ +0.75, and the second grader is the interesting one — it fails a run that invents an answer instead of saying nothing was stored, which the baseline did once in two runs.03-neg-plain-questionasks for a haiku. Δ 0.00 on purpose: it is the guard that having these tools loaded does not make an agent call them at a prompt that has nothing to do with them.
A negative case is not padding. A plugin that fires on everything is a
regression this suite is meant to go red on, and tool_used with min: 0,
max: 0 and arm: both is the shape that catches it.
bench/ is how every number here was measured.
python3 bench/fixture.py <transcript.jsonl> --target 120000 # a disposable copy
python3 bench/run.py --session <id> --replicates 3 baseline terse
python3 bench/report.py # per-arm cost and size
python3 bench/summarise_runs.py --days 7 # a week of real compactions
python3 bench/verify.py setup <transcript.jsonl> --target 120000
python3 bench/verify.py run all && python3 bench/verify.py table
python3 bench/baseline.py run --source <transcript.jsonl> --target 120000 --label large
python3 bench/baseline.py table
-
bench/prompts/*.txtare the arms.baseline.txtis Claude Code's own compaction instruction, lifted verbatim out ofpretty-v2.1.270.js. The variant is read off disk inside the hook and never enters the transcript, so two arms differ only in what the fork was asked. -
baseline.txtis still current. 2.1.274 ships three compaction prompts where 2.1.270 shipped one, and the full-history prompt among them is byte-identical to this file (diff, exit 0). The two that are new are committed beside it:recent274.txtsummarises only the tail because the engine now keeps earlier messages intact, andhandoff274.txtis written to sit at the start of a continuing session. Neither replacesbaseline.txtas the arm to measure against, because neither is what a full compaction runs. -
bench/compare274.pyputs all four side by side — the three stock prompts and the plugin — over one fixture in one session, andbench/page274.pyrenders the row as a single self-contained HTML page:python3 bench/compare274.py setup <transcript.jsonl> --target 60000 python3 bench/compare274.py run python3 bench/compare274.py page --out /tmp/compaction.htmlOne caveat the page repeats, because it bounds what the run proves:
$.model.forkappends one user message to the whole session transcript and offers no way to hand it only the tail, so therecent274arm reads the same conversation the others do. It measures what that prompt's instructions produce, not what the engine's kept-tail path produces. -
None of the three reaches a
session.compacthook. The 2.1.274 declarations hand the hook the whole transcript undermessagesand carry no kept-tail field, so the engine's prompt split is internal to the path this plugin replaces and changes nothing about what it does. -
bench/summarise_runs.pyreads a week of real compactions and asks a model nothing: cost distribution, which parts failed and why, how deep lineages went, what sessions re-ran after a compaction, and every feedback note. Costs are the plugin's own priced figures, never re-derived, so this and the plugin cannot disagree about what a compaction cost. A row the plugin could not price recordsnulland a reason, and is counted as unpriced rather than as zero: a total with unpriced rows in it is a floor, not a cost. -
bench/verify.pydrives real terminal sessions over a pty against an isolatedCLAUDE_CONFIG_DIRand data dir, and asserts on the rows the plugin stored rather than on what the terminal printed. A toast on screen is not evidence that anything was stored.
Five things about driving a session unattended, each of which cost a run:
- An unattended TUI stops on three dialogs in sequence, each silently eating
every keystroke meant for the prompt. The auth chooser wants
hasCompletedOnboardingin the config dir's.claude.json; the trust dialog wants aprojects[<cwd>]entry withhasTrustDialogAccepted; the bypassPermissions warning is suppressed by no settings key at all and has to be answered — and onlyESC O Bmoves its selection, because the TUI puts the terminal in application cursor mode. - A session spawned from inside another inherits
CLAUDE_CODE_CHILD_SESSIONand writes no transcript at all, so there is nothing to resume and nothing for the plugin to read. Unset it, or setCLAUDE_CODE_FORCE_SESSION_PERSISTENCE=1. - A session's transcript lives under its own
CLAUDE_CONFIG_DIR. WithCLAUDE_CONFIG_DIR=/x,--resume <id>reads/x/projects/<encoded-cwd>/<id>.jsonl, not~/.claude/projects/. The encoding turns every character that is not a letter or digit into-, and the bench derives the directory from the cwd it spawns with rather than naming it, because a name guessed for one checkout answered "No conversation found" from a worktree of this repo (2026-09-17). - The cwd is the checkout the plugin sits inside, found by walking up to the
nearest directory with
plugins/and a.git. Two parents up was wrong from a worktree of this repo: it landed inplugins/compact-handoff/.claude, whose CLAUDE.md lookup walked up to the checkout's@AGENTS.mdand raised engine 2.1.274's "Allow external CLAUDE.md file imports?" dialog, which thehasClaudeMdExternalIncludes*keys seeded for that cwd did not suppress. The bench answers that dialog too if it appears, before the bypass warning. /compactleaves its own text in the input box, so the next thing typed compacts a second time with instructions. The first run of these checks recorded a row dispositionedinstructedon the word/quit.
A session.compact hook is handed the whole transcript, and what it answers
is the conversation afterwards. Return { messages } without calling next
and the engine's summariser never runs.
The handoff is written inside the event, by $.model.fork({ prompt }): one
tool-less completion over the main thread's own cache-safe transcript snapshot,
which is how the engine's own compaction reads a conversation. Nothing has to be
marshalled in and the prompt cache is already warm. It answers { text, usage },
or null when there is no warm transcript.
That works despite the ten second dispatch budget, because a host op's time in
flight does not count against it. Every $ call crosses the ops bridge, and
the bridge pauses the budget timer while the op is out. The ten seconds are the
hook's own compute, not its wall clock:
| probe | result |
|---|---|
tool.call, $.clock.sleep loop on a file nobody writes |
elapsedMs: 10052, aborted |
tool.call, $.process.run(["sleep","30"]) |
elapsedMs: 30018, aborted: false |
turn.complete, $.model.fork |
elapsedMs: 21768, aborted: false |
a real /compact forking inline |
elapsedMs: 21116, replaced, aborted: false |
The first row is the trap that hid this: $.clock.sleep is a local timer, not a
host op, so that loop really did burn the budget. The engine's constant is
var BSe = 1e4, a literal with no env var, flag or settings key behind it, and
plugin handlers never carry the budgetMs override the engine's own handlers
use. Ten seconds cannot be raised. It just does not mean what it looks like.
A precompute is declined on purpose. The engine can build a compaction
before it needs one and keep it for the compaction that comes. Waving that
through was a hole: an automatic compaction could be served an engine summary
this plugin had approved, and no row would ever say so. The precompute trigger
now answers { skip }, which forces the engine to dispatch a real event when it
actually compacts.
- A function that takes
$may not share its name with anything else in the file. A helperpath($, relative)beside a localconst path = ...fails with "the function handed $pathis declared more than once in this file". $.env.gettakes a literal name, so the variables a module reads can be listed. A constant holding the name fails with "got the variable LIVE_ENV".