Skip to content

Commit fcd4eb2

Browse files
committed
pstack: tighten eval skill preflights
1 parent 8d53b5c commit fcd4eb2

1 file changed

Lines changed: 1 addition & 1 deletion

File tree

  • pstack/skills/poteto-mode/playbooks

pstack/skills/poteto-mode/playbooks/eval.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,7 @@ Evals are blinded, one-shot bakeoffs for deciding whether to promote or reject a
2020
1. **Frame.** State the variant and the promote-or-reject claim. Write a judge-only rubric with 3-6 concrete criteria. Grade task success and the intended behavioral shape. Never make a turn-1 skill load, a particular file read, a citation, or "did the skill trigger?" a pass condition.
2121
2. **Author an organic prompt set.** Include at least one task where the behavior should apply. If the variant changes a description, routing, sticky behavior, or when-to-apply rule, include at least one task where it should not engage and add false-positive cost to the rubric. Write what a user would type, with no leakage of what is measured. If the task prompt itself is the target, write matched current and proposed versions here; otherwise every arm gets the same prompt.
2222
3. **Build comparison arms.** Before editing, snapshot any prior skill contents the control will need. Variant gets the proposed skill, structure, or prompt. Control gets the current version. For skill presence or content changes, the control may be skill-absent when absence is the realistic baseline, otherwise the snapshotted prior skill. Never plant the ablation-target skill into a skill-absent control. Hold the project skeleton, model mix, and every non-target input constant. The control is another sanitized label. Promote only when the variant beats the control on the rubric.
23-
4. **Set up isolated trials.** Fresh per-trial workspace with only that arm's variant and organic-task context. Identical project skeleton across arms. A fresh workspace does not clear skills from workspace `.cursor/skills/`, user `~/.cursor/skills/`, or plugin installs: for a skill-absent arm, use a workspace-local isolation (or disable that is restored before any other arm runs). Never apply a shared user/plugin disable that also strips the skill from the variant arm. Preflight resolved sources and fail setup if the skill remains visible on a skill-absent arm or missing on a variant arm. Record each trial's workspace path and transcript ID as orchestrator-only metadata. Cheap deterministic preflights aid synthesis only; they never replace the blinded rubric.
23+
4. **Set up isolated trials.** Fresh per-trial workspace with only that arm's variant and organic-task context. Identical project skeleton across arms. A fresh workspace does not clear skills from workspace `.cursor/skills/`, user `~/.cursor/skills/`, or plugin installs: for a skill-absent arm, prefer workspace-local isolation. If a user/plugin copy must be disabled, record its prior state and restore it in a guaranteed cleanup path before any other arm runs and on preflight failure, abort, or bakeoff completion. Never apply a shared user/plugin disable that also strips the skill from the variant arm. For skill presence or content changes, preflight resolved sources and fail setup if the ablation-target skill remains visible on a skill-absent arm, the proposed skill is missing on the variant, or a snapshotted-prior control does not resolve exactly that snapshot. Record each trial's workspace path and transcript ID as orchestrator-only metadata. Cheap deterministic preflights aid synthesis only; they never replace the blinded rubric.
2424
5. **Run 2-3 one-shot trials per prompt and arm.** Launch each runner directly in its recorded workspace with that arm's isolated context. Fan out in parallel with no shared grounding and no candidate-visible files across workspaces. Match model and trial pairings across arms. Ask only for the organic task output, not a graft rationale. Missing output fails the trial. No retries, coaching, or repair. When budget binds, prefer 2 trials on fewer models over 1 on many.
2525
6. **Spawn one blinded judge** on a different model family after every trial finishes. In one pass, score every output by randomized sanitized label against the same rubric. Mark each criterion and output pass or fail. Do not run the arena pick/graft workflow. This bakeoff ends at arm-level scoring.
2626
7. **Inspect transcripts after scoring to explain how, not to decide pass or fail.** Read only the recorded transcript for each trial from that workspace's transcript directory (normally `~/.cursor/projects/<trial-workspace-slug>/agent-transcripts/`), using the session or transcript ID from setup. Derive the slug from the recorded workspace path. Do not glob across `~/.cursor/projects/*/` or open unregistered workspaces. Transcripts verify isolation and explain the output. They are not a pass gate.

0 commit comments

Comments
 (0)