Skip to content

feat(bot): /agentdiff approve — in-PR re-baselining with branded identity - #35

Merged
lostmartian merged 8 commits into
mainfrom
feat/approve-bot
Aug 31, 2026
Merged

lostmartian merged 8 commits into
mainfrom
feat/approve-bot

Conversation

@lostmartian

Copy link
Copy Markdown
Collaborator

What

Pillar 3, pulled forward (no more partials). Reviewers comment /agentdiff approve on a flagged PR → the approve workflow verifies the commenter has write access → downloads the candidate artifact from the check run → agentdiff approve BASELINE CANDIDATE re-baselines (envelope: rolling-window append; strict: replace) → commits → the check re-runs and flips PASSED. Policy D3: loop violations are never blessable — the CLI refuses with the specific loop findings; path drift and cost spikes are human judgments.

Also closes the product-side partials:

  • Branded identity — the bot authenticates as a GitHub App (agentdiff[bot] with AgentDiff's logo, avatar shipped in context/logos/) via actions/create-github-app-token; app-token pushes re-trigger CI (GITHUB_TOKEN pushes never do, so the PASSED flip is real). Graceful fallback to github-actions[bot] without the App.
  • Human-first PR verdict — "No infinite loops. No cost spikes. Trajectory within budget — safe to merge." / "Blocked by: tool_loop." leads the comment; the TDI/loops/cost table moves into a collapsible Gate details block.
  • Error Recovery Cascade is a default hard gate (PRD failure class): RSR ≥ 3× baseline blocks CI; explicit --max-recovery-ratio still wins.

Why

The PRD's three pain points must be fully closed in 0.5.0. Baseline maintenance fatigue's remaining loop (leave PR → checkout → update → commit → push) is exactly what this kills; CodeRabbit/Greptile-style PR-native conversation without their hosted backend — the engine runs in the user's own Actions (local-first moat intact).

How

  • ci/approve.py: approve_candidate() — gates the candidate, refuses loop-family codes (tool_loop, tool_repeats, loops), rotates envelope/strict baseline, writes atomically
  • cli.py: approve subcommand (--scenario --runs --pr), refuses exit 1 with findings; posts confirmation comment
  • init_wizard.py: approve-workflow template (issue_comment trigger, command+bot filters, collaborators/permission gate, concurrency guard, artifact download by run-id, app-token selection with fallback); check workflow template now uploads the candidate artifact; init --with-approve
  • reporters/pr.py: verdict-first layout; evaluate_gate/compare_envelope: recovery-cascade default 3.0

Testing

  • 8 new tests (test_approve_bot.py): blessable drift/cost, loop refusal + baseline untouched, envelope rotation window, strict replace, artifact upload, workflow guardrails (YAML-validated), flag-only emission
  • Full suite 423 green; lint clean; build green; generated workflows parse as valid YAML (pyyaml dev-dep added)

Checklist

  • make lint + full suite green
  • CHANGELOG [Unreleased] entries
  • No non-goal violations (no hosted backend, no LLM judge, local-first)

Links to context/ROADMAP.md Phase M / 0.5.0 (SPEC-0.5.0 Pillar 3 + D3/D5 supersede). Stacked on #34 → #33 → #32.

…illar 2)

Severity-aware gate evaluation shared by CLI, assertions, and suites:
HARD violations block CI (exit 1); SOFT warnings render everywhere but
never flip the exit code. New cyclical-tool-loop invariant (identical
inputs + stagnant outputs, non-consecutive included, on by default) and
opt-in max-tool-repeats cap. Path drift renders as a non-blocking note.
Provenance (G7) and threshold flagging (G6) cover the new knobs.
…, commutative equivalence (Pillar 1)

Baselines capture N >= 2 runs into a versioned envelope artifact
(schema 2.0.0, additive — AgentTrace schema untouched). Statistical
compare judges a candidate against normal variance: min-TDI-of-N
alignment, step-count and cost bands (mean ± k·sigma, cost ceiling =
max of relative cap and variance band), divergence ceiling. Hard
invariants flow through. v1 single-trace baselines load as strict
envelopes — full back-compat.

Topological equivalence: independent same-work reorders ([A->B] vs
[B->A]) merge to matched_commutative with zero TDI penalty; dependent
swaps, changed args, and changed outcomes stay real divergence.

[scenario.*] config (PRD v0.5 spec), record --runs N, rolling-window
envelope rotation, N=3/5 benchmarks (cost linear in N).
Detects the agent framework (LangGraph, CrewAI, OpenAI Agents SDK,
OpenTelemetry/OpenInference, or generic) from installed packages and
writes a production-ready agentdiff.toml (v0.5 statistical spec) plus
a GitHub Actions gate workflow. Idempotent (--force to overwrite);
--adapter/--scenario/--runs parameterize the output. Next-steps output
closes the loop: record envelope -> commit -> open a PR.
…tity (Pillar 3)

Reviewers comment '/agentdiff approve' on a flagged PR; the approve
workflow verifies write access, downloads the check run's candidate
artifact, re-baselines via 'agentdiff approve', commits, and the check
flips green. Loops are never blessable (policy D3) — path drift and
cost spikes are human judgments.

Branded identity: agentdiff[bot] via GitHub App token (logo avatar,
re-triggers CI) with github-actions[bot] fallback. init --with-approve
emits the bot workflow; the check workflow now uploads the candidate
artifact.

Product-side positioning: PR comment leads with a human verdict (no
infinite loops / no cost spikes) and demotes the metric table to a
collapsible details block. Recovery cascade (RSR >= 3x) is now a
default hard gate per the PRD failure classes.
@lostmartian
lostmartian changed the base branch from feat/agentdiff-init to main August 31, 2026 18:32
@lostmartian
lostmartian merged commit b966258 into main Aug 31, 2026
7 checks passed
@lostmartian
lostmartian deleted the feat/approve-bot branch August 31, 2026 18:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant