Skip to content

feat(benchmark): add cross-framework A/B benchmark mode - #14

Merged
lostmartian merged 1 commit into
mainfrom
feat/ab-benchmark-mode
Aug 22, 2026
Merged

lostmartian merged 1 commit into
mainfrom
feat/ab-benchmark-mode

Conversation

@lostmartian

Copy link
Copy Markdown
Collaborator

What

Implements G3 — cross-framework A/B benchmark mode: run_benchmark([BenchmarkCase(name, agent_a, agent_b)]) compares two agents head-to-head on the same task, scoring each case on four deterministic dimensions (steps, WEI, tokens, latency), majority-wins with explicit ties. Renders via summary() (CI logs) or to_markdown() (paste-ready table with overall tally).

Why

Roadmap WS-G3. With direct CrewAI/LangGraph ingestion landed (#12/#13), comparing a CrewAI run against a LangGraph run of the same task is now a natural ask — regression diffing answers "did we break?", A/B answers "which agent runs this leaner?".

How

  • Both sides are peers: each case loads both traces, diffs them for TDI/structure context (compare()), and separately scores per-side efficiency dimensions from the traces themselves.
  • Winner = majority of decided dimensions; equal values never count as wins; ties are first-class outcomes.
  • Deliberately structural-efficiency only — no semantic judgment, honoring the LLM-judge non-goal (documented in module docstring).
  • Mirrors the suite runner's execution contract: opt-in workers=N thread pool, input-order results, load/compare failures become contained ERROR outcomes.

Testing

  • 9 new tests: identical-agent tie, majority verdicts both directions, partial dimension splits, error containment, summary/markdown content incl. exact table row + tally, parallel==sequential equivalence with order preservation, empty/invalid workers, and a true cross-framework case diffing the real CrewAI fixture against the real LangGraph fixture.
  • Full suite: 300 passed; lint clean.

Checklist

  • Exactly one logical change in this PR
  • make lint passes
  • uv run python -m pytest is green (300 passed)
  • Docs updated if user-facing (public API docstring)
  • CHANGELOG.md entry added under [Unreleased]
  • Conventional title

@lostmartian
lostmartian merged commit a785e48 into main Aug 22, 2026
@lostmartian
lostmartian deleted the feat/ab-benchmark-mode branch August 22, 2026 20:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant