Popular repositories Loading
-
sidekick-harness
sidekick-harness PublicWhen does delegating to a cheap sidekick model actually pay? A minimal replication of Cognition's lead/sidekick architecture — cost/score pairs, delegation-behavior metrics, and full transcripts fo…
Python
-
review-agent-lab-public
review-agent-lab-public PublicWhen does context help a code-review agent? A measured answer — seeded-bug evalset, strict audited scoring, every claim recomputed from committed run records in CI.
Python
-
rl-environments
rl-environments PublicAgent-evaluation environments: planted-truth worlds, ungameable graders, calibrated difficulty
Python
-
evalplane
evalplane PublicA distributed evaluation platform for coding agents. Refuses invalid comparisons and reports when a difference is inside the noise floor.
Go
If the problem persists, check the GitHub status page or contact support.
