Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
-
Updated
Jul 16, 2025 - Python
Official repository for the paper "LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code"
Rules and checks that make Claude Code look around a change, not just at the lines it writes: ten questions before code, a reviewer that didn't write it, proof before done, bugs fixed as a class.
Evidence-grounded AI agent for Java code repair with LLM-guided patches and human-in-the-loop review.
Exploring and improving the quality of ChatGPT-generated code for LeetCode programming tasks.
基于 AI Agent 服务自动化修复系统:Agent 自动读取错误日志,定位 Bug,生成补丁,运行测试,提交 PR,并通知开发者 Review。AI-powered auto-fix agent for web services: analyzes logs, patches code, runs tests, and creates pull requests automatically.
Repository-level automated code repair agent using SWE-Bench dataset
Trusted autonomy T&E runtime that links mission needs, hazards, scenarios, telemetry, evidence, verification reports, and hash-chained ledgers so AI/autonomous decisions can be reviewed instead of merely trusted.
Multilingual software-engineering benchmark with pinned Docker environments & reproducible agent trajectories · 多语言软件工程评测基准
A reliability layer for AI-built systems: detect failures (tests or runtime drift), reproduce, repair one ticket at a time behind an approval gate, and prove the fix. The safety boundaries most AI agents skip.
Gymnasium RL environment for training LLM agents to autonomously debug and fix Python code with secure sandboxing and test-driven feedback.
A minimal lab for improving programs, agents and model parameters. Real demos, frozen evaluations, inspectable evidence. Zero runtime dependencies.
🦑 CT 11 — Secure Self-Learning Repair Agent. Droste Fusion. 4 Engines. Reflection Engine. 9/10 Benchmark.
Run broken Python code → it fixes itself. Local-first Python runtime repair.
Enforce authority before AI agents change code. Source-available control plane for federated workload identity, delegated least privilege, pre-tool MCP/API enforcement, evidence-gated human approval, revocation, revision-bound verification, and cryptographically chained receipts. AI proposes. Humans decide.
When Free Executors Cost More: The Free-Executor Paradox in Iterative LLM Code-Repair Loops (paper + reproducibility kit)
PyPatch— OpenEnv RL environment where AI agents debug & fix buggy Python code across 3 difficulty levels.
5-agent LangGraph system that autonomously resolves GitHub issues, benchmarked across single-agent baseline, paid API, and open-weight HPC configurations on SWE-bench Lite
Self-healing Python developer tool — uses AST surgery to extract and LLM-repair the exact failing function, then verifies in a closed loop.
Self-healing code reasoning engine. Detective → QA → Patcher closed loop on SWE-bench, with an RL layer that turns every reasoning trace into DPO training data.
To associate your repository with the code-repair topic, visit your repo's landing page and select "manage topics."