-
🎓 AI undergraduate, graduating in 2027. I work mostly at the boundary between agent systems, reinforcement learning, and embodied intelligence.
-
🧠 Current focus: making AI systems more reliable across time — memory, replay, evidence, verification, recovery, and explicit human authority.
-
🧪 Currently building RL Sentinel, chat-distiller, and Deep Native. Recent work includes RL reliability, Context Gateway 0.4, and adaptive coding-agent runtime experiments.
-
🤖 Robotics work: Go2W MoRA Navigation, a MuJoCo hierarchical navigation project using PPO, behavior cloning, DAgger, curriculum learning, ablations, and reproducible evaluation.
-
💡 Working rule: evidence over claims. A result is more useful when it can be replayed, inspected, and honestly bounded.
| Agent Systems & Research Infrastructure | ||
|---|---|---|
| Project | Stars | Tech / Focus |
| chat-distiller Versioned memory and knowledge for long-running agents. |
|
|
| Deep Native Evidence-first execution and adaptive runtime for coding agents. |
|
|
| FlowCredit Research Auditable research memory with grounded evidence and versioned belief change. |
|
|
| FlowCredit Evidence-aware risk infrastructure for AI-native systems. |
|
|
| Reinforcement Learning & Robotics | ||
|---|---|---|
| Project | Stars | Technologies |
| RL Sentinel Chronological replay and reliability analysis for RL experiments. |
|
|
| Go2W MoRA Navigation Hierarchical navigation for Unitree Go2W in MuJoCo. |
|
|


