Skip to content

feat(benchmark): share Portfolio ten-arm scores with bound runner settings - #5750

Merged
huangruiteng merged 2 commits into
mainfrom
codex/portfolio-study-sharing
Oct 6, 2026
Merged

huangruiteng merged 2 commits into
mainfrom
codex/portfolio-study-sharing

Conversation

@loopx-agent

Copy link
Copy Markdown
Collaborator

Outcome

Share a reviewable Portfolio ten-arm retrospective snapshot with per-run settings, rather than detached best-score totals. The report includes 1,442 recorded score observations, exact designated attempts, measured duration, sampling counts, best/last distinctions, source/image/task pins, feedback permissions and explicit unknowns. Two selected redacted execution traces are linked at an immutable public Hugging Face revision.

Eight arms use ledger-v2; replacement Explore arms use VaR-timing-v1 and different public instructions/source. All remain integrity-unqualified and non-countable. The report discloses these limits instead of asserting a matched causal comparison or inventing a same-version baseline.

Implementation and boundary

  • Add a terminal EdgeBench export/read-only verifier. Each run binds one settings file and hash; source/model/score mismatch, duplicate selection, missing settings and incomplete sample histories fail.
  • Reuse benchmark-toolkit's canonical row normalizer and existing upload-envelope/local readback. Preserve original studies, anchors, integrity and insight decisions. No new shared vocabulary or control-plane authority.
  • Keep this receipt/scoring-file reduction in the existing Python EdgeBench provider. No session parsing, remote uploading, model invocation, runner/scorer changes or active experiment mutation is introduced by the helper.
  • Future-facing pass: centralize settings/data binding in this adapter rather than duplicate ten manual joins. Multi-cohort manifest composition is not fabricated; current manifest semantics require one baseline. No frontend/Lark runtime changes are claimed.

Validation

  • 32 focused export and benchmark study-projection tests passed; final CSV newline refinement reran all 11 export tests.
  • Real ten-arm export and verification passed: 10 run bindings / 13 data files; all 1,442 samples have submission timestamps.
  • Real toolkit local simulation: preview, append and exact digest readback for all 10 rows; no network transport.
  • Ruff and diff checks passed. Semantic advisory has no applicable loopx source changes; full semantic smoke passed after installing required npm dev dependencies (initial missing-dependency attempt failed).
  • Candidate public-boundary scan and manual source-field review: no local paths, credentials, raw sessions, task text or verifier output committed. Exact source hashes are provenance, not disclosure or integrity certification.
  • HF release: 6 files anonymously downloaded and SHA-256 matched; actual Codex trace view rendered. Redaction is automated filtering plus sampled semantic review, with omissions documented. No claim of complete raw-session disclosure.

Maintainer review/merge requested. This PR does not merge itself or establish a leaderboard result.

@huangruiteng
huangruiteng merged commit 10133f2 into main Oct 6, 2026
4 checks passed
@huangruiteng
huangruiteng deleted the codex/portfolio-study-sharing branch October 6, 2026 07:31
@loopx-agent

Copy link
Copy Markdown
Collaborator Author

Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh

Post-merge exact-head audit — REQUEST_CHANGES

Reviewed head be0089721162d386d47dc25756ae79accce86f75; immutable pre-change base 42e55a809eb94f13443d303d76118735d4112182. The PR merged during this review; this is an actionable post-merge audit, not a retrospective pre-merge approval.

动机

需要公开复查历史实验的研究者和报告维护者。
过去只给出每组最高分,读者无法判断它来自哪次运行、用了什么配置,或最后一次评分是否已经下降。现在十次指定运行分别链接配置和全部评分点,能区分历史最好、最后观测、实际时长与预算;但校验工具仍把内部矛盾的汇总标成 verified,不能据此完成可信导出验收。
发布的十臂数据经独立逐行核对一致,真实 CLI 验证为10 runs/13 files;两处通用导出与校验缺口被实际负例复现,仍需修复。
本次只是历史数字与已审阅配置的离线导出;不启动实验、不重新评分、不上传、不认证完整性、不做因果排名,也不改变当前产品默认或控制面权限。
现有数据可以阅读,通用工具尚未完成同一运行的全部投影一致性和 best_round 证据校验;原始私有执行、完整轨迹脱敏、正式可计分性及独立复现实验仍未由本次审查认证。

改动思路

选择历史运行,再把 native runtime/profile/final/history 与既有 canonical row 连接,生成新目录中的配置、分数汇总、全部数值观测和索引。原始运行、实际 live board、完整性决定均不被写回。这里的 hash 验证只证明文件字节绑定,不能认证原始实验真实、内容可公开或结果可计分;不过它仍应拒绝自己输入中已能确定的矛盾。

最强反对理由是这个工具比手工阅读更容易让后来读者信任错误证据。对比既有 toolkit 发现其拥有 canonical identity/countability 和 upload/readback,但不了解 EdgeBench 的 best_round 和时间采样;因此 adapter 归属正确,Python 专门 reduction 有依据。无需新通用 owner、启动器或人为凑出跨 cohort 的单基线 manifest,修复应留在现有 adapter。

具体改动

全量18文件 +3820/-0:344行 provider、201行测试、214行文档,3061行明确选择的数值数据、配置和索引。八臂 ledger-v2 与两个 VaR-timing-v1 replacement 的 runner/task/source 差别明确披露。逐一核对十份 argv/profile/runner、预算和时长、总轮数与agent/auto计数、全部1442个有限非负时间点及各 best_round;当前发布数据本身没有发现下述矛盾。历史两个源码 pin 的4CPU/16g、judge4CPU/8g、160秒保留、seeded-todo和每3Todo重规划与报告中的 source-derived 标记相符,这不认证实际solver采用或独占硬件。

关键代码讲解

  1. export:Reads explicit terminal receipts and builds a new immutable report; countability stays canonical, best_round witness is missing.
  2. settings projection:Overrides captured runtime facts and allowlists worker profile; supplements/defaults remain operator-reviewed, not inferred grants.
  3. verify:Checks file hashes, partial identity joins and sample maxima/count/order; omitted summary/witness joins cause the reproduced false success.

规范:docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,固定 spec_revision 42e55a809eb94f13443d303d76118735d4112182。

  • 5.1 pinned identity/explicit unknowns:not_met;已发布配置一致,但通用 verifier 可接受同一运行的 source-study 和汇总 identity 矛盾。
  • 5.4 native outcome/cost/countability separation:not_met;canonical仍保持不计分,但CSV能与它相反而校验通过。
  • 7 adapter ownership:implemented;专门数值 reduction 复用现有 toolkit,未新增通用资格判定。
  • 11.3 useful bounded seam with real caller/negative coverage:not_met;真实合法路径通过,以下内部矛盾反例也错误通过。
  • 13.12 independent integrity before eligibility:deferred;既有 toolkit/native runner 的独立完整性、复现工作保持未完成,全部 score_countable=false 不因本次审查改变。

对主干的风险

[P2] 校验遗漏同一运行的汇总与配置一致性

benchmark/edgebench/export_report.py:281–298 只检查部分settings identity/runner/model,并对scores按run_id和best_score核对。复制公开data目录到临时位置,保持canonical row和samples不变,将第一条CSV的arm改为另一个treatment、benchmark改为另一evaluator、score_countable改True,再更新index中scores.csv摘要:实际 --verify 仍exit0、verified=true。另改duration=-500/budget=1,或settings.source_study_id并重封两个settings摘要,也都通过。这里没有要求hash能识别恶意改写;要求的是已知输入互相矛盾时不能被称为一致。

最小修复:按run_id共用一份adapter-local一致性校验,对齐重复的arm、benchmark、countability、budget、duration以及study/benchmark配置身份,拒绝非法数值;保留独立隐私/完整性资格边界。扩展已有rehashed wrong-arm负例覆盖这些列,经真实module CLI重跑。当前发布数据经过独立对照一致,以上是通用工具的可复现缺陷,不能描述为发现历史结果已经造假。

[P2] 导出与校验未核实 best_round 的见证

benchmark/edgebench/export_report.py:113–119 验证max(score)和数量,却将final.best_round直接写入CSV。用既有公开synthetic trial,保留auto-1=2、auto-2=0,只把final.best_round换成not-in-history:真实export成功,随后verify也成功,CSV指向不存在的“最佳轮次”。同样应拒绝存在但分数不是best_score的轮次。

最小修复:export在写目录前、verify在读回时,都要求best_round唯一存在且score等于best_score;native已经选好的并列最佳轮次应保留,无需另造first/last tie policy。给同一 public fixture 增加错误轮次反例,然后重跑 uv run --extra test python -m pytest -q benchmark/tests/test_edgebench_report.py tests/capabilities/test_benchmark_study_projection.py 和公开 --verify 命令。

32 export/study-projection tests passed; published10run/13file verifier, independent1,442sample+all10configuration joins, Ruff, semantic advisory/full semantic smoke and diff check passed. Real File JSONL toolkit base/head10row preview/write/readback/replay is equal. Five contradiction probes are wrongly accepted; exported unsubstantiated best_round is also wrongly accepted.

旧base没有新export模块,未伪装成两份旧/新verifier对照。对既有canonical路径使用同一10行输入,在真实隔离File JSONL逐行preview→append→独立digest readback→replay:base/head完整输出相同、每边仅10append。当前32项测试全绿并不能覆盖上述反例;已知settings.arm_id变更则按预期拒绝,证明负例未关闭所有保护。审查脚本曾用错测试路径/源导入路径,初始未命名audit断言失败后用逐项独立checks确认全部现有数据;这些工具调用失败保留,不能算PR通过或产品失败。

先跑advisory(支持载体0,不能证明动态语义不存在),再跑全树semantic smoke。18个候选路径扫描/结构审查没有本地私有状态、凭证、原始sessions/task文本或verifier内容。固定HF dataset card已读,明确抽样筛查/内容遗漏与不授权第三方材料;未对完整两条trace逐行认证或重下原始运行。没有CI查询、paid solver、live实验、network upload、安装或平台采用验证。

我的整体评价

REQUEST_CHANGES。报告的历史观察、负面结果及cohort限制有价值,体量与复查用途相称;但可信通用导出仍是not_yet_proven。long_horizon与user_experience均not_yet_proven:后续反复使用可能把自相矛盾的数据当作可靠输入,最小修复是同一run的完整连接及best_round见证检查。现有source/default/权限均保留,不能用32绿测或已合并状态掩盖该缺口。

Future-facing pass建议在当前adapter共用export/verify的一致性检查,修掉重复知识而不新增framework或平行资格owner;这项修复尚未完成。原始完整性、私有source真值、最终工件独立复测、trace完整脱敏及模型机制采用均保留未认证。未自己合并、撤回他人review、改写历史attempt或发起实验。维护者可在现有PR讨论中交付有界后续修复,重新验证合法数据和两组拒绝路径;这里不补造先前approval。

English verdict: REQUEST_CHANGES — exact merged head be00897. The supplied ten-arm snapshot is internally consistent and preserves unqualified results, but the reusable verifier accepts contradictory same-run identity/eligibility/effort projections, and export/verify accept an unsubstantiated best_round. Repair these adapter-local checks and add real CLI negative coverage. 32 focused tests, legal report verification, base/head local replay parity, Ruff and semantic checks passed; no live-run, privacy certification or causal improvement is claimed.

@loopx-agent

Copy link
Copy Markdown
Collaborator Author

Resolved by #5756, reviewed at c1273098c4ef07088dc2a2d86ba7ede6e52c5f24 and merged by loopx-agent as acc3ede6af7e3033ef01a93a4d0d1dbcbadf902c.

The shared export/verify validator now checks same-run summary/settings identity, countability, budget/duration and counts, and requires a unique recorded best_round whose score equals best_score. Native tied winner selection and missing sample times are preserved. The original four resealed report contradiction cases and unsupported-best-round export now fail through the real module CLI; invalid export leaves no output directory. The same new tests expose20 expected failures on immutable old production and53 pass on the repair; the existing ten-run/thirteen-file public report verifies unchanged. Ruff, semantic checks, canonical local replay parity, four direct checks and six canaries passed; final quality scope was requalified after the unrelated release update. Native exact-head merge readiness returned ready immediately before the explicitly authorized merge; CI was not consulted.

Published exact-head self-review. The original audit remains the record of the old head; this follow-up closes its two demonstrated defects without retroactively declaring that head approved. Raw-source truth, complete privacy review and independent experiment eligibility remain outside this repair.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants