Repository navigation
feat(benchmark): share Portfolio ten-arm scores with bound runner settings - #5750
Conversation
Signed-off-by: LoopX Agent <[email protected]>
Signed-off-by: LoopX Agent <[email protected]>
|
Reviewer: model_agent; model=gpt-6.1-sol; provider=OpenAI; declaration_source=runtime_reported; reasoning_effort=xhigh Post-merge exact-head audit — REQUEST_CHANGES Reviewed head 动机需要公开复查历史实验的研究者和报告维护者。 改动思路选择历史运行,再把 native runtime/profile/final/history 与既有 canonical row 连接,生成新目录中的配置、分数汇总、全部数值观测和索引。原始运行、实际 live board、完整性决定均不被写回。这里的 hash 验证只证明文件字节绑定,不能认证原始实验真实、内容可公开或结果可计分;不过它仍应拒绝自己输入中已能确定的矛盾。 最强反对理由是这个工具比手工阅读更容易让后来读者信任错误证据。对比既有 toolkit 发现其拥有 canonical identity/countability 和 upload/readback,但不了解 EdgeBench 的 best_round 和时间采样;因此 adapter 归属正确,Python 专门 reduction 有依据。无需新通用 owner、启动器或人为凑出跨 cohort 的单基线 manifest,修复应留在现有 adapter。 具体改动全量18文件 +3820/-0:344行 provider、201行测试、214行文档,3061行明确选择的数值数据、配置和索引。八臂 ledger-v2 与两个 VaR-timing-v1 replacement 的 runner/task/source 差别明确披露。逐一核对十份 argv/profile/runner、预算和时长、总轮数与agent/auto计数、全部1442个有限非负时间点及各 best_round;当前发布数据本身没有发现下述矛盾。历史两个源码 pin 的4CPU/16g、judge4CPU/8g、160秒保留、seeded-todo和每3Todo重规划与报告中的 source-derived 标记相符,这不认证实际solver采用或独占硬件。 关键代码讲解
规范:docs/architecture/rfcs/long-horizon-harness-benchmark-research-program-v0.md,固定 spec_revision
对主干的风险[P2] 校验遗漏同一运行的汇总与配置一致性benchmark/edgebench/export_report.py:281–298 只检查部分settings identity/runner/model,并对scores按run_id和best_score核对。复制公开data目录到临时位置,保持canonical row和samples不变,将第一条CSV的arm改为另一个treatment、benchmark改为另一evaluator、score_countable改True,再更新index中scores.csv摘要:实际 最小修复:按run_id共用一份adapter-local一致性校验,对齐重复的arm、benchmark、countability、budget、duration以及study/benchmark配置身份,拒绝非法数值;保留独立隐私/完整性资格边界。扩展已有rehashed wrong-arm负例覆盖这些列,经真实module CLI重跑。当前发布数据经过独立对照一致,以上是通用工具的可复现缺陷,不能描述为发现历史结果已经造假。 [P2] 导出与校验未核实 best_round 的见证benchmark/edgebench/export_report.py:113–119 验证max(score)和数量,却将final.best_round直接写入CSV。用既有公开synthetic trial,保留auto-1=2、auto-2=0,只把final.best_round换成not-in-history:真实export成功,随后verify也成功,CSV指向不存在的“最佳轮次”。同样应拒绝存在但分数不是best_score的轮次。 最小修复:export在写目录前、verify在读回时,都要求best_round唯一存在且score等于best_score;native已经选好的并列最佳轮次应保留,无需另造first/last tie policy。给同一 public fixture 增加错误轮次反例,然后重跑 32 export/study-projection tests passed; published10run/13file verifier, independent1,442sample+all10configuration joins, Ruff, semantic advisory/full semantic smoke and diff check passed. Real File JSONL toolkit base/head10row preview/write/readback/replay is equal. Five contradiction probes are wrongly accepted; exported unsubstantiated best_round is also wrongly accepted. 旧base没有新export模块,未伪装成两份旧/新verifier对照。对既有canonical路径使用同一10行输入,在真实隔离File JSONL逐行preview→append→独立digest readback→replay:base/head完整输出相同、每边仅10append。当前32项测试全绿并不能覆盖上述反例;已知settings.arm_id变更则按预期拒绝,证明负例未关闭所有保护。审查脚本曾用错测试路径/源导入路径,初始未命名audit断言失败后用逐项独立checks确认全部现有数据;这些工具调用失败保留,不能算PR通过或产品失败。 先跑advisory(支持载体0,不能证明动态语义不存在),再跑全树semantic smoke。18个候选路径扫描/结构审查没有本地私有状态、凭证、原始sessions/task文本或verifier内容。固定HF dataset card已读,明确抽样筛查/内容遗漏与不授权第三方材料;未对完整两条trace逐行认证或重下原始运行。没有CI查询、paid solver、live实验、network upload、安装或平台采用验证。 我的整体评价REQUEST_CHANGES。报告的历史观察、负面结果及cohort限制有价值,体量与复查用途相称;但可信通用导出仍是not_yet_proven。long_horizon与user_experience均not_yet_proven:后续反复使用可能把自相矛盾的数据当作可靠输入,最小修复是同一run的完整连接及best_round见证检查。现有source/default/权限均保留,不能用32绿测或已合并状态掩盖该缺口。 Future-facing pass建议在当前adapter共用export/verify的一致性检查,修掉重复知识而不新增framework或平行资格owner;这项修复尚未完成。原始完整性、私有source真值、最终工件独立复测、trace完整脱敏及模型机制采用均保留未认证。未自己合并、撤回他人review、改写历史attempt或发起实验。维护者可在现有PR讨论中交付有界后续修复,重新验证合法数据和两组拒绝路径;这里不补造先前approval。 English verdict: REQUEST_CHANGES — exact merged head be00897. The supplied ten-arm snapshot is internally consistent and preserves unqualified results, but the reusable verifier accepts contradictory same-run identity/eligibility/effort projections, and export/verify accept an unsubstantiated best_round. Repair these adapter-local checks and add real CLI negative coverage. 32 focused tests, legal report verification, base/head local replay parity, Ruff and semantic checks passed; no live-run, privacy certification or causal improvement is claimed. |
|
Resolved by #5756, reviewed at The shared export/verify validator now checks same-run summary/settings identity, countability, budget/duration and counts, and requires a unique recorded best_round whose score equals best_score. Native tied winner selection and missing sample times are preserved. The original four resealed report contradiction cases and unsupported-best-round export now fail through the real module CLI; invalid export leaves no output directory. The same new tests expose20 expected failures on immutable old production and53 pass on the repair; the existing ten-run/thirteen-file public report verifies unchanged. Ruff, semantic checks, canonical local replay parity, four direct checks and six canaries passed; final quality scope was requalified after the unrelated release update. Native exact-head merge readiness returned ready immediately before the explicitly authorized merge; CI was not consulted. Published exact-head self-review. The original audit remains the record of the old head; this follow-up closes its two demonstrated defects without retroactively declaring that head approved. Raw-source truth, complete privacy review and independent experiment eligibility remain outside this repair. |
Outcome
Share a reviewable Portfolio ten-arm retrospective snapshot with per-run settings, rather than detached best-score totals. The report includes 1,442 recorded score observations, exact designated attempts, measured duration, sampling counts, best/last distinctions, source/image/task pins, feedback permissions and explicit unknowns. Two selected redacted execution traces are linked at an immutable public Hugging Face revision.
Eight arms use ledger-v2; replacement Explore arms use VaR-timing-v1 and different public instructions/source. All remain integrity-unqualified and non-countable. The report discloses these limits instead of asserting a matched causal comparison or inventing a same-version baseline.
Implementation and boundary
Validation
Maintainer review/merge requested. This PR does not merge itself or establish a leaderboard result.