chore(release): prepare 1.3.0 long-horizon semantic control quality - #5682
Conversation
Signed-off-by: LoopX Agent <[email protected]>
Signed-off-by: LoopX Agent <[email protected]>
Signed-off-by: LoopX Agent <[email protected]>
loopx-agent
left a comment
There was a problem hiding this comment.
Reviewer: model_agent; gpt-6.1-sol; OpenAI; runtime_reported; reasoning_effort=xhigh
Approval conclusion (author-owned PR; GitHub blocks formal self-approval)
动机
维护者准备命名版本,用户通过版本、帮助和开发者书确认安装来源。准备 1.3.0 时,旧元数据仍标为 1.2.4,帮助分类检查遗漏现有配置备份命令;修改后版本对齐且帮助预检通过。干净候选报告 1.3.0,版本、manpage、开发者书及错误 tag 拒绝检查通过。本 PR 不发布 tag、不移动 stable、不替代完整资格验证,也不声称 benchmark 收益已成立。回归修复仍在 #5533,最终合并提交的完整验证、真实模型资格、个人飞书指南和发布产物 readback 尚待完成。
改动思路
本次评审基于完整 head 5c80a2551d36fb5fe6a0f466537c2da997490b6d,对照不可变主线 aada23d751c2a7e352d97430cd951e08335bcd8e。只复用既有版本源、帮助分类和文档生成路径,不引入另一份发布状态或新 CLI。版本只改 package identity;配置备份原本已经注册,加入既有命令专属帮助集合,让 parser 与 manpage 的集合检查重新闭合。同步主线解决旧界面语言断言,合并后 PR 差异仍为八个文件、九行新增和八行删除。
具体改动
规范依据:docs/product/release-readiness.md,spec_revision aada23d751c2a7e352d97430cd951e08335bcd8e。named-version-contract 已实现:loopx.__version__ 与 pyproject.toml 同为 1.3.0,错误 tag 会拒绝。documentation-preflight 已实现:manpage、四处开发者书版本锚点和帮助分类对齐,实际预检通过。compatibility-gate 属于最终晋级的 deferred 项,由维护者合并 #5533 后冻结提交执行,不能拿这次准备检查替代。
关键代码讲解
loopx/__init__.py::__version__ 是公开版本来源,pyproject.toml 镜像该值,既有 release artifact validator 负责匹配 tag;实际传入 v1.2.4 得到预期 exit 2,没有创建任何 ref。loopx/help_surface.py::MANPAGE_COMMAND_HELP_ONLY 在原集合中添加现有 configuration-backup,既有 parser census 检查所有顶层命令必须属于手册组或该集合。没有删命令、改参数或打开备份执行权限。man/loopx.1 只改变生成版本,四处 book checkpoint 只改变当前版本基线;历史示例没有批量重写。语义 advisory 识别了该集合扩展,决定复用既有本地 owner;全树语义检查通过。
对主干的风险
版本/help/manpage/developer-book smokes、source-built Chat verify,以及此前失败的界面语言检查均在此 head 通过。实际默认 CLI 的同 fixture 主线/候选对照有 96 行,零 candidate-only,动作签名与结构无差异;诊断与 Turn plan 的旧字符上限失败在两边一致,当前预算修正在 #5533,不能称这个 head 的完整 premerge 全绿。真实 Git baseline 与此 head 的 maintainability 检查还出现同样三个非本 PR 改动的 debt finding,保留为主线资格缺口。归因采用精确源码与失败身份,不以相同测试数量替代证据。同步后的新 GUI 包已按源码重建;没有宣称正式发布包、PyPI、Windows 或真实模型资格已通过。
语义与 CI 对齐
既有集合采用精确 membership,不是文字猜测或 substring denylist。CLI 可见性不授予执行、访问、费用或 actor 生命周期权限。本次没有 optional-capability 行为改变,实际同输入对照覆盖默认路径;并未借“未执行某功能”证明隔离。当前配置的评审依赖本地证据,不等待 CI;原构建失败和独立基线失败仍保留,最终发布门不因此豁免。
我的整体评价
APPROVE,交付判断为 justified_increment:这个独立、可回滚的准备步骤解决版本和文档预检一致性;long_horizon 与 user_experience 为 preserved,执行、调度、结算和授权路径没有改变。未来重构检查无需新增抽象,现有 owner 已是最小完整边界。最大的剩余风险是把候选准备误称已发布;因此本次不自合并,不宣称完成 release。最终合并、完整资格和发布 readback 仍由当前发布工作继续处理。
English verdict: APPROVE - 5c80a25; coherent 1.3.0 metadata/help/documentation preparation, with focused checks and paired CLI evidence. Baseline qualification failures and final release gates remain explicit; no release or maintainer merge is performed.
|
Owner-authorized release preparation integration at The exact-head COMMENTED approval conclusion remains valid. The refreshed LoopX gate returns Final integrated-source full regression, install/upgrade/host, live-model qualification and artifact/guide readback remain separate from this version/preparation merge. Existing failed receipts are retained, and no tag, release or stable promotion is claimed. CI is not consulted under the resolved managed policy. The adjacent refactor pass uses the existing help-only command classification; no new owner or capability is introduced. |
Release benchmark boundary review / 发布 benchmark 边界审查Reviewed source: The benchmark tree is unchanged from the previously reviewed 827 source. This release operation starts no benchmark jobs and submits no scores or leaderboard entries. The published EdgeBench study retains its pins, confounders, score corrections, actual policy failures and open questions. Its worker model is gpt-6.1-sol/xhigh; Astra describes engineering and semantic review. No matched v1.2.4-versus-v1.3.0 statistical/outcome uplift is claimed. Related runtime/adapter changes have published exact-head reviews. The future-facing pass keeps the existing runner/control-plane ownership and requires no additional abstraction. Verdict: benchmark manual boundary accepted. Release publication remains held pending final exact-source pytest and actual-model qualification plus artifact readbacks. This review does not waive any release gate. 同一冻结源码的全范围风险验证为 19 项选择检查、5 项直接检查通过,515 项公开 smoke 通过。原始 benchmark 人工 hold 保留,由本条审查单独处理;benchmark 单测 95 通过、48 因缺少 Harbor/SForge 跳过,跳过环境不算资格化。发布不启动新实验、不提交分数,不宣称匹配版本间的统计或 outcome 提升。完整 pytest、真实模型组合与发布产物回读仍是独立发布门槛。 |
v1.3.0 frozen-source benchmark boundary review / 冻结来源 benchmark 边界审查Source: Reviewed the shipped v1.2.4-to-source benchmark range and the subsequent 363-line adapter delta. The new effective-Turn cadence requires explicit opt-in on LoopX heartbeat/Explore profiles, refuses incompatible profiles or mixed units, retains the omitted legacy three-completed-Todo cadence, and reads back the exact unit/value. TypeScript remains the settlement/count owner. No scoring, task, judge, default trial budget, submission, leaderboard or experiment-promotion rule changes in this delta; rollback uses a new trial with the option omitted, never rewrites an active matched trial. The related ownership pass removes the reversed capability import and preserves the original context function. Validation: 19 selected canaries and five direct checks pass; all 515 full-public smokes and receipt health pass with zero failures/timeouts/source side effects. Benchmark tests: 96 passed, 62 skipped. Skips reflect absent optional Harbor/SForge dependencies: those runner environments are unqualified. The actual native Codex resume test uses disposable local Responses HTTP and proves same-Session resume; it is distinct from paid-model qualification. Installed CLI and packaged Goal UI prove invalid-value refusal, preview/apply, persisted reload readback and override removal. This release operation launches no benchmark job and claims no matched v1.2.4-to-v1.3.0 outcome uplift. The linked public study retains its original 中文结论:批准 benchmark-sensitive 人工审查项,保留原始一项 hold 与独立解决记录。 显式 Turn 周期只桥接到原 TypeScript owner;省略参数的旧默认、拒绝边界、评分/任务/judge/试验预算均保持。515 项公开 smoke 及 19+5 canary 通过;96 项 benchmark 测试通过、62 项可选环境跳过且未资格化。不启动新 benchmark,也不宣称版本间效果提升。完整 pytest、原生模型失败定位及制品/个人指南回读仍需完成。 |
v1.3.0 frozen-source benchmark review / 冻结来源边界审查Source: Rechecked the v1.2.4-to-source scope and exact tree parity with the previously reviewed benchmark boundary: the benchmark tree is unchanged. The subsequent macOS Chat persistence repair and deterministic wait tests do not alter scoring, tasks, judge, submissions, trial budgets or promotion rules. Effective-Turn cadence remains an explicit LoopX heartbeat/Explore opt-in; mixed units and unsupported profiles are refused, omitted legacy cadence is preserved, and TypeScript owns counting/settlement. Rollback starts a new trial with the option omitted; it never rewrites an active experiment. Current-source validation: 19 selected + 5 direct checks, 515/515 full-public smokes and receipt health all pass, with zero failures/timeouts/tracked side effects. 96 benchmark tests pass; 62 optional Harbor/SForge tests skip, so those provider environments remain unqualified. The real native Codex resume fixture uses disposable local Responses HTTP and the same Session. Separate paid DS Flash/high qualification passes 22 scenarios ×2, 44 actors and six contrasts with zero failures/skips; its native supplement passes two Turns/three settled spends. These receipts prove the stated acceptance, not matched-version benchmark uplift. The historical native assertion failure remains unattributed and is not reclassified by later passes. No benchmark job was launched by this release operation. The public study retains 中文结论:批准 benchmark_sensitive 人工项,原 hold 保留,独立审查后剩余为零。 当前 benchmark tree 与既有独立审查完全一致,后续 Chat 安装保留和等待测试修复不改变评分、任务、judge、试验预算或晋升权。当前19+5 canary、515公开smoke及健康通过;96项benchmark通过、62项可选provider跳过且未资格化。DS high22×2/44actors/6对照及2nativeTurns/3settledspends通过,历史原生断言失败仍未归属。未启动新benchmark,不宣称配对版本涨分;完整pytest、汇总及制品/个人指南回读仍是发布条件。 |
LoopX 1.3.0 release preparation
The release improves continuing work, evidence-driven replanning and scoped recovery/settlement, informed by the latest long-horizon study and Astra-assisted engineering. Study workers remain
gpt-6.1-sol/xhigh; no paired version outcome uplift is claimed.Release Decision
Who should upgrade: Operators of continuing Codex work, evidence-driven replanning and workspace Chat can upgrade to v1.3.0. Users satisfied with v1.2.4 may keep their current installation.
What this release solves: Useful history could be crowded out by repeated observations; an already-known replan could demand a failed round trip; task steps, permission failures and settlement hints could point to the wrong recovery. This release collects bounded fixes to those paths, alongside clearer workspace Chat and explicit configuration recovery.
Breaking changes: No intentional breaking migration. Ordinary host-declared project Chat now defaults to
workspace_write; chooseworkspace_readexplicitly when needed. Existing read-only App bindings retain their grant. New generated settlement commands carry correctly placed global--format json; direct CLI defaults and stored receipts are unchanged. Managed execution namesdeepseek-flash; explicit historical model settings remain respected. Canonical new-Goal creation and Explore execution remain opt-in. Existing Goals are not migrated by a device preference. Lifecycle-only registry metadata and stored workspace approval no longer authorize a Codex workspace-write resume; an admitted runtime source is required.How to verify: Expect
loopx 1.3.0, a healthy owning installation, and a current scoped status/diagnostic readback after upgrading. Investigate unavailable or blocked results before starting work.loopx --version loopx doctor loopx --format json status loopx diagnose --goal-id "$GOAL_ID"Contributors: @Inference1, @Duang777 and eight others; see the linked, annotated community credits below.
State Kernel & Control Plane
Explicit settled work-Turn review supports open Todos with legacy defaults preserved; typed selection, complete settlement and durable recovery keep authority explicit.
Capabilities & Workflows
Scoped Chat, configuration recovery, opt-in canonical Goals, governed delegation and original Session continuity. Saved macOS Chat storage/timeouts survive restart (#5714).
Quality & Testing
Shared typed owners, regression repairs and exact-source release qualification preserve counterexamples and failed/untested boundaries.
Benchmarks & Integrations
Opt-in Explore/GitHub evidence, DS profiles and scoped finance/DSH. Ego remains staged; native Bot adoption is unqualified. No matched-version uplift is claimed.
Community Contributors
Optional Capability Activation & Use
Explore Harness
Activation: In Goal settings → Capability Center choose evidence or planning. CLI planning opt-in is shown below; the default is off.
Validation: Read turn context and summary for the same Goal/Agent; recorded nodes alone do not prove adoption.
Disable / rollback: Set
--explore-mode off --execute; retained evidence stays readable.Authority boundary: Analysis and planning grant no worker spawning, claims, quota spend, merge or external publication.
Docs: Versioned guide
TurnEnvelope and captured decisions
Activation: Opt in per guard invocation with
--turn-envelope; add--decision-output-dironly with an explicit Turn id and a new directory whose parent exists.Validation: Read the returned capture and verify Goal/Agent/Turn, original source hash and
ok; observe rejection and incomplete publication honestly.Disable / rollback: Omit both options. Delete only no-longer-needed private captures through ordinary file management.
Authority boundary: Saved decisions are private observations, not fresh admission; selection, leases, cancellation and mutation-time checks remain mandatory. Frontend/Lark do not consume these files.
Docs: Versioned guide
Ordinary workspace Chat
Activation: Run
loopx chat; in the App steward conversation choose an already granted workspace in Scope. Host-declared roots now default toworkspace_write; use the explicit read-only command below when required.Validation: Read the selected scope, effective grant and returned result in the same conversation. A changed grant creates a new context; old history remains.
Disable / rollback: Restart the service with
--project-workspace-grant workspace_read, or revoke the host workspace grant; new work on the old context is rejected.Authority boundary: Workspace writes follow AGENTS.md and the actual sandbox. No hidden Goal, portfolio grant, peer delegation or external-send authority is created.
Docs: Versioned guide
Owner private Lark conversations
Activation: In Settings → Lark explicitly verify the selected App and personal owner, choose a workspace, executor, grant and ordinary Chat or steward role, then connect.
/delegate --tokens N objectiverequires an original-source confirmation./agents,/agent TARGET_REFand/projectuse separately authorized attached-Agent targets. In ordinary project/steward conversations send a PNG/JPEG/GIF/WebP image or image/text post under that same binding; no extra image switch is required. At most four images, 5 MiB each and 12 MiB total.Validation: In the bound private conversation send
/statusand/help; verify the App/source, role, workspace grant, original Session/Turn and queue. Inspect the same binding in Settings. Check the original conversation and caption, same Session/Turn, and conclusion returned to the bound private steward. A download failure, unsupported mixed media, control-command image or attached-host image must return an explicit non-execution notice; do not treat its text as executed.Disable / rollback: Disconnect that exact App binding in Settings → Lark;
/stoptargets the current ordinary Turn and/stop-commissiontargets the bound commission. Remove exact Agent target grants separately.Authority boundary: Personal credentials, source/owner and listener identity remain explicit. Ordinary Chat creates no Goal; commission confirmation does not grant arbitrary writes, automatic heartbeat or canonical task acceptance. Live Lark/mobile qualification is separate. The receiving App alone downloads its canonical message resources. Image bytes and keys stay private; grants are checked again after download and before return.
Docs: Versioned guide
Configuration checkpoints
Activation: Settings → Capability Center → Configuration backup and recovery downloads a private checkpoint. CLI export uses a new destination, preview first, then
--execute.Validation: Run
configuration-backup verifyon the exact file; compare source scope and digest. Integrity verification does not certify privacy.Disable / rollback: An isolated restored checkpoint can be removed without affecting live settings. Revert any adopted setting through its existing revision-checked editor or machine-config/configure-goal transaction.
Authority boundary: Capture and isolated restore copy no credential store, Host session, grant, provider selection, live registry, fence, lease or scheduler. Backups remain private, including secrets already embedded in configuration.
Docs: Versioned guide
Canonical new-Goal creation
Activation: In Device defaults → New Goal authority opt in to canonical creation and choose File/SQLite plus soft_claim/hard_lease. The same v1 document uses the existing preview/apply transaction; default remains off.
Validation: Inspect machine settings, bootstrap a new empty project, then read its Todos and native authority receipt; existing Goals are not retargeted.
Disable / rollback: Preview/apply
canonical_creation=falseto disable future creation, or remove the goal_storage namespace with its exact removal-plan revision. Existing Goals retain storage and fences.Authority boundary: Storage and execution policy are independent of tools, accounts, network, scheduling and migration authority. Missing authority cannot be recreated as empty by forced bootstrap.
Docs: Versioned guide
Governed delegation stop
Activation: Use an already configured exact binding, registered requester and operation;
delegation stop --executeexplicitly requests stop. CLI/MCP and the App team surface share that owner.Validation:
delegation readdistinguishes acknowledgement, proved native process drain and settlement. Missing supervision remains unknown and requires the returned recovery.Disable / rollback: Remove the exact operator binding to revoke new execution; use stop on an active operation. A stopped/settled receipt is not a resume grant.
Authority boundary: Do not infer Host exit, lease release or non-execution from absent records. Windows/unsupported stopping refuses before launch-side cancellation effects.
Docs: Versioned guide
PR review queue ownership and direction
Activation: Capability Center configures additional owner logins and forward/reverse direction; Goal CLI can set exact per-Agent direction. Current-session request intake precedes generic queue discovery.
Validation: Inspect configure-goal and the read-only pr-review packet; compare effective accounts, Agent and direction. Queue ownership is separate from GitHub author identity.
Disable / rollback: Use
--clear-pr-review-owner-logins --execute,--pr-review-agent-order AGENT=inherit, or--clear-pr-review-configuration --execute. Remove owner_logins from Goal/device settings before downgrading.Authority boundary: These settings grant no GitHub review, comment, dismissal, merge, cross-Agent write or scheduler authority; review depth and CI policy retain their owner.
Docs: Versioned guide
Public GitHub evidence
Activation: Opt in per plan with
--public-githuband a full commit SHA source;execute --executeauthorizes anonymous source reads.Validation: Read back the exact plan and execution receipt. Parent admit/reject and downstream research-ledger coverage are separate explicit operations.
Disable / rollback: Omit
--public-githuband--execute; no persistent provider switch is installed. Retire evidence only under the admitted downstream-coverage rules.Authority boundary: No token, cookie, private repository, mutable branch source, raw-page persistence, automatic admission or outbound message permission. Complete frontend initiation remains open.
Docs: Versioned guide
Managed DeepSeek model selection
Activation: Managed execution now names deepseek-flash@high. Explicit LOOPX_TURN_MODEL/DSH_MODEL or
--dsh-modeloverrides are still honored; configure credentials through the owning provider.Validation: Read
turn run-once --helpand the returned runtime profile/actual provider identity; a model alias alone is not live qualification.Disable / rollback: Set LOOPX_TURN_MODEL to the previous explicitly selected model and restart the owning host. Do not change an existing session silently; omitted configuration restores the shipped managed default.
Authority boundary: A model setting grants no API credential, paid invocation, workspace, permission, Goal or provider promotion authority. CPA routing remains separately installed/operator-owned.
Docs: Versioned guide
Finance evidence assessment
Activation: Finance Value Discovery is a separately installed source extension at version 0.8.5; install from the v1.3.0 checkout and register/enable its matching manifest. Select assess-period or assess-cash per command; no automatic assessment switch exists.
Validation: Run extension doctor, then assess-period on a reviewed local input. Producer accuracy, period eligibility, source authenticity and investment truth remain distinct. Use the supplied six-row cash example below; inspect arithmetic residuals, unknowns, source-column/unit/sign and measurement-kind failures. App/Lark cash presentation and production source adoption remain unqualified.
Disable / rollback: Disable loopx-finance-value-discovery with the command below. Retain private account/position observations outside public research projections.
Authority boundary: The package performs deterministic evidence assessments; it grants no investment advice, order, transfer, signing or execution authority. Missing source/timezone/binding evidence stays explicit.
Docs: Versioned guide
EdgeBench native trials
Activation: Explicit research-only invocation from a v1.3.0 checkout; Linux, Docker, pinned EdgeBench/SForge/Harbor, selected Codex binary and authorized model access are prerequisites. Use the selected worker and feedback profile. Only heartbeat-resume/explore accept
--replan-after-turns 3; Harbor uses explicitreplan_after_turns: 3.Validation: Run --help before the reviewed trial; read session/model, native terminal result, evaluator completion and isolation receipts. Registration and sampling counts do not prove countable outcomes. Its receipt must read
replan_after_effective_turns: 3.Disable / rollback: Do not start another trial; stop the exact owned trial/controller and TLS relay through their existing cancellation lifecycle. Preserve private logs and incomplete results. Omit the Turn option in a new trial to restore legacy three-Todo cadence;never alter an active matched trial.
Authority boundary: A release launches no benchmark jobs. The native judge owns task/scoring; blind policy requires separately qualified credential/network/submission isolation. Raw evidence and secrets remain private.
Docs: Versioned guide
DSH LoopX plug-in
Activation: Install the separately published 0.1.1-beta.6 into the web profile below. Check DSH compatibility first; the qualified current host is 0.2.0-rc.2. Loading prepares the isolated CLI/skills and GoalBar; the passive Driver activates only after this exact Session invokes the installed loopx skill.
Validation: Check dsh --version, then resolve the exact live Session binding. Require status=bound and one Goal/Agent pair before using GoalBar. Installed files alone do not establish a live binding.
Disable / rollback: Remove dsh-loopx-plugin from the same profile and restart DSH. To roll back, install a retained previous tarball supported by that host; preserve LoopX state.
Authority boundary: Installation may prepare the isolated CLI and skills, but grants no Goal binding, quota spend, model call or execution by itself. LoopX owns Goal/Todo/Agent decisions; no credentials are bundled. GoalBar and Driver use the authenticated local Session boundary.
Docs: Versioned guide
Already bound steward inbox execution
Activation: Use an existing reviewed task binding and requester. Configure its existing execution file below, then add the exact {goal_id, agent_id, requester_agent_id, binding_id} row to execution_bindings in the same verified private sources[channel] policy; retain sender_ids. Chat may select that authorized binding in context_handoff. No new task or host is provisioned.
Validation: Read the receiver inbox and exact delegation operation. Delivery, launch submission, receiver adoption, native result, canonical acceptance and original-audience return remain separate. runtime_unverified and refused admission retain their recovery.
Disable / rollback: Remove that exact execution_bindings row from the same source policy to revoke future launches and replays. Inspect and stop any already launched exact operation with delegation stop --execute; stopping the manager does not stop its independent worker.
Authority boundary: Sender/read/context grants do not grant execution. Canonical preflight, registered requester, provider provenance and original Turn remain required. No new scheduling, cross-host launch, arbitrary command, protected write or publication authority; operator-editor/discovery completion remains staged.
Docs: Versioned guide
Independent macOS service and Chat Codex homes
Activation: On macOS use the script from a v1.3.0 checkout and one qualified installed CLI owner. Set CODEX_HOME for service execution and LOOPX_CHAT_CODEX_HOME for Chat explicitly when changing them; omitted selections preserve existing bindings on install/restart. To change Chat storage/timeouts, set LOOPX_CHAT_RUNTIME_ROOT, LOOPX_CHAT_IDLE_TIMEOUT_SECONDS and LOOPX_CHAT_HARD_TIMEOUT_SECONDS before install/restart; omitted values retain saved settings, and invalid input preserves both prior plists.
Validation: Run service status, then inspect the actual running Chat identity and each selected home. Synthetic source checks do not prove login-after-reboot or provider acceptance.
Disable / rollback: Run stop to unload the two owned services or uninstall to remove their two plists. Reinstall the previous package through the same installation owner and restart to roll back; project state and credential homes are retained.
Authority boundary: Choosing a model home does not copy credentials between profiles or grant Goal, workspace-write, provider-call or external-message authority. New scan paths remain discovery inputs, bounded by Core authorization.
Docs: Versioned guide
Rendered public-source MCP reader
Activation: Explicitly add a unique
[mcp_servers.loopx_ego_source_read]to the actual execution host:commandis the installed environment Python,args = ["-m", "loopx.extensions.ego_source_reader"],startup_timeout_sec = 30,tool_timeout_sec = 40. Its env must supply the fourLOOPX_EGO_READ_*values below. Reserve one existing Ego Page and public origins only; restart the idle host and resume its original Session. The adapter remains a staged integration; native Bot/channel promotion is unqualified.Validation: Read the stdio tool inventory below, then call
read_public_urlandread_public_imagefrom that same host with an allowed URL/index. Inspect requested/observed URL, digest, truncation and actual image content; an out-of-origin call must be refused. A text result does not qualify visual understanding or complete article coverage.Disable / rollback: Remove only that uniquely named MCP entry and restart the idle host. Restore the prior private config/package through the same installation owner when rolling back; preserve other MCP entries and Session bindings.
Authority boundary: Origin setup grants no private/admin-site permission, sandbox change, workspace write, note edit, publication or delegation. Page content is untrusted. Rendered image regions may be occluded; keep the Page exclusively reserved and respect user takeover.
Docs: Versioned guide
Canonical retirement with an existing successor
Activation: For already promoted canonical Goals, opt in per original guarded Turn with
todo supersede --successor-todo-id; use the original Goal/Agent/Todo/Turn and existing successor. Carry the actual lease proof required by that Goal. This is a state mutation without an additional execute flag; it does not enable canonical authority.Validation: Read
todo listfor the same Goal after the operation. Verify the original is superseded, links point to existing work, the original lease is released, and successor owner/scope/status/due time are unchanged. Reuse the exact original Turn and unchanged intent after an ambiguous response; replay is not a new lease.Disable / rollback: Omit the new option to retain the existing prelinked/generated recovery. Missing, self, cross-Goal or mixed generated/existing links refuse before mutation. Unpromoted Markdown requires its existing migration path. Retirement is terminal: do not reset canonical state to undo it; create explicitly authorized corrective work if required.
Authority boundary: The operation grants no new scope, changes no successor requirements or monitor due time, closes no Goal and certifies no deliverable. Original writeback and one quota spend remain separate; Chat/Lark actions and the remaining frontend migration journey are unchanged.
Docs: Versioned guide
Single-material ranked moves and receipt coverage
Activation: Opt in per SDK caller to plan_material_single_move and material_rerank_receipt_chunks. An explicitly activated project source may pass MaterialProjectScope instead of goal_id to existing inventory and intake/rollback builders; exactly one owner is required. Optionally install only the project-local loopx-material skill. Both synthetic previews below change no source; real effects require the original source adapter.
Validation: Run the previews to verify complete ranking, unchanged relative order, 100/67 receipt coverage and project-scoped metadata without a Goal. Read project-local skill status. Actual project intake/rollback requires provider.verify_project_scope to resolve current Core caller, audience, exact profile/store, workspace grant, gate expiry and revocation before access/publication, plus the source transaction fence. Returned metadata is not permission or an applied move.
Disable / rollback: Uninstall the project-local skill with the exact command below and stop invoking the helpers or passing project_scope. Restore a real store only through its verified backup, source adapter, owner gate and rollback contract. Skill removal does not roll back source data.
Authority boundary: References select existing scope and grant no write, private access, scoring change, cutover or automatic reranking. Project intake creates no Goal and borrows no steward authority. Migration, rebuild and Explore keep the Goal route; only existing SDK/provider intake is extended. This experimental slice adds no frontend/Lark/CLI writer, so the broader project-material journey remains partial.
Docs: Versioned guide
Effective-work-Turn review cadence
Activation: Opt in in Capability Center → Goal review cadence by choosing
Settled work Turnsand a count from 1–5, for the device default or selected Goal. The CLI below explicitly selects three settled work Turns; the legacy completed-Todo default remains unchanged.Validation: Read configure-goal for the same Goal, then quota should-run for its Agent. The shared TypeScript owner counts distinct settled work Turns after the accepted review checkpoint; duplicate retry, poll, unspent writeback and missing settlement do not count.
Disable / rollback: Clear the effective-Turn Goal override with the command below to restore the current device default. Use the revision-checked device editor to choose completed Todos, or remove todo_replan_cadence to restore the legacy capability default. Preserve existing state and receipts.
Authority boundary: The threshold creates a review obligation; it grants no host scheduling, execution, quota, workspace access or external-write authority. Existing v0 configuration retains its completed-Todo meaning until an explicit editor apply migrates it. An accepted negative work result can count.
Docs: Versioned guide
Qualification
Personal illustrated release guide; existing Feishu access applies.
Qualified release source:
e1dd9e519c3057fda4b3765d0c93e7ae793608bf, tree92ecd2e76d8bb72cfc7d87487363edecc20990da. Eight exact-source lanes passed, including pytest 18,077 with subtests/133 skipped, real wheel upgrade/Chat/three hosts,515 public smokes and DS high22×2. Packages/PyPI, attested product/builder identity, desktop assets/signed feed and stable installation were read back. Fixture recovery passed all46 packaged browser scenarios and team-evidence return. Native desktop builds passed; the maintainer published verified CI bytes after cancelling queued publishers. First DMG failure and historic DS assertion remain unattributed. Windows execution, Apple notarization, live Lark/Bot and missing Harbor/SForge/PostgreSQL/NoKV remain unqualified.Turn/diagnostic presentation ceilings are15,000/45,000 JSON characters after identical14,647/44,126 outputs exceeded14,600/44,000. Monitor retains2,000 and all settlement assertions with a physically stable fixture. These are not token/fee budgets or promotion criteria. No matched version outcome uplift is claimed.
中文摘要
升级决策
**谁需要升级:**需要持续 Codex 工作、证据驱动重规划或普通 workspace Chat 的用户可升级到v1.3.0;当前满足需求的v1.2.4用户也可保留现有安装。
**解决了什么:**早期有效结果被重复观察淹没、已知 replan 仍要求失败重入、任务步骤与权限/结算提示错位;本次集合这些有界修复,并改善 workspace 对话与配置恢复。
**是否有破坏性变更:**无主动 breaking migration。host-declared project Chat 新默认
workspace_write,需只读时显式选择workspace_read,旧 App 只读 grant 保留。新生成结算命令自带位置正确的 JSON 参数,历史命令与直接 CLI 默认不变。managed 模型命名改为deepseek-flash,明确旧配置保持优先。canonical 新 Goal 与 Explore 执行保持 opt-in,不迁移已有 Goal。**如何验证:**正式升级后期望
loopx 1.3.0、健康安装及正确作用域状态。使用上方loopx --version、loopx doctor、loopx --format json status与loopx diagnose --goal-id "$GOAL_ID",遇到 unavailable/blocked 先恢复再工作。贡献者: @Inference1、@Duang777等十位;具体贡献与PR链接见下方社区贡献者。
状态内核与控制面
已结算工作Turn复核开放Todo,保留旧默认;类型化选择、完整结算与持久恢复保持原权限。
能力与工作流
普通Scope Chat、配置恢复、opt-in规范Goal、受治理委派与原Session续接。macOS Chat保存的存储/超时重启后保留(#5714)。
质量与测试
复用类型化owner、修复回归并按精确源码资格验证;反例、失败与未测边界保留。
基准与集成
按需Explore/GitHub证据、DS profile及finance/DSH;Ego分阶段、原生Bot未资格化。blog worker仍是gpt-6.1-sol/xhigh,Astra代表工程/语义审查,不宣称配对涨分。
社区贡献者
可选能力启用与使用
Explore Harness
启用: 在 Goal 设置 → 能力中心选择 evidence 或 planning;下方命令启用 planning,默认关闭。
验证: 读取同一 Goal/Agent 的 turn-context 与 summary;节点存在不等于已采用。
停用 / 回退: 执行
--explore-mode off --execute;保留既有证据。权限边界: 分析与规划不授予 spawn、claim、扣额、合并或外发权限。
文档: 固定版本指南
TurnEnvelope and captured decisions
启用: 每次 guard 显式加
--turn-envelope;保存完整 decision 还需明确 Turn id、已存在父目录与全新目标目录。验证: 读取 capture,核对 Goal/Agent/Turn、原始 source hash 与
ok;拒绝或不完整保存不可当成功。停用 / 回退: 省略两个参数;仅通过普通文件管理删除不再需要的私有 capture。
权限边界: 旧观察不提供新准入;选择、lease、取消与写入时校验仍有效。frontend/Lark 不消费这些文件。
文档: 固定版本指南
Ordinary workspace Chat
启用: 运行
loopx chat,在管家对话的 Scope 选择已授权 workspace。宿主声明的目录默认workspace_write;需要只读时用下方命令。验证: 核对 Scope、生效 grant 与同一对话的结果;grant 改变创建新上下文,旧历史保留。
停用 / 回退: 以
--project-workspace-grant workspace_read重启服务,或撤销宿主 workspace grant;旧上下文的新工作被拒绝。权限边界: 写入遵循 AGENTS.md 与实际 sandbox;不创建隐藏 Goal,不借 portfolio、peer 委派或外发权限。
文档: 固定版本指南
Owner private Lark conversations
启用: 在设置 → Lark 核验 App 与个人 owner,选择 workspace、executor、grant 和普通 Chat/管家角色后连接。
/delegate --tokens N objective需原来源确认;/agents、/agent TARGET_REF、/project使用另行授权目标。 在同一普通项目/管家私聊发送 PNG/JPEG/GIF/WebP 图片或图文,不需要额外图片开关;最多 4 张、每张 5 MiB、总计 12 MiB。验证: 在绑定私聊发送
/status、/help,核对 App/source、角色、workspace grant、原 Session/Turn 与队列;设置回读同一绑定。 核对原对话、caption、同一 Session/Turn,以及回到绑定私聊的结论;下载失败、不支持的混合媒体、控制命令图片或附着宿主图片应明确告知未执行,不能把其中的文字当成功执行。停用 / 回退: 在设置 → Lark 断开精确 App;
/stop停止普通 Turn,/stop-commission停止绑定委托;另行撤销精确 Agent grant。权限边界: 个人凭据、source/owner、listener 身份仍明确。普通 Chat 不创建 Goal,委托确认不授予任意写入、自动 heartbeat 或规范任务验收。真实 Lark/手机资格另行验收。 接收 App 仅下载其规范消息的资源;图片字节和资源 key 保持私有,下载后和返回前再次检查 grants。
文档: 固定版本指南
Configuration checkpoints
启用: 设置 → 能力中心 → 配置备份与恢复可下载私有 checkpoint;CLI export 使用新目标,先预览再
--execute。验证: 对精确文件运行
configuration-backup verify,核对来源范围与摘要;完整性不证明隐私安全。停用 / 回退: 未采用的隔离 checkpoint 可删除而不影响 live settings;已采用设置通过已有 revision 校验 editor 或 machine-config/configure-goal 回退。
权限边界: 不复制 credential store、Host session、grant、live registry、fence、lease 或 scheduler;配置本来包含的秘密仍属私有。
文档: 固定版本指南
Canonical new-Goal creation
启用: 设备默认 → 新 Goal 的权威存储显式启用 canonical creation,选择 File/SQLite 及 soft_claim/hard_lease;v1 document 采用已有 preview/apply,默认关闭。
验证: inspect 后 bootstrap 新空项目,再读 Todos 和原生 authority receipt;已有 Goal 不改目标。
停用 / 回退: 预览/应用
canonical_creation=false关闭未来创建,或按精确 removal-plan revision 删除 goal_storage namespace;已有存储与 fence 保持。权限边界: 存储/执行策略不授权工具、账号、网络、调度或迁移;丢失的 authority 不能以 forced bootstrap 当空数据重建。
文档: 固定版本指南
Governed delegation stop
启用: 使用已配置的精确 binding、注册 requester 与 operation,
delegation stop --execute显式请求停止;CLI/MCP 与 App 团队表面复用 owner。验证:
delegation read区分确认、已证明的原生进程 drain 与结算;监督缺失保持 unknown,执行返回的恢复路径。停用 / 回退: 移除精确 operator binding 撤销新执行,对 active operation 请求 stop;停止/结算回执不是 resume grant。
权限边界: 缺少记录不证明 Host 退出、lease 释放或未执行;不支持的停止平台在 launch-side 取消效果之前拒绝。
文档: 固定版本指南
PR review queue ownership and direction
启用: 能力中心配置额外 owner 登录名与 forward/reverse;Goal CLI 可设置精确 Agent 方向。当前会话请求先于普通队列发现。
验证: 读取 configure-goal 与只读 pr-review packet,核对实际账号、Agent 和方向;queue owner 不等于 GitHub author。
停用 / 回退: 用
--clear-pr-review-owner-logins --execute、--pr-review-agent-order AGENT=inherit或--clear-pr-review-configuration --execute;降级前移除 Goal/设备 owner_logins。权限边界: 不授予 GitHub review、comment、dismiss、merge、跨 Agent 写入或调度权限;review 深度和 CI policy 沿用已有 owner。
文档: 固定版本指南
Public GitHub evidence
启用: 每个 plan 使用
--public-github和完整 SHA 来源;execute --execute只授权匿名来源读取。验证: 回读精确 plan/execution receipt;父采纳/拒绝与研究账本覆盖另行显式操作。
停用 / 回退: 省略
--public-github与--execute;没有持久 provider switch。证据退休仍遵循采纳与下游覆盖规则。权限边界: 不使用 token、cookie、私有仓库、可变分支、原始页面持久化、自动采纳或外发权限;完整 frontend 发起仍开放。
文档: 固定版本指南
Managed DeepSeek model selection
启用: managed execution 使用 deepseek-flash@high;显式 LOOPX_TURN_MODEL/DSH_MODEL 或
--dsh-model仍优先,凭据由 provider 配置。验证: 读
turn run-once --help及返回的 runtime profile/实际 provider 身份;alias 本身不证明 live qualification。停用 / 回退: 将 LOOPX_TURN_MODEL 设置为之前明确选择的模型,重启 owning host;不要静默修改旧 session;省略配置恢复 shipped managed default。
权限边界: 模型设置不提供 API 凭据、付费调用、workspace、Goal 或 provider 晋升权限;CPA route 仍为独立 operator-owned 部署。
文档: 固定版本指南
Finance evidence assessment
启用: Finance Value Discovery 是单独安装的源码 extension,版本 0.8.5;从 v1.3.0 checkout 安装并登记/启用匹配 manifest。每次明确选择 assess-period 或 assess-cash,不自动运行评估。
验证: 运行 extension doctor,再对已审阅本地输入 assess-period;producer 精度、期间资格、来源真实性与投资真值分开。 下方六行现金示例检查算术 residual、unknown、来源列/单位/符号和 measurement-kind 失败;App/Lark 现金呈现与真实来源采纳仍未获资格。
停用 / 回退: 用下方命令 disable loopx-finance-value-discovery;私有账户/持仓观察不放公开研究投影。
权限边界: 只做确定性证据评估,不授予投资建议、下单、转账、签名或执行权限;来源/timezone/binding 缺失保持明确。
文档: 固定版本指南
EdgeBench native trials
启用: 仅在 v1.3.0 checkout 中显式启动研究:需要 Linux、Docker、固定 EdgeBench/SForge/Harbor、选定 Codex 与已授权模型。明确 worker 和反馈 profile。 仅heartbeat-resume/explore接受
--replan-after-turns 3,Harbor显式设置replan_after_turns: 3。验证: 先 --help 再审阅试验命令;回读 session/model、终态、评测完成与隔离回执。注册/采样计数不证明有效 outcome。 回执须显示
replan_after_effective_turns: 3。停用 / 回退: 不再启动新试验;按已有 cancellation lifecycle 停止精确 trial/controller 与 TLS relay,保留私有日志和未完成结果。 新trial省略Turn选项可恢复旧三Todo节奏,不能改正在匹配的trial。
权限边界: 发布不会启动 benchmark;judge 拥有任务/评分,blind 还需凭据、网络、提交资格隔离验收;原始证据与秘密保持私有。
文档: 固定版本指南
DSH LoopX plug-in
启用: 按下方命令将独立发布的 0.1.1-beta.6 安装到 web profile,先核对 DSH 兼容范围;当前已资格验证宿主为 0.2.0-rc.2。加载准备隔离 CLI/skills 和 GoalBar;Driver 仅在这个精确 Session 调用已安装 loopx skill 后激活。
验证: 读 dsh --version,再解析精确 live Session 的绑定,要求 status=bound 且唯一 Goal/Agent 对。文件已安装不等于 Session 已绑定。
停用 / 回退: 从同一 profile remove dsh-loopx-plugin 后重启 DSH。回退时安装预先保留且与宿主兼容的旧 tarball,保留 LoopX 状态。
权限边界: 安装可以准备隔离 CLI 与 skills,本身不授予 Goal 绑定、扣额、模型调用或执行权限。Goal/Todo/Agent 决策仍由 LoopX 管理;包不含凭据,GoalBar/Driver 使用已认证本地 Session 边界。
文档: 固定版本指南
Already bound steward inbox execution
启用: 使用已审阅的已有 task binding 与 requester;下方配置现有 execution 文件,再将精确 {goal_id, agent_id, requester_agent_id, binding_id} 条目加入同一已核验私有 sources[channel] policy 的 execution_bindings,保留 sender_ids。Chat 可在 context_handoff 选择此授权绑定,不创建 task 或 host。
验证: 回读 receiver inbox 和精确 delegation operation;送达、提交启动、接收采纳、原生结果、规范验收与原 audience 返回分开。runtime_unverified 或准入拒绝保持明确恢复路径。
停用 / 回退: 删除同一 source policy 的精确 execution_bindings 条目,撤销后续启动与回放;已启动 operation 用 delegation stop --execute 检查/停止。停止管家不停止独立 worker。
权限边界: sender/read/context grant 不提供执行权;canonical preflight、注册 requester、provider provenance 与原 Turn 仍必需。不授予新调度、跨 host 启动、任意命令、保护写入或发布;operator editor/发现完整交付仍分阶段。
文档: 固定版本指南
Independent macOS service and Chat Codex homes
启用: 在 macOS 使用 v1.3.0 checkout 的脚本和一个已资格验证的 installed CLI owner。需要改变时分别显式设置执行用 CODEX_HOME 与 Chat 用 LOOPX_CHAT_CODEX_HOME;省略选择时 install/restart 保留已有绑定。 改变 Chat 存储/超时前显式设置 LOOPX_CHAT_RUNTIME_ROOT、LOOPX_CHAT_IDLE_TIMEOUT_SECONDS 和 LOOPX_CHAT_HARD_TIMEOUT_SECONDS;省略保留已存值,非法输入保留两份原 plist。
验证: 运行 service status,再读实际运行的 Chat identity 和各自选择的 home;合成源码检查不证明登录后重启或真实 provider 接受。
停用 / 回退: 运行 stop 卸载两个 owned service,或 uninstall 删除其两个 plist;通过同一安装 owner 恢复之前的包并 restart 回退,保留项目状态与 credential home。
权限边界: 选择模型 home 不复制不同 profile 凭据,不授予 Goal、workspace 写入、provider 调用或外发权限;scan path 只负责发现,仍由 Core 授权约束。
文档: 固定版本指南
Rendered public-source MCP reader
启用: 在实际执行宿主中显式添加唯一
[mcp_servers.loopx_ego_source_read]:command为安装环境的 Python,args = ["-m", "loopx.extensions.ego_source_reader"]、startup_timeout_sec = 30、tool_timeout_sec = 40;env 必须配置下方四个LOOPX_EGO_READ_*值。保留一个既有 Ego Page,仅允许任务需要的公共 origin;重启空闲宿主后恢复原 Session。适配器属于分阶段集成,原生 Bot/原渠道晋升未资格化。验证: 先读下方 stdio 工具清单,再由原宿主对允许的 URL/index 调用
read_public_url与read_public_image;核对 requested/observed URL、digest、truncation 和真实 ImageContent,并确认越界 URL 被拒绝。正文读取不等于图片理解或文章完整覆盖。停用 / 回退: 仅删除该唯一 MCP 配置项并重启空闲宿主;回退时由同一安装 owner 恢复私有备份/旧包,保留其他 MCP 项与 Session 绑定。
权限边界: origin 配置不授予私有/管理站点访问、sandbox 变更、workspace 写入、笔记编辑、发布或委派;页面内容是非可信证据。图片是可能被遮挡的渲染区域,Page 必须独占保留并遵守用户接管。
文档: 固定版本指南
Canonical retirement with an existing successor
启用: 仅在已晋升 canonical 的 Goal 中,对原 guard Turn 每次显式使用
todo supersede --successor-todo-id;填写原 Goal/Agent/Todo/Turn 和既有 successor,并携带该 Goal 要求的真实 lease proof。这是直接状态写入,不需额外 execute flag,也不启用 canonical 权威。验证: 操作后对同一 Goal 读取
todo list:核对前项已 superseded、既有 successor 链、原 lease 已释放,successor 的 owner/scope/status/due time 未改。响应含糊时复用原 Turn 与完全相同 intent;receipt replay 不代表新 lease。停用 / 回退: 不加新选项可保留旧的预先关联/生成 successor 流程;缺失、自关联、跨 Goal 或混合 intent 在写前拒绝。未晋升 Markdown 走既有迁移流程。退休是终态,不通过重置权威状态回滚;如需纠正,创建明确授权的修复工作。
权限边界: 不增授范围,不改 successor 要求或 Monitor 到期时间,不关闭 Goal,也不认证产物;原 writeback 与唯一 quota spend 仍独立。Chat/Lark 操作与尚未完成的前端迁移旅程保持原边界。
文档: 固定版本指南
Single-material ranked moves and receipt coverage
启用: 每个 SDK caller 显式调用 plan_material_single_move 与 material_rerank_receipt_chunks;已明确激活的项目 source 可在既有 inventory、intake/rollback builder 中传 MaterialProjectScope 代替 goal_id,必须二选一。可仅安装项目本地 loopx-material skill。下方两个合成 preview 不改变来源;真实效果仍需原 source adapter。
验证: 运行 preview 核对完整排名、其他条目相对顺序、100/67 回执覆盖,以及无需 Goal 的 project scope 元数据,再读 skill status。真实 project intake/rollback 的 provider.verify_project_scope 必须在访问/发布前解析当前 Core caller、audience、精确 profile/store、workspace grant、gate 到期与撤销状态,并保留来源 transaction fence。元数据不等于权限或已应用操作。
停用 / 回退: 按下方命令卸载项目本地 skill,停止调用 helpers 或传 project_scope;真实 store 只经既有已验证 backup、source adapter、owner gate 和 rollback 合同恢复。卸载 skill 不回退来源数据。
权限边界: 引用只选择已有 scope,不增授写入、私有访问、评分变更、cutover 或自动重排。项目 intake 不创建 Goal、不借用管家权限。migration、rebuild 和 Explore 保留 Goal 路径;本次只扩展既有 SDK/provider intake。实验性 slice 未增加 frontend/Lark/CLI writer,更广的项目材料旅程仍为 partial。
文档: 固定版本指南
Effective-work-Turn review cadence
启用: 在能力中心 → Goal 复核周期显式选择“已结算工作 Turn”和 1–5 次,可设置设备默认或指定 Goal。下方 CLI 选择三次已结算工作 Turn;原完成 Todo 默认节奏保持。
验证: 读取同一 Goal 的 configure-goal,再读取其 Agent 的 quota should-run。共享 TypeScript owner 在已接受复盘检查点后按唯一已结算工作 Turn 计数;重复 retry、poll、未扣额 writeback 和缺失结算均不计数。
停用 / 回退: 用下方 clear 命令移除 Goal 的有效 Turn 覆盖,恢复当前设备默认。用带 revision 校验的设备编辑器改回已完成 Todo,或移除 todo_replan_cadence 恢复旧能力默认;保留既有状态与回执。
权限边界: 阈值产生复盘义务,不授予 host 调度、执行、quota、workspace 访问或外部写入权限。v0 配置读取时仍按完成 Todo 解释,只有明确 apply 才迁移;已接受的负面工作结果也可以计数。
文档: Versioned guide
发布验证
个人图文发布指南,沿用现有飞书访问范围。
发布源码为
e1dd9e519c3057fda4b3765d0c93e7ae793608bf,tree为92ecd2e76d8bb72cfc7d87487363edecc20990da。八项精确资格通过,含pytest及子测试18,077通过/133跳过、实际wheel升级/Chat/三宿主、515public smoke及DS high22×2。包/PyPI、已证明的产品与构建身份、桌面/签名feed及stable安装已回读。夹具恢复保留46个浏览器场景及团队证据返回。原生构建通过;取消排队publisher后由维护者发布已核验CI bytes。首次DMG与历史DS断言原因未归属。Windows运行、Apple公证、真实Lark/Bot及缺失Harbor/SForge/PostgreSQL/NoKV未验证。Turn/diagnostic展示阈值为15,000/45,000 JSON字符,相同输出14,647/44,126超过旧14,600/44,000;monitor保留2,000及全部结算断言,fixture使用稳定物理目录。它们不是token/费用预算或晋升标准。不宣称匹配版本outcome提升。