Skip to content

v0.2.0: multilingual ADR generation, CHECK hardening, and ADR search/relationships - #2

Merged
SHcommit merged 62 commits into
developfrom
feature/v0.2.0-multilingual-and-check-confidence
Aug 31, 2026
Merged

SHcommit merged 62 commits into
developfrom
feature/v0.2.0-multilingual-and-check-confidence

Conversation

@SHcommit

Copy link
Copy Markdown
Owner

Summary

  • Repository-configured deterministic ADR generation across eight canonical
    locales (en/ko/ja/zh/fr/es/de/pt-BR), with portable Unicode
    filenames and an approved-semantic-slug fallback.
  • Hardened CHECK against several false-clean evidence gaps (incomplete git
    diff, renames, Unicode paths, ignored/deleted paths, conflicting diff
    modes) and promoted its confidence classification to a stable
    confidence field (VERIFIED/VIOLATED/UNVERIFIABLE) on every finding.
  • Added a deterministic policy-exception mechanism (adr.py exception):
    schema-validated, annotate-only — an exception is recorded and visible on
    a finding, never a silent pass.
  • Corrected repository ADR provenance/quality and superseded ADR-0003 with
    ADR-0006 (localization decision); this session's own design decisions are
    dogfooded as ADR-0007..0010.
  • New: adr.py search (title+body keyword, tags, status, --id,
    governed-path matching, deterministic ranking, --limit/total/
    truncated) and a Relationships section in the generated index
    (supersession chains + related lists with titles, localized across all 8
    catalogs), plus validate.py relationship-integrity checks (broken link,
    one-sided supersession, supersession cycle). Design rationale — grounded
    in how npryce/adr-tools and thomvaill/log4brains handle this — is in
    README's "Why storage stays flat" section and
    docs/superpowers/specs/2026-08-31-adr-search-and-relationships-design.md.
  • Fixed several smaller correctness issues found along the way:
    sync_version.py silent-drift/crash/non-ASCII bugs, manifest-description
    drift, an unused Codex adapter install step, and a related.py matching
    gap (title-only keyword, and a dropped non-list-field guard).

Test plan

  • python3 -m pytest -q → 386 passed
  • python3 scripts/sync_version.py --check → exit 0
  • adr.py validate --dir docs/decisions → 10 ADRs, no errors (including
    the new relationship-integrity/cycle checks)
  • adr.py index --dir docs/decisions regenerates cleanly, byte-stable
  • Every newly documented CLI command was run in a scratch repository
    before being written into README/SKILL.md
  • CI (5-leg OS/Python matrix + version-drift job) — verify green on this PR

Release readiness detail (predates the search feature; re-verify before
treating its GO/NO-GO judgment as current) is in
docs/adr-toolkit-v0.2.0-readiness-report.md. Version bump, release branch,
and tag are intentionally not part of this PR — separate owner approval
required per AGENTS.md's Git Flow policy.

🤖 Generated with Claude Code

https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF

SHcommit and others added 30 commits August 30, 2026 19:09
Adds repository-configured eight-locale ADR generation with strict
config/schema validation, portable Unicode filenames with an approved
semantic ASCII slug, closes several CHECK false-clean gaps, and makes
repository-scoped commands resolve relative paths from --root instead
of the caller's CWD. Includes updated docs, the dogfooded ADR history
(ADR-0006 superseding ADR-0003), and the Korean v0.2.0 readiness and
enterprise-adoption reports with local release evidence.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
A manifest that keeps its file but loses its "version" key was silently
skipped by --check, so CI would report no drift even after the tracked
key vanished. A manifest path outside the repo root crashed
require_known_paths() with an unrelated ValueError instead of its
intended SystemExit. JSON writes escaped non-ASCII manifest content
instead of preserving it, which regresses this project's own
multilingual metadata.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
sync_version.py --check reads repository content only; its result
cannot vary by OS or Python interpreter. Running it identically on all
5 pytest matrix legs wasted CI minutes for no additional coverage.
Split it into its own single-run job instead.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
related, index, validate, and check each redefined the same SKIP_FILES
set and re-walked the ADR directory with their own glob/skip/parse
loop. Extract core.adr_directory.iter_adr_files as the single source
of truth for "which files in this directory are candidate ADRs,"
while each command keeps its own bad-filename and frontmatter-error
handling, since those differ deliberately (validate reports a bad
filename as an error; the others skip it silently).

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Re-verified end to end against Codex CLI 0.151.0 with no
adapters/codex/skills/ symlink present at all: codex plugin marketplace
add + codex plugin add + preflight --json all still succeed, because
registering the repo root resolves .claude-plugin/plugin.json as the
plugin manifest, whose sibling skills/ is the real package. Install
step 1 (create that symlink) was therefore dead documentation from an
earlier design plan. Removed it from the install flow and added a
section explaining that .codex-plugin/plugin.json and its symlink are
structural-only, kept for consistency with the other adapters rather
than exercised by the verified path.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Manifest descriptions had drifted the same way versions once did:
.claude-plugin/plugin.json was missing "and existing decisions" that
SKILL.md and the three adapter manifests all carried. Extend
sync_version.py to treat SKILL.md's frontmatter description as
canonical and sync it into every duplicating manifest, with --check
enforcing it the same way version drift already is. Fixes the one
real drift this uncovered.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Every subcommand parsed --json but main() always printed JSON
regardless, matching its own "JSON-only-stdout contract" comment and
every existing test/doc caller. Rather than inventing a human-readable
mode or dropping the flag as a breaking change, made that existing
behavior the deliberate, documented contract: centralized the 14
duplicated --json argument definitions into one helper with clear
--help text, added a regression test proving output is JSON with or
without the flag, and added --json to the three doc examples that had
omitted it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Item 1 (shared ADR directory iterator) shipped in 2f2a8de. Annotated
the remaining four with their actual current state: item 2 blocked on
the repo going public, item 5 blocked on a second repository existing,
and items 3-4 genuinely not started, with the concrete gap named (no
VERIFIED-equivalent finding kind exists yet; register_exception is a
label with no backing schema).

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
README, SKILL.md, and conflict-rules.md already documented a mapping
from each finding's kind to a VERIFIED/VIOLATED/UNVERIFIABLE
confidence value, but it existed only in prose — the agent had to
re-derive it from docs on every run. Compute it once in check.py and
attach it to every finding directly as `confidence`, using the exact
mapping those docs already specified (related->VERIFIED,
verified_violation->VIOLATED, review_required and
no_applicable_constraint->UNVERIFIABLE). Updated the docs to describe
the field instead of a derivation the agent had to perform, and synced
quickstart's example output to real command output.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
register_exception existed only as a resolution label with no backing
schema or storage. Add scripts/core/exceptions.py (validation,
expiry, scope matching against globs), a new `adr.py exception`
command that validates a draft (adr_id, rule_id, owner, reason, scope,
expiry) and writes docs/decisions/exceptions/NNNN.json as EXC-NNNN,
and schemas/exception.schema.json documenting the same shape.

CHECK loads active (non-expired, schema-valid) exceptions and
annotates a matching verified_violation finding's `exception` field —
it never suppresses or downgrades the finding, keeping kind/confidence
exactly what the structural evidence says, consistent with this
project's "no false governance confidence" principle. A malformed
exception file degrades to a BAD_EXCEPTION warning rather than
aborting the run, matching the existing ADR/constraints pattern.

Updated README, SKILL.md, and conflict-rules.md to document the
command and field; verified the documented command end to end in a
scratch repository.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
The last handoff predated 8 commits (sync_version hardening, CI
dedup, ADR directory centralization, Codex adapter doc fix, manifest
description ownership, --json contract, CHECK confidence field,
exception schema) and still cited a 290-test baseline against the
current 327. Recorded what's done, what's left (two owner-gated P1
items and the P0 release gate), and confirmed project-roadmap.md has
nothing actionable right now.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
quick_validate.py's rejection of user-invocable/version comes from
Codex CLI's own local skill-creator authoring tool, never from the
actual plugin install/discovery path (independently verified working).
skill-creator targets simple Codex-only skills with a narrower schema;
adr-toolkit is deliberately cross-harness, so there's no real gap to
close. Recorded the decision and moved the item to Done rather than
leaving it open pending a fix that isn't needed.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Both origin/SHcommit/feat-plan-adr-toolkit and
origin/feat/adr-toolkit-mvp-implement were fully merged into
origin/develop with no open PR referencing either. Owner approved
deletion; deleted via git push origin --delete and confirmed gone.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Dogfoods the toolkit's own RECORD workflow: scored each candidate
decision with adr.py significance before writing anything.

- ADR-0007 (10, recommended): promote CHECK's kind-to-confidence
  mapping to a stable output field.
- ADR-0008 (12, recommended): deterministic CHECK policy exceptions,
  schema-validated and annotate-only, never suppressing a violation.
- ADR-0009 (7, recommended): --json is a documented no-op; CLI output
  is always JSON.
- ADR-0010 (4, optional): Codex skill-creator's quick_validate.py
  incompatibility is not this project's problem, with the verification
  evidence recorded so the question doesn't need re-investigating.

Regenerated docs/decisions/README.md; full suite (327 tests) and
validate both pass with all 10 ADRs.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
SHcommit and others added 28 commits August 31, 2026 02:03
Moves the directory-boundary prefix-matching primitive CHECK's
affected_paths_overlap already used out of rules/conflict.py's private
_path_under into core/globs.py's public path_under(), so the upcoming
search command's --path filter can reuse it without core depending on
rules (nothing in core currently imports from rules).

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
matches_keyword fixes the title-only search gap (matches body too).
matches_tags_any/matches_paths_exact extract related.py's existing
set-intersection logic unchanged. path_governed_by is new: does a real
file path fall under an ADR's governed scope, using the same
directory-boundary prefix + glob logic CHECK uses -- a different
question from matches_paths_exact's frontmatter-literal-list check.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Fixed tiers (exact id > exact title > title substring > tag-only >
body substring), id as tie-break. No numeric score is exposed as a
public field -- this is purely an internal sort key.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
…ps plan

Recorded completed commits, exact remaining task list, and pointed at
the plan file's location (gitignored, filesystem-only) plus a
fallback (the committed spec's own sequencing) in case a fresh session
needs to resume without it.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
resolve() turns each ADR's related/supersedes/superseded_by frontmatter
into a flat edge list, so index.py and validate.py both read the same
structured data instead of each re-deriving it from raw frontmatter.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
missing_targets finds a related/supersedes/superseded_by reference to
an ADR id that doesn't exist. supersession_mismatches finds a
one-sided supersession edit -- exactly what adr.py supersede's atomic
write normally prevents, but a hand-edit could still produce.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Migrated onto core/query.py's shared matching primitives. Combination
policy is unchanged (path OR tag OR keyword -- related wants a broad
net before drafting); all pre-existing tests pass unmodified.

Also fixes a regression the migration would otherwise have introduced
and one already latent in path_governed_by: extracting related.py's
matching logic dropped its _as_list() guard against a malformed ADR
carrying a non-list value for affected_paths/tags (e.g. a plain string
instead of a YAML list), which would silently decompose into
individual characters and produce nonsense matches. Restored the guard
in core/query.py, applied consistently to all three matchers.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
No filters given returns every ADR (browse mode). Different filter
fields combine with AND; multiple values in --tags combine with OR.
--path uses path_governed_by (directory-boundary + glob), not
related.py's exact frontmatter-list matching.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Results are best-match-first via core/query.py's rank_key. --limit
truncates the already-ranked list (weakest matches dropped first);
total is always the untruncated count; truncated is total > count.
query echoes the effective filters back for harness introspection.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Recorded two real bugs found and fixed mid-task (a dropped non-list
guard, and this repo's minimal frontmatter parser not supporting
YAML flow-style lists) so the next session doesn't rediscover them.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Supersession chains and related lists now render with titles, not
just IDs, built from core/relationships.py's resolved edge list. An
ADR with no relationships is omitted from the section entirely.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Same key-set-equality contract test_locale.py already enforces for
every other generated-index string.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
BROKEN_SUPERSESSION_LINK for a dangling target, SUPERSESSION_MISMATCH
for a one-sided edit -- exactly what adr.py supersede's atomic write
normally prevents. Both are errors, matching BROKEN_RELATED_LINK's
existing severity; validate.py has no warnings mechanism and this
doesn't add one for two checks. Verified against this repository's
real 10 ADRs: ok, no errors.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Every shown command was run in a scratch repository first; example
output is real, not hand-written.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
improvements.md gets the full Done entry, project-roadmap.md's "ADR
navigation and scale" item is split into a done bullet (search +
relationship visibility) and three still-deferred ones (rendered
graph, 500+-scale sharding/real search index, semantic retrieval),
and handoff.md reflects the feature as complete with the P0 release
gate as the only remaining open item.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
find_cycles() walks only "supersedes" edges (not "related", which is
symmetric and has no logical cycle concept) via DFS, returning each
cycle's node path. validate.py reports SUPERSESSION_CYCLE, matching
the existing error-only contract. Verified against this repository's
real 10 ADRs: ok, no errors.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
--id combines with other filters via the same AND policy as
keyword/tags/status/path, rather than a separate short-circuit path,
keeping search's combination semantics uniform across every filter.
Verified against a real ADR in this repository.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
README's Search section documented the command but not the design
rationale. Added it: checked how npryce/adr-tools (no search, expects
grep) and log4brains (1.5k+ stars, most-adopted actively maintained
ADR tool, still flat storage + search/graph layered on top) handle
this before deciding. The corollary this repo follows is that
retrieval, not storage layout, is where a growing ADR set gets harder
to use -- folder sharding, a rendered relationship graph, and a real
search index stay deliberately deferred in project-roadmap.md until
ADR count actually demonstrates the need.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
…ault

Windows CI failed test_uncommitted_mode_preserves_unicode_untracked_path_and_content:
a Unicode filename round-tripped through subprocess.run(..., text=True)
without capture as garbage on windows-latest (both 3.9 and 3.12), while
macOS/Ubuntu passed. Root cause: without an explicit encoding=, Python
decodes subprocess output using the platform's preferred locale
encoding, which is UTF-8 on macOS/Linux but not guaranteed on Windows
-- git itself always writes UTF-8. Added encoding="utf-8" to every git
subprocess.run() call in diff.py and git_paths.py. Pre-existing bug
from earlier CHECK-correctness work in this branch; only surfaced now
because this branch's CI had never run before this PR.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Owner-approved minor release version bump. skills/adr-toolkit/VERSION
is the source of truth; scripts/sync_version.py propagated it to
every manifest that duplicates it (.claude-plugin/plugin.json,
adapters/gemini-cli/gemini-extension.json, SKILL.md's frontmatter).

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
Recorded everything completed this session (ADR search/relationships,
CHECK confidence/exceptions, --json contract, sync_version hardening,
Windows encoding fix, v0.2.0 bump) in changelog.md, which already
serves as this project's durable human-readable history per AGENTS.md.
improvements.md's Done section and handoff.md's long commit-by-commit
walkthrough are redundant with that plus git log, so both are trimmed
to current/active state only -- nothing is lost, it's consolidated
into the file whose stated purpose is exactly this record.

Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01CeDZgjQakpnGGtsWoC7ijF
@SHcommit
SHcommit merged commit 3a3a482 into develop Aug 31, 2026
6 checks passed
SHcommit added a commit that referenced this pull request Aug 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant