Skip to content

About

Show prompt changes as readable diffs with affected tests and variables.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

18 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Prompt Version Diff

Compare two versions of a structured prompt the way a reviewer needs to read them: by named section, by variable, by output contract, and by which test cases each change touches. Whitespace-only edits can be ignored; a removed required variable and a weakened instruction cannot.

  • Repository: edilec/prompt-version-diff
  • Area: Prompt & Agent Workflows
  • License: MIT
  • Dependencies: none, at runtime and in development. Node 22 or newer.

Why it exists

A line diff of a prompt tells you which characters moved. It does not tell you that must became should in the grounding section, that the template stopped accepting repository, that the output format changed from JSON to Markdown, or that three test cases now point at something that no longer exists. Those are the changes that break callers, and they are exactly the ones a reflowed paragraph hides.

This tool reads both versions as structures rather than as text, matches sections and variables by name rather than by position, and reports the differences that change the contract. Reformatting is quiet; a weakened instruction is loud.

What it does not do

  • It does not diff prose. A section whose wording changed is reported as changed, at its pointer. The tool does not say which words moved, does not produce a hunk, and makes no judgement about whether the new wording is better. Use an ordinary diff alongside it.
  • It does not call a model, and never reaches the network. No provider call, no telemetry, no fetch of any kind, in the tool or in its tests.
  • It does not judge whether a change is safe. error means "breaks the contract a caller relied on". Intent is the reviewer's to supply.
  • It does not run the test cases. testCases entries are references. The tool reports which of them a change touches; it does not execute anything.
  • It does not compare nested schema structure. An output property is compared by its declared type string and by whether it is required. A property whose value carries no string type is reported as not comparable and the run is incomplete.
  • It does not match a renamed section to its old self. A rename reads as one section removed and one added, which is the honest reading: the tool has no evidence that the two are the same section.
  • It does not parse Markdown or free-form prompt files. Input is JSON.
  • It does not walk directories and it never writes to its inputs. Exactly two files are passed explicitly; the only file it writes is the one named by --out, and that destination is refused unless it resolves inside the write root and is neither of the two versions.

Quick start

# an additive revision: exits 0
node bin/prompt-version-diff.mjs --root examples \
  examples/v1-release-notes.json examples/v2-additive.json

# the same, human-readable
node bin/prompt-version-diff.mjs --root examples --format text \
  examples/v1-release-notes.json examples/v2-additive.json

# a revision with breaking changes: exits 1
node bin/prompt-version-diff.mjs --root examples --format text \
  examples/v1-release-notes.json examples/v3-breaking.json

# report every text difference, whitespace included
node bin/prompt-version-diff.mjs --root examples --no-ignore-whitespace \
  examples/v1-release-notes.json examples/v2-additive.json

# everything the repository checks before a commit
npm run check

The input document

Both versions use the same shape.

{
  "id": "release-notes-writer",
  "version": "1.4.0",
  "sections": [
    { "name": "objective", "authority": "must", "text": "Produce release notes." },
    { "name": "tone", "authority": "should", "text": "Write for engineers." }
  ],
  "variables": [
    { "name": "repository", "required": true },
    { "name": "since", "required": false, "default": "the previous tag" }
  ],
  "outputContract": {
    "format": "json",
    "properties": { "title": { "type": "string" } },
    "required": ["title"]
  },
  "testCases": [{ "id": "tc-title", "covers": ["section:objective", "output:title"] }]
}
  • id must be present in both versions and must match, unless --allow-id-mismatch is given. Two different prompts are not two versions.
  • sections are matched by name. Every section must declare a text string and an authority. Duplicate names make sections uncomparable.
  • authority is one of info, may, should, must, weakest first. A step down the ladder is an error; a step up is a warning. Both are reported.
  • variables are matched by name. required defaults to false.
  • outputContract must be present and must declare a format. Properties are compared by their declared type.
  • testCases[].covers names targets in the form section:<name>, variable:<name>, output:<property> or contract:format. The list is deduplicated and walked in code-unit order, so a target named twice produces one finding rather than two identical ones, and the findings a single covers list produces do not depend on the order the list happens to be written in.

Whitespace, and the one thing it must never hide

--ignore-whitespace is the default. Under it, a section whose text differs only in whitespace produces a section-whitespace-only-change info, is not counted as a change, and does not fail the run. --no-ignore-whitespace reports it as section-text-changed instead.

Authority is compared independently of text, and always. A section whose text moved only in whitespace and whose authority fell from must to should still produces section-authority-weakened and still exits 1. The suite pins exactly that combination, because folding the two comparisons together — for instance by skipping a section whose normalised text is unchanged — is the plausible implementation that silently loses the finding.

Rules

Sixty-two rules, each with a stable kebab-case id and a severity fixed in one frozen catalog (src/rules.mjs). The full table is in docs/rules.md; it is generated from the catalog and the suite asserts the two agree in both directions.

Severity decides the exit code: any error fails the run. Rule severities are pinned in the suite behaviourally — each one is driven through the real report path as a real pair of versions, and the assertion is on the resulting status, not on a second copy of the table.

rule severity what it catches
section-authority-weakened error must became should, or should became may
section-authority-strengthened warning a step up the same ladder
section-removed error a named section is gone
section-text-changed warning a text change that is more than whitespace
section-whitespace-only-change info a text change that is only whitespace
required-variable-removed error the new version no longer accepts a variable callers pass
required-variable-added error a caller that satisfied the old contract no longer satisfies this one
variable-requirement-tightened error an optional variable became required
output-format-changed error the declared output format is different
output-property-removed error a declared output property is gone
output-property-type-changed error a declared output property changed type
test-case-orphaned error a test case names a target the new version removed
test-case-impacted warning a test case names a target that changed
change-without-test-coverage warning a change no test case names
section-name-duplicated error sections cannot be matched by name, so none were compared

Output

stdout carries the JSON report and nothing else, so it can be piped straight into a parser. --format text swaps stdout to the human summary instead. Operational diagnostics go to stderr.

{
  "schemaVersion": "1",
  "tool": "prompt-version-diff",
  "status": "fail",
  "summary": {
    "checked": 1,
    "errors": 9,
    "warnings": 7,
    "infos": 0,
    "inputs": 2,
    "changes": 16,
    "findingsTotal": 16
  },
  "findings": [
    {
      "ruleId": "section-authority-weakened",
      "severity": "error",
      "message": "Section \"grounding\" fell from \"must\" to \"should\".",
      "location": { "file": "v3-breaking.json", "pointer": "/sections/1/authority" },
      "suggestion": "Confirm the weaker instruction is intended; callers relying on the stronger one are affected."
    }
  ]
}

A finding is located in the document that holds its evidence: a removal in the earlier version, an addition or a change in the later one. summary counts every finding produced, including any the finding limit kept out of findings.

Findings sort by (location.file, location.pointer, ruleId) using UTF-16 code units — never localeCompare or Intl.Collator, whose ICU data differs between Node builds — so two machines produce the same bytes. Several findings from one covers list share all three keys, so that sort leaves them in emission order; the covers list is itself ordered by code unit for the same reason, and the suite pins the emitted order with targets whose document, collation and code-unit orders all disagree. A findings-truncated notice, when present, is appended last because it describes the report rather than a document. Findings that describe the run rather than a document are located at ".".

Exit codes

code meaning
0 no breaking change between the two versions
1 at least one breaking change between the two versions
2 invalid usage, or a comparison could not be made

Exit 2 has two shapes, and the difference is deliberate:

situation stdout status
invalid usage, unknown option, not exactly two files empty no report
a version that could not be read, decoded, parsed or keyed a report incomplete

A comparison the tool could not make is never reported as an absence of difference. If the sections of either version cannot be matched by name, no section finding is emitted at all — not "all sections unchanged", not "every section added" — and the run is incomplete. Each collection is independent, so a removed required variable is still reported when the sections are unusable.

Limits

Every limit is enforced from the command line, and exceeding one produces a finding naming it and an incomplete run — never a silent truncation and never a pass.

option default exceeding it
--max-bytes 1000000 input-too-large; the version is not read
--max-items 1000 too-many-items; that collection is not compared
--max-findings 500 findings-truncated; the counts still include the dropped findings
--time-budget-ms 10000 time-budget-exceeded; comparisons the budget never reached contribute nothing and checked stays 0

An unknown option is a usage error, never a no-op, so a typo such as --no-ignore-witespace cannot silently keep ignoring whitespace.

Safety properties

Each of these is covered by a test that fails when the guard is removed.

  • Parse failures never reproduce the document. V8 embeds a window of the input in its JSON.parse error messages, so error.message is never interpolated. parseFailureDetail keeps the offset and drops the quoted span, recognises the quoting shape before looking for an offset (a document whose own text reads at position 1 defeats the other order), and ends with a backstop that refuses any detail still carrying a double quote.
  • Every untrusted string is sanitized, not only evidence. Section names, coverage targets, paths, pointers, messages and excerpts all pass the same filter, which removes C0, DEL, C1, U+2028/U+2029 and the bidi controls. Identifiers and format/type labels are matched by that rendered form: an invisible-only identifier or a collision after sanitization is incomplete, not a fabricated add/remove or a vacuous pass. Unknown non-string authority values are reported without attempting unsafe string coercion.
  • The write destination is checked before anything is written. --out is refused if it is a symlink (realpath would follow it, and following is the dangerous act), if its parent resolves outside the write root (a lexical prefix check passes for root/link/out where link leaves the root), if it is not a regular file, or if it is the same file as one of the two versions by device and inode (a hard link shares no path with its input, so only identity sees it). All three destroyed a file here and exited 0 reporting success before the guard was wired in. The write root is the working directory unless --out-root declares another; --out-root without --out is a usage error rather than an option accepted and ignored, and an --out-root that does not exist is refused the same way rather than ending the run in an unhandled error.
  • Inputs are confined by real path. A symlink planted inside the declared root that points out of it is refused, its content is never read, and its destination is never printed.
  • Encoding validity is decided by a strict decoder, never inferred from decoded text, so a document legitimately containing U+FFFD still parses.

Repository layout

  • src/ — library. rules.mjs is the frozen catalog, diff.mjs the preparation and comparison, report.mjs the envelope, ordering and budget, inputs.mjs the reader and confinement, text.mjs the sanitizer and comparator, parse-failure.mjs the parse-error helper, write-guard.mjs the --out destination check.
  • bin/prompt-version-diff.mjs — the CLI.
  • examples/ — one base version, one additive revision, one breaking revision.
  • test/ — node:test, no dev dependencies.
  • docs/rules.md — the generated rule table.

License

MIT. See LICENSE.

About

Show prompt changes as readable diffs with affected tests and variables.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages