Skip to content

mctest: Add better options for test validity based on sim / target errorbars - #2776

Merged
willend merged 13 commits into
mccode-dev:mainfrom
willend:mctest-sigma
Oct 9, 2026
Merged

willend merged 13 commits into
mccode-dev:mainfrom
willend:mctest-sigma

Conversation

@willend

@willend willend commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Free-form text area

Please describe what your PR is adding in terms of features or bugfixes:

mctest: judge test values by their Monte Carlo error bars; %Test: / %TestScan: tags

Why

mctest has always accepted a test value within a fixed ±20 % of its target. That is too strict for noisy, low-statistics tests and blind to how precise a run actually is. This PR lets mctest use the error bar every McStas/McXtrace run already prints, and lets a whole test run be scaled to fewer or more particles. Background: mccode-dev/mcstas-ess-instr-citool#4.

Nothing changes unless the new options are used, and no instrument files are touched. The full explanation, written for statistics-minded readers, is in tools/Python/mctest-statistics.md.

New mctest options

Option Effect
--nsigma n accept a value within n × the combined error bar of test and target
--pvalue p the same threshold as a two-sided Gaussian p-value (5.7e-7 ≡ 5σ)
--statfactor F scale the ncount of every test (and of each scan point) by F, e.g. 0.1 for quick runs or 100 for reference runs

How a value is judged

With Δ = I_test − I_target:

Target line Error bars usable Accepted when
no _ERR (all instruments today) yes within ±20 % or |Δ| ≤ n·√(1+F)·ERR_test
with _ERR yes |Δ| ≤ n·√(ERR_test² + ERR_target²)
any no within ±20 %
  • Unrecorded target error: assumed to come from the unscaled ncount, i.e. σ_target = σ_test·√F, which is √2·σ_test by default.
  • Usable error bars: at least 100 effective events, N_eff = (I/ERR)², the Kish effective sample size of the weighted particles. A run with only a handful of rays, a zero result, or a missing ERR falls back to ±20 %, so a huge error bar can never wave a bad value through.
  • Scans: every point of a scan is judged this way.
  • Today's targets: while they carry no _ERR, these options can only widen acceptance compared with ±20 %.

ERR is read from the monitor file, from the stdout Detector: …_ERR= line (NeXus), or from a scan's mccode.dat. A target may state its own error bar in the same format: NAME_ERR=…, or NAME_ERR={…} for scans.

%Test: / %TestScan: tags

These are the future names of %Example: / %Scan:. Both spellings are accepted everywhere, so instruments can switch gradually:

  • mctest, mcviewtest: run and report both spellings.
  • mcdoc: renders the lines as Test: / TestScan:.
  • mcstas-pygen: translates %Test: lines into instr.add_test, in file order with %Example:. %TestScan: lines have no add_test equivalent and are left alone.
  • mccode-git-diff-code: treats both as test definitions, including a scan's _ERR={…} block.
  • PR template: asks contributors for a %Test: (or %TestScan:) line instead of %Example:.

Also in this PR

  • mcviewtest: colours each cell by the same rule mctest used, from the threshold and factor stored with the results.
  • CI: the mcstas/mcxtrace conda testsuite workflows run with --nsigma 5. Basictests are unchanged.
  • Docs:
    • new mctest-statistics.md;
    • mctest.md, mcrun.md, codegen.md and the overview brought up to date. The old %Scan: example lacked the {…} the parser requires, so it never worked as written.
  • devel/bin/mccode-tooldocs-html: regenerates the tools/Python/*.html doc pages from their .md sources (mcdoc.md → mcdoc_tool.html). It reproduces all existing pages byte for byte. It needs python-markdown, now a dependency.

Choosing the threshold

A full suite judges a few thousand values once scan points are counted. At 3σ that gives several spurious failures per run; at 5σ, far below one. CI therefore uses 5.

Testing (macOS arm64, mccode-dev built from this branch)

Run Outcome
Basic-test instruments, --nsigma 5, MPI (McStas + McXtrace) all pass; extremes BNL_H8 −4.2σ, ESS_BEER_MCPL +4.6σ, Test_KDSource_3/4 −13σ/−42σ at 97 % of target (pass via ±20 %)
--statfactor 10 all pass, incl. BNL_H8_simple (100.8 %), which is flaky at default statistics
--statfactor 0.1 5 failures, all below the N_eff limit: too little statistics, reported as such
-n100 starved run falls back to ±20 % and fails; without the gate a value at 28 % of target (−2.2σ) would have passed
Targets with _ERR (scratch copies) pure error-bar test, BNL_H8 fails at −6.5σ as intended
NeXus output identical verdicts
mcstas-pygen %Test: gives the same code as %Example:; mixed tags kept in file order; %TestScan: skipped; generated script compiles; no compiler warnings

Unit checks cover every acceptance mode, mctest/mcviewtest agreement, parsing of all four tags, and the diff tool.

Known limitations

  • SPLIT instruments: ERR understates the real scatter there (2–3× measured on SSRL/LLB), so targets with _ERR on such instruments would fail spuriously.
  • File-input tests (MCPL, KDSource) don't change with --statfactor; for them only the assumed target error changes.
  • Recording _ERR: record it only together with a fresh, high-statistics target value. Some targets are a few per cent off for non-statistical reasons and would fail a pure error-bar test.
  • Not run locally: Linux, Windows, macOS x86_64 (CI covers them). --statfactor was exercised on McStas instruments only.

Declaration of use of AI-tools

  • Please add a checkmark here if you used AI-tools during the work for this contribution
  • Furter, please describe how / where and for what the tools were used:

^ 🤖 Generated with Claude Code


Development OS / boundary conditions

Please describe what OS you developed and tested your additions on, and if any special dependencies are required:


PR Checklist for contributing to McStas/McXtrace

For a coherent and useful contribution to McStas/McXtrace, please fill in relevant parts of the checklist:

  • My contribution contains something else

    • Explanation is added in free form text above or below the checklist

willend and others added 13 commits October 9, 2026 13:28
The simulation's own error bars (ERR) are now read alongside I: from the
monitor file's "# values:" line, the stdout "Detector: NAME_I=...
NAME_ERR=..." line (NeXus), or the NAME_ERR column of a scan's mccode.dat.

With --sigma S a test value (each point of a %Scan) is accepted if
|I - I_target| <= S * err, where
- err = sqrt(ERR_test^2 + ERR_target^2) when the target line gives its
  error bar (NAME_ERR=... for %Example, NAME_ERR={...} for %Scan, as in
  the simulation's own Detector: line); judged by the error bars alone
- otherwise the target is assumed to have the test run's error bar,
  err = sqrt(2) * ERR_test, and the value is also accepted within 20%.
Without --sigma the 20% rule is unchanged.

Results store testerr/targeterr/sigma, and mcviewtest applies the same
rule (older results keep the 20% rule). mccode-git-diff-code includes a
scan's _ERR={...} block in the test definition. No instrument files
are changed.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Accept test values within 20% or within 5 x the combined error bar
(sqrt(2) x the test run's ERR, as no targets carry _ERR yet), to see how
the error-bar based acceptance behaves across all platforms.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Both spellings are accepted, so instruments can move to the new tags
gradually:
- mctest: %Test: lines are tests like %Example:, %TestScan: lines scans
  like %Scan: (incl. NAME_ERR targets and multi-line value lists)
- mccode-git-diff-code: %Test:/%TestScan: lines are test definitions too
- mcdoc (mccodelib parse_header/format_examples): the test section also
  starts at %Test:/%TestScan:, and the lines are labelled "Test:" and
  "TestScan:" (scan lines were labelled "Scan:" before); Tomography's
  README updated to the new label

Checked: all four tags parse in mctest, mcdoc and the diff tool; mcdoc
output for all 452 example instrument headers is unchanged apart from
the scan label.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
A value from very few rays has a huge error bar, so S x ERR would accept
almost anything (e.g. ISIS_CRISP at -n100: 28% of target, but only -2.2
sigma). --sigma now only uses error bars with ERR/|I| <= 25%
(SIGMA_MAX_RELERR, ~1/sqrt(N) for N >~ 16 unweighted events), for the
test value and, when given, the target's _ERR; I = 0 or a missing ERR
also count as unusable. Otherwise the value is judged by the 20% rule.
The log says when this happens. mcviewtest applies the same rule.

Checked: basictest instruments at -n1e6 --mpi=auto unchanged (all pass
on "20% or 5 sigma"); at -n100 the starved values fall back to 20% and
fail.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
- --sigma is renamed --nsigma: it is a maximum z-score, not a sigma
  value (no alias, it was not in use yet). CI testsuites use --nsigma 5.
- --pvalue P: the same threshold as a two-sided Gaussian p-value
  (0.0027 = 3 sigma, 5.7e-7 = 5 sigma), converted with
  statistics.NormalDist; cannot be combined with --nsigma.
- Error bars are only used from >= 100 effective events,
  N_eff = (I/ERR)^2 (was ERR/I <= 25%, ~16 events), matching the ~100
  event floor proposed for --statfactor; below that the 20% rule
  applies. Noisy tests like BNL_H8_simple (N_eff ~19 at 1e6) are thus
  judged by the 20% rule.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
--statfactor F multiplies the ncount of every test by F: the general
ncount (-n, default 1e6) for single-value tests, and the per-point ncount
of each scan (the scan line's -n, else at most 1e5). Run timeouts grow
with F, and the factor is part of the result directory's name.

With --nsigma/--pvalue, a target without NAME_ERR is assumed to come from
the unscaled ncount, i.e. sigma_target = sigma_test * sqrt(F), so the
combined error bar is sqrt(1 + F) * ERR_test (sqrt(2) for F = 1, as
before). The test run's own ERR follows the scaled statistics by itself.
mcviewtest applies the same rule from the stored statfactor.

Checked on the McStas basic-test instruments + SE_example scans:
F = 10 all pass (incl. BNL_H8_simple, which fails 20% at 1e6); F = 0.1
gives 5 failures that fall below 100 effective events and fail 20%.
File-input tests (MCPL, KDSource) do not change with F, so for them only
the assumed target error changes.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Explains for statistics-savvy readers how mctest's --nsigma, --pvalue and
--statfactor judge test values against their targets: where I and ERR
come from (checked against mcestimate_error), the decision table, the
sqrt(1+f) assumption for unrecorded target errors, the N_eff >= 100 gate
(Kish effective sample size), the n-sigma / p-value equivalence and the
multiple-comparisons argument for 5 sigma in CI, known limitations
(SPLIT-correlated weights, systematic target drift, file-input tests, MPI
random streams, Gaussian approximation), and measured results on the CI
basic tests at statfactor 0.1, 1 and 10. Linked from mctest.md.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
mctest.md:
- %Test:/%TestScan: are the documented test tags (older names %Example:
  and %Scan: still accepted)
- the scan example uses the {}-enclosed target list the parser requires
  (the old example without braces would not have been recognised)
- a value or scan point passes within 20%, or with --nsigma/--pvalue
  within the statistical tolerance; --nsigma row gives sqrt(1+F) with
  --statfactor F; --skipnontest/--noscans/--strict rows use the new tags
- mcviewtest: one row per test/scan, cells coloured by the rule mctest
  used, per-point tooltip on scan cells
mcrun.md: --scan_split is ignored (with a warning) for an explicit
multi-process --mpi.
README.md: link to the statistical acceptance page.

The .html twins are regenerated from the .md files with python-markdown
(tables, toc, fenced_code) and the existing page head, which reproduces
the committed mcplot.html exactly; mctest.html and mcrun.html had not
followed the earlier %Scan:/--scan_split doc changes. Adds
mctest-statistics.html.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Converts tools/Python/NAME.md to NAME.html with python-markdown (tables,
toc, fenced_code), keeping each page's existing <head>; mcdoc.md becomes
mcdoc_tool.html, and links to .md files are rewritten to the .html names.
With no arguments all tool docs are converted. Running it on the current
tree reproduces all 9 committed .html pages byte for byte.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
mcstas-pygen turns test lines into McStasScript instr.add_test calls.
It now accepts %Test: as well as %Example:, in file order when both are
used. %TestScan: lines (parameter scans) have no add_test equivalent and
are left alone; "%Test:" does not match them. Warnings, the generated
comment, the --no-tests help text and codegen.md/.html name both tags.

Checked: BNL_H8 with %Test: instead of %Example: generates identical
code; a header mixing %Example:/%Test:/%TestScan: gives the three
single-value tests in file order and the generated script compiles; no
compiler warnings for pygen.c.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@willend
willend merged commit 31a2e7f into mccode-dev:main Oct 9, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant