Repository navigation
mctest: Add better options for test validity based on sim / target errorbars - #2776
Merged
Merged
Conversation
The simulation's own error bars (ERR) are now read alongside I: from the
monitor file's "# values:" line, the stdout "Detector: NAME_I=...
NAME_ERR=..." line (NeXus), or the NAME_ERR column of a scan's mccode.dat.
With --sigma S a test value (each point of a %Scan) is accepted if
|I - I_target| <= S * err, where
- err = sqrt(ERR_test^2 + ERR_target^2) when the target line gives its
error bar (NAME_ERR=... for %Example, NAME_ERR={...} for %Scan, as in
the simulation's own Detector: line); judged by the error bars alone
- otherwise the target is assumed to have the test run's error bar,
err = sqrt(2) * ERR_test, and the value is also accepted within 20%.
Without --sigma the 20% rule is unchanged.
Results store testerr/targeterr/sigma, and mcviewtest applies the same
rule (older results keep the 20% rule). mccode-git-diff-code includes a
scan's _ERR={...} block in the test definition. No instrument files
are changed.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Accept test values within 20% or within 5 x the combined error bar (sqrt(2) x the test run's ERR, as no targets carry _ERR yet), to see how the error-bar based acceptance behaves across all platforms. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Both spellings are accepted, so instruments can move to the new tags gradually: - mctest: %Test: lines are tests like %Example:, %TestScan: lines scans like %Scan: (incl. NAME_ERR targets and multi-line value lists) - mccode-git-diff-code: %Test:/%TestScan: lines are test definitions too - mcdoc (mccodelib parse_header/format_examples): the test section also starts at %Test:/%TestScan:, and the lines are labelled "Test:" and "TestScan:" (scan lines were labelled "Scan:" before); Tomography's README updated to the new label Checked: all four tags parse in mctest, mcdoc and the diff tool; mcdoc output for all 452 example instrument headers is unchanged apart from the scan label. Co-Authored-By: Claude Opus 5.5 <[email protected]>
A value from very few rays has a huge error bar, so S x ERR would accept almost anything (e.g. ISIS_CRISP at -n100: 28% of target, but only -2.2 sigma). --sigma now only uses error bars with ERR/|I| <= 25% (SIGMA_MAX_RELERR, ~1/sqrt(N) for N >~ 16 unweighted events), for the test value and, when given, the target's _ERR; I = 0 or a missing ERR also count as unusable. Otherwise the value is judged by the 20% rule. The log says when this happens. mcviewtest applies the same rule. Checked: basictest instruments at -n1e6 --mpi=auto unchanged (all pass on "20% or 5 sigma"); at -n100 the starved values fall back to 20% and fail. Co-Authored-By: Claude Opus 5.5 <[email protected]>
- --sigma is renamed --nsigma: it is a maximum z-score, not a sigma value (no alias, it was not in use yet). CI testsuites use --nsigma 5. - --pvalue P: the same threshold as a two-sided Gaussian p-value (0.0027 = 3 sigma, 5.7e-7 = 5 sigma), converted with statistics.NormalDist; cannot be combined with --nsigma. - Error bars are only used from >= 100 effective events, N_eff = (I/ERR)^2 (was ERR/I <= 25%, ~16 events), matching the ~100 event floor proposed for --statfactor; below that the 20% rule applies. Noisy tests like BNL_H8_simple (N_eff ~19 at 1e6) are thus judged by the 20% rule. Co-Authored-By: Claude Opus 5.5 <[email protected]>
--statfactor F multiplies the ncount of every test by F: the general ncount (-n, default 1e6) for single-value tests, and the per-point ncount of each scan (the scan line's -n, else at most 1e5). Run timeouts grow with F, and the factor is part of the result directory's name. With --nsigma/--pvalue, a target without NAME_ERR is assumed to come from the unscaled ncount, i.e. sigma_target = sigma_test * sqrt(F), so the combined error bar is sqrt(1 + F) * ERR_test (sqrt(2) for F = 1, as before). The test run's own ERR follows the scaled statistics by itself. mcviewtest applies the same rule from the stored statfactor. Checked on the McStas basic-test instruments + SE_example scans: F = 10 all pass (incl. BNL_H8_simple, which fails 20% at 1e6); F = 0.1 gives 5 failures that fall below 100 effective events and fail 20%. File-input tests (MCPL, KDSource) do not change with F, so for them only the assumed target error changes. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Explains for statistics-savvy readers how mctest's --nsigma, --pvalue and --statfactor judge test values against their targets: where I and ERR come from (checked against mcestimate_error), the decision table, the sqrt(1+f) assumption for unrecorded target errors, the N_eff >= 100 gate (Kish effective sample size), the n-sigma / p-value equivalence and the multiple-comparisons argument for 5 sigma in CI, known limitations (SPLIT-correlated weights, systematic target drift, file-input tests, MPI random streams, Gaussian approximation), and measured results on the CI basic tests at statfactor 0.1, 1 and 10. Linked from mctest.md. Co-Authored-By: Claude Opus 5.5 <[email protected]>
mctest.md:
- %Test:/%TestScan: are the documented test tags (older names %Example:
and %Scan: still accepted)
- the scan example uses the {}-enclosed target list the parser requires
(the old example without braces would not have been recognised)
- a value or scan point passes within 20%, or with --nsigma/--pvalue
within the statistical tolerance; --nsigma row gives sqrt(1+F) with
--statfactor F; --skipnontest/--noscans/--strict rows use the new tags
- mcviewtest: one row per test/scan, cells coloured by the rule mctest
used, per-point tooltip on scan cells
mcrun.md: --scan_split is ignored (with a warning) for an explicit
multi-process --mpi.
README.md: link to the statistical acceptance page.
The .html twins are regenerated from the .md files with python-markdown
(tables, toc, fenced_code) and the existing page head, which reproduces
the committed mcplot.html exactly; mctest.html and mcrun.html had not
followed the earlier %Scan:/--scan_split doc changes. Adds
mctest-statistics.html.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Converts tools/Python/NAME.md to NAME.html with python-markdown (tables, toc, fenced_code), keeping each page's existing <head>; mcdoc.md becomes mcdoc_tool.html, and links to .md files are rewritten to the .html names. With no arguments all tool docs are converted. Running it on the current tree reproduces all 9 committed .html pages byte for byte. Co-Authored-By: Claude Opus 5.5 <[email protected]>
mcstas-pygen turns test lines into McStasScript instr.add_test calls. It now accepts %Test: as well as %Example:, in file order when both are used. %TestScan: lines (parameter scans) have no add_test equivalent and are left alone; "%Test:" does not match them. Warnings, the generated comment, the --no-tests help text and codegen.md/.html name both tags. Checked: BNL_H8 with %Test: instead of %Example: generates identical code; a header mixing %Example:/%Test:/%TestScan: gives the three single-value tests in file order and the generated script compiles; no compiler warnings for pygen.c. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Free-form text area
Please describe what your PR is adding in terms of features or bugfixes:
mctest: judge test values by their Monte Carlo error bars;
%Test:/%TestScan:tagsWhy
mctest has always accepted a test value within a fixed ±20 % of its target. That is too strict for noisy, low-statistics tests and blind to how precise a run actually is. This PR lets mctest use the error bar every McStas/McXtrace run already prints, and lets a whole test run be scaled to fewer or more particles. Background: mccode-dev/mcstas-ess-instr-citool#4.
Nothing changes unless the new options are used, and no instrument files are touched. The full explanation, written for statistics-minded readers, is in
tools/Python/mctest-statistics.md.New mctest options
--nsigma n--pvalue p5.7e-7≡ 5σ)--statfactor F0.1for quick runs or100for reference runsHow a value is judged
With Δ = I_test − I_target:
_ERR(all instruments today)_ERRN_eff = (I/ERR)², the Kish effective sample size of the weighted particles. A run with only a handful of rays, a zero result, or a missing ERR falls back to ±20 %, so a huge error bar can never wave a bad value through._ERR, these options can only widen acceptance compared with ±20 %.ERR is read from the monitor file, from the stdout
Detector: …_ERR=line (NeXus), or from a scan'smccode.dat. A target may state its own error bar in the same format:NAME_ERR=…, orNAME_ERR={…}for scans.%Test:/%TestScan:tagsThese are the future names of
%Example:/%Scan:. Both spellings are accepted everywhere, so instruments can switch gradually:Test:/TestScan:.%Test:lines intoinstr.add_test, in file order with%Example:.%TestScan:lines have noadd_testequivalent and are left alone.mccode-git-diff-code: treats both as test definitions, including a scan's_ERR={…}block.%Test:(or%TestScan:) line instead of%Example:.Also in this PR
--nsigma 5. Basictests are unchanged.mctest-statistics.md;mctest.md,mcrun.md,codegen.mdand the overview brought up to date. The old%Scan:example lacked the{…}the parser requires, so it never worked as written.devel/bin/mccode-tooldocs-html: regenerates thetools/Python/*.htmldoc pages from their.mdsources (mcdoc.md→mcdoc_tool.html). It reproduces all existing pages byte for byte. It needs python-markdown, now a dependency.Choosing the threshold
A full suite judges a few thousand values once scan points are counted. At 3σ that gives several spurious failures per run; at 5σ, far below one. CI therefore uses 5.
Testing (macOS arm64,
mccode-devbuilt from this branch)--nsigma 5, MPI (McStas + McXtrace)--statfactor 10--statfactor 0.1-n100starved run_ERR(scratch copies)%Test:gives the same code as%Example:; mixed tags kept in file order;%TestScan:skipped; generated script compiles; no compiler warningsUnit checks cover every acceptance mode, mctest/mcviewtest agreement, parsing of all four tags, and the diff tool.
Known limitations
SPLITinstruments: ERR understates the real scatter there (2–3× measured on SSRL/LLB), so targets with_ERRon such instruments would fail spuriously.--statfactor; for them only the assumed target error changes._ERR: record it only together with a fresh, high-statistics target value. Some targets are a few per cent off for non-statistical reasons and would fail a pure error-bar test.--statfactorwas exercised on McStas instruments only.Declaration of use of AI-tools
^ 🤖 Generated with Claude Code
Development OS / boundary conditions
Please describe what OS you developed and tested your additions on, and if any special dependencies are required:
PR Checklist for contributing to McStas/McXtrace
For a coherent and useful contribution to McStas/McXtrace, please fill in relevant parts of the checklist:
My contribution contains something else