Python: benchmark harness¶
Note
This page documents python/oracles/ — the benchmark, verification, and
trust-ledger internals used to develop and audit spectrafit-core. It is
not part of the stable public API described in
Python: core API; it may change without notice between
releases. For the narrative tour of this package (what a contributor
touches to add a case or a model), see
Benchmark engine. For the
project's central methodological claim — that it verifies its numbers but
had never verified its own self-description — see
The self-auditing benchmark.
python/oracles/ is 107 modules and roughly 358 public symbols — an order of
magnitude larger than the spectrafit_core surface on the previous page. It
is grouped below by concern, not dumped as one flat symbol list, and each
group carries a short note on what the code actually does before the
generated API. show_source is disabled throughout: read the linked files
directly on GitLab/GitHub for implementation detail; this page is for
signatures and docstrings.
Verification wires¶
The numerical trust ledger. Each wire_w* function in oracles.audit.wires
computes one independently-checkable property of a benchmark run (W1–W11) and
returns a list of WireResult records — "pass", "warn", "fail",
"skipped", or "gap", never a silent pass-by-absence. oracles.audit.nist
is wire W8's evidence emitter: it re-runs 22 of NIST's 27 StRD nonlinear
regression datasets against spectrafit and returns a structured
NistValidation. See
The self-auditing benchmark
for how these wires roll up into a credibility rung, and
NIST StRD Validation /
NIST StRD reference for what W8 means and the full
per-dataset agreement table.
oracles.audit.wires
¶
One function per wire. Each returns a list of WireResult records.
W1/W3/W4/W6 read the pytest lastfailed cache for their status via a TRI-STATE
helper: an absent cache means the test never ran, so the wire reports "skipped"
(no pass-by-absence) rather than inflating the rung with an unbacked "pass".
W2a/W2b/W2c accept an optional audit_records parameter; when present they
RECOMPUTE from the full undecimated arrays in the audit.json sidecar instead of
relying on a stale cache entry. When audit_records is None they return
"skipped" so the rung is not inflated.
Note
_SUBJECT_BACKEND (used by W2c) names the subject under test in the
benchmark harness — the backend whose \(\kappa(J)\) capability W2c verifies.
Oracle backends (lmfit, jax) that do not expose \(\kappa\) are a disclosed
per-backend limitation, not a capability gap in the subject.
wire_w1_synth_invariants()
¶
Wire W1: hypothesis-driven invariants on the synthetic data generator.
wire_w2a_metric_identity(audit_records=None)
¶
W2a: recompute \(r^2\)/rmse/reduced-\(\chi^2\) and compare to stored.
The three metrics checked here correspond exactly to the three W2a-backed claims:
metric.r2, metric.rmse, and metric.red_chi2 (which also covers
metric.chi2 because \(\chi^2 = \chi^2_{\mathrm{red}} \times \mathrm{dof}\) — same formula, same check).
All three must agree with their stored counterparts within \(1 \times 10^{-6}\) for W2a to pass.
Note
This is a metric-IDENTITY check: each metric is recomputed under the SAME definition the engine stored it with. The engine stores reduced-\(\chi^2\) UNWEIGHTED (SSR/dof), so the identity recompute is unweighted too — passing the sidecar's \(\sigma\) here re-weighted it by \(1/\sigma^2\), which diverged ~\(1/\sigma^2\) on every noisy case and went to \(\infty\) for the noiseless (\(\sigma=0\)) optfn landscapes (the Track-0 root cause). The \(\sigma\)-weighted "true" reduced-\(\chi^2\) is a separate, additive metric — not this identity oracle.
wire_w2b_coverage(audit_records=None)
¶
W2b: recompute \(1\sigma\) pull coverage from the MC ensemble in the sidecar.
wire_w2c_jacobian_kappa(audit_records=None)
¶
W2c: verify Jacobian conditioning \(\kappa(J)\) per (case, backend).
Four outcomes (in priority order):
skipped— no audit sidecar present; the wire cannot recompute.fail— at least one \(\kappa\) WAS computed (by any backend) but is non-finite (inf/nan): a real numerical failure that caps the rung.pass— the subject (spectrafit) exposes a finite \(\kappa(J)\) for every one of its entries AND no backend has a non-finite computed \(\kappa\). Oracle backends (lmfit/jax) that do not expose \(\kappa\) at all are a disclosed per-backend limitation, not a capability gap — their absence is reported indetailsbut does NOT change the pass.gap— no non-finite \(\kappa\), AND the subject has no entries with a finite \(\kappa\) (subject does not expose \(\kappa\) either). This is a genuine subject capability gap.
wire_w2d_solver_output_oracle()
¶
Wire W2d: spectrafit's solver output matches an independent oracle.
Fitted parameters AND covariance agree with scipy.optimize.least_squares
(method='lm') within tolerance — the end-to-end solver-output check
(Invariant V, V3).
Backed by tests/audit/test_audit_solver_output_oracle.py; the value side
of the LM solve (params + per-parameter \(\sigma\) from the covariance) is verified
against a second, independent LM implementation on the same data.
Status is read from the pytest lastfailed cache (the shared pattern with
W1/W3/W4/W6): a "pass" means the audit suite ran and this test was not among
the failures. It therefore reports "pass" only when the suite actually ran;
a fresh checkout with no cache reports "skipped" (non-capping), never a
false pass from an absent cache.
wire_w3_results_roundtrip(current_run_id=None)
¶
Wire W3: results.json parses and re-emits to canonical form.
The roundtrip test is parametrized per run, so its lastfailed entries are
run-specific (…canonical_roundtrip[benchmark/<run>]). Scope to the CURRENT
run id so a stale failure for an OLDER run on disk cannot pin the current
run's trust red — the Track-0 ② conflation where run_031 passed but the wire
went red on runs 029/030. When no run id is supplied (legacy callers), fall
back to the file-level fragment.
wire_w4_api_schema()
¶
Wire W4: every /api response validates against the BenchReport contract.
wire_w5_render_fidelity()
¶
Wire W5: Playwright JSON-vs-render differential (CI-only, always skipped here).
W5 is a Playwright e2e test; controller-side it is "skipped" unless we invoke Playwright from here. Conservative: mark "skipped" so it does not inflate the rung — Task 15's poe task / CI job will provide ground truth.
wire_w6_gate_state_parity()
¶
Wire W6: gate_state is a closed Literal and manifest carries it.
wire_w7_inference_validity(*, has_inference_cases=None)
¶
W7: the seeded inference is reproducible — same seed → identical CI.
Evidence-conditional (no pass-by-absence, mirroring W10/W11): a run with no
non-baseline backend (e.g. a single-backend/self-baseline run — see
oracles.cli._resolve_run_backends) never populates inference.cases, so
there is nothing for this wire's claims (inference.speedup_ci,
inference.delta_r2_ci) to validate against. has_inference_cases=False
reports "skipped" instead of running the synthetic self-check below, so the
write-time claim-integrity guard (_assert_audited_claims_resolve) never
expects those claims to resolve for such a run. has_inference_cases=None (the
direct-call/test default) preserves prior behavior exactly — always runs the
self-check, for callers with no run context to report against.
wire_w8_nist_certified_validation(nist_validation=None)
¶
W8: independent external replication against NIST StRD certified values.
Re-runs every NIST StRD fit registered in oracles.audit.nist._RECIPES
(22 of the 27 datasets in the external collection) and compares the
recovered parameters to NIST's extended-precision certified values:
pass— every dataset reproduces the certified values to \(\geq\) the sig-fig threshold. This is the independent differential validation + external replication evidence that earns RUNG_5.fail— the harness ran but at least one dataset fell short of the threshold (a real recovery regression).skipped— the harness could not run (spectrafit_core unavailable, fixture import error), ORnist_validationwas explicitly passed asNone(e.g. run_audit computed it before the loop but the call failed).
nist_validation — when provided by the caller (e.g. run_audit
pre-computed it so the TrustBlock and W8 share a single object), the wire
derives its status from that block without a second run_nist_validation()
call. Divergence between the W8 WireResult and TrustBlock.nist_validation is
therefore impossible. Pass None (the default) to run the harness inline,
preserving direct-call / test back-compat.
wire_w9_nested_adequacy(nested_adequacy=None)
¶
W9: nested-model adequacy V&V — BIC recovers the true model order.
Verifies that model-selection criteria applied to a known generative case recover the true peak order m*:
pass—nested_adequacy.recovered_true_order_bicisTrue. BIC is the consistent model-order estimator; it governs the wire verdict. AIC and the LRT p-value are included in the evidence string so any AIC/BIC disagreement is visible without changing the verdict.fail— the V&V was present butrecovered_true_order_bicisFalse(BIC selected the wrong order).skipped—nested_adequacy is None; evidence absent ⇒ the rung must not be inflated (no pass-by-absence, mirror W8).
nested_adequacy — when provided by the caller (e.g. run_audit
pre-computed it from results.json), the wire derives its status from
that block without a second computation call. Pass None (the default)
to return skipped immediately (e.g. on a fresh checkout where the
V&V has not yet been run).
wire_w10_sigma_calibration(calibration=None)
¶
W10: \(\sigma\)-calibration quality — practical-equivalence gate on pull coverage.
Verifies that the subject's reported \(\sigma\) values are calibrated: the fraction of \((\theta_{\mathrm{est}} - \theta_{\mathrm{true}}) / \sigma_{\mathrm{est}}\) pulls that fall within \(\pm 1\sigma\) should lie close to the nominal 68.27 % coverage.
Gate criterion (A2 fix — CI-inclusion TOST for a proportion): passed iff the Clopper–Pearson CI of empirical coverage lies entirely within \([\mathrm{nominal} - \mathrm{margin}, \mathrm{nominal} + \mathrm{margin}]\) (default margin = 0.03 = \(\pm 3\) pp). This replaces the old point-null binomial gate that had extreme power at large n and rejected a practically-negligible 1.5 pp deviation.
Honest strict diagnostic retained
The strict point-null binomial p-value (H0: coverage = nominal) is always computed and reported in the evidence string so the dashboard can honestly show "coverage 0.668 vs 0.6827, strict binomial p < 0.001, slightly optimistic but within \(\pm 0.03\) equivalence band." Only the binary gate verdict changes; the diagnostic does not disappear.
pass—calibration.passed is True(CI within equivalence band).fail—calibration.passed is False(CI outside equivalence band).skipped—calibration is NoneORcalibration.skipped is True; the rung must not be inflated (no pass-by-absence).
calibration — when provided by the caller (e.g. run_audit
pre-computed it from results.json), the wire derives its status from
that block. Pass None (the default) to return skipped immediately.
wire_w11_speed_inference(speed_inference=None)
¶
W11: speed inference — bootstrap CI on geomean speedup vs baseline.
Verifies that the best-performing non-baseline backend per case is
significantly faster than the baseline solver (= spectrafit on the current
roster; see compute_inference I2 comment in inference_report.py for
the exact selection criterion): the bootstrap CI on the geometric mean of
per-case speedup ratios must exclude 1.0. Secondary diagnostics: sign test
and Wilcoxon signed-rank p-values.
pass—speed_inference.passed is True(CI excludes 1.0).fail—speed_inference.passed is False(CI includes 1.0).skipped—speed_inference is NoneORspeed_inference.skipped is True; the rung must not be inflated (no pass-by-absence).
speed_inference — when provided by the caller (e.g. run_audit
pre-computed it from results.json), the wire derives its status from
that block. Pass None (the default) to return skipped immediately.
oracles.audit.nist
¶
NIST StRD certified-value validation emitter (wire W8 evidence).
Re-runs 22 of NIST's 27 StRD nonlinear-regression datasets — Gauss1, Gauss2,
Gauss3, Lanczos1, Lanczos2, Lanczos3, BoxBOD, Misra1a, Misra1b, MGH17,
Bennett5, MGH09, Eckerle4, Roszman1, DanWood, Kirby2, Hahn1, Thurber, Rat42,
Rat43, Chwirut1, Chwirut2 — and returns a structured
:class:~oracles.trust_ledger.NistValidation. The original ten also have a
dedicated scenario test apiece under tests/scenario/nist_strd/; the
remaining twelve are exercised here and via kernel-correctness numpy-oracle
tests only. Each fit starts from NIST's published START2 guess, recovers the
parameters via spectrafit's LM solver, projects them back to the NIST
parameterization, and records the significant-figure agreement against the
certified values.
The fits are tiny (\(\leq\) 250 points, ~0.4 s total) and deterministic; this is the cheap, independent external-replication oracle that earns the honest RUNG_5.
The build/project recipes intentionally mirror the scenario tests' _build_graph
/ _project_to_nist helpers verbatim (same parameterization mapping, same
START2 guess) so this emitter and the green scenario tests assert the same fit.
Bennett5 note: Bennett5 is NIST "Higher" difficulty and may not converge to
the certified values from START2 via the LM solver. Its _NistRecipe entry
is included for completeness; if it does not pass the sig-fig threshold the
NistDataset.passed flag will be False for that entry. NistValidation.passed
is a strict all() over every recipe dataset with no exclusion in this
production path — a Bennett5 regression fails the overall validation and caps
the W8 wire (and therefore the RUNG_5 unlock) exactly like any other dataset
would. The unit test suite (tests/audit/test_nist_validation.py) separately
tracks a narrower _OPTIONAL_DATASETS subset for its own "mandatory datasets
pass" assertion, but that exclusion is local to the test and is never applied
to the value this module (or the W8 wire) actually returns.
MGH09 note: MGH09 is also NIST "Higher" difficulty (Kowalik–Osborne rational
function). It is included for kernel-correctness evidence (the MGH09_RATIONAL
kernel and parity oracle are verified), but LM-solver convergence to the certified
values is not guaranteed from either NIST start. Like Bennett5, it is in the test
suite's _OPTIONAL_DATASETS set, not excluded from this module's own
NistValidation.passed.
Threshold & denominator: NIST_SIGFIG_THRESHOLD (4.0) is the minimum
significant-figure agreement required for a dataset to "pass". The scenario
tests assert 1e-3 relative (~3 sig figs) on parameters; this emitter holds
to the stricter \(\ge\)4 sig figs (1e-4 relative) that the RSS/\(\chi^2\)
assertions use, which the actual fits clear by ~6 figures of headroom.
NIST_STRD_TOTAL (27) is the size of the external NIST StRD
nonlinear-regression universe (Lower/Average/Higher difficulty) — the
canonical denominator for "N of M datasets reproduced"
(https://www.itl.nist.gov/div898/strd/nls/nls_main.shtml). It lives here
(the validation source of truth) and is emitted on the contract so the UI
never hardcodes it.
run_nist_validation(threshold=NIST_SIGFIG_THRESHOLD)
¶
Fit the eight NIST StRD datasets and return their certified-value agreement.
Each fit is tiny and deterministic; the whole sweep runs in ~0.3 s. Returns a
:class:NistValidation whose passed is True iff every dataset recovers the
NIST certified values to \(\geq\) threshold significant figures.
Contract and trust ledger¶
oracles.bench_contract is the frozen BenchReport Pydantic contract — the
single source of truth for the JSON the benchmark engine emits and the
web/ React app consumes, camelCase on the wire. oracles.trust_ledger
defines the TrustBlock/WireResult/CredibilityRung types that carry wire
outcomes and NIST validation as an optional, additive slice of that contract.
oracles.bench_contract
¶
Frozen data contract for the benchmark report (the BENCH payload).
This module is the single source of truth for the JSON the benchmark engine
emits and the web/ React app consumes. Field names serialize to camelCase
via an alias generator, so the TypeScript types generated from this schema
(web/src/contract.ts) line up 1:1 with the React chart code.
Conventions
- Per-solver maps are
dict[str, X]keyed by solver id ("spectrafit"/"lmfit"/"jax"). - Peak parameters use the design's short keys:
a(amplitude),c(center),s(sigma) — this is the report contract, distinct from the solver's canonicalamplitude/center/sigmanames. - Convergence history carries a
history_sourcemap ("real"|"reconstructed") so the UI can honestly label proxy curves.
Schema version window
SUPPORTED_SCHEMA (below) is the single source of truth for which payload
versions a current build renders — the web SUPPORTED_SCHEMA set is pinned
to it by tests/unit/benchmark/test_supported_schema_window.py, which
fails if the two drift (the guard that closes the latent-red failure class
where the web gate silently lagged a schema bump). It is a window, not a
single value, on purpose: an older version stays accepted while its payloads
can still be migrated/rendered, so a transition can widen the window here
independently of bumping SCHEMA_VERSION — one authoritative place
instead of a hand-edited TS literal. Invariants enforced by the guard test:
SCHEMA_VERSION is in the window, and every windowed version has a
migration path to SCHEMA_VERSION in oracles.migrate.MIGRATIONS.
Emit with model.model_dump(by_alias=True). The TypeScript types the web app
consumes (web/src/contract.ts / web/src/openapi.gen.ts) are generated by
openapi-typescript from the OpenAPI schema the FastAPI app (:mod:oracles.api)
publishes for these same models — Pydantic is the single source of truth.
GateState = Literal['pass', 'warn', 'fail']
module-attribute
¶
Three-axis gate state. Wire format = lowercase to match spc-bench gate
--json output (cli.py lines 218-262). The web binding
(web/src/shell/EvidenceVerdict.tsx) currently uppercases all three states
on render — that's a presentation decision; the wire stays lowercase.
Note: pinned by the same pattern as the Rust-side ModelTypeStr::as_str()
parity test (model_type_as_str_matches_serde_wire_for_every_variant) —
tests/parity/test_canonical_wire_strings.py guards this wire string
against producer(Python)/consumer(TS) drift, the bug class Top-10 #7 was
raised to prevent.
GATE_STATES = ('pass', 'warn', 'fail')
module-attribute
¶
Exhaustive ordered roster — GATE_STATES[GATE_RANK[s]] == s for every
state. Used by tests + the worst-of aggregator.
GATE_RANK = {'pass': 0, 'warn': 1, 'fail': 2}
module-attribute
¶
Worst-of rank: max(rank) wins when aggregating 3-axis gate states into a
single headline. fail always wins, then warn, then pass.
KNOWN_SOLVER_IDS = frozenset({'spectrafit', 'lmfit', 'jax', 'scipy-ls-lm', 'scipy-ls-trf', 'scipy-ls-dogbox'})
module-attribute
¶
Closed roster of currently registered backend ids. New backends are a
deliberate one-line addition here PLUS a register_backend(...) entry in
the bench backends module — see Top-10 #2 (deferred). A new bench run with
an unknown id round-trips through the contract because SolverId = str,
but the parity test in tests/parity/test_canonical_wire_strings.py
flags the orphan so the maintainer adds it to the roster intentionally.
SolverMeta
pydantic-model
¶
Bases: _Contract
Solver legend entry (id, label, and theme color tokens).
Re-exported from oracles.bench_contract for BenchReport.solvers.
Show JSON schema:
{
"additionalProperties": false,
"description": "Solver legend entry (id, label, and theme color tokens).\n\nRe-exported from oracles.bench_contract for BenchReport.solvers.",
"properties": {
"id": {
"description": "Stable backend/solver key (e.g. `spectrafit`, `lmfit`, `scipy-ls-trf`) used to join this entry against per-backend results elsewhere in the report.",
"title": "Id",
"type": "string"
},
"label": {
"description": "Human-readable display name shown in legends and tables.",
"title": "Label",
"type": "string"
},
"color": {
"description": "Primary theme color token (CSS `var(--c-...)` reference, with an optional CSS fallback) used for this backend's main chart marks.",
"title": "Color",
"type": "string"
},
"soft": {
"description": "Muted/soft variant of `color` (CSS `var(--c-...-soft)` reference) used for secondary chart elements \u2014 bands, fills, and de-emphasized series \u2014 that must read as the same backend without competing with the primary marks.",
"title": "Soft",
"type": "string"
}
},
"required": [
"id",
"label",
"color",
"soft"
],
"title": "SolverMeta",
"type": "object"
}
Fields:
id
pydantic-field
¶
Stable backend/solver key (e.g. spectrafit, lmfit, scipy-ls-trf) used to join this entry against per-backend results elsewhere in the report.
label
pydantic-field
¶
Human-readable display name shown in legends and tables.
color
pydantic-field
¶
Primary theme color token (CSS var(--c-...) reference, with an optional CSS fallback) used for this backend's main chart marks.
soft
pydantic-field
¶
Muted/soft variant of color (CSS var(--c-...-soft) reference) used for secondary chart elements — bands, fills, and de-emphasized series — that must read as the same backend without competing with the primary marks.
TrustBlock
pydantic-model
¶
Bases: _Contract
Aggregate trust evidence attached to a BenchReport.
Show JSON schema:
{
"$defs": {
"CredibilityRung": {
"description": "V&V maturity ladder (inspired by ASME V&V credibility levels).\n\nThe live wire set earns RUNG_2..RUNG_4. RUNG_1 / RUNG_5 are reserved\nend-stops (see the per-member comments), part of the public contract and\npinned by an enum-value test \u2014 reserved, not dead.",
"enum": [
1,
2,
3,
4,
5
],
"title": "CredibilityRung",
"type": "integer"
},
"NistDataset": {
"additionalProperties": false,
"description": "Per-dataset NIST StRD certified-value validation result.",
"properties": {
"name": {
"description": "StRD problem name, e.g. 'Gauss1'.",
"title": "Name",
"type": "string"
},
"model": {
"description": "Human-readable model description.",
"title": "Model",
"type": "string"
},
"nParams": {
"title": "Nparams",
"type": "integer"
},
"params": {
"items": {
"$ref": "#/$defs/NistParam"
},
"title": "Params",
"type": "array"
},
"minSigFigs": {
"description": "Worst (minimum) per-parameter sig-fig agreement.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff min_sig_figs $\\ge$ the validation threshold.",
"title": "Passed",
"type": "boolean"
}
},
"required": [
"name",
"model",
"nParams",
"params",
"minSigFigs",
"passed"
],
"title": "NistDataset",
"type": "object"
},
"NistParam": {
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
},
"NistValidation": {
"additionalProperties": false,
"description": "Aggregate NIST StRD certified-value validation (the W8 evidence block).\n\nIndependent external replication against NIST's extended-precision certified\nvalues \u2014 the evidence RUNG_5 was reserved for. Additive on TrustBlock.",
"properties": {
"thresholdSigFigs": {
"description": "Required minimum significant-figure agreement.",
"title": "Thresholdsigfigs",
"type": "number"
},
"datasets": {
"items": {
"$ref": "#/$defs/NistDataset"
},
"title": "Datasets",
"type": "array"
},
"minSigFigs": {
"description": "Worst min_sig_figs across all datasets.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff every dataset agrees to $\\ge$ threshold sig figs.",
"title": "Passed",
"type": "boolean"
},
"totalAvailable": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Size of the external NIST StRD nonlinear-regression universe (the denominator for 'N of M' coverage). Emitted by the validation builder so the UI never hardcodes the total. Additive \u2014 None for payloads written before this field existed.",
"title": "Totalavailable"
}
},
"required": [
"thresholdSigFigs",
"datasets",
"minSigFigs",
"passed"
],
"title": "NistValidation",
"type": "object"
},
"WireResult": {
"additionalProperties": false,
"description": "One verification wire's outcome.\n\nNote:\n ``status=\"gap\"`` is distinct from ``\"fail\"``: it means the property\n could not be verified because the backend does NOT expose the input\n (a disclosed CAPABILITY gap, e.g. $\\kappa(J)$ is not surfaced by\n spectrafit/lmfit/jax), not because a computed value was wrong. Like\n ``\"skipped\"``, a ``\"gap\"`` does not cap the credibility rung; only a\n genuine ``\"fail\"`` does (see ``runner._compute_rung``).",
"properties": {
"wireId": {
"description": "Audit wire id: W1, W2a-W2d, or W3..W11 from the value-stream diagram.",
"title": "Wireid",
"type": "string"
},
"name": {
"description": "Short machine-readable wire name (e.g. `synth_invariants`).",
"title": "Name",
"type": "string"
},
"status": {
"description": "Wire outcome: `pass`, `warn`, or `fail` for a computed verdict; `skipped` when the underlying test never ran; `gap` for a disclosed capability gap (see the class Note).",
"enum": [
"pass",
"warn",
"fail",
"skipped",
"gap"
],
"title": "Status",
"type": "string"
},
"evidence": {
"description": "One-line evidence statement.",
"title": "Evidence",
"type": "string"
},
"details": {
"additionalProperties": {
"anyOf": [
{
"type": "number"
},
{
"type": "integer"
},
{
"type": "string"
},
{
"type": "boolean"
},
{
"type": "null"
}
]
},
"description": "Optional structured detail (counts, thresholds, sample sizes, ...) supporting `evidence`; empty when the one-line statement is self-sufficient.",
"title": "Details",
"type": "object"
}
},
"required": [
"wireId",
"name",
"status",
"evidence"
],
"title": "WireResult",
"type": "object"
}
},
"additionalProperties": false,
"description": "Aggregate trust evidence attached to a BenchReport.",
"properties": {
"rung": {
"$ref": "#/$defs/CredibilityRung",
"description": "Aggregate V&V credibility rung (see `CredibilityRung`), capped by the worst genuine `fail` among `wires` \u2014 a `gap` or `skipped` wire never lowers it."
},
"wires": {
"description": "Every verification wire's outcome making up this ledger.",
"items": {
"$ref": "#/$defs/WireResult"
},
"title": "Wires",
"type": "array"
},
"nClaimsAudited": {
"description": "Number of report claims this audit run actually checked.",
"title": "Nclaimsaudited",
"type": "integer"
},
"nClaimsTotal": {
"description": "Total number of claims the report makes, whether or not they were audited this run.",
"title": "Nclaimstotal",
"type": "integer"
},
"nistValidation": {
"anyOf": [
{
"$ref": "#/$defs/NistValidation"
},
{
"type": "null"
}
],
"default": null,
"description": "NIST StRD certified-value validation (W8). Additive \u2014 Pydantic fills None for payloads that predate the A7 external-validation wire."
}
},
"required": [
"rung",
"wires",
"nClaimsAudited",
"nClaimsTotal"
],
"title": "TrustBlock",
"type": "object"
}
Fields:
-
rung(CredibilityRung) -
wires(list[WireResult]) -
n_claims_audited(int) -
n_claims_total(int) -
nist_validation(NistValidation | None)
Validators:
-
_rung5_requires_nist_and_inference
rung
pydantic-field
¶
Aggregate V&V credibility rung (see CredibilityRung), capped by the worst genuine fail among wires — a gap or skipped wire never lowers it.
wires
pydantic-field
¶
Every verification wire's outcome making up this ledger.
n_claims_audited
pydantic-field
¶
Number of report claims this audit run actually checked.
n_claims_total
pydantic-field
¶
Total number of claims the report makes, whether or not they were audited this run.
nist_validation = None
pydantic-field
¶
NIST StRD certified-value validation (W8). Additive — Pydantic fills None for payloads that predate the A7 external-validation wire.
ExprEdge
pydantic-model
¶
Bases: _Base
One tied/shared-parameter expression edge.
Wire form (camelCase): {"targetNode": "p1", "targetParam": "sigma",
"expression": "p0.sigma"}. Typed (not a bare dict[str, str]) so the
inner keys camelize like every other contract field — the web reads
edge.targetNode / edge.targetParam. Accepts the engine's snake_case
dicts via populate_by_name on :class:_Base.
Show JSON schema:
{
"additionalProperties": false,
"description": "One tied/shared-parameter expression edge.\n\nWire form (camelCase): ``{\"targetNode\": \"p1\", \"targetParam\": \"sigma\",\n\"expression\": \"p0.sigma\"}``. Typed (not a bare ``dict[str, str]``) so the\ninner keys camelize like every other contract field \u2014 the web reads\n``edge.targetNode`` / ``edge.targetParam``. Accepts the engine's snake_case\ndicts via ``populate_by_name`` on :class:`_Base`.",
"properties": {
"targetNode": {
"title": "Targetnode",
"type": "string"
},
"targetParam": {
"title": "Targetparam",
"type": "string"
},
"expression": {
"title": "Expression",
"type": "string"
}
},
"required": [
"targetNode",
"targetParam",
"expression"
],
"title": "ExprEdge",
"type": "object"
}
Fields:
-
target_node(str) -
target_param(str) -
expression(str)
Point2D
pydantic-model
¶
Bases: _Base
A single (x, y) point used by line / ecdf / scaling / warmup series.
Show JSON schema:
Fields:
-
x(float) -
y(float)
PeakACS
pydantic-model
¶
Bases: _Base
Peak parameters in the report's short form: amplitude / center / sigma.
Show JSON schema:
{
"additionalProperties": false,
"description": "Peak parameters in the report's short form: amplitude / center / sigma.",
"properties": {
"a": {
"title": "A",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
},
"s": {
"title": "S",
"type": "number"
}
},
"required": [
"a",
"c",
"s"
],
"title": "PeakACS",
"type": "object"
}
Fields:
-
a(float) -
c(float) -
s(float)
CategoryMeta
pydantic-model
¶
Bases: _Base
Benchmark category definition (id, label, case count, hue token).
Show JSON schema:
{
"additionalProperties": false,
"description": "Benchmark category definition (id, label, case count, hue token).",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"label": {
"title": "Label",
"type": "string"
},
"n": {
"title": "N",
"type": "integer"
},
"hue": {
"title": "Hue",
"type": "string"
}
},
"required": [
"id",
"label",
"n",
"hue"
],
"title": "CategoryMeta",
"type": "object"
}
Fields:
-
id(str) -
label(str) -
n(int) -
hue(str)
SolverFit
pydantic-model
¶
Bases: _Base
One solver's fit of the featured case: params + curve + residuals.
Show JSON schema:
{
"$defs": {
"PeakACS": {
"additionalProperties": false,
"description": "Peak parameters in the report's short form: amplitude / center / sigma.",
"properties": {
"a": {
"title": "A",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
},
"s": {
"title": "S",
"type": "number"
}
},
"required": [
"a",
"c",
"s"
],
"title": "PeakACS",
"type": "object"
}
},
"additionalProperties": false,
"description": "One solver's fit of the featured case: params + curve + residuals.",
"properties": {
"params": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Params",
"type": "array"
},
"curve": {
"items": {
"type": "number"
},
"title": "Curve",
"type": "array"
},
"resid": {
"items": {
"type": "number"
},
"title": "Resid",
"type": "array"
}
},
"required": [
"params",
"curve",
"resid"
],
"title": "SolverFit",
"type": "object"
}
Fields:
-
params(list[PeakACS]) -
curve(list[float]) -
resid(list[float])
Peak
pydantic-model
¶
Bases: _Base
A single peak contribution curve (featured solver), labeled g1/g2/….
Show JSON schema:
{
"additionalProperties": false,
"description": "A single peak contribution curve (featured solver), labeled g1/g2/\u2026.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"y": {
"items": {
"type": "number"
},
"title": "Y",
"type": "array"
}
},
"required": [
"label",
"y"
],
"title": "Peak",
"type": "object"
}
Fields:
-
label(str) -
y(list[float])
TimingDist
pydantic-model
¶
Bases: _Base
Per-solver runtime distribution over timing repetitions (milliseconds).
Show JSON schema:
{
"additionalProperties": false,
"description": "Per-solver runtime distribution over timing repetitions (milliseconds).",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"mean": {
"title": "Mean",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
},
"p95": {
"title": "P95",
"type": "number"
},
"iqr": {
"title": "Iqr",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
}
},
"required": [
"raw",
"median",
"mean",
"p5",
"p25",
"p75",
"p95",
"iqr",
"cv"
],
"title": "TimingDist",
"type": "object"
}
Fields:
-
raw(list[float]) -
median(float) -
mean(float) -
p5(float) -
p25(float) -
p75(float) -
p95(float) -
iqr(float) -
cv(float)
AccuracyDist
pydantic-model
¶
Bases: _Base
Per-solver reduced-\(\chi^2\) distribution over repetitions.
Show JSON schema:
{
"additionalProperties": false,
"description": "Per-solver reduced-$\\chi^2$ distribution over repetitions.",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
}
},
"required": [
"raw",
"median",
"p5",
"p25",
"p75"
],
"title": "AccuracyDist",
"type": "object"
}
Fields:
-
raw(list[float]) -
median(float) -
p5(float) -
p25(float) -
p75(float)
Summary
pydantic-model
¶
Bases: _Base
Headline KPI block per solver for the featured case.
Show JSON schema:
{
"$defs": {
"CI": {
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
}
},
"additionalProperties": false,
"description": "Headline KPI block per solver for the featured case.",
"properties": {
"r2": {
"title": "R2",
"type": "number"
},
"chi2": {
"title": "Chi2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"rmse": {
"title": "Rmse",
"type": "number"
},
"mae": {
"title": "Mae",
"type": "number"
},
"nIter": {
"title": "Niter",
"type": "integer"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"iqrMs": {
"title": "Iqrms",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
},
"speedup": {
"title": "Speedup",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"aic": {
"title": "Aic",
"type": "number"
},
"bic": {
"title": "Bic",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
},
"speedupCi": {
"anyOf": [
{
"$ref": "#/$defs/CI"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"r2",
"chi2",
"redChi2",
"rmse",
"mae",
"nIter",
"medMs",
"iqrMs",
"cv",
"speedup",
"success",
"aic",
"bic",
"dAIC",
"dBIC"
],
"title": "Summary",
"type": "object"
}
Fields:
-
r2(float) -
chi2(float) -
red_chi2(float) -
rmse(float) -
mae(float) -
n_iter(int) -
med_ms(float) -
iqr_ms(float) -
cv(float) -
speedup(float) -
success(bool) -
aic(float) -
bic(float) -
d_aic(float) -
d_bic(float) -
speedup_ci(CI | None)
WarmupPt
pydantic-model
¶
Bases: _Base
One cold/hot amortization point: per-run ms at a cumulative run count.
Show JSON schema:
{
"additionalProperties": false,
"description": "One cold/hot amortization point: per-run ms at a cumulative run count.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"perRun": {
"title": "Perrun",
"type": "number"
}
},
"required": [
"n",
"perRun"
],
"title": "WarmupPt",
"type": "object"
}
Fields:
-
n(int) -
per_run(float)
Warmup
pydantic-model
¶
Bases: _Base
Cold/hot amortization for one solver (JAX compile cost amortizing 1→100).
Show JSON schema:
{
"$defs": {
"Point2D": {
"additionalProperties": false,
"description": "A single (x, y) point used by line / ecdf / scaling / warmup series.",
"properties": {
"x": {
"title": "X",
"type": "number"
},
"y": {
"title": "Y",
"type": "number"
}
},
"required": [
"x",
"y"
],
"title": "Point2D",
"type": "object"
},
"WarmupPt": {
"additionalProperties": false,
"description": "One cold/hot amortization point: per-run ms at a cumulative run count.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"perRun": {
"title": "Perrun",
"type": "number"
}
},
"required": [
"n",
"perRun"
],
"title": "WarmupPt",
"type": "object"
}
},
"additionalProperties": false,
"description": "Cold/hot amortization for one solver (JAX compile cost amortizing 1\u2192100).",
"properties": {
"curve": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Curve",
"type": "array"
},
"pts": {
"items": {
"$ref": "#/$defs/WarmupPt"
},
"title": "Pts",
"type": "array"
},
"hotThroughput": {
"title": "Hotthroughput",
"type": "number"
},
"coldMs": {
"title": "Coldms",
"type": "number"
},
"hotMs": {
"title": "Hotms",
"type": "number"
}
},
"required": [
"curve",
"pts",
"hotThroughput",
"coldMs",
"hotMs"
],
"title": "Warmup",
"type": "object"
}
Fields:
Uncertainty
pydantic-model
¶
Bases: _Base
Parameter-pull diagnostics: pulls, coverage, reported \(\sigma\).
Pulls are \((\mathrm{est}-\mathrm{truth})/\sigma\). coverage is None when no
valid \(\sigma\) was available for any MC sample (e.g. backends like jax that set
every param_stderr entry to None).
A None coverage is distinct from a genuine 0.0 (all pulls \(|p| \ge 1\sigma\)).
Show JSON schema:
{
"additionalProperties": false,
"description": "Parameter-pull diagnostics: pulls, coverage, reported $\\sigma$.\n\nPulls are $(\\mathrm{est}-\\mathrm{truth})/\\sigma$. ``coverage`` is ``None`` when no\nvalid $\\sigma$ was available for any MC sample (e.g. backends like jax that set\nevery ``param_stderr`` entry to ``None``).\nA ``None`` coverage is distinct from a genuine ``0.0`` (all pulls $|p| \\ge 1\\sigma$).",
"properties": {
"pulls": {
"items": {
"type": "number"
},
"title": "Pulls",
"type": "array"
},
"coverage": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"title": "Coverage"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"pulls",
"coverage",
"sigma"
],
"title": "Uncertainty",
"type": "object"
}
Fields:
-
pulls(list[float]) -
coverage(float | None) -
sigma(list[float])
SpreadPt
pydantic-model
¶
Bases: _Base
A mean \(\pm\) sd point versus run count (param-spread / stability series).
Show JSON schema:
{
"additionalProperties": false,
"description": "A mean $\\pm$ sd point versus run count (param-spread / stability series).",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"mean": {
"title": "Mean",
"type": "number"
},
"sd": {
"title": "Sd",
"type": "number"
}
},
"required": [
"n",
"mean",
"sd"
],
"title": "SpreadPt",
"type": "object"
}
Fields:
-
n(int) -
mean(float) -
sd(float)
StabilityEntry
pydantic-model
¶
Bases: _Base
One backend's metric stability versus run count (mean \(\pm\) sd per metric).
Show JSON schema:
{
"$defs": {
"SpreadPt": {
"additionalProperties": false,
"description": "A mean $\\pm$ sd point versus run count (param-spread / stability series).",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"mean": {
"title": "Mean",
"type": "number"
},
"sd": {
"title": "Sd",
"type": "number"
}
},
"required": [
"n",
"mean",
"sd"
],
"title": "SpreadPt",
"type": "object"
}
},
"additionalProperties": false,
"description": "One backend's metric stability versus run count (mean $\\pm$ sd per metric).",
"properties": {
"r2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "R2",
"type": "array"
},
"rmse": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Rmse",
"type": "array"
},
"redChi2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Redchi2",
"type": "array"
},
"iters": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Iters",
"type": "array"
}
},
"required": [
"r2",
"rmse",
"redChi2",
"iters"
],
"title": "StabilityEntry",
"type": "object"
}
Fields:
NdPeak
pydantic-model
¶
Bases: _Base
A recovered N-D Gaussian peak: amplitude + per-axis center and sigma.
center and sigma each have n_dims entries (indexed center_0.. /
sigma_0.. from the parametric gaussian_nd kernel).
Show JSON schema:
{
"additionalProperties": false,
"description": "A recovered N-D Gaussian peak: amplitude + per-axis center and sigma.\n\n``center`` and ``sigma`` each have ``n_dims`` entries (indexed `center_0..` /\n`sigma_0..` from the parametric ``gaussian_nd`` kernel).",
"properties": {
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"items": {
"type": "number"
},
"title": "Center",
"type": "array"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "NdPeak",
"type": "object"
}
Fields:
-
amplitude(float) -
center(list[float]) -
sigma(list[float])
Projection
pydantic-model
¶
Bases: _Base
A 2-D projection (marginal slice) of a higher-D fit: axis-pair + matrix.
Show JSON schema:
{
"additionalProperties": false,
"description": "A 2-D projection (marginal slice) of a higher-D fit: axis-pair + matrix.",
"properties": {
"labels": {
"maxItems": 2,
"minItems": 2,
"prefixItems": [
{
"type": "string"
},
{
"type": "string"
}
],
"title": "Labels",
"type": "array"
},
"matrix": {
"items": {
"items": {
"type": "number"
},
"type": "array"
},
"title": "Matrix",
"type": "array"
}
},
"required": [
"labels",
"matrix"
],
"title": "Projection",
"type": "object"
}
Fields:
-
labels(tuple[str, str]) -
matrix(list[list[float]])
MultiDim
pydantic-model
¶
Bases: _Base
Genuinely N-dimensional (\(\ge\)3-D) fit showcase.
A synthetic N-D Gaussian is recovered by a real least-squares solve with the
parametric gaussian_nd kernel (D inferred from the node's indexed
center_<i> parameters). peaks are the fitted parameters and
r_squared the recovery quality. Full N-D obs/model/residual grids are
intentionally NOT stored — they do not scale past 2-D — so projections
carries 2-D marginal slices of the fitted model for any visualization.
SYNTHETIC — no experimental data. Rendered in the production UI as
part of Evidence's "Native showcases" section (multidim-showcase panel,
web/src/panels/registry.tsx).
Show JSON schema:
{
"$defs": {
"NdPeak": {
"additionalProperties": false,
"description": "A recovered N-D Gaussian peak: amplitude + per-axis center and sigma.\n\n``center`` and ``sigma`` each have ``n_dims`` entries (indexed `center_0..` /\n`sigma_0..` from the parametric ``gaussian_nd`` kernel).",
"properties": {
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"items": {
"type": "number"
},
"title": "Center",
"type": "array"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "NdPeak",
"type": "object"
},
"Projection": {
"additionalProperties": false,
"description": "A 2-D projection (marginal slice) of a higher-D fit: axis-pair + matrix.",
"properties": {
"labels": {
"maxItems": 2,
"minItems": 2,
"prefixItems": [
{
"type": "string"
},
{
"type": "string"
}
],
"title": "Labels",
"type": "array"
},
"matrix": {
"items": {
"items": {
"type": "number"
},
"type": "array"
},
"title": "Matrix",
"type": "array"
}
},
"required": [
"labels",
"matrix"
],
"title": "Projection",
"type": "object"
}
},
"additionalProperties": false,
"description": "Genuinely N-dimensional ($\\ge$3-D) fit showcase.\n\nA synthetic N-D Gaussian is recovered by a real least-squares solve with the\nparametric ``gaussian_nd`` kernel (D inferred from the node's indexed\n``center_<i>`` parameters). ``peaks`` are the *fitted* parameters and\n``r_squared`` the recovery quality. Full N-D obs/model/residual grids are\nintentionally NOT stored \u2014 they do not scale past 2-D \u2014 so ``projections``\ncarries 2-D marginal slices of the fitted model for any visualization.\n\n**SYNTHETIC \u2014 no experimental data.** Rendered in the production UI as\npart of Evidence's \"Native showcases\" section (``multidim-showcase`` panel,\n``web/src/panels/registry.tsx``).",
"properties": {
"nDims": {
"title": "Ndims",
"type": "integer"
},
"shape": {
"items": {
"type": "integer"
},
"title": "Shape",
"type": "array"
},
"nPoints": {
"title": "Npoints",
"type": "integer"
},
"rSquared": {
"title": "Rsquared",
"type": "number"
},
"peaks": {
"items": {
"$ref": "#/$defs/NdPeak"
},
"title": "Peaks",
"type": "array"
},
"projections": {
"items": {
"$ref": "#/$defs/Projection"
},
"title": "Projections",
"type": "array"
},
"source": {
"default": "spectrafit-core",
"title": "Source",
"type": "string"
},
"dataProvenance": {
"default": "synthetic",
"enum": [
"measured",
"synthetic"
],
"title": "Dataprovenance",
"type": "string"
}
},
"required": [
"nDims",
"shape",
"nPoints",
"rSquared",
"peaks",
"projections"
],
"title": "MultiDim",
"type": "object"
}
Fields:
-
n_dims(int) -
shape(list[int]) -
n_points(int) -
r_squared(float) -
peaks(list[NdPeak]) -
projections(list[Projection]) -
source(str) -
data_provenance(Literal['measured', 'synthetic'])
GlobalFitSlice
pydantic-model
¶
Bases: _Base
Dataset slice in a global fit with jointly-fitted shared model.
The coord field carries the axis coordinate for this slice (e.g. a time
value, a temperature, a sample index) — the name is axis-neutral because the
global-fit machinery is not specific to time.
Show JSON schema:
{
"additionalProperties": false,
"description": "Dataset slice in a global fit with jointly-fitted shared model.\n\nThe ``coord`` field carries the axis coordinate for this slice (e.g. a time\nvalue, a temperature, a sample index) \u2014 the name is axis-neutral because the\nglobal-fit machinery is not specific to time.",
"properties": {
"coord": {
"title": "Coord",
"type": "number"
},
"obs": {
"items": {
"type": "number"
},
"title": "Obs",
"type": "array"
},
"model": {
"items": {
"type": "number"
},
"title": "Model",
"type": "array"
}
},
"required": [
"coord",
"obs",
"model"
],
"title": "GlobalFitSlice",
"type": "object"
}
Fields:
-
coord(float) -
obs(list[float]) -
model(list[float])
PeakTrace
pydantic-model
¶
Bases: _Base
A shared peak's amplitude evolving across the dataset axis.
Show JSON schema:
{
"additionalProperties": false,
"description": "A shared peak's amplitude evolving across the dataset axis.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"amplitude": {
"items": {
"type": "number"
},
"title": "Amplitude",
"type": "array"
}
},
"required": [
"label",
"center",
"sigma",
"amplitude"
],
"title": "PeakTrace",
"type": "object"
}
Fields:
-
label(str) -
center(float) -
sigma(float) -
amplitude(list[float])
GlobalFit
pydantic-model
¶
Bases: _Base
Shared-model multi-spectrum global fit (lmfit pattern).
SYNTHETIC placeholder — no experimental data. Uses a GlobalFitGraph
fit over synthetic Gaussian peaks. Rendered in the production UI as part of
Evidence's "Native showcases" section (global-fit-showcase panel,
web/src/panels/registry.tsx). Peak centers/widths are shared across every
dataset slice (global analysis) while each slice's amplitudes vary; traces
are the recovered per-peak amplitude traces and slices carry the per-dataset
observed + jointly-fitted curves on the shared x grid.
The dataset_axis / axis_label names are deliberately axis-neutral: the
global-fit machinery in spectrafit-core (GlobalFitGraph) is not specific to
time — here the axis incidentally represents time, but the contract does not
claim that.
Show JSON schema:
{
"$defs": {
"GlobalFitSlice": {
"additionalProperties": false,
"description": "Dataset slice in a global fit with jointly-fitted shared model.\n\nThe ``coord`` field carries the axis coordinate for this slice (e.g. a time\nvalue, a temperature, a sample index) \u2014 the name is axis-neutral because the\nglobal-fit machinery is not specific to time.",
"properties": {
"coord": {
"title": "Coord",
"type": "number"
},
"obs": {
"items": {
"type": "number"
},
"title": "Obs",
"type": "array"
},
"model": {
"items": {
"type": "number"
},
"title": "Model",
"type": "array"
}
},
"required": [
"coord",
"obs",
"model"
],
"title": "GlobalFitSlice",
"type": "object"
},
"PeakTrace": {
"additionalProperties": false,
"description": "A shared peak's amplitude evolving across the dataset axis.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"amplitude": {
"items": {
"type": "number"
},
"title": "Amplitude",
"type": "array"
}
},
"required": [
"label",
"center",
"sigma",
"amplitude"
],
"title": "PeakTrace",
"type": "object"
}
},
"additionalProperties": false,
"description": "Shared-model multi-spectrum global fit (lmfit pattern).\n\n**SYNTHETIC placeholder \u2014 no experimental data.** Uses a ``GlobalFitGraph``\nfit over synthetic Gaussian peaks. Rendered in the production UI as part of\nEvidence's \"Native showcases\" section (``global-fit-showcase`` panel,\n``web/src/panels/registry.tsx``). Peak centers/widths are shared across every\ndataset slice (global analysis) while each slice's amplitudes vary; ``traces``\nare the recovered per-peak amplitude traces and ``slices`` carry the per-dataset\nobserved + jointly-fitted curves on the shared ``x`` grid.\n\nThe ``dataset_axis`` / ``axis_label`` names are deliberately axis-neutral: the\nglobal-fit machinery in spectrafit-core (``GlobalFitGraph``) is not specific to\ntime \u2014 here the axis incidentally represents time, but the contract does not\nclaim that.",
"properties": {
"x": {
"items": {
"type": "number"
},
"title": "X",
"type": "array"
},
"datasetAxis": {
"items": {
"type": "number"
},
"title": "Datasetaxis",
"type": "array"
},
"slices": {
"items": {
"$ref": "#/$defs/GlobalFitSlice"
},
"title": "Slices",
"type": "array"
},
"traces": {
"items": {
"$ref": "#/$defs/PeakTrace"
},
"title": "Traces",
"type": "array"
},
"source": {
"default": "spectrafit-core",
"title": "Source",
"type": "string"
},
"xLabel": {
"default": "x",
"title": "Xlabel",
"type": "string"
},
"axisLabel": {
"default": "dataset",
"title": "Axislabel",
"type": "string"
},
"dataProvenance": {
"default": "synthetic",
"enum": [
"measured",
"synthetic"
],
"title": "Dataprovenance",
"type": "string"
}
},
"required": [
"x",
"datasetAxis",
"slices",
"traces"
],
"title": "GlobalFit",
"type": "object"
}
Fields:
-
x(list[float]) -
dataset_axis(list[float]) -
slices(list[GlobalFitSlice]) -
traces(list[PeakTrace]) -
source(str) -
x_label(str) -
axis_label(str) -
data_provenance(Literal['measured', 'synthetic'])
BackendProfile
pydantic-model
¶
Bases: _Base
All of ONE backend's featured-case metrics, grouped (no parallel dicts).
Element types match the per-panel charts exactly: scaling/ecdf_* keep the
(x, y) Point2D axes, param_err is per-parameter, param_spread/stability
keep the mean \(\pm\) sd SpreadPt bands.
Show JSON schema:
{
"$defs": {
"AccuracyDist": {
"additionalProperties": false,
"description": "Per-solver reduced-$\\chi^2$ distribution over repetitions.",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
}
},
"required": [
"raw",
"median",
"p5",
"p25",
"p75"
],
"title": "AccuracyDist",
"type": "object"
},
"CI": {
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
},
"PeakACS": {
"additionalProperties": false,
"description": "Peak parameters in the report's short form: amplitude / center / sigma.",
"properties": {
"a": {
"title": "A",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
},
"s": {
"title": "S",
"type": "number"
}
},
"required": [
"a",
"c",
"s"
],
"title": "PeakACS",
"type": "object"
},
"Point2D": {
"additionalProperties": false,
"description": "A single (x, y) point used by line / ecdf / scaling / warmup series.",
"properties": {
"x": {
"title": "X",
"type": "number"
},
"y": {
"title": "Y",
"type": "number"
}
},
"required": [
"x",
"y"
],
"title": "Point2D",
"type": "object"
},
"SolverFit": {
"additionalProperties": false,
"description": "One solver's fit of the featured case: params + curve + residuals.",
"properties": {
"params": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Params",
"type": "array"
},
"curve": {
"items": {
"type": "number"
},
"title": "Curve",
"type": "array"
},
"resid": {
"items": {
"type": "number"
},
"title": "Resid",
"type": "array"
}
},
"required": [
"params",
"curve",
"resid"
],
"title": "SolverFit",
"type": "object"
},
"SpreadPt": {
"additionalProperties": false,
"description": "A mean $\\pm$ sd point versus run count (param-spread / stability series).",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"mean": {
"title": "Mean",
"type": "number"
},
"sd": {
"title": "Sd",
"type": "number"
}
},
"required": [
"n",
"mean",
"sd"
],
"title": "SpreadPt",
"type": "object"
},
"StabilityEntry": {
"additionalProperties": false,
"description": "One backend's metric stability versus run count (mean $\\pm$ sd per metric).",
"properties": {
"r2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "R2",
"type": "array"
},
"rmse": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Rmse",
"type": "array"
},
"redChi2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Redchi2",
"type": "array"
},
"iters": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Iters",
"type": "array"
}
},
"required": [
"r2",
"rmse",
"redChi2",
"iters"
],
"title": "StabilityEntry",
"type": "object"
},
"Summary": {
"additionalProperties": false,
"description": "Headline KPI block per solver for the featured case.",
"properties": {
"r2": {
"title": "R2",
"type": "number"
},
"chi2": {
"title": "Chi2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"rmse": {
"title": "Rmse",
"type": "number"
},
"mae": {
"title": "Mae",
"type": "number"
},
"nIter": {
"title": "Niter",
"type": "integer"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"iqrMs": {
"title": "Iqrms",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
},
"speedup": {
"title": "Speedup",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"aic": {
"title": "Aic",
"type": "number"
},
"bic": {
"title": "Bic",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
},
"speedupCi": {
"anyOf": [
{
"$ref": "#/$defs/CI"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"r2",
"chi2",
"redChi2",
"rmse",
"mae",
"nIter",
"medMs",
"iqrMs",
"cv",
"speedup",
"success",
"aic",
"bic",
"dAIC",
"dBIC"
],
"title": "Summary",
"type": "object"
},
"TimingDist": {
"additionalProperties": false,
"description": "Per-solver runtime distribution over timing repetitions (milliseconds).",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"mean": {
"title": "Mean",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
},
"p95": {
"title": "P95",
"type": "number"
},
"iqr": {
"title": "Iqr",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
}
},
"required": [
"raw",
"median",
"mean",
"p5",
"p25",
"p75",
"p95",
"iqr",
"cv"
],
"title": "TimingDist",
"type": "object"
},
"Uncertainty": {
"additionalProperties": false,
"description": "Parameter-pull diagnostics: pulls, coverage, reported $\\sigma$.\n\nPulls are $(\\mathrm{est}-\\mathrm{truth})/\\sigma$. ``coverage`` is ``None`` when no\nvalid $\\sigma$ was available for any MC sample (e.g. backends like jax that set\nevery ``param_stderr`` entry to ``None``).\nA ``None`` coverage is distinct from a genuine ``0.0`` (all pulls $|p| \\ge 1\\sigma$).",
"properties": {
"pulls": {
"items": {
"type": "number"
},
"title": "Pulls",
"type": "array"
},
"coverage": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"title": "Coverage"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"pulls",
"coverage",
"sigma"
],
"title": "Uncertainty",
"type": "object"
},
"Warmup": {
"additionalProperties": false,
"description": "Cold/hot amortization for one solver (JAX compile cost amortizing 1\u2192100).",
"properties": {
"curve": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Curve",
"type": "array"
},
"pts": {
"items": {
"$ref": "#/$defs/WarmupPt"
},
"title": "Pts",
"type": "array"
},
"hotThroughput": {
"title": "Hotthroughput",
"type": "number"
},
"coldMs": {
"title": "Coldms",
"type": "number"
},
"hotMs": {
"title": "Hotms",
"type": "number"
}
},
"required": [
"curve",
"pts",
"hotThroughput",
"coldMs",
"hotMs"
],
"title": "Warmup",
"type": "object"
},
"WarmupPt": {
"additionalProperties": false,
"description": "One cold/hot amortization point: per-run ms at a cumulative run count.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"perRun": {
"title": "Perrun",
"type": "number"
}
},
"required": [
"n",
"perRun"
],
"title": "WarmupPt",
"type": "object"
}
},
"additionalProperties": false,
"description": "All of ONE backend's featured-case metrics, grouped (no parallel dicts).\n\nElement types match the per-panel charts exactly: `scaling`/`ecdf_*` keep the\n(x, y) `Point2D` axes, `param_err` is per-parameter, `param_spread`/`stability`\nkeep the mean $\\pm$ sd `SpreadPt` bands.",
"properties": {
"fit": {
"$ref": "#/$defs/SolverFit"
},
"conv": {
"items": {
"type": "number"
},
"title": "Conv",
"type": "array"
},
"grad": {
"items": {
"type": "number"
},
"title": "Grad",
"type": "array"
},
"convEff": {
"anyOf": [
{
"items": {
"type": "number"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"description": "Per-iteration convergence efficiency (cost reduction per step) derived from `conv`. Populated ONLY when `history_source == 'real'` (the subject's measured cost history); `None` for backends whose history is `reconstructed`, since a proxy-derived efficiency must not be presented as a measured series. Additive field \u2014 old payloads validate with None.",
"title": "Conveff"
},
"historySource": {
"enum": [
"real",
"reconstructed"
],
"title": "Historysource",
"type": "string"
},
"timing": {
"$ref": "#/$defs/TimingDist"
},
"accuracy": {
"$ref": "#/$defs/AccuracyDist"
},
"summary": {
"$ref": "#/$defs/Summary"
},
"paramErr": {
"items": {
"type": "number"
},
"title": "Paramerr",
"type": "array"
},
"ecdfResid": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Ecdfresid",
"type": "array"
},
"ecdfTime": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Ecdftime",
"type": "array"
},
"warmup": {
"$ref": "#/$defs/Warmup"
},
"scaling": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Scaling",
"type": "array"
},
"uncertainty": {
"$ref": "#/$defs/Uncertainty"
},
"paramSpread": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Paramspread",
"type": "array"
},
"stability": {
"$ref": "#/$defs/StabilityEntry"
},
"jacobianConditionNumber": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "$\\kappa(J)$ at convergence. Structural property of the case, not the solver. $\\kappa \\ge$ 1e6 \u2192 ill-conditioned (explains backend disagreement); $\\kappa \\le$ 1e4 \u2192 well-posed. Carried through BackendOutcome; backends that don't expose a Jacobian leave this None (scipy-ls backends populate it from raw.jac).",
"title": "Jacobianconditionnumber"
},
"thetaDistance": {
"anyOf": [
{
"items": {
"type": "number"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"description": "Scale-normalized per-iteration distance of the parameter vector to the synthetic ground truth, $d_k = \\lVert(\\theta_k - \\theta_{\\mathrm{true}})/s\\rVert_2$ ($s$ = per-parameter magnitude scale). Same length as `conv`. The REAL convergence-to-truth metric \u2014 distinct from the $\\chi^2$ descent in `conv`. None for non-synthetic cases (no $\\theta_{\\mathrm{true}}$) or backends without a parameter trajectory (only spectrafit's faer LM records it). Additive field \u2014 old payloads validate with None.",
"title": "Thetadistance"
}
},
"required": [
"fit",
"conv",
"grad",
"historySource",
"timing",
"accuracy",
"summary",
"paramErr",
"ecdfResid",
"ecdfTime",
"warmup",
"scaling",
"uncertainty",
"paramSpread",
"stability"
],
"title": "BackendProfile",
"type": "object"
}
Fields:
-
fit(SolverFit) -
conv(list[float]) -
grad(list[float]) -
conv_eff(list[float] | None) -
history_source(Literal['real', 'reconstructed']) -
timing(TimingDist) -
accuracy(AccuracyDist) -
summary(Summary) -
param_err(list[float]) -
ecdf_resid(list[Point2D]) -
ecdf_time(list[Point2D]) -
warmup(Warmup) -
scaling(list[Point2D]) -
uncertainty(Uncertainty) -
param_spread(list[SpreadPt]) -
stability(StabilityEntry) -
jacobian_condition_number(float | None) -
theta_distance(list[float] | None)
conv_eff = None
pydantic-field
¶
Per-iteration convergence efficiency (cost reduction per step) derived from conv. Populated ONLY when history_source == 'real' (the subject's measured cost history); None for backends whose history is reconstructed, since a proxy-derived efficiency must not be presented as a measured series. Additive field — old payloads validate with None.
jacobian_condition_number = None
pydantic-field
¶
\(\kappa(J)\) at convergence. Structural property of the case, not the solver. \(\kappa \ge\) 1e6 → ill-conditioned (explains backend disagreement); \(\kappa \le\) 1e4 → well-posed. Carried through BackendOutcome; backends that don't expose a Jacobian leave this None (scipy-ls backends populate it from raw.jac).
theta_distance = None
pydantic-field
¶
Scale-normalized per-iteration distance of the parameter vector to the synthetic ground truth, \(d_k = \lVert(\theta_k - \theta_{\mathrm{true}})/s\rVert_2\) (\(s\) = per-parameter magnitude scale). Same length as conv. The REAL convergence-to-truth metric — distinct from the \(\chi^2\) descent in conv. None for non-synthetic cases (no \(\theta_{\mathrm{true}}\)) or backends without a parameter trajectory (only spectrafit's faer LM records it). Additive field — old payloads validate with None.
SelectionStats
pydantic-model
¶
Bases: _Base
Frozen contract view of oracles.nested.SelectionStats.
Statistics for one nested-model comparison (full minus reduced): LRT / F-test / information criteria. All statistics are "full minus reduced" oriented: negative \(\Delta\)AIC/\(\Delta\)BIC means the full model is preferred.
Show JSON schema:
{
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.SelectionStats``.\n\nStatistics for one nested-model comparison (full minus reduced):\nLRT / F-test / information criteria. All statistics are \"full minus\nreduced\" oriented: negative $\\Delta$AIC/$\\Delta$BIC means the full model is\npreferred.",
"properties": {
"lrtStat": {
"title": "Lrtstat",
"type": "number"
},
"lrtP": {
"title": "Lrtp",
"type": "number"
},
"fStat": {
"title": "Fstat",
"type": "number"
},
"fP": {
"title": "Fp",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
}
},
"required": [
"lrtStat",
"lrtP",
"fStat",
"fP",
"dAIC",
"dBIC"
],
"title": "SelectionStats",
"type": "object"
}
Fields:
-
lrt_stat(float) -
lrt_p(float) -
f_stat(float) -
f_p(float) -
d_aic(float) -
d_bic(float)
NestedAdequacy
pydantic-model
¶
Bases: _Base
Frozen contract view of oracles.nested.NestedAdequacy.
Result of a nested-order V&V against a known generative order: checks that model-selection criteria (LRT / F / AIC / BIC) recover the true order m* by comparing the (m-1) reduced, m true, and (m*+1) over-fitted models.
Both AIC and BIC verdicts are reported separately. On the featured
real tri-Gaussian, AIC over-selects the 4-peak model
(over_not_preferred_aic = False) while BIC correctly recovers the
true 3-peak order (over_not_preferred_bic = True).
nested_adequacy: NestedAdequacy | None = None on Featured is an
additive-minor field — old payloads that lack it validate with None.
Show JSON schema:
{
"$defs": {
"SelectionStats": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.SelectionStats``.\n\nStatistics for one nested-model comparison (full minus reduced):\nLRT / F-test / information criteria. All statistics are \"full minus\nreduced\" oriented: negative $\\Delta$AIC/$\\Delta$BIC means the full model is\npreferred.",
"properties": {
"lrtStat": {
"title": "Lrtstat",
"type": "number"
},
"lrtP": {
"title": "Lrtp",
"type": "number"
},
"fStat": {
"title": "Fstat",
"type": "number"
},
"fP": {
"title": "Fp",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
}
},
"required": [
"lrtStat",
"lrtP",
"fStat",
"fP",
"dAIC",
"dBIC"
],
"title": "SelectionStats",
"type": "object"
}
},
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.NestedAdequacy``.\n\nResult of a nested-order V&V against a known generative order: checks\nthat model-selection criteria (LRT / F / AIC / BIC) recover the true\norder m* by comparing the (m*-1) reduced, m* true, and (m*+1)\nover-fitted models.\n\nBoth AIC and BIC verdicts are reported separately. On the featured\nreal tri-Gaussian, AIC over-selects the 4-peak model\n(``over_not_preferred_aic = False``) while BIC correctly recovers the\ntrue 3-peak order (``over_not_preferred_bic = True``).\n\n``nested_adequacy: NestedAdequacy | None = None`` on ``Featured`` is an\nadditive-minor field \u2014 old payloads that lack it validate with ``None``.",
"properties": {
"trueOrder": {
"title": "Trueorder",
"type": "integer"
},
"reducedRejected": {
"title": "Reducedrejected",
"type": "boolean"
},
"overNotPreferredAic": {
"title": "Overnotpreferredaic",
"type": "boolean"
},
"overNotPreferredBic": {
"title": "Overnotpreferredbic",
"type": "boolean"
},
"selectedOrderAic": {
"title": "Selectedorderaic",
"type": "integer"
},
"selectedOrderBic": {
"title": "Selectedorderbic",
"type": "integer"
},
"recoveredTrueOrderAic": {
"title": "Recoveredtrueorderaic",
"type": "boolean"
},
"recoveredTrueOrderBic": {
"title": "Recoveredtrueorderbic",
"type": "boolean"
},
"reducedVsTrue": {
"$ref": "#/$defs/SelectionStats"
},
"trueVsOver": {
"$ref": "#/$defs/SelectionStats"
}
},
"required": [
"trueOrder",
"reducedRejected",
"overNotPreferredAic",
"overNotPreferredBic",
"selectedOrderAic",
"selectedOrderBic",
"recoveredTrueOrderAic",
"recoveredTrueOrderBic",
"reducedVsTrue",
"trueVsOver"
],
"title": "NestedAdequacy",
"type": "object"
}
Fields:
-
true_order(int) -
reduced_rejected(bool) -
over_not_preferred_aic(bool) -
over_not_preferred_bic(bool) -
selected_order_aic(int) -
selected_order_bic(int) -
recovered_true_order_aic(bool) -
recovered_true_order_bic(bool) -
reduced_vs_true(SelectionStats) -
true_vs_over(SelectionStats)
Featured
pydantic-model
¶
Bases: _Base
One analyzed case rendered in full detail across all panels.
Per-backend metrics live in one profiles map (keyed by solver id); the
remaining fields are case-level (shared across backends). A report carries a
list of these (BenchReport.analyzed), one per deep-dived case, selectable
in the UI; the id is unique within the list.
Show JSON schema:
{
"$defs": {
"AccuracyDist": {
"additionalProperties": false,
"description": "Per-solver reduced-$\\chi^2$ distribution over repetitions.",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
}
},
"required": [
"raw",
"median",
"p5",
"p25",
"p75"
],
"title": "AccuracyDist",
"type": "object"
},
"BackendProfile": {
"additionalProperties": false,
"description": "All of ONE backend's featured-case metrics, grouped (no parallel dicts).\n\nElement types match the per-panel charts exactly: `scaling`/`ecdf_*` keep the\n(x, y) `Point2D` axes, `param_err` is per-parameter, `param_spread`/`stability`\nkeep the mean $\\pm$ sd `SpreadPt` bands.",
"properties": {
"fit": {
"$ref": "#/$defs/SolverFit"
},
"conv": {
"items": {
"type": "number"
},
"title": "Conv",
"type": "array"
},
"grad": {
"items": {
"type": "number"
},
"title": "Grad",
"type": "array"
},
"convEff": {
"anyOf": [
{
"items": {
"type": "number"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"description": "Per-iteration convergence efficiency (cost reduction per step) derived from `conv`. Populated ONLY when `history_source == 'real'` (the subject's measured cost history); `None` for backends whose history is `reconstructed`, since a proxy-derived efficiency must not be presented as a measured series. Additive field \u2014 old payloads validate with None.",
"title": "Conveff"
},
"historySource": {
"enum": [
"real",
"reconstructed"
],
"title": "Historysource",
"type": "string"
},
"timing": {
"$ref": "#/$defs/TimingDist"
},
"accuracy": {
"$ref": "#/$defs/AccuracyDist"
},
"summary": {
"$ref": "#/$defs/Summary"
},
"paramErr": {
"items": {
"type": "number"
},
"title": "Paramerr",
"type": "array"
},
"ecdfResid": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Ecdfresid",
"type": "array"
},
"ecdfTime": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Ecdftime",
"type": "array"
},
"warmup": {
"$ref": "#/$defs/Warmup"
},
"scaling": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Scaling",
"type": "array"
},
"uncertainty": {
"$ref": "#/$defs/Uncertainty"
},
"paramSpread": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Paramspread",
"type": "array"
},
"stability": {
"$ref": "#/$defs/StabilityEntry"
},
"jacobianConditionNumber": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "$\\kappa(J)$ at convergence. Structural property of the case, not the solver. $\\kappa \\ge$ 1e6 \u2192 ill-conditioned (explains backend disagreement); $\\kappa \\le$ 1e4 \u2192 well-posed. Carried through BackendOutcome; backends that don't expose a Jacobian leave this None (scipy-ls backends populate it from raw.jac).",
"title": "Jacobianconditionnumber"
},
"thetaDistance": {
"anyOf": [
{
"items": {
"type": "number"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"description": "Scale-normalized per-iteration distance of the parameter vector to the synthetic ground truth, $d_k = \\lVert(\\theta_k - \\theta_{\\mathrm{true}})/s\\rVert_2$ ($s$ = per-parameter magnitude scale). Same length as `conv`. The REAL convergence-to-truth metric \u2014 distinct from the $\\chi^2$ descent in `conv`. None for non-synthetic cases (no $\\theta_{\\mathrm{true}}$) or backends without a parameter trajectory (only spectrafit's faer LM records it). Additive field \u2014 old payloads validate with None.",
"title": "Thetadistance"
}
},
"required": [
"fit",
"conv",
"grad",
"historySource",
"timing",
"accuracy",
"summary",
"paramErr",
"ecdfResid",
"ecdfTime",
"warmup",
"scaling",
"uncertainty",
"paramSpread",
"stability"
],
"title": "BackendProfile",
"type": "object"
},
"CI": {
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
},
"ExprEdge": {
"additionalProperties": false,
"description": "One tied/shared-parameter expression edge.\n\nWire form (camelCase): ``{\"targetNode\": \"p1\", \"targetParam\": \"sigma\",\n\"expression\": \"p0.sigma\"}``. Typed (not a bare ``dict[str, str]``) so the\ninner keys camelize like every other contract field \u2014 the web reads\n``edge.targetNode`` / ``edge.targetParam``. Accepts the engine's snake_case\ndicts via ``populate_by_name`` on :class:`_Base`.",
"properties": {
"targetNode": {
"title": "Targetnode",
"type": "string"
},
"targetParam": {
"title": "Targetparam",
"type": "string"
},
"expression": {
"title": "Expression",
"type": "string"
}
},
"required": [
"targetNode",
"targetParam",
"expression"
],
"title": "ExprEdge",
"type": "object"
},
"GlobalFit": {
"additionalProperties": false,
"description": "Shared-model multi-spectrum global fit (lmfit pattern).\n\n**SYNTHETIC placeholder \u2014 no experimental data.** Uses a ``GlobalFitGraph``\nfit over synthetic Gaussian peaks. Rendered in the production UI as part of\nEvidence's \"Native showcases\" section (``global-fit-showcase`` panel,\n``web/src/panels/registry.tsx``). Peak centers/widths are shared across every\ndataset slice (global analysis) while each slice's amplitudes vary; ``traces``\nare the recovered per-peak amplitude traces and ``slices`` carry the per-dataset\nobserved + jointly-fitted curves on the shared ``x`` grid.\n\nThe ``dataset_axis`` / ``axis_label`` names are deliberately axis-neutral: the\nglobal-fit machinery in spectrafit-core (``GlobalFitGraph``) is not specific to\ntime \u2014 here the axis incidentally represents time, but the contract does not\nclaim that.",
"properties": {
"x": {
"items": {
"type": "number"
},
"title": "X",
"type": "array"
},
"datasetAxis": {
"items": {
"type": "number"
},
"title": "Datasetaxis",
"type": "array"
},
"slices": {
"items": {
"$ref": "#/$defs/GlobalFitSlice"
},
"title": "Slices",
"type": "array"
},
"traces": {
"items": {
"$ref": "#/$defs/PeakTrace"
},
"title": "Traces",
"type": "array"
},
"source": {
"default": "spectrafit-core",
"title": "Source",
"type": "string"
},
"xLabel": {
"default": "x",
"title": "Xlabel",
"type": "string"
},
"axisLabel": {
"default": "dataset",
"title": "Axislabel",
"type": "string"
},
"dataProvenance": {
"default": "synthetic",
"enum": [
"measured",
"synthetic"
],
"title": "Dataprovenance",
"type": "string"
}
},
"required": [
"x",
"datasetAxis",
"slices",
"traces"
],
"title": "GlobalFit",
"type": "object"
},
"GlobalFitSlice": {
"additionalProperties": false,
"description": "Dataset slice in a global fit with jointly-fitted shared model.\n\nThe ``coord`` field carries the axis coordinate for this slice (e.g. a time\nvalue, a temperature, a sample index) \u2014 the name is axis-neutral because the\nglobal-fit machinery is not specific to time.",
"properties": {
"coord": {
"title": "Coord",
"type": "number"
},
"obs": {
"items": {
"type": "number"
},
"title": "Obs",
"type": "array"
},
"model": {
"items": {
"type": "number"
},
"title": "Model",
"type": "array"
}
},
"required": [
"coord",
"obs",
"model"
],
"title": "GlobalFitSlice",
"type": "object"
},
"MultiDim": {
"additionalProperties": false,
"description": "Genuinely N-dimensional ($\\ge$3-D) fit showcase.\n\nA synthetic N-D Gaussian is recovered by a real least-squares solve with the\nparametric ``gaussian_nd`` kernel (D inferred from the node's indexed\n``center_<i>`` parameters). ``peaks`` are the *fitted* parameters and\n``r_squared`` the recovery quality. Full N-D obs/model/residual grids are\nintentionally NOT stored \u2014 they do not scale past 2-D \u2014 so ``projections``\ncarries 2-D marginal slices of the fitted model for any visualization.\n\n**SYNTHETIC \u2014 no experimental data.** Rendered in the production UI as\npart of Evidence's \"Native showcases\" section (``multidim-showcase`` panel,\n``web/src/panels/registry.tsx``).",
"properties": {
"nDims": {
"title": "Ndims",
"type": "integer"
},
"shape": {
"items": {
"type": "integer"
},
"title": "Shape",
"type": "array"
},
"nPoints": {
"title": "Npoints",
"type": "integer"
},
"rSquared": {
"title": "Rsquared",
"type": "number"
},
"peaks": {
"items": {
"$ref": "#/$defs/NdPeak"
},
"title": "Peaks",
"type": "array"
},
"projections": {
"items": {
"$ref": "#/$defs/Projection"
},
"title": "Projections",
"type": "array"
},
"source": {
"default": "spectrafit-core",
"title": "Source",
"type": "string"
},
"dataProvenance": {
"default": "synthetic",
"enum": [
"measured",
"synthetic"
],
"title": "Dataprovenance",
"type": "string"
}
},
"required": [
"nDims",
"shape",
"nPoints",
"rSquared",
"peaks",
"projections"
],
"title": "MultiDim",
"type": "object"
},
"NdPeak": {
"additionalProperties": false,
"description": "A recovered N-D Gaussian peak: amplitude + per-axis center and sigma.\n\n``center`` and ``sigma`` each have ``n_dims`` entries (indexed `center_0..` /\n`sigma_0..` from the parametric ``gaussian_nd`` kernel).",
"properties": {
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"items": {
"type": "number"
},
"title": "Center",
"type": "array"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "NdPeak",
"type": "object"
},
"NestedAdequacy": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.NestedAdequacy``.\n\nResult of a nested-order V&V against a known generative order: checks\nthat model-selection criteria (LRT / F / AIC / BIC) recover the true\norder m* by comparing the (m*-1) reduced, m* true, and (m*+1)\nover-fitted models.\n\nBoth AIC and BIC verdicts are reported separately. On the featured\nreal tri-Gaussian, AIC over-selects the 4-peak model\n(``over_not_preferred_aic = False``) while BIC correctly recovers the\ntrue 3-peak order (``over_not_preferred_bic = True``).\n\n``nested_adequacy: NestedAdequacy | None = None`` on ``Featured`` is an\nadditive-minor field \u2014 old payloads that lack it validate with ``None``.",
"properties": {
"trueOrder": {
"title": "Trueorder",
"type": "integer"
},
"reducedRejected": {
"title": "Reducedrejected",
"type": "boolean"
},
"overNotPreferredAic": {
"title": "Overnotpreferredaic",
"type": "boolean"
},
"overNotPreferredBic": {
"title": "Overnotpreferredbic",
"type": "boolean"
},
"selectedOrderAic": {
"title": "Selectedorderaic",
"type": "integer"
},
"selectedOrderBic": {
"title": "Selectedorderbic",
"type": "integer"
},
"recoveredTrueOrderAic": {
"title": "Recoveredtrueorderaic",
"type": "boolean"
},
"recoveredTrueOrderBic": {
"title": "Recoveredtrueorderbic",
"type": "boolean"
},
"reducedVsTrue": {
"$ref": "#/$defs/SelectionStats"
},
"trueVsOver": {
"$ref": "#/$defs/SelectionStats"
}
},
"required": [
"trueOrder",
"reducedRejected",
"overNotPreferredAic",
"overNotPreferredBic",
"selectedOrderAic",
"selectedOrderBic",
"recoveredTrueOrderAic",
"recoveredTrueOrderBic",
"reducedVsTrue",
"trueVsOver"
],
"title": "NestedAdequacy",
"type": "object"
},
"Peak": {
"additionalProperties": false,
"description": "A single peak contribution curve (featured solver), labeled g1/g2/\u2026.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"y": {
"items": {
"type": "number"
},
"title": "Y",
"type": "array"
}
},
"required": [
"label",
"y"
],
"title": "Peak",
"type": "object"
},
"PeakACS": {
"additionalProperties": false,
"description": "Peak parameters in the report's short form: amplitude / center / sigma.",
"properties": {
"a": {
"title": "A",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
},
"s": {
"title": "S",
"type": "number"
}
},
"required": [
"a",
"c",
"s"
],
"title": "PeakACS",
"type": "object"
},
"PeakTrace": {
"additionalProperties": false,
"description": "A shared peak's amplitude evolving across the dataset axis.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"amplitude": {
"items": {
"type": "number"
},
"title": "Amplitude",
"type": "array"
}
},
"required": [
"label",
"center",
"sigma",
"amplitude"
],
"title": "PeakTrace",
"type": "object"
},
"Point2D": {
"additionalProperties": false,
"description": "A single (x, y) point used by line / ecdf / scaling / warmup series.",
"properties": {
"x": {
"title": "X",
"type": "number"
},
"y": {
"title": "Y",
"type": "number"
}
},
"required": [
"x",
"y"
],
"title": "Point2D",
"type": "object"
},
"Projection": {
"additionalProperties": false,
"description": "A 2-D projection (marginal slice) of a higher-D fit: axis-pair + matrix.",
"properties": {
"labels": {
"maxItems": 2,
"minItems": 2,
"prefixItems": [
{
"type": "string"
},
{
"type": "string"
}
],
"title": "Labels",
"type": "array"
},
"matrix": {
"items": {
"items": {
"type": "number"
},
"type": "array"
},
"title": "Matrix",
"type": "array"
}
},
"required": [
"labels",
"matrix"
],
"title": "Projection",
"type": "object"
},
"SelectionStats": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.SelectionStats``.\n\nStatistics for one nested-model comparison (full minus reduced):\nLRT / F-test / information criteria. All statistics are \"full minus\nreduced\" oriented: negative $\\Delta$AIC/$\\Delta$BIC means the full model is\npreferred.",
"properties": {
"lrtStat": {
"title": "Lrtstat",
"type": "number"
},
"lrtP": {
"title": "Lrtp",
"type": "number"
},
"fStat": {
"title": "Fstat",
"type": "number"
},
"fP": {
"title": "Fp",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
}
},
"required": [
"lrtStat",
"lrtP",
"fStat",
"fP",
"dAIC",
"dBIC"
],
"title": "SelectionStats",
"type": "object"
},
"SolverFit": {
"additionalProperties": false,
"description": "One solver's fit of the featured case: params + curve + residuals.",
"properties": {
"params": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Params",
"type": "array"
},
"curve": {
"items": {
"type": "number"
},
"title": "Curve",
"type": "array"
},
"resid": {
"items": {
"type": "number"
},
"title": "Resid",
"type": "array"
}
},
"required": [
"params",
"curve",
"resid"
],
"title": "SolverFit",
"type": "object"
},
"SpreadPt": {
"additionalProperties": false,
"description": "A mean $\\pm$ sd point versus run count (param-spread / stability series).",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"mean": {
"title": "Mean",
"type": "number"
},
"sd": {
"title": "Sd",
"type": "number"
}
},
"required": [
"n",
"mean",
"sd"
],
"title": "SpreadPt",
"type": "object"
},
"StabilityEntry": {
"additionalProperties": false,
"description": "One backend's metric stability versus run count (mean $\\pm$ sd per metric).",
"properties": {
"r2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "R2",
"type": "array"
},
"rmse": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Rmse",
"type": "array"
},
"redChi2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Redchi2",
"type": "array"
},
"iters": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Iters",
"type": "array"
}
},
"required": [
"r2",
"rmse",
"redChi2",
"iters"
],
"title": "StabilityEntry",
"type": "object"
},
"Summary": {
"additionalProperties": false,
"description": "Headline KPI block per solver for the featured case.",
"properties": {
"r2": {
"title": "R2",
"type": "number"
},
"chi2": {
"title": "Chi2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"rmse": {
"title": "Rmse",
"type": "number"
},
"mae": {
"title": "Mae",
"type": "number"
},
"nIter": {
"title": "Niter",
"type": "integer"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"iqrMs": {
"title": "Iqrms",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
},
"speedup": {
"title": "Speedup",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"aic": {
"title": "Aic",
"type": "number"
},
"bic": {
"title": "Bic",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
},
"speedupCi": {
"anyOf": [
{
"$ref": "#/$defs/CI"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"r2",
"chi2",
"redChi2",
"rmse",
"mae",
"nIter",
"medMs",
"iqrMs",
"cv",
"speedup",
"success",
"aic",
"bic",
"dAIC",
"dBIC"
],
"title": "Summary",
"type": "object"
},
"TimingDist": {
"additionalProperties": false,
"description": "Per-solver runtime distribution over timing repetitions (milliseconds).",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"mean": {
"title": "Mean",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
},
"p95": {
"title": "P95",
"type": "number"
},
"iqr": {
"title": "Iqr",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
}
},
"required": [
"raw",
"median",
"mean",
"p5",
"p25",
"p75",
"p95",
"iqr",
"cv"
],
"title": "TimingDist",
"type": "object"
},
"Uncertainty": {
"additionalProperties": false,
"description": "Parameter-pull diagnostics: pulls, coverage, reported $\\sigma$.\n\nPulls are $(\\mathrm{est}-\\mathrm{truth})/\\sigma$. ``coverage`` is ``None`` when no\nvalid $\\sigma$ was available for any MC sample (e.g. backends like jax that set\nevery ``param_stderr`` entry to ``None``).\nA ``None`` coverage is distinct from a genuine ``0.0`` (all pulls $|p| \\ge 1\\sigma$).",
"properties": {
"pulls": {
"items": {
"type": "number"
},
"title": "Pulls",
"type": "array"
},
"coverage": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"title": "Coverage"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"pulls",
"coverage",
"sigma"
],
"title": "Uncertainty",
"type": "object"
},
"Warmup": {
"additionalProperties": false,
"description": "Cold/hot amortization for one solver (JAX compile cost amortizing 1\u2192100).",
"properties": {
"curve": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Curve",
"type": "array"
},
"pts": {
"items": {
"$ref": "#/$defs/WarmupPt"
},
"title": "Pts",
"type": "array"
},
"hotThroughput": {
"title": "Hotthroughput",
"type": "number"
},
"coldMs": {
"title": "Coldms",
"type": "number"
},
"hotMs": {
"title": "Hotms",
"type": "number"
}
},
"required": [
"curve",
"pts",
"hotThroughput",
"coldMs",
"hotMs"
],
"title": "Warmup",
"type": "object"
},
"WarmupPt": {
"additionalProperties": false,
"description": "One cold/hot amortization point: per-run ms at a cumulative run count.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"perRun": {
"title": "Perrun",
"type": "number"
}
},
"required": [
"n",
"perRun"
],
"title": "WarmupPt",
"type": "object"
}
},
"additionalProperties": false,
"description": "One analyzed case rendered in full detail across all panels.\n\nPer-backend metrics live in one `profiles` map (keyed by solver id); the\nremaining fields are case-level (shared across backends). A report carries a\n*list* of these (``BenchReport.analyzed``), one per deep-dived case, selectable\nin the UI; the `id` is unique within the list.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"name": {
"title": "Name",
"type": "string"
},
"category": {
"title": "Category",
"type": "string"
},
"x": {
"items": {
"type": "number"
},
"title": "X",
"type": "array"
},
"ref": {
"items": {
"type": "number"
},
"title": "Ref",
"type": "array"
},
"guess": {
"items": {
"type": "number"
},
"title": "Guess",
"type": "array"
},
"truth": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Truth",
"type": "array"
},
"guessParams": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Guessparams",
"type": "array"
},
"noise": {
"title": "Noise",
"type": "number"
},
"baseline": {
"title": "Baseline",
"type": "number"
},
"profiles": {
"additionalProperties": {
"$ref": "#/$defs/BackendProfile"
},
"title": "Profiles",
"type": "object"
},
"peaks": {
"items": {
"$ref": "#/$defs/Peak"
},
"title": "Peaks",
"type": "array"
},
"paramNames": {
"items": {
"type": "string"
},
"title": "Paramnames",
"type": "array"
},
"corr": {
"items": {
"items": {
"type": "number"
},
"type": "array"
},
"title": "Corr",
"type": "array"
},
"Ngrid": {
"items": {
"type": "integer"
},
"title": "Ngrid",
"type": "array"
},
"schedule": {
"items": {
"type": "integer"
},
"title": "Schedule",
"type": "array"
},
"runsSched": {
"items": {
"type": "integer"
},
"title": "Runssched",
"type": "array"
},
"crossN": {
"title": "Crossn",
"type": "number"
},
"multidim": {
"anyOf": [
{
"$ref": "#/$defs/MultiDim"
},
{
"type": "null"
}
],
"default": null
},
"globalFit": {
"anyOf": [
{
"$ref": "#/$defs/GlobalFit"
},
{
"type": "null"
}
],
"default": null
},
"nestedAdequacy": {
"anyOf": [
{
"$ref": "#/$defs/NestedAdequacy"
},
{
"type": "null"
}
],
"default": null,
"description": "Nested-model adequacy V&V result for this case. Populated when the benchmark engine runs the nested-order oracle (Task 4.3). None for cases without a known generative order or when the oracle was not executed. Additive-minor field \u2014 old payloads without nestedAdequacy validate with None."
},
"modelSourceFile": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Relative path to the Rust kernel source for this case's model, e.g. \"crates/spectrafit-models/src/gaussian.rs\". None for landscape/optfn cases without a dedicated kernel. Additive field \u2014 old payloads validate with None.",
"title": "Modelsourcefile"
},
"modelFormula": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "LaTeX formula string for the featured case's model shape. Sourced from PeakModel.formula_latex in the oracle registry. None when no formula is registered. Additive field.",
"title": "Modelformula"
},
"fixedParams": {
"additionalProperties": {
"items": {
"type": "string"
},
"type": "array"
},
"description": "Per-node fixed parameter names, e.g. {\"p0\": [\"center\"]} means node p0's center was held at its guess value during the fit. Empty for unconstrained cases. Additive-minor field \u2014 old payloads validate with {}.",
"title": "Fixedparams",
"type": "object"
},
"exprEdges": {
"description": "Tied/shared-parameter expression edges (camelCase keys on the wire: targetNode / targetParam / expression). Empty for unconstrained cases. Additive-minor field.",
"items": {
"$ref": "#/$defs/ExprEdge"
},
"title": "Expredges",
"type": "array"
}
},
"required": [
"id",
"name",
"category",
"x",
"ref",
"guess",
"truth",
"noise",
"baseline",
"profiles",
"peaks",
"paramNames",
"corr",
"Ngrid",
"schedule",
"runsSched",
"crossN"
],
"title": "Featured",
"type": "object"
}
Fields:
-
id(str) -
name(str) -
category(str) -
x(list[float]) -
ref(list[float]) -
guess(list[float]) -
truth(list[PeakACS]) -
guess_params(list[PeakACS]) -
noise(float) -
baseline(float) -
profiles(dict[SolverId, BackendProfile]) -
peaks(list[Peak]) -
param_names(list[str]) -
corr(list[list[float]]) -
n_grid(list[int]) -
schedule(list[int]) -
runs_sched(list[int]) -
cross_n(float) -
multidim(MultiDim | None) -
global_fit(GlobalFit | None) -
nested_adequacy(NestedAdequacy | None) -
model_source_file(str | None) -
model_formula(str | None) -
fixed_params(dict[str, list[str]]) -
expr_edges(list[ExprEdge])
nested_adequacy = None
pydantic-field
¶
Nested-model adequacy V&V result for this case. Populated when the benchmark engine runs the nested-order oracle (Task 4.3). None for cases without a known generative order or when the oracle was not executed. Additive-minor field — old payloads without nestedAdequacy validate with None.
model_source_file = None
pydantic-field
¶
Relative path to the Rust kernel source for this case's model, e.g. "crates/spectrafit-models/src/gaussian.rs". None for landscape/optfn cases without a dedicated kernel. Additive field — old payloads validate with None.
model_formula = None
pydantic-field
¶
LaTeX formula string for the featured case's model shape. Sourced from PeakModel.formula_latex in the oracle registry. None when no formula is registered. Additive field.
fixed_params
pydantic-field
¶
Per-node fixed parameter names, e.g. {"p0": ["center"]} means node p0's center was held at its guess value during the fit. Empty for unconstrained cases. Additive-minor field — old payloads validate with {}.
expr_edges
pydantic-field
¶
Tied/shared-parameter expression edges (camelCase keys on the wire: targetNode / targetParam / expression). Empty for unconstrained cases. Additive-minor field.
SuiteMetric
pydantic-model
¶
Bases: _Base
Single-point per-solver metrics for a suite row.
Show JSON schema:
{
"additionalProperties": false,
"description": "Single-point per-solver metrics for a suite row.",
"properties": {
"speedup": {
"title": "Speedup",
"type": "number"
},
"r2": {
"title": "R2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"paramErr": {
"title": "Paramerr",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"convergenceEfficiency": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "Mean cost reduction per iteration: $(\\mathrm{cost}_0 - \\mathrm{cost}_n)$ / n_iter. Populated when a real cost history exists (history_source=='real'); None for reconstructed histories. Additive field \u2014 old payloads validate with None.",
"title": "Convergenceefficiency"
},
"illConditioned": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"description": "True when $\\kappa(J) \\ge$ 1e6 (ill-conditioned). None when $\\kappa(J)$ was not computed. Additive field.",
"title": "Illconditioned"
},
"redChi2Weighted": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "$\\sigma$-weighted reduced-$\\chi^2$: $\\sum((y-\\hat{y})/\\sigma)^2$ / dof. None when $\\sigma = 0$ (noiseless case); see metric_undefined_reason. Additive field \u2014 old payloads validate with None.",
"title": "Redchi2Weighted"
},
"metricUndefinedReason": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Human-readable explanation when red_chi2_weighted is None, e.g. \"sigma=0 (noiseless)\". Additive field.",
"title": "Metricundefinedreason"
}
},
"required": [
"speedup",
"r2",
"redChi2",
"medMs",
"paramErr",
"success"
],
"title": "SuiteMetric",
"type": "object"
}
Fields:
-
speedup(float) -
r2(float) -
red_chi2(float) -
med_ms(float) -
param_err(float) -
success(bool) -
convergence_efficiency(float | None) -
ill_conditioned(bool | None) -
red_chi2_weighted(float | None) -
metric_undefined_reason(str | None)
Validators:
-
_reject_non_finite_required→speedup,r2,red_chi2,med_ms,param_err -
_reject_non_finite_optional→convergence_efficiency,red_chi2_weighted
convergence_efficiency = None
pydantic-field
¶
Mean cost reduction per iteration: \((\mathrm{cost}_0 - \mathrm{cost}_n)\) / n_iter. Populated when a real cost history exists (history_source=='real'); None for reconstructed histories. Additive field — old payloads validate with None.
ill_conditioned = None
pydantic-field
¶
True when \(\kappa(J) \ge\) 1e6 (ill-conditioned). None when \(\kappa(J)\) was not computed. Additive field.
red_chi2_weighted = None
pydantic-field
¶
\(\sigma\)-weighted reduced-\(\chi^2\): \(\sum((y-\hat{y})/\sigma)^2\) / dof. None when \(\sigma = 0\) (noiseless case); see metric_undefined_reason. Additive field — old payloads validate with None.
metric_undefined_reason = None
pydantic-field
¶
Human-readable explanation when red_chi2_weighted is None, e.g. "sigma=0 (noiseless)". Additive field.
SuiteCase
pydantic-model
¶
Bases: _Base
One row in the all-cases suite (regression map / win-rate tables).
Show JSON schema:
{
"$defs": {
"SuiteMetric": {
"additionalProperties": false,
"description": "Single-point per-solver metrics for a suite row.",
"properties": {
"speedup": {
"title": "Speedup",
"type": "number"
},
"r2": {
"title": "R2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"paramErr": {
"title": "Paramerr",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"convergenceEfficiency": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "Mean cost reduction per iteration: $(\\mathrm{cost}_0 - \\mathrm{cost}_n)$ / n_iter. Populated when a real cost history exists (history_source=='real'); None for reconstructed histories. Additive field \u2014 old payloads validate with None.",
"title": "Convergenceefficiency"
},
"illConditioned": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"description": "True when $\\kappa(J) \\ge$ 1e6 (ill-conditioned). None when $\\kappa(J)$ was not computed. Additive field.",
"title": "Illconditioned"
},
"redChi2Weighted": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "$\\sigma$-weighted reduced-$\\chi^2$: $\\sum((y-\\hat{y})/\\sigma)^2$ / dof. None when $\\sigma = 0$ (noiseless case); see metric_undefined_reason. Additive field \u2014 old payloads validate with None.",
"title": "Redchi2Weighted"
},
"metricUndefinedReason": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Human-readable explanation when red_chi2_weighted is None, e.g. \"sigma=0 (noiseless)\". Additive field.",
"title": "Metricundefinedreason"
}
},
"required": [
"speedup",
"r2",
"redChi2",
"medMs",
"paramErr",
"success"
],
"title": "SuiteMetric",
"type": "object"
}
},
"additionalProperties": false,
"description": "One row in the all-cases suite (regression map / win-rate tables).",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"name": {
"title": "Name",
"type": "string"
},
"category": {
"title": "Category",
"type": "string"
},
"difficulty": {
"title": "Difficulty",
"type": "number"
},
"m": {
"additionalProperties": {
"$ref": "#/$defs/SuiteMetric"
},
"title": "M",
"type": "object"
},
"winner": {
"title": "Winner",
"type": "string"
},
"regression": {
"title": "Regression",
"type": "boolean"
},
"winnerReason": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Short human-readable explanation of WHY the winner won, contrasting key signals ($\\kappa(J)$, n_iter, convergence_efficiency) for the winner vs the baseline. None when the required signals are absent. Additive field \u2014 old payloads validate with None.",
"title": "Winnerreason"
}
},
"required": [
"id",
"name",
"category",
"difficulty",
"m",
"winner",
"regression"
],
"title": "SuiteCase",
"type": "object"
}
Fields:
-
id(str) -
name(str) -
category(str) -
difficulty(float) -
m(dict[SolverId, SuiteMetric]) -
winner(str) -
regression(bool) -
winner_reason(str | None)
winner_reason = None
pydantic-field
¶
Short human-readable explanation of WHY the winner won, contrasting key signals (\(\kappa(J)\), n_iter, convergence_efficiency) for the winner vs the baseline. None when the required signals are absent. Additive field — old payloads validate with None.
PinnedBaseline
pydantic-model
¶
Bases: _Base
The .spectrafit_reports/perf_baseline.json sidecar surfaced to the web.
A subset of :func:oracles.reports.write_perf_baseline's on-disk shape
(which also writes schema_version, category, and baseline_solver_id
— the fields :func:oracles.cli._gate_self_perf_check uses server-side to
refuse a cross-context comparison) — only the four fields the web needs to
render the self-vs-self signal are carried here.
Show JSON schema:
{
"additionalProperties": false,
"description": "The ``.spectrafit_reports/perf_baseline.json`` sidecar surfaced to the web.\n\nA **subset** of :func:`oracles.reports.write_perf_baseline`'s on-disk shape\n(which also writes ``schema_version``, ``category``, and ``baseline_solver_id``\n\u2014 the fields :func:`oracles.cli._gate_self_perf_check` uses server-side to\nrefuse a cross-context comparison) \u2014 only the four fields the web needs to\nrender the self-vs-self signal are carried here.",
"properties": {
"runId": {
"title": "Runid",
"type": "string"
},
"recordedAt": {
"title": "Recordedat",
"type": "string"
},
"geomeanSpeedupVsBaseline": {
"title": "Geomeanspeedupvsbaseline",
"type": "number"
},
"nCases": {
"title": "Ncases",
"type": "integer"
}
},
"required": [
"runId",
"recordedAt",
"geomeanSpeedupVsBaseline",
"nCases"
],
"title": "PinnedBaseline",
"type": "object"
}
Fields:
-
run_id(str) -
recorded_at(str) -
geomean_speedup_vs_baseline(float) -
n_cases(int)
ManifestSignals
pydantic-model
¶
Bases: _Base
Headline numbers from manifest.json surfaced via the contract.
Mirrors the keys :func:oracles.reports._headline writes; pinned
is populated from :func:oracles.reports.read_perf_baseline (None
when no baseline is currently pinned). Surfaced to the web so the
gate-state UI can render real values instead of telling users to run
spc-bench show-baseline. Additive 1.1 -> 1.2 bump — the field on
:class:BenchReport is optional so old payloads on disk validate
without going through migrate.py.
Show JSON schema:
{
"$defs": {
"PinnedBaseline": {
"additionalProperties": false,
"description": "The ``.spectrafit_reports/perf_baseline.json`` sidecar surfaced to the web.\n\nA **subset** of :func:`oracles.reports.write_perf_baseline`'s on-disk shape\n(which also writes ``schema_version``, ``category``, and ``baseline_solver_id``\n\u2014 the fields :func:`oracles.cli._gate_self_perf_check` uses server-side to\nrefuse a cross-context comparison) \u2014 only the four fields the web needs to\nrender the self-vs-self signal are carried here.",
"properties": {
"runId": {
"title": "Runid",
"type": "string"
},
"recordedAt": {
"title": "Recordedat",
"type": "string"
},
"geomeanSpeedupVsBaseline": {
"title": "Geomeanspeedupvsbaseline",
"type": "number"
},
"nCases": {
"title": "Ncases",
"type": "integer"
}
},
"required": [
"runId",
"recordedAt",
"geomeanSpeedupVsBaseline",
"nCases"
],
"title": "PinnedBaseline",
"type": "object"
}
},
"additionalProperties": false,
"description": "Headline numbers from ``manifest.json`` surfaced via the contract.\n\nMirrors the keys :func:`oracles.reports._headline` writes; ``pinned``\nis populated from :func:`oracles.reports.read_perf_baseline` (``None``\nwhen no baseline is currently pinned). Surfaced to the web so the\ngate-state UI can render real values instead of telling users to run\n``spc-bench show-baseline``. Additive 1.1 -> 1.2 bump \u2014 the field on\n:class:`BenchReport` is optional so old payloads on disk validate\nwithout going through ``migrate.py``.",
"properties": {
"geomeanSpeedupVsBaseline": {
"title": "Geomeanspeedupvsbaseline",
"type": "number"
},
"maxAbsDeltaR2": {
"title": "Maxabsdeltar2",
"type": "number"
},
"spectrafitWinRate": {
"title": "Spectrafitwinrate",
"type": "number"
},
"regressions": {
"title": "Regressions",
"type": "integer"
},
"pinned": {
"anyOf": [
{
"$ref": "#/$defs/PinnedBaseline"
},
{
"type": "null"
}
],
"default": null
},
"harmonicMeanSpeedupVsBaseline": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "Harmonic mean of per-case speedup values vs the baseline solver. Complements the geometric mean per Eeckhout (2024): the harmonic mean is the correct aggregate for equal-time comparisons and is always $\\le$ the geometric mean for positively-skewed speedup distributions. Additive field \u2014 old payloads on disk validate with None.",
"title": "Harmonicmeanspeedupvsbaseline"
},
"gateState": {
"anyOf": [
{
"enum": [
"pass",
"warn",
"fail"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Aggregate gate verdict. Single source of truth; the web UI must render this rather than recomputing from regression flags.",
"title": "Gatestate"
},
"nonfiniteDr2CaseIds": {
"description": "Case ids whose $|\\Delta r^2|$ vs the baseline was non-finite (NaN/Inf). The accuracy gate fails when this is non-empty; carried as a list so _sanitize (which coerces non-finite floats) cannot erase the signal.",
"items": {
"type": "string"
},
"title": "Nonfinitedr2Caseids",
"type": "array"
},
"saturatedCategories": {
"description": "Categories where every supported backend hits $r^2 \\ge$ 0.999 on every case. Differential below the noise floor; UI should mark these explicitly.",
"items": {
"type": "string"
},
"title": "Saturatedcategories",
"type": "array"
},
"sanitizedValuePaths": {
"description": "JSONPath-ish locators ('$.suite[3].m.jax.r2') of non-finite floats the presentation sanitizer coerced to 0.0 when this artifact was written. Non-empty means some rendered 0.0 values are suppressions, not measurements. Carried as a list so _sanitize cannot erase the signal; additive field \u2014 old payloads validate with [].",
"items": {
"type": "string"
},
"title": "Sanitizedvaluepaths",
"type": "array"
}
},
"required": [
"geomeanSpeedupVsBaseline",
"maxAbsDeltaR2",
"spectrafitWinRate",
"regressions"
],
"title": "ManifestSignals",
"type": "object"
}
Fields:
-
geomean_speedup_vs_baseline(float) -
max_abs_delta_r2(float) -
spectrafit_win_rate(float) -
regressions(int) -
pinned(PinnedBaseline | None) -
harmonic_mean_speedup_vs_baseline(float | None) -
gate_state(GateState | None) -
nonfinite_dr2_case_ids(list[str]) -
saturated_categories(list[str]) -
sanitized_value_paths(list[str])
Validators:
-
_harmonic_le_geomean
harmonic_mean_speedup_vs_baseline = None
pydantic-field
¶
Harmonic mean of per-case speedup values vs the baseline solver. Complements the geometric mean per Eeckhout (2024): the harmonic mean is the correct aggregate for equal-time comparisons and is always \(\le\) the geometric mean for positively-skewed speedup distributions. Additive field — old payloads on disk validate with None.
gate_state = None
pydantic-field
¶
Aggregate gate verdict. Single source of truth; the web UI must render this rather than recomputing from regression flags.
nonfinite_dr2_case_ids
pydantic-field
¶
Case ids whose \(|\Delta r^2|\) vs the baseline was non-finite (NaN/Inf). The accuracy gate fails when this is non-empty; carried as a list so _sanitize (which coerces non-finite floats) cannot erase the signal.
saturated_categories
pydantic-field
¶
Categories where every supported backend hits \(r^2 \ge\) 0.999 on every case. Differential below the noise floor; UI should mark these explicitly.
sanitized_value_paths
pydantic-field
¶
JSONPath-ish locators ('$.suite[3].m.jax.r2') of non-finite floats the presentation sanitizer coerced to 0.0 when this artifact was written. Non-empty means some rendered 0.0 values are suppressions, not measurements. Carried as a list so _sanitize cannot erase the signal; additive field — old payloads validate with [].
PanelLayout
pydantic-model
¶
Bases: _Base
Grid sizing intent for one panel — hints, not CSS.
Show JSON schema:
{
"additionalProperties": false,
"description": "Grid sizing intent for one panel \u2014 hints, not CSS.",
"properties": {
"wide": {
"default": false,
"title": "Wide",
"type": "boolean"
},
"height": {
"default": 230,
"maximum": 640,
"minimum": 120,
"title": "Height",
"type": "integer"
}
},
"title": "PanelLayout",
"type": "object"
}
Fields:
-
wide(bool) -
height(int)
PanelSpec
pydantic-model
¶
Bases: _Base
One report panel as data: which series, which chart, which slot.
The web app renders these generically (PanelRenderer); adding a panel is
one record in oracles.panels.DEFAULT_PANELS, not a new React block.
source keys into the web-side series-builder registry — it is a contract
string, pinned both sides by the default-panels fixture test.
Additive 1.2 -> 1.3 bump together with BenchReport.panels.
Note
chart_kind is one of :data:ChartKind, mirroring the components
exported by web/src/charts/index.tsx one-to-one; the web-side
dispatch (panelRegistry.tsx) carries the mirror table and a
vitest sync test pins the two sides together. "band" maps to the
Line component's band-series mode.
Show JSON schema:
{
"$defs": {
"PanelLayout": {
"additionalProperties": false,
"description": "Grid sizing intent for one panel \u2014 hints, not CSS.",
"properties": {
"wide": {
"default": false,
"title": "Wide",
"type": "boolean"
},
"height": {
"default": 230,
"maximum": 640,
"minimum": 120,
"title": "Height",
"type": "integer"
}
},
"title": "PanelLayout",
"type": "object"
}
},
"additionalProperties": false,
"description": "One report panel as data: which series, which chart, which slot.\n\nThe web app renders these generically (PanelRenderer); adding a panel is\none record in ``oracles.panels.DEFAULT_PANELS``, not a new React block.\n``source`` keys into the web-side series-builder registry \u2014 it is a contract\nstring, pinned both sides by the default-panels fixture test.\nAdditive 1.2 -> 1.3 bump together with ``BenchReport.panels``.\n\nNote:\n ``chart_kind`` is one of :data:`ChartKind`, mirroring the components\n exported by ``web/src/charts/index.tsx`` one-to-one; the web-side\n dispatch (``panelRegistry.tsx``) carries the mirror table and a\n vitest sync test pins the two sides together. ``\"band\"`` maps to the\n Line component's band-series mode.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"title": {
"title": "Title",
"type": "string"
},
"desc": {
"default": "",
"title": "Desc",
"type": "string"
},
"chartKind": {
"enum": [
"line",
"scatter",
"ecdf",
"violin",
"box",
"heatmap",
"field",
"lollipop",
"dumbbell",
"sparkline",
"band"
],
"title": "Chartkind",
"type": "string"
},
"source": {
"title": "Source",
"type": "string"
},
"layout": {
"$ref": "#/$defs/PanelLayout"
}
},
"required": [
"id",
"title",
"chartKind",
"source"
],
"title": "PanelSpec",
"type": "object"
}
Fields:
-
id(str) -
title(str) -
desc(str) -
chart_kind(ChartKind) -
source(str) -
layout(PanelLayout)
CI
pydantic-model
¶
Bases: _Base
A point estimate with its confidence interval.
Show JSON schema:
{
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
}
Fields:
-
lo(float) -
point(float) -
hi(float)
CaseInference
pydantic-model
¶
Bases: _Base
Per-case inference: intervals that make each comparison falsifiable.
Show JSON schema:
{
"$defs": {
"CI": {
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
}
},
"additionalProperties": false,
"description": "Per-case inference: intervals that make each comparison falsifiable.",
"properties": {
"caseId": {
"title": "Caseid",
"type": "string"
},
"speedupCi": {
"$ref": "#/$defs/CI"
},
"deltaR2Ci": {
"$ref": "#/$defs/CI"
}
},
"required": [
"caseId",
"speedupCi",
"deltaR2Ci"
],
"title": "CaseInference",
"type": "object"
}
Fields:
EquivalenceResult
pydantic-model
¶
Bases: _Base
Per-category TOST: 'saturated' earned, not asserted.
Show JSON schema:
{
"additionalProperties": false,
"description": "Per-category TOST: 'saturated' earned, not asserted.",
"properties": {
"category": {
"title": "Category",
"type": "string"
},
"equivalent": {
"title": "Equivalent",
"type": "boolean"
},
"margin": {
"title": "Margin",
"type": "number"
},
"diff": {
"title": "Diff",
"type": "number"
}
},
"required": [
"category",
"equivalent",
"margin",
"diff"
],
"title": "EquivalenceResult",
"type": "object"
}
Fields:
-
category(str) -
equivalent(bool) -
margin(float) -
diff(float)
InferenceConfig
pydantic-model
¶
Bases: _Base
Pre-registered inference parameters (written to the manifest, not tuned).
The four new fields are the Bonferroni-split family \(\alpha\) values and the calibration-test hyper-parameters of the pre-registered inference protocol. They carry defaults so old payloads that lack them validate without a migrator.
Note
calibration_margin is the pre-registered equivalence band for the
\(\sigma\)-calibration gate (A2 fix): the gate uses a CI-inclusion
TOST, passing iff the Clopper–Pearson CI of empirical coverage lies
entirely within [coverage_nominal - calibration_margin,
coverage_nominal + calibration_margin].
Show JSON schema:
{
"additionalProperties": false,
"description": "Pre-registered inference parameters (written to the manifest, not tuned).\n\nThe four new fields are the Bonferroni-split family $\\alpha$ values and the\ncalibration-test hyper-parameters of the pre-registered inference protocol.\nThey carry defaults so old payloads that lack them validate without a migrator.\n\nNote:\n ``calibration_margin`` is the pre-registered equivalence band for the\n $\\sigma$-calibration gate (A2 fix): the gate uses a CI-inclusion\n TOST, passing iff the Clopper\u2013Pearson CI of empirical coverage lies\n entirely within ``[coverage_nominal - calibration_margin,\n coverage_nominal + calibration_margin]``.",
"properties": {
"equivalenceMargin": {
"title": "Equivalencemargin",
"type": "number"
},
"bootstrapB": {
"title": "Bootstrapb",
"type": "integer"
},
"seed": {
"title": "Seed",
"type": "integer"
},
"fdrQ": {
"title": "Fdrq",
"type": "number"
},
"alphaCalibration": {
"default": 0.025,
"title": "Alphacalibration",
"type": "number"
},
"alphaSpeed": {
"default": 0.025,
"title": "Alphaspeed",
"type": "number"
},
"coverageNominal": {
"default": 0.6827,
"title": "Coveragenominal",
"type": "number"
},
"minPulls": {
"default": 20,
"title": "Minpulls",
"type": "integer"
},
"calibrationMargin": {
"default": 0.03,
"title": "Calibrationmargin",
"type": "number"
}
},
"required": [
"equivalenceMargin",
"bootstrapB",
"seed",
"fdrQ"
],
"title": "InferenceConfig",
"type": "object"
}
Fields:
-
equivalence_margin(float) -
bootstrap_b(int) -
seed(int) -
fdr_q(float) -
alpha_calibration(float) -
alpha_speed(float) -
coverage_nominal(float) -
min_pulls(int) -
calibration_margin(float)
alpha_calibration = 0.025
pydantic-field
¶
Info:
Pre-registered calibration \(\alpha\) — Bonferroni-split family-wise 0.05
across the two primaries (this field and alpha_speed).
CalibrationResult
pydantic-model
¶
Bases: _Base
Frozen contract view of oracles.inference.CalibrationStat (Task 5.3).
Result of a binomial calibration test on parameter pulls — each pull is
(fitted value − true value) / its standard error. All fields mirror
CalibrationStat exactly
so the engine can pass the stat object straight into model_validate.
When fewer than InferenceConfig.min_pulls pulls were collected, the
engine still emits a populated record with skipped=True, passed=False,
coverage=0.0, binomial_p=1.0 — it is never None for insufficient data.
None on InferenceBlock only appears for old payloads that predate
the field (additive-minor schema upgrade); the web layer must check
skipped to distinguish "test ran but data was insufficient" from
"old payload that never had the field".
binomial_p is retained as an honest strict diagnostic: it reports the
point-null test of H0: coverage = nominal. At large n this has extreme power
and will reject even a negligible deviation; it is NOT the gate criterion.
passed uses a practical-equivalence rule: the Clopper–Pearson CI must lie
entirely within [nominal - equivalence_margin, nominal + equivalence_margin].
equivalence_margin carries the pre-registered band (default 0.03 = \(\pm\)3 pp).
Additive field — old payloads without it default to 0.03 (matching the engine
default) so validation is lossless.
Show JSON schema:
{
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.inference.CalibrationStat`` (Task 5.3).\n\nResult of a binomial calibration test on parameter pulls \u2014 each pull is\n(fitted value \u2212 true value) / its standard error. All fields mirror\n``CalibrationStat`` exactly\nso the engine can pass the stat object straight into ``model_validate``.\n\nWhen fewer than ``InferenceConfig.min_pulls`` pulls were collected, the\nengine still emits a **populated** record with ``skipped=True, passed=False,\ncoverage=0.0, binomial_p=1.0`` \u2014 it is never ``None`` for insufficient data.\n``None`` on ``InferenceBlock`` only appears for old payloads that predate\nthe field (additive-minor schema upgrade); the web layer must check\n``skipped`` to distinguish \"test ran but data was insufficient\" from\n\"old payload that never had the field\".\n\n``binomial_p`` is retained as an *honest strict diagnostic*: it reports the\npoint-null test of H0: coverage = nominal. At large n this has extreme power\nand will reject even a negligible deviation; it is NOT the gate criterion.\n\n``passed`` uses a practical-equivalence rule: the Clopper\u2013Pearson CI must lie\nentirely within ``[nominal - equivalence_margin, nominal + equivalence_margin]``.\n``equivalence_margin`` carries the pre-registered band (default 0.03 = $\\pm$3 pp).\nAdditive field \u2014 old payloads without it default to 0.03 (matching the engine\ndefault) so validation is lossless.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"coverage": {
"title": "Coverage",
"type": "number"
},
"coverageCiLo": {
"title": "Coveragecilo",
"type": "number"
},
"coverageCiHi": {
"title": "Coveragecihi",
"type": "number"
},
"nominal": {
"title": "Nominal",
"type": "number"
},
"binomialP": {
"title": "Binomialp",
"type": "number"
},
"ksStat": {
"title": "Ksstat",
"type": "number"
},
"ksP": {
"title": "Ksp",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"equivalenceMargin": {
"default": 0.03,
"title": "Equivalencemargin",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"n",
"coverage",
"coverageCiLo",
"coverageCiHi",
"nominal",
"binomialP",
"ksStat",
"ksP",
"alpha",
"passed",
"skipped"
],
"title": "CalibrationResult",
"type": "object"
}
Fields:
-
n(int) -
coverage(float) -
coverage_ci_lo(float) -
coverage_ci_hi(float) -
nominal(float) -
binomial_p(float) -
ks_stat(float) -
ks_p(float) -
alpha(float) -
equivalence_margin(float) -
passed(bool) -
skipped(bool)
SpeedInferenceResult
pydantic-model
¶
Bases: _Base
Frozen contract view of oracles.inference.SpeedStat (Task 5.3).
Result of a geomean-speedup significance test: bootstrap CI on the
geometric mean of per-case speedup ratios plus sign-test and Wilcoxon
signed-rank p-values. All fields mirror SpeedStat exactly.
When fewer than 2 valid per-case speedups were available, the engine still
emits a populated record with skipped=True, passed=False,
geomean_speedup=0.0, ci_lo=0.0, ci_hi=0.0 — it is never None for
insufficient data. None on InferenceBlock only appears for old
payloads that predate the field (additive-minor schema upgrade); the web
layer must check skipped to distinguish "test ran but data was
insufficient" from "old payload that never had the field".
Show JSON schema:
{
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.inference.SpeedStat`` (Task 5.3).\n\nResult of a geomean-speedup significance test: bootstrap CI on the\ngeometric mean of per-case speedup ratios plus sign-test and Wilcoxon\nsigned-rank p-values. All fields mirror ``SpeedStat`` exactly.\n\nWhen fewer than 2 valid per-case speedups were available, the engine still\nemits a **populated** record with ``skipped=True, passed=False,\ngeomean_speedup=0.0, ci_lo=0.0, ci_hi=0.0`` \u2014 it is never ``None`` for\ninsufficient data. ``None`` on ``InferenceBlock`` only appears for old\npayloads that predate the field (additive-minor schema upgrade); the web\nlayer must check ``skipped`` to distinguish \"test ran but data was\ninsufficient\" from \"old payload that never had the field\".",
"properties": {
"geomeanSpeedup": {
"title": "Geomeanspeedup",
"type": "number"
},
"ciLo": {
"title": "Cilo",
"type": "number"
},
"ciHi": {
"title": "Cihi",
"type": "number"
},
"excludesOne": {
"title": "Excludesone",
"type": "boolean"
},
"signP": {
"title": "Signp",
"type": "number"
},
"wilcoxonP": {
"title": "Wilcoxonp",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"geomeanSpeedup",
"ciLo",
"ciHi",
"excludesOne",
"signP",
"wilcoxonP",
"alpha",
"passed",
"skipped"
],
"title": "SpeedInferenceResult",
"type": "object"
}
Fields:
-
geomean_speedup(float) -
ci_lo(float) -
ci_hi(float) -
excludes_one(bool) -
sign_p(float) -
wilcoxon_p(float) -
alpha(float) -
passed(bool) -
skipped(bool)
InferenceBlock
pydantic-model
¶
Bases: _Base
Top-level inference attached to BenchReport (additive 1.3 -> 1.4).
calibration and speed_inference are additive-minor fields (Task 5.3):
old payloads that lack them validate with None; the gate axis is inert
when None (no-pass-by-absence rule).
Show JSON schema:
{
"$defs": {
"CI": {
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
},
"CalibrationResult": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.inference.CalibrationStat`` (Task 5.3).\n\nResult of a binomial calibration test on parameter pulls \u2014 each pull is\n(fitted value \u2212 true value) / its standard error. All fields mirror\n``CalibrationStat`` exactly\nso the engine can pass the stat object straight into ``model_validate``.\n\nWhen fewer than ``InferenceConfig.min_pulls`` pulls were collected, the\nengine still emits a **populated** record with ``skipped=True, passed=False,\ncoverage=0.0, binomial_p=1.0`` \u2014 it is never ``None`` for insufficient data.\n``None`` on ``InferenceBlock`` only appears for old payloads that predate\nthe field (additive-minor schema upgrade); the web layer must check\n``skipped`` to distinguish \"test ran but data was insufficient\" from\n\"old payload that never had the field\".\n\n``binomial_p`` is retained as an *honest strict diagnostic*: it reports the\npoint-null test of H0: coverage = nominal. At large n this has extreme power\nand will reject even a negligible deviation; it is NOT the gate criterion.\n\n``passed`` uses a practical-equivalence rule: the Clopper\u2013Pearson CI must lie\nentirely within ``[nominal - equivalence_margin, nominal + equivalence_margin]``.\n``equivalence_margin`` carries the pre-registered band (default 0.03 = $\\pm$3 pp).\nAdditive field \u2014 old payloads without it default to 0.03 (matching the engine\ndefault) so validation is lossless.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"coverage": {
"title": "Coverage",
"type": "number"
},
"coverageCiLo": {
"title": "Coveragecilo",
"type": "number"
},
"coverageCiHi": {
"title": "Coveragecihi",
"type": "number"
},
"nominal": {
"title": "Nominal",
"type": "number"
},
"binomialP": {
"title": "Binomialp",
"type": "number"
},
"ksStat": {
"title": "Ksstat",
"type": "number"
},
"ksP": {
"title": "Ksp",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"equivalenceMargin": {
"default": 0.03,
"title": "Equivalencemargin",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"n",
"coverage",
"coverageCiLo",
"coverageCiHi",
"nominal",
"binomialP",
"ksStat",
"ksP",
"alpha",
"passed",
"skipped"
],
"title": "CalibrationResult",
"type": "object"
},
"CaseInference": {
"additionalProperties": false,
"description": "Per-case inference: intervals that make each comparison falsifiable.",
"properties": {
"caseId": {
"title": "Caseid",
"type": "string"
},
"speedupCi": {
"$ref": "#/$defs/CI"
},
"deltaR2Ci": {
"$ref": "#/$defs/CI"
}
},
"required": [
"caseId",
"speedupCi",
"deltaR2Ci"
],
"title": "CaseInference",
"type": "object"
},
"EquivalenceResult": {
"additionalProperties": false,
"description": "Per-category TOST: 'saturated' earned, not asserted.",
"properties": {
"category": {
"title": "Category",
"type": "string"
},
"equivalent": {
"title": "Equivalent",
"type": "boolean"
},
"margin": {
"title": "Margin",
"type": "number"
},
"diff": {
"title": "Diff",
"type": "number"
}
},
"required": [
"category",
"equivalent",
"margin",
"diff"
],
"title": "EquivalenceResult",
"type": "object"
},
"InferenceConfig": {
"additionalProperties": false,
"description": "Pre-registered inference parameters (written to the manifest, not tuned).\n\nThe four new fields are the Bonferroni-split family $\\alpha$ values and the\ncalibration-test hyper-parameters of the pre-registered inference protocol.\nThey carry defaults so old payloads that lack them validate without a migrator.\n\nNote:\n ``calibration_margin`` is the pre-registered equivalence band for the\n $\\sigma$-calibration gate (A2 fix): the gate uses a CI-inclusion\n TOST, passing iff the Clopper\u2013Pearson CI of empirical coverage lies\n entirely within ``[coverage_nominal - calibration_margin,\n coverage_nominal + calibration_margin]``.",
"properties": {
"equivalenceMargin": {
"title": "Equivalencemargin",
"type": "number"
},
"bootstrapB": {
"title": "Bootstrapb",
"type": "integer"
},
"seed": {
"title": "Seed",
"type": "integer"
},
"fdrQ": {
"title": "Fdrq",
"type": "number"
},
"alphaCalibration": {
"default": 0.025,
"title": "Alphacalibration",
"type": "number"
},
"alphaSpeed": {
"default": 0.025,
"title": "Alphaspeed",
"type": "number"
},
"coverageNominal": {
"default": 0.6827,
"title": "Coveragenominal",
"type": "number"
},
"minPulls": {
"default": 20,
"title": "Minpulls",
"type": "integer"
},
"calibrationMargin": {
"default": 0.03,
"title": "Calibrationmargin",
"type": "number"
}
},
"required": [
"equivalenceMargin",
"bootstrapB",
"seed",
"fdrQ"
],
"title": "InferenceConfig",
"type": "object"
},
"SpeedInferenceResult": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.inference.SpeedStat`` (Task 5.3).\n\nResult of a geomean-speedup significance test: bootstrap CI on the\ngeometric mean of per-case speedup ratios plus sign-test and Wilcoxon\nsigned-rank p-values. All fields mirror ``SpeedStat`` exactly.\n\nWhen fewer than 2 valid per-case speedups were available, the engine still\nemits a **populated** record with ``skipped=True, passed=False,\ngeomean_speedup=0.0, ci_lo=0.0, ci_hi=0.0`` \u2014 it is never ``None`` for\ninsufficient data. ``None`` on ``InferenceBlock`` only appears for old\npayloads that predate the field (additive-minor schema upgrade); the web\nlayer must check ``skipped`` to distinguish \"test ran but data was\ninsufficient\" from \"old payload that never had the field\".",
"properties": {
"geomeanSpeedup": {
"title": "Geomeanspeedup",
"type": "number"
},
"ciLo": {
"title": "Cilo",
"type": "number"
},
"ciHi": {
"title": "Cihi",
"type": "number"
},
"excludesOne": {
"title": "Excludesone",
"type": "boolean"
},
"signP": {
"title": "Signp",
"type": "number"
},
"wilcoxonP": {
"title": "Wilcoxonp",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"geomeanSpeedup",
"ciLo",
"ciHi",
"excludesOne",
"signP",
"wilcoxonP",
"alpha",
"passed",
"skipped"
],
"title": "SpeedInferenceResult",
"type": "object"
}
},
"additionalProperties": false,
"description": "Top-level inference attached to BenchReport (additive 1.3 -> 1.4).\n\n``calibration`` and ``speed_inference`` are additive-minor fields (Task 5.3):\nold payloads that lack them validate with ``None``; the gate axis is inert\nwhen ``None`` (no-pass-by-absence rule).",
"properties": {
"config": {
"$ref": "#/$defs/InferenceConfig"
},
"cases": {
"items": {
"$ref": "#/$defs/CaseInference"
},
"title": "Cases",
"type": "array"
},
"equivalence": {
"items": {
"$ref": "#/$defs/EquivalenceResult"
},
"title": "Equivalence",
"type": "array"
},
"winnerStability": {
"additionalProperties": {
"type": "number"
},
"title": "Winnerstability",
"type": "object"
},
"calibration": {
"anyOf": [
{
"$ref": "#/$defs/CalibrationResult"
},
{
"type": "null"
}
],
"default": null
},
"speedInference": {
"anyOf": [
{
"$ref": "#/$defs/SpeedInferenceResult"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"config",
"cases",
"equivalence",
"winnerStability"
],
"title": "InferenceBlock",
"type": "object"
}
Fields:
-
config(InferenceConfig) -
cases(list[CaseInference]) -
equivalence(list[EquivalenceResult]) -
winner_stability(dict[str, float]) -
calibration(CalibrationResult | None) -
speed_inference(SpeedInferenceResult | None)
BenchReport
pydantic-model
¶
Bases: _Base
Complete benchmark report payload consumed by the web app.
Show JSON schema:
{
"$defs": {
"AccuracyDist": {
"additionalProperties": false,
"description": "Per-solver reduced-$\\chi^2$ distribution over repetitions.",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
}
},
"required": [
"raw",
"median",
"p5",
"p25",
"p75"
],
"title": "AccuracyDist",
"type": "object"
},
"BackendProfile": {
"additionalProperties": false,
"description": "All of ONE backend's featured-case metrics, grouped (no parallel dicts).\n\nElement types match the per-panel charts exactly: `scaling`/`ecdf_*` keep the\n(x, y) `Point2D` axes, `param_err` is per-parameter, `param_spread`/`stability`\nkeep the mean $\\pm$ sd `SpreadPt` bands.",
"properties": {
"fit": {
"$ref": "#/$defs/SolverFit"
},
"conv": {
"items": {
"type": "number"
},
"title": "Conv",
"type": "array"
},
"grad": {
"items": {
"type": "number"
},
"title": "Grad",
"type": "array"
},
"convEff": {
"anyOf": [
{
"items": {
"type": "number"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"description": "Per-iteration convergence efficiency (cost reduction per step) derived from `conv`. Populated ONLY when `history_source == 'real'` (the subject's measured cost history); `None` for backends whose history is `reconstructed`, since a proxy-derived efficiency must not be presented as a measured series. Additive field \u2014 old payloads validate with None.",
"title": "Conveff"
},
"historySource": {
"enum": [
"real",
"reconstructed"
],
"title": "Historysource",
"type": "string"
},
"timing": {
"$ref": "#/$defs/TimingDist"
},
"accuracy": {
"$ref": "#/$defs/AccuracyDist"
},
"summary": {
"$ref": "#/$defs/Summary"
},
"paramErr": {
"items": {
"type": "number"
},
"title": "Paramerr",
"type": "array"
},
"ecdfResid": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Ecdfresid",
"type": "array"
},
"ecdfTime": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Ecdftime",
"type": "array"
},
"warmup": {
"$ref": "#/$defs/Warmup"
},
"scaling": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Scaling",
"type": "array"
},
"uncertainty": {
"$ref": "#/$defs/Uncertainty"
},
"paramSpread": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Paramspread",
"type": "array"
},
"stability": {
"$ref": "#/$defs/StabilityEntry"
},
"jacobianConditionNumber": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "$\\kappa(J)$ at convergence. Structural property of the case, not the solver. $\\kappa \\ge$ 1e6 \u2192 ill-conditioned (explains backend disagreement); $\\kappa \\le$ 1e4 \u2192 well-posed. Carried through BackendOutcome; backends that don't expose a Jacobian leave this None (scipy-ls backends populate it from raw.jac).",
"title": "Jacobianconditionnumber"
},
"thetaDistance": {
"anyOf": [
{
"items": {
"type": "number"
},
"type": "array"
},
{
"type": "null"
}
],
"default": null,
"description": "Scale-normalized per-iteration distance of the parameter vector to the synthetic ground truth, $d_k = \\lVert(\\theta_k - \\theta_{\\mathrm{true}})/s\\rVert_2$ ($s$ = per-parameter magnitude scale). Same length as `conv`. The REAL convergence-to-truth metric \u2014 distinct from the $\\chi^2$ descent in `conv`. None for non-synthetic cases (no $\\theta_{\\mathrm{true}}$) or backends without a parameter trajectory (only spectrafit's faer LM records it). Additive field \u2014 old payloads validate with None.",
"title": "Thetadistance"
}
},
"required": [
"fit",
"conv",
"grad",
"historySource",
"timing",
"accuracy",
"summary",
"paramErr",
"ecdfResid",
"ecdfTime",
"warmup",
"scaling",
"uncertainty",
"paramSpread",
"stability"
],
"title": "BackendProfile",
"type": "object"
},
"CI": {
"additionalProperties": false,
"description": "A point estimate with its confidence interval.",
"properties": {
"lo": {
"title": "Lo",
"type": "number"
},
"point": {
"title": "Point",
"type": "number"
},
"hi": {
"title": "Hi",
"type": "number"
}
},
"required": [
"lo",
"point",
"hi"
],
"title": "CI",
"type": "object"
},
"CalibrationResult": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.inference.CalibrationStat`` (Task 5.3).\n\nResult of a binomial calibration test on parameter pulls \u2014 each pull is\n(fitted value \u2212 true value) / its standard error. All fields mirror\n``CalibrationStat`` exactly\nso the engine can pass the stat object straight into ``model_validate``.\n\nWhen fewer than ``InferenceConfig.min_pulls`` pulls were collected, the\nengine still emits a **populated** record with ``skipped=True, passed=False,\ncoverage=0.0, binomial_p=1.0`` \u2014 it is never ``None`` for insufficient data.\n``None`` on ``InferenceBlock`` only appears for old payloads that predate\nthe field (additive-minor schema upgrade); the web layer must check\n``skipped`` to distinguish \"test ran but data was insufficient\" from\n\"old payload that never had the field\".\n\n``binomial_p`` is retained as an *honest strict diagnostic*: it reports the\npoint-null test of H0: coverage = nominal. At large n this has extreme power\nand will reject even a negligible deviation; it is NOT the gate criterion.\n\n``passed`` uses a practical-equivalence rule: the Clopper\u2013Pearson CI must lie\nentirely within ``[nominal - equivalence_margin, nominal + equivalence_margin]``.\n``equivalence_margin`` carries the pre-registered band (default 0.03 = $\\pm$3 pp).\nAdditive field \u2014 old payloads without it default to 0.03 (matching the engine\ndefault) so validation is lossless.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"coverage": {
"title": "Coverage",
"type": "number"
},
"coverageCiLo": {
"title": "Coveragecilo",
"type": "number"
},
"coverageCiHi": {
"title": "Coveragecihi",
"type": "number"
},
"nominal": {
"title": "Nominal",
"type": "number"
},
"binomialP": {
"title": "Binomialp",
"type": "number"
},
"ksStat": {
"title": "Ksstat",
"type": "number"
},
"ksP": {
"title": "Ksp",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"equivalenceMargin": {
"default": 0.03,
"title": "Equivalencemargin",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"n",
"coverage",
"coverageCiLo",
"coverageCiHi",
"nominal",
"binomialP",
"ksStat",
"ksP",
"alpha",
"passed",
"skipped"
],
"title": "CalibrationResult",
"type": "object"
},
"CaseInference": {
"additionalProperties": false,
"description": "Per-case inference: intervals that make each comparison falsifiable.",
"properties": {
"caseId": {
"title": "Caseid",
"type": "string"
},
"speedupCi": {
"$ref": "#/$defs/CI"
},
"deltaR2Ci": {
"$ref": "#/$defs/CI"
}
},
"required": [
"caseId",
"speedupCi",
"deltaR2Ci"
],
"title": "CaseInference",
"type": "object"
},
"CategoryMeta": {
"additionalProperties": false,
"description": "Benchmark category definition (id, label, case count, hue token).",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"label": {
"title": "Label",
"type": "string"
},
"n": {
"title": "N",
"type": "integer"
},
"hue": {
"title": "Hue",
"type": "string"
}
},
"required": [
"id",
"label",
"n",
"hue"
],
"title": "CategoryMeta",
"type": "object"
},
"CredibilityRung": {
"description": "V&V maturity ladder (inspired by ASME V&V credibility levels).\n\nThe live wire set earns RUNG_2..RUNG_4. RUNG_1 / RUNG_5 are reserved\nend-stops (see the per-member comments), part of the public contract and\npinned by an enum-value test \u2014 reserved, not dead.",
"enum": [
1,
2,
3,
4,
5
],
"title": "CredibilityRung",
"type": "integer"
},
"EquivalenceResult": {
"additionalProperties": false,
"description": "Per-category TOST: 'saturated' earned, not asserted.",
"properties": {
"category": {
"title": "Category",
"type": "string"
},
"equivalent": {
"title": "Equivalent",
"type": "boolean"
},
"margin": {
"title": "Margin",
"type": "number"
},
"diff": {
"title": "Diff",
"type": "number"
}
},
"required": [
"category",
"equivalent",
"margin",
"diff"
],
"title": "EquivalenceResult",
"type": "object"
},
"ExprEdge": {
"additionalProperties": false,
"description": "One tied/shared-parameter expression edge.\n\nWire form (camelCase): ``{\"targetNode\": \"p1\", \"targetParam\": \"sigma\",\n\"expression\": \"p0.sigma\"}``. Typed (not a bare ``dict[str, str]``) so the\ninner keys camelize like every other contract field \u2014 the web reads\n``edge.targetNode`` / ``edge.targetParam``. Accepts the engine's snake_case\ndicts via ``populate_by_name`` on :class:`_Base`.",
"properties": {
"targetNode": {
"title": "Targetnode",
"type": "string"
},
"targetParam": {
"title": "Targetparam",
"type": "string"
},
"expression": {
"title": "Expression",
"type": "string"
}
},
"required": [
"targetNode",
"targetParam",
"expression"
],
"title": "ExprEdge",
"type": "object"
},
"Featured": {
"additionalProperties": false,
"description": "One analyzed case rendered in full detail across all panels.\n\nPer-backend metrics live in one `profiles` map (keyed by solver id); the\nremaining fields are case-level (shared across backends). A report carries a\n*list* of these (``BenchReport.analyzed``), one per deep-dived case, selectable\nin the UI; the `id` is unique within the list.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"name": {
"title": "Name",
"type": "string"
},
"category": {
"title": "Category",
"type": "string"
},
"x": {
"items": {
"type": "number"
},
"title": "X",
"type": "array"
},
"ref": {
"items": {
"type": "number"
},
"title": "Ref",
"type": "array"
},
"guess": {
"items": {
"type": "number"
},
"title": "Guess",
"type": "array"
},
"truth": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Truth",
"type": "array"
},
"guessParams": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Guessparams",
"type": "array"
},
"noise": {
"title": "Noise",
"type": "number"
},
"baseline": {
"title": "Baseline",
"type": "number"
},
"profiles": {
"additionalProperties": {
"$ref": "#/$defs/BackendProfile"
},
"title": "Profiles",
"type": "object"
},
"peaks": {
"items": {
"$ref": "#/$defs/Peak"
},
"title": "Peaks",
"type": "array"
},
"paramNames": {
"items": {
"type": "string"
},
"title": "Paramnames",
"type": "array"
},
"corr": {
"items": {
"items": {
"type": "number"
},
"type": "array"
},
"title": "Corr",
"type": "array"
},
"Ngrid": {
"items": {
"type": "integer"
},
"title": "Ngrid",
"type": "array"
},
"schedule": {
"items": {
"type": "integer"
},
"title": "Schedule",
"type": "array"
},
"runsSched": {
"items": {
"type": "integer"
},
"title": "Runssched",
"type": "array"
},
"crossN": {
"title": "Crossn",
"type": "number"
},
"multidim": {
"anyOf": [
{
"$ref": "#/$defs/MultiDim"
},
{
"type": "null"
}
],
"default": null
},
"globalFit": {
"anyOf": [
{
"$ref": "#/$defs/GlobalFit"
},
{
"type": "null"
}
],
"default": null
},
"nestedAdequacy": {
"anyOf": [
{
"$ref": "#/$defs/NestedAdequacy"
},
{
"type": "null"
}
],
"default": null,
"description": "Nested-model adequacy V&V result for this case. Populated when the benchmark engine runs the nested-order oracle (Task 4.3). None for cases without a known generative order or when the oracle was not executed. Additive-minor field \u2014 old payloads without nestedAdequacy validate with None."
},
"modelSourceFile": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Relative path to the Rust kernel source for this case's model, e.g. \"crates/spectrafit-models/src/gaussian.rs\". None for landscape/optfn cases without a dedicated kernel. Additive field \u2014 old payloads validate with None.",
"title": "Modelsourcefile"
},
"modelFormula": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "LaTeX formula string for the featured case's model shape. Sourced from PeakModel.formula_latex in the oracle registry. None when no formula is registered. Additive field.",
"title": "Modelformula"
},
"fixedParams": {
"additionalProperties": {
"items": {
"type": "string"
},
"type": "array"
},
"description": "Per-node fixed parameter names, e.g. {\"p0\": [\"center\"]} means node p0's center was held at its guess value during the fit. Empty for unconstrained cases. Additive-minor field \u2014 old payloads validate with {}.",
"title": "Fixedparams",
"type": "object"
},
"exprEdges": {
"description": "Tied/shared-parameter expression edges (camelCase keys on the wire: targetNode / targetParam / expression). Empty for unconstrained cases. Additive-minor field.",
"items": {
"$ref": "#/$defs/ExprEdge"
},
"title": "Expredges",
"type": "array"
}
},
"required": [
"id",
"name",
"category",
"x",
"ref",
"guess",
"truth",
"noise",
"baseline",
"profiles",
"peaks",
"paramNames",
"corr",
"Ngrid",
"schedule",
"runsSched",
"crossN"
],
"title": "Featured",
"type": "object"
},
"GlobalFit": {
"additionalProperties": false,
"description": "Shared-model multi-spectrum global fit (lmfit pattern).\n\n**SYNTHETIC placeholder \u2014 no experimental data.** Uses a ``GlobalFitGraph``\nfit over synthetic Gaussian peaks. Rendered in the production UI as part of\nEvidence's \"Native showcases\" section (``global-fit-showcase`` panel,\n``web/src/panels/registry.tsx``). Peak centers/widths are shared across every\ndataset slice (global analysis) while each slice's amplitudes vary; ``traces``\nare the recovered per-peak amplitude traces and ``slices`` carry the per-dataset\nobserved + jointly-fitted curves on the shared ``x`` grid.\n\nThe ``dataset_axis`` / ``axis_label`` names are deliberately axis-neutral: the\nglobal-fit machinery in spectrafit-core (``GlobalFitGraph``) is not specific to\ntime \u2014 here the axis incidentally represents time, but the contract does not\nclaim that.",
"properties": {
"x": {
"items": {
"type": "number"
},
"title": "X",
"type": "array"
},
"datasetAxis": {
"items": {
"type": "number"
},
"title": "Datasetaxis",
"type": "array"
},
"slices": {
"items": {
"$ref": "#/$defs/GlobalFitSlice"
},
"title": "Slices",
"type": "array"
},
"traces": {
"items": {
"$ref": "#/$defs/PeakTrace"
},
"title": "Traces",
"type": "array"
},
"source": {
"default": "spectrafit-core",
"title": "Source",
"type": "string"
},
"xLabel": {
"default": "x",
"title": "Xlabel",
"type": "string"
},
"axisLabel": {
"default": "dataset",
"title": "Axislabel",
"type": "string"
},
"dataProvenance": {
"default": "synthetic",
"enum": [
"measured",
"synthetic"
],
"title": "Dataprovenance",
"type": "string"
}
},
"required": [
"x",
"datasetAxis",
"slices",
"traces"
],
"title": "GlobalFit",
"type": "object"
},
"GlobalFitSlice": {
"additionalProperties": false,
"description": "Dataset slice in a global fit with jointly-fitted shared model.\n\nThe ``coord`` field carries the axis coordinate for this slice (e.g. a time\nvalue, a temperature, a sample index) \u2014 the name is axis-neutral because the\nglobal-fit machinery is not specific to time.",
"properties": {
"coord": {
"title": "Coord",
"type": "number"
},
"obs": {
"items": {
"type": "number"
},
"title": "Obs",
"type": "array"
},
"model": {
"items": {
"type": "number"
},
"title": "Model",
"type": "array"
}
},
"required": [
"coord",
"obs",
"model"
],
"title": "GlobalFitSlice",
"type": "object"
},
"InferenceBlock": {
"additionalProperties": false,
"description": "Top-level inference attached to BenchReport (additive 1.3 -> 1.4).\n\n``calibration`` and ``speed_inference`` are additive-minor fields (Task 5.3):\nold payloads that lack them validate with ``None``; the gate axis is inert\nwhen ``None`` (no-pass-by-absence rule).",
"properties": {
"config": {
"$ref": "#/$defs/InferenceConfig"
},
"cases": {
"items": {
"$ref": "#/$defs/CaseInference"
},
"title": "Cases",
"type": "array"
},
"equivalence": {
"items": {
"$ref": "#/$defs/EquivalenceResult"
},
"title": "Equivalence",
"type": "array"
},
"winnerStability": {
"additionalProperties": {
"type": "number"
},
"title": "Winnerstability",
"type": "object"
},
"calibration": {
"anyOf": [
{
"$ref": "#/$defs/CalibrationResult"
},
{
"type": "null"
}
],
"default": null
},
"speedInference": {
"anyOf": [
{
"$ref": "#/$defs/SpeedInferenceResult"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"config",
"cases",
"equivalence",
"winnerStability"
],
"title": "InferenceBlock",
"type": "object"
},
"InferenceConfig": {
"additionalProperties": false,
"description": "Pre-registered inference parameters (written to the manifest, not tuned).\n\nThe four new fields are the Bonferroni-split family $\\alpha$ values and the\ncalibration-test hyper-parameters of the pre-registered inference protocol.\nThey carry defaults so old payloads that lack them validate without a migrator.\n\nNote:\n ``calibration_margin`` is the pre-registered equivalence band for the\n $\\sigma$-calibration gate (A2 fix): the gate uses a CI-inclusion\n TOST, passing iff the Clopper\u2013Pearson CI of empirical coverage lies\n entirely within ``[coverage_nominal - calibration_margin,\n coverage_nominal + calibration_margin]``.",
"properties": {
"equivalenceMargin": {
"title": "Equivalencemargin",
"type": "number"
},
"bootstrapB": {
"title": "Bootstrapb",
"type": "integer"
},
"seed": {
"title": "Seed",
"type": "integer"
},
"fdrQ": {
"title": "Fdrq",
"type": "number"
},
"alphaCalibration": {
"default": 0.025,
"title": "Alphacalibration",
"type": "number"
},
"alphaSpeed": {
"default": 0.025,
"title": "Alphaspeed",
"type": "number"
},
"coverageNominal": {
"default": 0.6827,
"title": "Coveragenominal",
"type": "number"
},
"minPulls": {
"default": 20,
"title": "Minpulls",
"type": "integer"
},
"calibrationMargin": {
"default": 0.03,
"title": "Calibrationmargin",
"type": "number"
}
},
"required": [
"equivalenceMargin",
"bootstrapB",
"seed",
"fdrQ"
],
"title": "InferenceConfig",
"type": "object"
},
"ManifestSignals": {
"additionalProperties": false,
"description": "Headline numbers from ``manifest.json`` surfaced via the contract.\n\nMirrors the keys :func:`oracles.reports._headline` writes; ``pinned``\nis populated from :func:`oracles.reports.read_perf_baseline` (``None``\nwhen no baseline is currently pinned). Surfaced to the web so the\ngate-state UI can render real values instead of telling users to run\n``spc-bench show-baseline``. Additive 1.1 -> 1.2 bump \u2014 the field on\n:class:`BenchReport` is optional so old payloads on disk validate\nwithout going through ``migrate.py``.",
"properties": {
"geomeanSpeedupVsBaseline": {
"title": "Geomeanspeedupvsbaseline",
"type": "number"
},
"maxAbsDeltaR2": {
"title": "Maxabsdeltar2",
"type": "number"
},
"spectrafitWinRate": {
"title": "Spectrafitwinrate",
"type": "number"
},
"regressions": {
"title": "Regressions",
"type": "integer"
},
"pinned": {
"anyOf": [
{
"$ref": "#/$defs/PinnedBaseline"
},
{
"type": "null"
}
],
"default": null
},
"harmonicMeanSpeedupVsBaseline": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "Harmonic mean of per-case speedup values vs the baseline solver. Complements the geometric mean per Eeckhout (2024): the harmonic mean is the correct aggregate for equal-time comparisons and is always $\\le$ the geometric mean for positively-skewed speedup distributions. Additive field \u2014 old payloads on disk validate with None.",
"title": "Harmonicmeanspeedupvsbaseline"
},
"gateState": {
"anyOf": [
{
"enum": [
"pass",
"warn",
"fail"
],
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Aggregate gate verdict. Single source of truth; the web UI must render this rather than recomputing from regression flags.",
"title": "Gatestate"
},
"nonfiniteDr2CaseIds": {
"description": "Case ids whose $|\\Delta r^2|$ vs the baseline was non-finite (NaN/Inf). The accuracy gate fails when this is non-empty; carried as a list so _sanitize (which coerces non-finite floats) cannot erase the signal.",
"items": {
"type": "string"
},
"title": "Nonfinitedr2Caseids",
"type": "array"
},
"saturatedCategories": {
"description": "Categories where every supported backend hits $r^2 \\ge$ 0.999 on every case. Differential below the noise floor; UI should mark these explicitly.",
"items": {
"type": "string"
},
"title": "Saturatedcategories",
"type": "array"
},
"sanitizedValuePaths": {
"description": "JSONPath-ish locators ('$.suite[3].m.jax.r2') of non-finite floats the presentation sanitizer coerced to 0.0 when this artifact was written. Non-empty means some rendered 0.0 values are suppressions, not measurements. Carried as a list so _sanitize cannot erase the signal; additive field \u2014 old payloads validate with [].",
"items": {
"type": "string"
},
"title": "Sanitizedvaluepaths",
"type": "array"
}
},
"required": [
"geomeanSpeedupVsBaseline",
"maxAbsDeltaR2",
"spectrafitWinRate",
"regressions"
],
"title": "ManifestSignals",
"type": "object"
},
"MultiDim": {
"additionalProperties": false,
"description": "Genuinely N-dimensional ($\\ge$3-D) fit showcase.\n\nA synthetic N-D Gaussian is recovered by a real least-squares solve with the\nparametric ``gaussian_nd`` kernel (D inferred from the node's indexed\n``center_<i>`` parameters). ``peaks`` are the *fitted* parameters and\n``r_squared`` the recovery quality. Full N-D obs/model/residual grids are\nintentionally NOT stored \u2014 they do not scale past 2-D \u2014 so ``projections``\ncarries 2-D marginal slices of the fitted model for any visualization.\n\n**SYNTHETIC \u2014 no experimental data.** Rendered in the production UI as\npart of Evidence's \"Native showcases\" section (``multidim-showcase`` panel,\n``web/src/panels/registry.tsx``).",
"properties": {
"nDims": {
"title": "Ndims",
"type": "integer"
},
"shape": {
"items": {
"type": "integer"
},
"title": "Shape",
"type": "array"
},
"nPoints": {
"title": "Npoints",
"type": "integer"
},
"rSquared": {
"title": "Rsquared",
"type": "number"
},
"peaks": {
"items": {
"$ref": "#/$defs/NdPeak"
},
"title": "Peaks",
"type": "array"
},
"projections": {
"items": {
"$ref": "#/$defs/Projection"
},
"title": "Projections",
"type": "array"
},
"source": {
"default": "spectrafit-core",
"title": "Source",
"type": "string"
},
"dataProvenance": {
"default": "synthetic",
"enum": [
"measured",
"synthetic"
],
"title": "Dataprovenance",
"type": "string"
}
},
"required": [
"nDims",
"shape",
"nPoints",
"rSquared",
"peaks",
"projections"
],
"title": "MultiDim",
"type": "object"
},
"NdPeak": {
"additionalProperties": false,
"description": "A recovered N-D Gaussian peak: amplitude + per-axis center and sigma.\n\n``center`` and ``sigma`` each have ``n_dims`` entries (indexed `center_0..` /\n`sigma_0..` from the parametric ``gaussian_nd`` kernel).",
"properties": {
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"items": {
"type": "number"
},
"title": "Center",
"type": "array"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "NdPeak",
"type": "object"
},
"NestedAdequacy": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.NestedAdequacy``.\n\nResult of a nested-order V&V against a known generative order: checks\nthat model-selection criteria (LRT / F / AIC / BIC) recover the true\norder m* by comparing the (m*-1) reduced, m* true, and (m*+1)\nover-fitted models.\n\nBoth AIC and BIC verdicts are reported separately. On the featured\nreal tri-Gaussian, AIC over-selects the 4-peak model\n(``over_not_preferred_aic = False``) while BIC correctly recovers the\ntrue 3-peak order (``over_not_preferred_bic = True``).\n\n``nested_adequacy: NestedAdequacy | None = None`` on ``Featured`` is an\nadditive-minor field \u2014 old payloads that lack it validate with ``None``.",
"properties": {
"trueOrder": {
"title": "Trueorder",
"type": "integer"
},
"reducedRejected": {
"title": "Reducedrejected",
"type": "boolean"
},
"overNotPreferredAic": {
"title": "Overnotpreferredaic",
"type": "boolean"
},
"overNotPreferredBic": {
"title": "Overnotpreferredbic",
"type": "boolean"
},
"selectedOrderAic": {
"title": "Selectedorderaic",
"type": "integer"
},
"selectedOrderBic": {
"title": "Selectedorderbic",
"type": "integer"
},
"recoveredTrueOrderAic": {
"title": "Recoveredtrueorderaic",
"type": "boolean"
},
"recoveredTrueOrderBic": {
"title": "Recoveredtrueorderbic",
"type": "boolean"
},
"reducedVsTrue": {
"$ref": "#/$defs/SelectionStats"
},
"trueVsOver": {
"$ref": "#/$defs/SelectionStats"
}
},
"required": [
"trueOrder",
"reducedRejected",
"overNotPreferredAic",
"overNotPreferredBic",
"selectedOrderAic",
"selectedOrderBic",
"recoveredTrueOrderAic",
"recoveredTrueOrderBic",
"reducedVsTrue",
"trueVsOver"
],
"title": "NestedAdequacy",
"type": "object"
},
"NistDataset": {
"additionalProperties": false,
"description": "Per-dataset NIST StRD certified-value validation result.",
"properties": {
"name": {
"description": "StRD problem name, e.g. 'Gauss1'.",
"title": "Name",
"type": "string"
},
"model": {
"description": "Human-readable model description.",
"title": "Model",
"type": "string"
},
"nParams": {
"title": "Nparams",
"type": "integer"
},
"params": {
"items": {
"$ref": "#/$defs/NistParam"
},
"title": "Params",
"type": "array"
},
"minSigFigs": {
"description": "Worst (minimum) per-parameter sig-fig agreement.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff min_sig_figs $\\ge$ the validation threshold.",
"title": "Passed",
"type": "boolean"
}
},
"required": [
"name",
"model",
"nParams",
"params",
"minSigFigs",
"passed"
],
"title": "NistDataset",
"type": "object"
},
"NistParam": {
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
},
"NistValidation": {
"additionalProperties": false,
"description": "Aggregate NIST StRD certified-value validation (the W8 evidence block).\n\nIndependent external replication against NIST's extended-precision certified\nvalues \u2014 the evidence RUNG_5 was reserved for. Additive on TrustBlock.",
"properties": {
"thresholdSigFigs": {
"description": "Required minimum significant-figure agreement.",
"title": "Thresholdsigfigs",
"type": "number"
},
"datasets": {
"items": {
"$ref": "#/$defs/NistDataset"
},
"title": "Datasets",
"type": "array"
},
"minSigFigs": {
"description": "Worst min_sig_figs across all datasets.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff every dataset agrees to $\\ge$ threshold sig figs.",
"title": "Passed",
"type": "boolean"
},
"totalAvailable": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Size of the external NIST StRD nonlinear-regression universe (the denominator for 'N of M' coverage). Emitted by the validation builder so the UI never hardcodes the total. Additive \u2014 None for payloads written before this field existed.",
"title": "Totalavailable"
}
},
"required": [
"thresholdSigFigs",
"datasets",
"minSigFigs",
"passed"
],
"title": "NistValidation",
"type": "object"
},
"PanelLayout": {
"additionalProperties": false,
"description": "Grid sizing intent for one panel \u2014 hints, not CSS.",
"properties": {
"wide": {
"default": false,
"title": "Wide",
"type": "boolean"
},
"height": {
"default": 230,
"maximum": 640,
"minimum": 120,
"title": "Height",
"type": "integer"
}
},
"title": "PanelLayout",
"type": "object"
},
"PanelSpec": {
"additionalProperties": false,
"description": "One report panel as data: which series, which chart, which slot.\n\nThe web app renders these generically (PanelRenderer); adding a panel is\none record in ``oracles.panels.DEFAULT_PANELS``, not a new React block.\n``source`` keys into the web-side series-builder registry \u2014 it is a contract\nstring, pinned both sides by the default-panels fixture test.\nAdditive 1.2 -> 1.3 bump together with ``BenchReport.panels``.\n\nNote:\n ``chart_kind`` is one of :data:`ChartKind`, mirroring the components\n exported by ``web/src/charts/index.tsx`` one-to-one; the web-side\n dispatch (``panelRegistry.tsx``) carries the mirror table and a\n vitest sync test pins the two sides together. ``\"band\"`` maps to the\n Line component's band-series mode.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"title": {
"title": "Title",
"type": "string"
},
"desc": {
"default": "",
"title": "Desc",
"type": "string"
},
"chartKind": {
"enum": [
"line",
"scatter",
"ecdf",
"violin",
"box",
"heatmap",
"field",
"lollipop",
"dumbbell",
"sparkline",
"band"
],
"title": "Chartkind",
"type": "string"
},
"source": {
"title": "Source",
"type": "string"
},
"layout": {
"$ref": "#/$defs/PanelLayout"
}
},
"required": [
"id",
"title",
"chartKind",
"source"
],
"title": "PanelSpec",
"type": "object"
},
"Peak": {
"additionalProperties": false,
"description": "A single peak contribution curve (featured solver), labeled g1/g2/\u2026.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"y": {
"items": {
"type": "number"
},
"title": "Y",
"type": "array"
}
},
"required": [
"label",
"y"
],
"title": "Peak",
"type": "object"
},
"PeakACS": {
"additionalProperties": false,
"description": "Peak parameters in the report's short form: amplitude / center / sigma.",
"properties": {
"a": {
"title": "A",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
},
"s": {
"title": "S",
"type": "number"
}
},
"required": [
"a",
"c",
"s"
],
"title": "PeakACS",
"type": "object"
},
"PeakTrace": {
"additionalProperties": false,
"description": "A shared peak's amplitude evolving across the dataset axis.",
"properties": {
"label": {
"title": "Label",
"type": "string"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"amplitude": {
"items": {
"type": "number"
},
"title": "Amplitude",
"type": "array"
}
},
"required": [
"label",
"center",
"sigma",
"amplitude"
],
"title": "PeakTrace",
"type": "object"
},
"PinnedBaseline": {
"additionalProperties": false,
"description": "The ``.spectrafit_reports/perf_baseline.json`` sidecar surfaced to the web.\n\nA **subset** of :func:`oracles.reports.write_perf_baseline`'s on-disk shape\n(which also writes ``schema_version``, ``category``, and ``baseline_solver_id``\n\u2014 the fields :func:`oracles.cli._gate_self_perf_check` uses server-side to\nrefuse a cross-context comparison) \u2014 only the four fields the web needs to\nrender the self-vs-self signal are carried here.",
"properties": {
"runId": {
"title": "Runid",
"type": "string"
},
"recordedAt": {
"title": "Recordedat",
"type": "string"
},
"geomeanSpeedupVsBaseline": {
"title": "Geomeanspeedupvsbaseline",
"type": "number"
},
"nCases": {
"title": "Ncases",
"type": "integer"
}
},
"required": [
"runId",
"recordedAt",
"geomeanSpeedupVsBaseline",
"nCases"
],
"title": "PinnedBaseline",
"type": "object"
},
"Point2D": {
"additionalProperties": false,
"description": "A single (x, y) point used by line / ecdf / scaling / warmup series.",
"properties": {
"x": {
"title": "X",
"type": "number"
},
"y": {
"title": "Y",
"type": "number"
}
},
"required": [
"x",
"y"
],
"title": "Point2D",
"type": "object"
},
"Projection": {
"additionalProperties": false,
"description": "A 2-D projection (marginal slice) of a higher-D fit: axis-pair + matrix.",
"properties": {
"labels": {
"maxItems": 2,
"minItems": 2,
"prefixItems": [
{
"type": "string"
},
{
"type": "string"
}
],
"title": "Labels",
"type": "array"
},
"matrix": {
"items": {
"items": {
"type": "number"
},
"type": "array"
},
"title": "Matrix",
"type": "array"
}
},
"required": [
"labels",
"matrix"
],
"title": "Projection",
"type": "object"
},
"SelectionStats": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.nested.SelectionStats``.\n\nStatistics for one nested-model comparison (full minus reduced):\nLRT / F-test / information criteria. All statistics are \"full minus\nreduced\" oriented: negative $\\Delta$AIC/$\\Delta$BIC means the full model is\npreferred.",
"properties": {
"lrtStat": {
"title": "Lrtstat",
"type": "number"
},
"lrtP": {
"title": "Lrtp",
"type": "number"
},
"fStat": {
"title": "Fstat",
"type": "number"
},
"fP": {
"title": "Fp",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
}
},
"required": [
"lrtStat",
"lrtP",
"fStat",
"fP",
"dAIC",
"dBIC"
],
"title": "SelectionStats",
"type": "object"
},
"SolverFit": {
"additionalProperties": false,
"description": "One solver's fit of the featured case: params + curve + residuals.",
"properties": {
"params": {
"items": {
"$ref": "#/$defs/PeakACS"
},
"title": "Params",
"type": "array"
},
"curve": {
"items": {
"type": "number"
},
"title": "Curve",
"type": "array"
},
"resid": {
"items": {
"type": "number"
},
"title": "Resid",
"type": "array"
}
},
"required": [
"params",
"curve",
"resid"
],
"title": "SolverFit",
"type": "object"
},
"SolverMeta": {
"additionalProperties": false,
"description": "Solver legend entry (id, label, and theme color tokens).\n\nRe-exported from oracles.bench_contract for BenchReport.solvers.",
"properties": {
"id": {
"description": "Stable backend/solver key (e.g. `spectrafit`, `lmfit`, `scipy-ls-trf`) used to join this entry against per-backend results elsewhere in the report.",
"title": "Id",
"type": "string"
},
"label": {
"description": "Human-readable display name shown in legends and tables.",
"title": "Label",
"type": "string"
},
"color": {
"description": "Primary theme color token (CSS `var(--c-...)` reference, with an optional CSS fallback) used for this backend's main chart marks.",
"title": "Color",
"type": "string"
},
"soft": {
"description": "Muted/soft variant of `color` (CSS `var(--c-...-soft)` reference) used for secondary chart elements \u2014 bands, fills, and de-emphasized series \u2014 that must read as the same backend without competing with the primary marks.",
"title": "Soft",
"type": "string"
}
},
"required": [
"id",
"label",
"color",
"soft"
],
"title": "SolverMeta",
"type": "object"
},
"SpeedInferenceResult": {
"additionalProperties": false,
"description": "Frozen contract view of ``oracles.inference.SpeedStat`` (Task 5.3).\n\nResult of a geomean-speedup significance test: bootstrap CI on the\ngeometric mean of per-case speedup ratios plus sign-test and Wilcoxon\nsigned-rank p-values. All fields mirror ``SpeedStat`` exactly.\n\nWhen fewer than 2 valid per-case speedups were available, the engine still\nemits a **populated** record with ``skipped=True, passed=False,\ngeomean_speedup=0.0, ci_lo=0.0, ci_hi=0.0`` \u2014 it is never ``None`` for\ninsufficient data. ``None`` on ``InferenceBlock`` only appears for old\npayloads that predate the field (additive-minor schema upgrade); the web\nlayer must check ``skipped`` to distinguish \"test ran but data was\ninsufficient\" from \"old payload that never had the field\".",
"properties": {
"geomeanSpeedup": {
"title": "Geomeanspeedup",
"type": "number"
},
"ciLo": {
"title": "Cilo",
"type": "number"
},
"ciHi": {
"title": "Cihi",
"type": "number"
},
"excludesOne": {
"title": "Excludesone",
"type": "boolean"
},
"signP": {
"title": "Signp",
"type": "number"
},
"wilcoxonP": {
"title": "Wilcoxonp",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"geomeanSpeedup",
"ciLo",
"ciHi",
"excludesOne",
"signP",
"wilcoxonP",
"alpha",
"passed",
"skipped"
],
"title": "SpeedInferenceResult",
"type": "object"
},
"SpreadPt": {
"additionalProperties": false,
"description": "A mean $\\pm$ sd point versus run count (param-spread / stability series).",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"mean": {
"title": "Mean",
"type": "number"
},
"sd": {
"title": "Sd",
"type": "number"
}
},
"required": [
"n",
"mean",
"sd"
],
"title": "SpreadPt",
"type": "object"
},
"StabilityEntry": {
"additionalProperties": false,
"description": "One backend's metric stability versus run count (mean $\\pm$ sd per metric).",
"properties": {
"r2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "R2",
"type": "array"
},
"rmse": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Rmse",
"type": "array"
},
"redChi2": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Redchi2",
"type": "array"
},
"iters": {
"items": {
"$ref": "#/$defs/SpreadPt"
},
"title": "Iters",
"type": "array"
}
},
"required": [
"r2",
"rmse",
"redChi2",
"iters"
],
"title": "StabilityEntry",
"type": "object"
},
"SuiteCase": {
"additionalProperties": false,
"description": "One row in the all-cases suite (regression map / win-rate tables).",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"name": {
"title": "Name",
"type": "string"
},
"category": {
"title": "Category",
"type": "string"
},
"difficulty": {
"title": "Difficulty",
"type": "number"
},
"m": {
"additionalProperties": {
"$ref": "#/$defs/SuiteMetric"
},
"title": "M",
"type": "object"
},
"winner": {
"title": "Winner",
"type": "string"
},
"regression": {
"title": "Regression",
"type": "boolean"
},
"winnerReason": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Short human-readable explanation of WHY the winner won, contrasting key signals ($\\kappa(J)$, n_iter, convergence_efficiency) for the winner vs the baseline. None when the required signals are absent. Additive field \u2014 old payloads validate with None.",
"title": "Winnerreason"
}
},
"required": [
"id",
"name",
"category",
"difficulty",
"m",
"winner",
"regression"
],
"title": "SuiteCase",
"type": "object"
},
"SuiteMetric": {
"additionalProperties": false,
"description": "Single-point per-solver metrics for a suite row.",
"properties": {
"speedup": {
"title": "Speedup",
"type": "number"
},
"r2": {
"title": "R2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"paramErr": {
"title": "Paramerr",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"convergenceEfficiency": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "Mean cost reduction per iteration: $(\\mathrm{cost}_0 - \\mathrm{cost}_n)$ / n_iter. Populated when a real cost history exists (history_source=='real'); None for reconstructed histories. Additive field \u2014 old payloads validate with None.",
"title": "Convergenceefficiency"
},
"illConditioned": {
"anyOf": [
{
"type": "boolean"
},
{
"type": "null"
}
],
"default": null,
"description": "True when $\\kappa(J) \\ge$ 1e6 (ill-conditioned). None when $\\kappa(J)$ was not computed. Additive field.",
"title": "Illconditioned"
},
"redChi2Weighted": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"default": null,
"description": "$\\sigma$-weighted reduced-$\\chi^2$: $\\sum((y-\\hat{y})/\\sigma)^2$ / dof. None when $\\sigma = 0$ (noiseless case); see metric_undefined_reason. Additive field \u2014 old payloads validate with None.",
"title": "Redchi2Weighted"
},
"metricUndefinedReason": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Human-readable explanation when red_chi2_weighted is None, e.g. \"sigma=0 (noiseless)\". Additive field.",
"title": "Metricundefinedreason"
}
},
"required": [
"speedup",
"r2",
"redChi2",
"medMs",
"paramErr",
"success"
],
"title": "SuiteMetric",
"type": "object"
},
"Summary": {
"additionalProperties": false,
"description": "Headline KPI block per solver for the featured case.",
"properties": {
"r2": {
"title": "R2",
"type": "number"
},
"chi2": {
"title": "Chi2",
"type": "number"
},
"redChi2": {
"title": "Redchi2",
"type": "number"
},
"rmse": {
"title": "Rmse",
"type": "number"
},
"mae": {
"title": "Mae",
"type": "number"
},
"nIter": {
"title": "Niter",
"type": "integer"
},
"medMs": {
"title": "Medms",
"type": "number"
},
"iqrMs": {
"title": "Iqrms",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
},
"speedup": {
"title": "Speedup",
"type": "number"
},
"success": {
"title": "Success",
"type": "boolean"
},
"aic": {
"title": "Aic",
"type": "number"
},
"bic": {
"title": "Bic",
"type": "number"
},
"dAIC": {
"title": "Daic",
"type": "number"
},
"dBIC": {
"title": "Dbic",
"type": "number"
},
"speedupCi": {
"anyOf": [
{
"$ref": "#/$defs/CI"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"r2",
"chi2",
"redChi2",
"rmse",
"mae",
"nIter",
"medMs",
"iqrMs",
"cv",
"speedup",
"success",
"aic",
"bic",
"dAIC",
"dBIC"
],
"title": "Summary",
"type": "object"
},
"TimingDist": {
"additionalProperties": false,
"description": "Per-solver runtime distribution over timing repetitions (milliseconds).",
"properties": {
"raw": {
"items": {
"type": "number"
},
"title": "Raw",
"type": "array"
},
"median": {
"title": "Median",
"type": "number"
},
"mean": {
"title": "Mean",
"type": "number"
},
"p5": {
"title": "P5",
"type": "number"
},
"p25": {
"title": "P25",
"type": "number"
},
"p75": {
"title": "P75",
"type": "number"
},
"p95": {
"title": "P95",
"type": "number"
},
"iqr": {
"title": "Iqr",
"type": "number"
},
"cv": {
"title": "Cv",
"type": "number"
}
},
"required": [
"raw",
"median",
"mean",
"p5",
"p25",
"p75",
"p95",
"iqr",
"cv"
],
"title": "TimingDist",
"type": "object"
},
"TrustBlock": {
"additionalProperties": false,
"description": "Aggregate trust evidence attached to a BenchReport.",
"properties": {
"rung": {
"$ref": "#/$defs/CredibilityRung",
"description": "Aggregate V&V credibility rung (see `CredibilityRung`), capped by the worst genuine `fail` among `wires` \u2014 a `gap` or `skipped` wire never lowers it."
},
"wires": {
"description": "Every verification wire's outcome making up this ledger.",
"items": {
"$ref": "#/$defs/WireResult"
},
"title": "Wires",
"type": "array"
},
"nClaimsAudited": {
"description": "Number of report claims this audit run actually checked.",
"title": "Nclaimsaudited",
"type": "integer"
},
"nClaimsTotal": {
"description": "Total number of claims the report makes, whether or not they were audited this run.",
"title": "Nclaimstotal",
"type": "integer"
},
"nistValidation": {
"anyOf": [
{
"$ref": "#/$defs/NistValidation"
},
{
"type": "null"
}
],
"default": null,
"description": "NIST StRD certified-value validation (W8). Additive \u2014 Pydantic fills None for payloads that predate the A7 external-validation wire."
}
},
"required": [
"rung",
"wires",
"nClaimsAudited",
"nClaimsTotal"
],
"title": "TrustBlock",
"type": "object"
},
"Uncertainty": {
"additionalProperties": false,
"description": "Parameter-pull diagnostics: pulls, coverage, reported $\\sigma$.\n\nPulls are $(\\mathrm{est}-\\mathrm{truth})/\\sigma$. ``coverage`` is ``None`` when no\nvalid $\\sigma$ was available for any MC sample (e.g. backends like jax that set\nevery ``param_stderr`` entry to ``None``).\nA ``None`` coverage is distinct from a genuine ``0.0`` (all pulls $|p| \\ge 1\\sigma$).",
"properties": {
"pulls": {
"items": {
"type": "number"
},
"title": "Pulls",
"type": "array"
},
"coverage": {
"anyOf": [
{
"type": "number"
},
{
"type": "null"
}
],
"title": "Coverage"
},
"sigma": {
"items": {
"type": "number"
},
"title": "Sigma",
"type": "array"
}
},
"required": [
"pulls",
"coverage",
"sigma"
],
"title": "Uncertainty",
"type": "object"
},
"Warmup": {
"additionalProperties": false,
"description": "Cold/hot amortization for one solver (JAX compile cost amortizing 1\u2192100).",
"properties": {
"curve": {
"items": {
"$ref": "#/$defs/Point2D"
},
"title": "Curve",
"type": "array"
},
"pts": {
"items": {
"$ref": "#/$defs/WarmupPt"
},
"title": "Pts",
"type": "array"
},
"hotThroughput": {
"title": "Hotthroughput",
"type": "number"
},
"coldMs": {
"title": "Coldms",
"type": "number"
},
"hotMs": {
"title": "Hotms",
"type": "number"
}
},
"required": [
"curve",
"pts",
"hotThroughput",
"coldMs",
"hotMs"
],
"title": "Warmup",
"type": "object"
},
"WarmupPt": {
"additionalProperties": false,
"description": "One cold/hot amortization point: per-run ms at a cumulative run count.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"perRun": {
"title": "Perrun",
"type": "number"
}
},
"required": [
"n",
"perRun"
],
"title": "WarmupPt",
"type": "object"
},
"WireResult": {
"additionalProperties": false,
"description": "One verification wire's outcome.\n\nNote:\n ``status=\"gap\"`` is distinct from ``\"fail\"``: it means the property\n could not be verified because the backend does NOT expose the input\n (a disclosed CAPABILITY gap, e.g. $\\kappa(J)$ is not surfaced by\n spectrafit/lmfit/jax), not because a computed value was wrong. Like\n ``\"skipped\"``, a ``\"gap\"`` does not cap the credibility rung; only a\n genuine ``\"fail\"`` does (see ``runner._compute_rung``).",
"properties": {
"wireId": {
"description": "Audit wire id: W1, W2a-W2d, or W3..W11 from the value-stream diagram.",
"title": "Wireid",
"type": "string"
},
"name": {
"description": "Short machine-readable wire name (e.g. `synth_invariants`).",
"title": "Name",
"type": "string"
},
"status": {
"description": "Wire outcome: `pass`, `warn`, or `fail` for a computed verdict; `skipped` when the underlying test never ran; `gap` for a disclosed capability gap (see the class Note).",
"enum": [
"pass",
"warn",
"fail",
"skipped",
"gap"
],
"title": "Status",
"type": "string"
},
"evidence": {
"description": "One-line evidence statement.",
"title": "Evidence",
"type": "string"
},
"details": {
"additionalProperties": {
"anyOf": [
{
"type": "number"
},
{
"type": "integer"
},
{
"type": "string"
},
{
"type": "boolean"
},
{
"type": "null"
}
]
},
"description": "Optional structured detail (counts, thresholds, sample sizes, ...) supporting `evidence`; empty when the one-line statement is self-sufficient.",
"title": "Details",
"type": "object"
}
},
"required": [
"wireId",
"name",
"status",
"evidence"
],
"title": "WireResult",
"type": "object"
}
},
"additionalProperties": false,
"description": "Complete benchmark report payload consumed by the web app.",
"properties": {
"schemaVersion": {
"default": "1.7",
"title": "Schemaversion",
"type": "string"
},
"solvers": {
"items": {
"$ref": "#/$defs/SolverMeta"
},
"title": "Solvers",
"type": "array"
},
"categories": {
"items": {
"$ref": "#/$defs/CategoryMeta"
},
"title": "Categories",
"type": "array"
},
"analyzed": {
"items": {
"$ref": "#/$defs/Featured"
},
"title": "Analyzed",
"type": "array"
},
"suite": {
"items": {
"$ref": "#/$defs/SuiteCase"
},
"title": "Suite",
"type": "array"
},
"baselineSolverId": {
"default": "lmfit",
"description": "Solver id whose median runtime defines the speedup baseline (x1.0). Was implicitly \"lmfit\" through hardcoded engine lookups (`by_name[\"lmfit\"]`) and a hardcoded web label; now an explicit contract field so adding a fourth backend or retiring lmfit doesn't ripple across every consumer. Additive 1.0 -> 1.1 bump: old payloads without this field default here.",
"title": "Baselinesolverid",
"type": "string"
},
"manifest": {
"anyOf": [
{
"$ref": "#/$defs/ManifestSignals"
},
{
"type": "null"
}
],
"default": null,
"description": "Manifest headline numbers (mirrors `reports._headline`) \u2014 exposed so the web gate-state UI can render geomean / max $|\\Delta r^2|$ / win-rate / pinned ratio without a sidecar fetch. Additive 1.1 -> 1.2 bump; old payloads on disk validate against the bumped schema because Pydantic fills the default."
},
"trustBlock": {
"anyOf": [
{
"$ref": "#/$defs/TrustBlock"
},
{
"type": "null"
}
],
"default": null,
"description": "V&V evidence block emitted by oracles.audit.runner. Additive \u2014 Pydantic fills None for pre-Plan-E payloads, so old runs still parse."
},
"panels": {
"description": "Declarative panel registry consumed by the web PanelRenderer. Additive 1.2 -> 1.3 bump: old payloads validate with [] and the web app renders its explicit \"report predates panel specs\" empty state (no silent fallback to a hardcoded layout).",
"items": {
"$ref": "#/$defs/PanelSpec"
},
"title": "Panels",
"type": "array"
},
"inference": {
"anyOf": [
{
"$ref": "#/$defs/InferenceBlock"
},
{
"type": "null"
}
],
"default": null
},
"gitCommit": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Git commit hash (short, $\\ge$7 chars) at benchmark run time. None when git is unavailable or the run is not in a git repo. Additive field \u2014 old payloads validate with None.",
"title": "Gitcommit"
},
"gitBranch": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"description": "Git branch name at benchmark run time. None when git is unavailable. Additive field.",
"title": "Gitbranch"
},
"runTimestampUnix": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Unix epoch seconds when the benchmark run started (int). None when not captured. Additive field.",
"title": "Runtimestampunix"
}
},
"required": [
"solvers",
"categories",
"analyzed",
"suite"
],
"title": "BenchReport",
"type": "object"
}
Fields:
-
schema_version(str) -
solvers(list[SolverMeta]) -
categories(list[CategoryMeta]) -
analyzed(list[Featured]) -
suite(list[SuiteCase]) -
baseline_solver_id(str) -
manifest(ManifestSignals | None) -
trust_block(TrustBlock | None) -
panels(list[PanelSpec]) -
inference(InferenceBlock | None) -
git_commit(str | None) -
git_branch(str | None) -
run_timestamp_unix(int | None)
baseline_solver_id = 'lmfit'
pydantic-field
¶
Solver id whose median runtime defines the speedup baseline (x1.0). Was implicitly "lmfit" through hardcoded engine lookups (by_name["lmfit"]) and a hardcoded web label; now an explicit contract field so adding a fourth backend or retiring lmfit doesn't ripple across every consumer. Additive 1.0 -> 1.1 bump: old payloads without this field default here.
manifest = None
pydantic-field
¶
Manifest headline numbers (mirrors reports._headline) — exposed so the web gate-state UI can render geomean / max \(|\Delta r^2|\) / win-rate / pinned ratio without a sidecar fetch. Additive 1.1 -> 1.2 bump; old payloads on disk validate against the bumped schema because Pydantic fills the default.
trust_block = None
pydantic-field
¶
V&V evidence block emitted by oracles.audit.runner. Additive — Pydantic fills None for pre-Plan-E payloads, so old runs still parse.
panels
pydantic-field
¶
Declarative panel registry consumed by the web PanelRenderer. Additive 1.2 -> 1.3 bump: old payloads validate with [] and the web app renders its explicit "report predates panel specs" empty state (no silent fallback to a hardcoded layout).
git_commit = None
pydantic-field
¶
Git commit hash (short, \(\ge\)7 chars) at benchmark run time. None when git is unavailable or the run is not in a git repo. Additive field — old payloads validate with None.
git_branch = None
pydantic-field
¶
Git branch name at benchmark run time. None when git is unavailable. Additive field.
run_timestamp_unix = None
pydantic-field
¶
Unix epoch seconds when the benchmark run started (int). None when not captured. Additive field.
oracles.trust_ledger
¶
Trust-ledger contract — single source of truth for what was audited.
The TrustBlock is embedded in BenchReport (optional, additive). It maps each
audited claim to a wire-id, the evidence sentence, and a pass/warn/fail status,
plus an aggregate credibility rung. The block is written to disk as
trust.json next to manifest.json AND inlined into results.json so a
single payload carries its own provenance.
CredibilityRung
¶
Bases: IntEnum
V&V maturity ladder (inspired by ASME V&V credibility levels).
The live wire set earns RUNG_2..RUNG_4. RUNG_1 / RUNG_5 are reserved end-stops (see the per-member comments), part of the public contract and pinned by an enum-value test — reserved, not dead.
WireResult
pydantic-model
¶
Bases: _Contract
One verification wire's outcome.
Note
status="gap" is distinct from "fail": it means the property
could not be verified because the backend does NOT expose the input
(a disclosed CAPABILITY gap, e.g. \(\kappa(J)\) is not surfaced by
spectrafit/lmfit/jax), not because a computed value was wrong. Like
"skipped", a "gap" does not cap the credibility rung; only a
genuine "fail" does (see runner._compute_rung).
Show JSON schema:
{
"additionalProperties": false,
"description": "One verification wire's outcome.\n\nNote:\n ``status=\"gap\"`` is distinct from ``\"fail\"``: it means the property\n could not be verified because the backend does NOT expose the input\n (a disclosed CAPABILITY gap, e.g. $\\kappa(J)$ is not surfaced by\n spectrafit/lmfit/jax), not because a computed value was wrong. Like\n ``\"skipped\"``, a ``\"gap\"`` does not cap the credibility rung; only a\n genuine ``\"fail\"`` does (see ``runner._compute_rung``).",
"properties": {
"wireId": {
"description": "Audit wire id: W1, W2a-W2d, or W3..W11 from the value-stream diagram.",
"title": "Wireid",
"type": "string"
},
"name": {
"description": "Short machine-readable wire name (e.g. `synth_invariants`).",
"title": "Name",
"type": "string"
},
"status": {
"description": "Wire outcome: `pass`, `warn`, or `fail` for a computed verdict; `skipped` when the underlying test never ran; `gap` for a disclosed capability gap (see the class Note).",
"enum": [
"pass",
"warn",
"fail",
"skipped",
"gap"
],
"title": "Status",
"type": "string"
},
"evidence": {
"description": "One-line evidence statement.",
"title": "Evidence",
"type": "string"
},
"details": {
"additionalProperties": {
"anyOf": [
{
"type": "number"
},
{
"type": "integer"
},
{
"type": "string"
},
{
"type": "boolean"
},
{
"type": "null"
}
]
},
"description": "Optional structured detail (counts, thresholds, sample sizes, ...) supporting `evidence`; empty when the one-line statement is self-sufficient.",
"title": "Details",
"type": "object"
}
},
"required": [
"wireId",
"name",
"status",
"evidence"
],
"title": "WireResult",
"type": "object"
}
Config:
extra:forbidalias_generator:to_camelpopulate_by_name:True
Fields:
-
wire_id(str) -
name(str) -
status(WireStatus) -
evidence(str) -
details(dict[str, float | int | str | bool | None])
wire_id
pydantic-field
¶
Audit wire id: W1, W2a-W2d, or W3..W11 from the value-stream diagram.
name
pydantic-field
¶
Short machine-readable wire name (e.g. synth_invariants).
status
pydantic-field
¶
Wire outcome: pass, warn, or fail for a computed verdict; skipped when the underlying test never ran; gap for a disclosed capability gap (see the class Note).
evidence
pydantic-field
¶
One-line evidence statement.
details
pydantic-field
¶
Optional structured detail (counts, thresholds, sample sizes, ...) supporting evidence; empty when the one-line statement is self-sufficient.
NistParam
pydantic-model
¶
Bases: _Contract
One parameter's certified-vs-fitted agreement for a NIST StRD dataset.
Show JSON schema:
{
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
}
Fields:
-
name(str) -
certified(float) -
fitted(float) -
sig_figs_agreed(float)
name
pydantic-field
¶
NIST parameter name, e.g. 'b1'.
certified
pydantic-field
¶
NIST certified value (10+ sig figs).
fitted
pydantic-field
¶
Spectrafit recovered value.
sig_figs_agreed
pydantic-field
¶
\(-\log_{10}\left(|\mathrm{fitted}-\mathrm{certified}|/|\mathrm{certified}|\right)\); \(\infty\)-capped for exact agreement.
NistDataset
pydantic-model
¶
Bases: _Contract
Per-dataset NIST StRD certified-value validation result.
Show JSON schema:
{
"$defs": {
"NistParam": {
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
}
},
"additionalProperties": false,
"description": "Per-dataset NIST StRD certified-value validation result.",
"properties": {
"name": {
"description": "StRD problem name, e.g. 'Gauss1'.",
"title": "Name",
"type": "string"
},
"model": {
"description": "Human-readable model description.",
"title": "Model",
"type": "string"
},
"nParams": {
"title": "Nparams",
"type": "integer"
},
"params": {
"items": {
"$ref": "#/$defs/NistParam"
},
"title": "Params",
"type": "array"
},
"minSigFigs": {
"description": "Worst (minimum) per-parameter sig-fig agreement.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff min_sig_figs $\\ge$ the validation threshold.",
"title": "Passed",
"type": "boolean"
}
},
"required": [
"name",
"model",
"nParams",
"params",
"minSigFigs",
"passed"
],
"title": "NistDataset",
"type": "object"
}
Fields:
-
name(str) -
model(str) -
n_params(int) -
params(list[NistParam]) -
min_sig_figs(float) -
passed(bool)
NistValidation
pydantic-model
¶
Bases: _Contract
Aggregate NIST StRD certified-value validation (the W8 evidence block).
Independent external replication against NIST's extended-precision certified values — the evidence RUNG_5 was reserved for. Additive on TrustBlock.
Show JSON schema:
{
"$defs": {
"NistDataset": {
"additionalProperties": false,
"description": "Per-dataset NIST StRD certified-value validation result.",
"properties": {
"name": {
"description": "StRD problem name, e.g. 'Gauss1'.",
"title": "Name",
"type": "string"
},
"model": {
"description": "Human-readable model description.",
"title": "Model",
"type": "string"
},
"nParams": {
"title": "Nparams",
"type": "integer"
},
"params": {
"items": {
"$ref": "#/$defs/NistParam"
},
"title": "Params",
"type": "array"
},
"minSigFigs": {
"description": "Worst (minimum) per-parameter sig-fig agreement.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff min_sig_figs $\\ge$ the validation threshold.",
"title": "Passed",
"type": "boolean"
}
},
"required": [
"name",
"model",
"nParams",
"params",
"minSigFigs",
"passed"
],
"title": "NistDataset",
"type": "object"
},
"NistParam": {
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
}
},
"additionalProperties": false,
"description": "Aggregate NIST StRD certified-value validation (the W8 evidence block).\n\nIndependent external replication against NIST's extended-precision certified\nvalues \u2014 the evidence RUNG_5 was reserved for. Additive on TrustBlock.",
"properties": {
"thresholdSigFigs": {
"description": "Required minimum significant-figure agreement.",
"title": "Thresholdsigfigs",
"type": "number"
},
"datasets": {
"items": {
"$ref": "#/$defs/NistDataset"
},
"title": "Datasets",
"type": "array"
},
"minSigFigs": {
"description": "Worst min_sig_figs across all datasets.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff every dataset agrees to $\\ge$ threshold sig figs.",
"title": "Passed",
"type": "boolean"
},
"totalAvailable": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Size of the external NIST StRD nonlinear-regression universe (the denominator for 'N of M' coverage). Emitted by the validation builder so the UI never hardcodes the total. Additive \u2014 None for payloads written before this field existed.",
"title": "Totalavailable"
}
},
"required": [
"thresholdSigFigs",
"datasets",
"minSigFigs",
"passed"
],
"title": "NistValidation",
"type": "object"
}
Fields:
-
threshold_sig_figs(float) -
datasets(list[NistDataset]) -
min_sig_figs(float) -
passed(bool) -
total_available(int | None)
threshold_sig_figs
pydantic-field
¶
Required minimum significant-figure agreement.
min_sig_figs
pydantic-field
¶
Worst min_sig_figs across all datasets.
passed
pydantic-field
¶
True iff every dataset agrees to \(\ge\) threshold sig figs.
total_available = None
pydantic-field
¶
Size of the external NIST StRD nonlinear-regression universe (the denominator for 'N of M' coverage). Emitted by the validation builder so the UI never hardcodes the total. Additive — None for payloads written before this field existed.
TrustBlock
pydantic-model
¶
Bases: _Contract
Aggregate trust evidence attached to a BenchReport.
Show JSON schema:
{
"$defs": {
"CredibilityRung": {
"description": "V&V maturity ladder (inspired by ASME V&V credibility levels).\n\nThe live wire set earns RUNG_2..RUNG_4. RUNG_1 / RUNG_5 are reserved\nend-stops (see the per-member comments), part of the public contract and\npinned by an enum-value test \u2014 reserved, not dead.",
"enum": [
1,
2,
3,
4,
5
],
"title": "CredibilityRung",
"type": "integer"
},
"NistDataset": {
"additionalProperties": false,
"description": "Per-dataset NIST StRD certified-value validation result.",
"properties": {
"name": {
"description": "StRD problem name, e.g. 'Gauss1'.",
"title": "Name",
"type": "string"
},
"model": {
"description": "Human-readable model description.",
"title": "Model",
"type": "string"
},
"nParams": {
"title": "Nparams",
"type": "integer"
},
"params": {
"items": {
"$ref": "#/$defs/NistParam"
},
"title": "Params",
"type": "array"
},
"minSigFigs": {
"description": "Worst (minimum) per-parameter sig-fig agreement.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff min_sig_figs $\\ge$ the validation threshold.",
"title": "Passed",
"type": "boolean"
}
},
"required": [
"name",
"model",
"nParams",
"params",
"minSigFigs",
"passed"
],
"title": "NistDataset",
"type": "object"
},
"NistParam": {
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
},
"NistValidation": {
"additionalProperties": false,
"description": "Aggregate NIST StRD certified-value validation (the W8 evidence block).\n\nIndependent external replication against NIST's extended-precision certified\nvalues \u2014 the evidence RUNG_5 was reserved for. Additive on TrustBlock.",
"properties": {
"thresholdSigFigs": {
"description": "Required minimum significant-figure agreement.",
"title": "Thresholdsigfigs",
"type": "number"
},
"datasets": {
"items": {
"$ref": "#/$defs/NistDataset"
},
"title": "Datasets",
"type": "array"
},
"minSigFigs": {
"description": "Worst min_sig_figs across all datasets.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff every dataset agrees to $\\ge$ threshold sig figs.",
"title": "Passed",
"type": "boolean"
},
"totalAvailable": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Size of the external NIST StRD nonlinear-regression universe (the denominator for 'N of M' coverage). Emitted by the validation builder so the UI never hardcodes the total. Additive \u2014 None for payloads written before this field existed.",
"title": "Totalavailable"
}
},
"required": [
"thresholdSigFigs",
"datasets",
"minSigFigs",
"passed"
],
"title": "NistValidation",
"type": "object"
},
"WireResult": {
"additionalProperties": false,
"description": "One verification wire's outcome.\n\nNote:\n ``status=\"gap\"`` is distinct from ``\"fail\"``: it means the property\n could not be verified because the backend does NOT expose the input\n (a disclosed CAPABILITY gap, e.g. $\\kappa(J)$ is not surfaced by\n spectrafit/lmfit/jax), not because a computed value was wrong. Like\n ``\"skipped\"``, a ``\"gap\"`` does not cap the credibility rung; only a\n genuine ``\"fail\"`` does (see ``runner._compute_rung``).",
"properties": {
"wireId": {
"description": "Audit wire id: W1, W2a-W2d, or W3..W11 from the value-stream diagram.",
"title": "Wireid",
"type": "string"
},
"name": {
"description": "Short machine-readable wire name (e.g. `synth_invariants`).",
"title": "Name",
"type": "string"
},
"status": {
"description": "Wire outcome: `pass`, `warn`, or `fail` for a computed verdict; `skipped` when the underlying test never ran; `gap` for a disclosed capability gap (see the class Note).",
"enum": [
"pass",
"warn",
"fail",
"skipped",
"gap"
],
"title": "Status",
"type": "string"
},
"evidence": {
"description": "One-line evidence statement.",
"title": "Evidence",
"type": "string"
},
"details": {
"additionalProperties": {
"anyOf": [
{
"type": "number"
},
{
"type": "integer"
},
{
"type": "string"
},
{
"type": "boolean"
},
{
"type": "null"
}
]
},
"description": "Optional structured detail (counts, thresholds, sample sizes, ...) supporting `evidence`; empty when the one-line statement is self-sufficient.",
"title": "Details",
"type": "object"
}
},
"required": [
"wireId",
"name",
"status",
"evidence"
],
"title": "WireResult",
"type": "object"
}
},
"additionalProperties": false,
"description": "Aggregate trust evidence attached to a BenchReport.",
"properties": {
"rung": {
"$ref": "#/$defs/CredibilityRung",
"description": "Aggregate V&V credibility rung (see `CredibilityRung`), capped by the worst genuine `fail` among `wires` \u2014 a `gap` or `skipped` wire never lowers it."
},
"wires": {
"description": "Every verification wire's outcome making up this ledger.",
"items": {
"$ref": "#/$defs/WireResult"
},
"title": "Wires",
"type": "array"
},
"nClaimsAudited": {
"description": "Number of report claims this audit run actually checked.",
"title": "Nclaimsaudited",
"type": "integer"
},
"nClaimsTotal": {
"description": "Total number of claims the report makes, whether or not they were audited this run.",
"title": "Nclaimstotal",
"type": "integer"
},
"nistValidation": {
"anyOf": [
{
"$ref": "#/$defs/NistValidation"
},
{
"type": "null"
}
],
"default": null,
"description": "NIST StRD certified-value validation (W8). Additive \u2014 Pydantic fills None for payloads that predate the A7 external-validation wire."
}
},
"required": [
"rung",
"wires",
"nClaimsAudited",
"nClaimsTotal"
],
"title": "TrustBlock",
"type": "object"
}
Fields:
-
rung(CredibilityRung) -
wires(list[WireResult]) -
n_claims_audited(int) -
n_claims_total(int) -
nist_validation(NistValidation | None)
Validators:
-
_rung5_requires_nist_and_inference
rung
pydantic-field
¶
Aggregate V&V credibility rung (see CredibilityRung), capped by the worst genuine fail among wires — a gap or skipped wire never lowers it.
wires
pydantic-field
¶
Every verification wire's outcome making up this ledger.
n_claims_audited
pydantic-field
¶
Number of report claims this audit run actually checked.
n_claims_total
pydantic-field
¶
Total number of claims the report makes, whether or not they were audited this run.
nist_validation = None
pydantic-field
¶
NIST StRD certified-value validation (W8). Additive — Pydantic fills None for payloads that predate the A7 external-validation wire.
TrustLedger
pydantic-model
¶
Bases: BaseModel
Persisted ledger written to trust.json.
Show JSON schema:
{
"$defs": {
"CredibilityRung": {
"description": "V&V maturity ladder (inspired by ASME V&V credibility levels).\n\nThe live wire set earns RUNG_2..RUNG_4. RUNG_1 / RUNG_5 are reserved\nend-stops (see the per-member comments), part of the public contract and\npinned by an enum-value test \u2014 reserved, not dead.",
"enum": [
1,
2,
3,
4,
5
],
"title": "CredibilityRung",
"type": "integer"
},
"NistDataset": {
"additionalProperties": false,
"description": "Per-dataset NIST StRD certified-value validation result.",
"properties": {
"name": {
"description": "StRD problem name, e.g. 'Gauss1'.",
"title": "Name",
"type": "string"
},
"model": {
"description": "Human-readable model description.",
"title": "Model",
"type": "string"
},
"nParams": {
"title": "Nparams",
"type": "integer"
},
"params": {
"items": {
"$ref": "#/$defs/NistParam"
},
"title": "Params",
"type": "array"
},
"minSigFigs": {
"description": "Worst (minimum) per-parameter sig-fig agreement.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff min_sig_figs $\\ge$ the validation threshold.",
"title": "Passed",
"type": "boolean"
}
},
"required": [
"name",
"model",
"nParams",
"params",
"minSigFigs",
"passed"
],
"title": "NistDataset",
"type": "object"
},
"NistParam": {
"additionalProperties": false,
"description": "One parameter's certified-vs-fitted agreement for a NIST StRD dataset.",
"properties": {
"name": {
"description": "NIST parameter name, e.g. 'b1'.",
"title": "Name",
"type": "string"
},
"certified": {
"description": "NIST certified value (10+ sig figs).",
"title": "Certified",
"type": "number"
},
"fitted": {
"description": "Spectrafit recovered value.",
"title": "Fitted",
"type": "number"
},
"sigFigsAgreed": {
"description": "$-\\log_{10}\\left(|\\mathrm{fitted}-\\mathrm{certified}|/|\\mathrm{certified}|\\right)$; $\\infty$-capped for exact agreement.",
"title": "Sigfigsagreed",
"type": "number"
}
},
"required": [
"name",
"certified",
"fitted",
"sigFigsAgreed"
],
"title": "NistParam",
"type": "object"
},
"NistValidation": {
"additionalProperties": false,
"description": "Aggregate NIST StRD certified-value validation (the W8 evidence block).\n\nIndependent external replication against NIST's extended-precision certified\nvalues \u2014 the evidence RUNG_5 was reserved for. Additive on TrustBlock.",
"properties": {
"thresholdSigFigs": {
"description": "Required minimum significant-figure agreement.",
"title": "Thresholdsigfigs",
"type": "number"
},
"datasets": {
"items": {
"$ref": "#/$defs/NistDataset"
},
"title": "Datasets",
"type": "array"
},
"minSigFigs": {
"description": "Worst min_sig_figs across all datasets.",
"title": "Minsigfigs",
"type": "number"
},
"passed": {
"description": "True iff every dataset agrees to $\\ge$ threshold sig figs.",
"title": "Passed",
"type": "boolean"
},
"totalAvailable": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"default": null,
"description": "Size of the external NIST StRD nonlinear-regression universe (the denominator for 'N of M' coverage). Emitted by the validation builder so the UI never hardcodes the total. Additive \u2014 None for payloads written before this field existed.",
"title": "Totalavailable"
}
},
"required": [
"thresholdSigFigs",
"datasets",
"minSigFigs",
"passed"
],
"title": "NistValidation",
"type": "object"
},
"TrustBlock": {
"additionalProperties": false,
"description": "Aggregate trust evidence attached to a BenchReport.",
"properties": {
"rung": {
"$ref": "#/$defs/CredibilityRung",
"description": "Aggregate V&V credibility rung (see `CredibilityRung`), capped by the worst genuine `fail` among `wires` \u2014 a `gap` or `skipped` wire never lowers it."
},
"wires": {
"description": "Every verification wire's outcome making up this ledger.",
"items": {
"$ref": "#/$defs/WireResult"
},
"title": "Wires",
"type": "array"
},
"nClaimsAudited": {
"description": "Number of report claims this audit run actually checked.",
"title": "Nclaimsaudited",
"type": "integer"
},
"nClaimsTotal": {
"description": "Total number of claims the report makes, whether or not they were audited this run.",
"title": "Nclaimstotal",
"type": "integer"
},
"nistValidation": {
"anyOf": [
{
"$ref": "#/$defs/NistValidation"
},
{
"type": "null"
}
],
"default": null,
"description": "NIST StRD certified-value validation (W8). Additive \u2014 Pydantic fills None for payloads that predate the A7 external-validation wire."
}
},
"required": [
"rung",
"wires",
"nClaimsAudited",
"nClaimsTotal"
],
"title": "TrustBlock",
"type": "object"
},
"WireResult": {
"additionalProperties": false,
"description": "One verification wire's outcome.\n\nNote:\n ``status=\"gap\"`` is distinct from ``\"fail\"``: it means the property\n could not be verified because the backend does NOT expose the input\n (a disclosed CAPABILITY gap, e.g. $\\kappa(J)$ is not surfaced by\n spectrafit/lmfit/jax), not because a computed value was wrong. Like\n ``\"skipped\"``, a ``\"gap\"`` does not cap the credibility rung; only a\n genuine ``\"fail\"`` does (see ``runner._compute_rung``).",
"properties": {
"wireId": {
"description": "Audit wire id: W1, W2a-W2d, or W3..W11 from the value-stream diagram.",
"title": "Wireid",
"type": "string"
},
"name": {
"description": "Short machine-readable wire name (e.g. `synth_invariants`).",
"title": "Name",
"type": "string"
},
"status": {
"description": "Wire outcome: `pass`, `warn`, or `fail` for a computed verdict; `skipped` when the underlying test never ran; `gap` for a disclosed capability gap (see the class Note).",
"enum": [
"pass",
"warn",
"fail",
"skipped",
"gap"
],
"title": "Status",
"type": "string"
},
"evidence": {
"description": "One-line evidence statement.",
"title": "Evidence",
"type": "string"
},
"details": {
"additionalProperties": {
"anyOf": [
{
"type": "number"
},
{
"type": "integer"
},
{
"type": "string"
},
{
"type": "boolean"
},
{
"type": "null"
}
]
},
"description": "Optional structured detail (counts, thresholds, sample sizes, ...) supporting `evidence`; empty when the one-line statement is self-sufficient.",
"title": "Details",
"type": "object"
}
},
"required": [
"wireId",
"name",
"status",
"evidence"
],
"title": "WireResult",
"type": "object"
}
},
"additionalProperties": false,
"description": "Persisted ledger written to ``trust.json``.",
"properties": {
"schema_version": {
"default": "1.0",
"description": "Ledger schema version, independent of the BenchReport schema.",
"title": "Schema Version",
"type": "string"
},
"run_id": {
"description": "The benchmark run id this ledger belongs to.",
"title": "Run Id",
"type": "string"
},
"block": {
"$ref": "#/$defs/TrustBlock",
"description": "The aggregate trust evidence itself."
}
},
"required": [
"run_id",
"block"
],
"title": "TrustLedger",
"type": "object"
}
Config:
extra:forbid
Fields:
-
schema_version(str) -
run_id(str) -
block(TrustBlock)
Statistics¶
Four modules that turn raw fit outcomes into the report's statistical claims.
oracles.metrics computes the frozen contract's distributional sub-objects
(timing, accuracy) from real repetitions and Monte-Carlo realizations —
nothing fabricated. oracles.inference holds the pure inferential-statistics
functions (confidence intervals, equivalence tests, stability scores) so
every comparison in the report carries its own uncertainty. oracles.nested
covers nested-model selection statistics for model-adequacy verification and
validation. oracles.stability aggregates a reps-ladder of benchmark runs
into per-backend convergence data — at which timing-repetition budget the
headline numbers stop moving.
oracles.metrics
¶
Statistics that turn fit outcomes into the frozen report contract sub-objects.
Everything here is computed from real measurements (timing repetitions, Monte-Carlo noise realizations, covariance at the solution); nothing is fabricated. The 1/\(\sqrt{n}\) stability projections are derived from the measured spread, not invented.
Note
The canonical quantile vectors (_TIMING_QUANTILES / _ACCURACY_QUANTILES)
are pinned by the frozen BenchReport contract: TimingDist exposes
{median, p5, p25, p75, p95} and AccuracyDist exposes
{median, p5, p25, p75}. The 0.5 slot feeds the median field, not a
(non-existent) p50 field — neither contract model declares p50.
Keeping the quantiles named (instead of inline literals) makes the
contract<->implementation link explicit and prevents an accidental tuple
drift between the two helpers. Unpack order at the call sites:
p5, p25, med, p75[, p95].
pcts(xs, qs)
¶
Linear-interpolated percentiles of xs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
xs
|
list[float]
|
Sample values. |
required |
qs
|
tuple[float, ...]
|
Quantiles to compute, each in |
required |
Returns:
| Type | Description |
|---|---|
list[float]
|
One interpolated value per entry in |
timing_dist(ms)
¶
Build a :class:TimingDist from raw per-rep milliseconds.
accuracy_dist(red_chi2_samples)
¶
Build an :class:AccuracyDist from reduced-\(\chi^2\) Monte-Carlo samples.
ecdf(values)
¶
Empirical CDF of values as (x, y) points (y in [0,1]); empty → [].
cov_to_corr(cov)
¶
Convert a covariance matrix to a correlation matrix (zeros if unavailable).
Note
_CORR_DIAG_FLOOR is the numerical floor applied to the covariance
diagonal before the sqrt that normalises off-diagonals into a
correlation matrix. Anything smaller would produce NaN/inf in the
division; the value matches the historical inline literal so the
pre-refactor JSON output is preserved byte-for-byte.
spread_vs_runs(samples, runs_sched)
¶
Mean and sd of samples subsampled per run count.
Exhibits real \(1/\sqrt{n}\) scaling. The sample sd needs \(\ge\)2 points; for
a single-point subsample sd is 0.0 (never nan from ddof=1 dividing by zero).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
list[float]
|
The full pool of measurements to subsample from. |
required |
runs_sched
|
list[int]
|
Cumulative-run counts to evaluate; each is clamped to
|
required |
Returns:
| Name | Type | Description |
|---|---|---|
One |
list[SpreadPt]
|
class: |
amortization_curve(cold_ms, hot_ms, schedule)
¶
Per-run time as the one-off cold cost amortizes over cumulative runs.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cold_ms
|
float
|
One-off cost paid on the first run (e.g. JIT warm-up). |
required |
hot_ms
|
float
|
Steady-state per-run cost paid on every subsequent run. |
required |
schedule
|
list[int]
|
Cumulative run counts to evaluate. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
One |
list[Point2D]
|
class: |
list[Point2D]
|
count |
|
list[Point2D]
|
|
pulls_from_mc(estimates, stderrs, true_params)
¶
Pull statistic and coverage across MC fits.
Pull = \((\mathrm{estimate}-\mathrm{truth})/\sigma\); coverage = fraction with |pull| < 1.
Skips MC samples where \(\sigma\) is None, \(\le\) 0, or the estimate key is absent
(see :func:_iter_valid_pulls for the skip rules).
The returned sigma vector is the last-seen \(\sigma\) per parameter — each MC
iteration that contributes a sample overwrites the previous \(\sigma\) for that key,
so the final entry is whichever sample happened to be processed last (output
is then sorted alphabetically by key). This is a diagnostic field, not a
statistical summary; consumers wanting mean/median \(\sigma\) should reduce across MC
repetitions themselves. The "last" choice is preserved verbatim from the
pre-refactor implementation so existing report-payload tests pin the same
value; revisiting whether this should instead be mean(sigma) or
median(sigma) is tracked as a separate Plan C3 concern and is
intentionally out of scope here.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
estimates
|
list[dict[str, float]]
|
Per-MC-iteration dicts of fitted parameter values. |
required |
stderrs
|
list[dict[str, float | None]]
|
Per-MC-iteration dicts of fitted parameter standard errors
(paired with |
required |
true_params
|
dict[str, float]
|
The ground-truth value for each parameter of interest. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
An |
Uncertainty
|
class: |
Uncertainty
|
the coverage fraction ( |
|
Uncertainty
|
last-seen \(\sigma\) per parameter sorted alphabetically by key |
|
Uncertainty
|
( |
Note
coverage is None when no valid \(\sigma\) was available (empty
pulls means every \(\sigma\) was None / non-positive). A genuine "0%
within \(1\sigma\)" failure requires at least one valid pull that
happened to be \(|p| \ge 1\); that produces 0.0, not None.
Downstream consumers (audit W2b, web renders) must treat None as
"\(\sigma\) not reported" rather than "0% coverage".
r2_of(y, fit)
¶
Coefficient of determination.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
y
|
Array
|
Observed data. |
required |
fit
|
Array
|
Model evaluated at the fitted parameters. |
required |
Returns:
| Type | Description |
|---|---|
float
|
\(R^2 = 1 - \mathrm{SS}_{\mathrm{res}}/\mathrm{SS}_{\mathrm{tot}}\); |
float
|
|
rmse_of(y, fit)
¶
Root-mean-square error. Canonical recomputation oracle for wire W2a.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
y
|
Array
|
Observed data. |
required |
fit
|
Array
|
Model evaluated at the fitted parameters. |
required |
Returns:
| Type | Description |
|---|---|
float
|
The root-mean-square residual between |
chi2_red_of(y, fit, sigma, dof)
¶
Reduced \(\chi^2\); unweighted if sigma is None.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
y
|
Array
|
Observed data. |
required |
fit
|
Array
|
Model evaluated at the fitted parameters. |
required |
sigma
|
Array | None
|
Per-point uncertainty, or |
required |
dof
|
int
|
Degrees of freedom to divide by (floored at 1). |
required |
Returns:
| Type | Description |
|---|---|
float
|
\(\chi^2_{\nu} = \sum w_i (y_i - \mathrm{fit}_i)^2 / \max(\mathrm{dof}, 1)\), |
float
|
with \(w_i = 1/\sigma_i^2\) (or \(1\) when |
oracles.inference
¶
Inferential statistics for the trustworthy benchmark.
Pure functions over plain sequences — no engine coupling. Every comparison the report makes carries an interval, an equivalence verdict, or a stability score, so the numbers state their own uncertainty rather than asserting it.
EquivalenceVerdict
pydantic-model
¶
Bases: BaseModel
Result of a two-one-sided-tests (TOST) equivalence test.
Show JSON schema:
{
"additionalProperties": false,
"description": "Result of a two-one-sided-tests (TOST) equivalence test.",
"properties": {
"equivalent": {
"title": "Equivalent",
"type": "boolean"
},
"margin": {
"title": "Margin",
"type": "number"
},
"diff": {
"title": "Diff",
"type": "number"
},
"p_lower": {
"title": "P Lower",
"type": "number"
},
"p_upper": {
"title": "P Upper",
"type": "number"
}
},
"required": [
"equivalent",
"margin",
"diff",
"p_lower",
"p_upper"
],
"title": "EquivalenceVerdict",
"type": "object"
}
Config:
extra:forbid
Fields:
-
equivalent(bool) -
margin(float) -
diff(float) -
p_lower(float) -
p_upper(float)
CalibrationStat
pydantic-model
¶
Bases: BaseModel
Result of a binomial calibration test on parameter pulls.
Show JSON schema:
{
"additionalProperties": false,
"description": "Result of a binomial calibration test on parameter pulls.",
"properties": {
"n": {
"title": "N",
"type": "integer"
},
"coverage": {
"title": "Coverage",
"type": "number"
},
"coverage_ci_lo": {
"title": "Coverage Ci Lo",
"type": "number"
},
"coverage_ci_hi": {
"title": "Coverage Ci Hi",
"type": "number"
},
"nominal": {
"title": "Nominal",
"type": "number"
},
"binomial_p": {
"title": "Binomial P",
"type": "number"
},
"ks_stat": {
"title": "Ks Stat",
"type": "number"
},
"ks_p": {
"title": "Ks P",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"equivalence_margin": {
"title": "Equivalence Margin",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"n",
"coverage",
"coverage_ci_lo",
"coverage_ci_hi",
"nominal",
"binomial_p",
"ks_stat",
"ks_p",
"alpha",
"equivalence_margin",
"passed",
"skipped"
],
"title": "CalibrationStat",
"type": "object"
}
Config:
extra:forbid
Fields:
-
n(int) -
coverage(float) -
coverage_ci_lo(float) -
coverage_ci_hi(float) -
nominal(float) -
binomial_p(float) -
ks_stat(float) -
ks_p(float) -
alpha(float) -
equivalence_margin(float) -
passed(bool) -
skipped(bool)
SpeedStat
pydantic-model
¶
Bases: BaseModel
Result of a geomean speedup significance test.
Show JSON schema:
{
"additionalProperties": false,
"description": "Result of a geomean speedup significance test.",
"properties": {
"geomean_speedup": {
"title": "Geomean Speedup",
"type": "number"
},
"ci_lo": {
"title": "Ci Lo",
"type": "number"
},
"ci_hi": {
"title": "Ci Hi",
"type": "number"
},
"excludes_one": {
"title": "Excludes One",
"type": "boolean"
},
"sign_p": {
"title": "Sign P",
"type": "number"
},
"wilcoxon_p": {
"title": "Wilcoxon P",
"type": "number"
},
"alpha": {
"title": "Alpha",
"type": "number"
},
"passed": {
"title": "Passed",
"type": "boolean"
},
"skipped": {
"title": "Skipped",
"type": "boolean"
}
},
"required": [
"geomean_speedup",
"ci_lo",
"ci_hi",
"excludes_one",
"sign_p",
"wilcoxon_p",
"alpha",
"passed",
"skipped"
],
"title": "SpeedStat",
"type": "object"
}
Config:
extra:forbid
Fields:
-
geomean_speedup(float) -
ci_lo(float) -
ci_hi(float) -
excludes_one(bool) -
sign_p(float) -
wilcoxon_p(float) -
alpha(float) -
passed(bool) -
skipped(bool)
bootstrap_ci(samples, *, stat, b, alpha, seed)
¶
Percentile-bootstrap (1-alpha) CI of stat over samples.
Single-sample input returns a degenerate point interval (lo == hi); a bootstrap over one value has no spread and must not raise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
Sequence[float]
|
The observed sample to resample from. |
required |
stat
|
Callable[[ndarray], float]
|
Statistic to compute on each resample (e.g. |
required |
b
|
int
|
Number of bootstrap resamples. |
required |
alpha
|
float
|
Significance level; the CI covers |
required |
seed
|
int
|
Random seed for reproducibility. |
required |
Returns:
| Type | Description |
|---|---|
CI
|
The |
speedup_ci(baseline_ms, subject_ms, *, b, alpha, seed)
¶
(lo, point, hi) CI for speedup = median(baseline)/median(subject).
Resamples each leg independently (the reps are unpaired across backends) and bootstraps the ratio of medians.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
baseline_ms
|
Sequence[float]
|
Per-repetition timings of the baseline backend. |
required |
subject_ms
|
Sequence[float]
|
Per-repetition timings of the subject backend. |
required |
b
|
int
|
Number of bootstrap resamples. |
required |
alpha
|
float
|
Significance level; the CI covers |
required |
seed
|
int
|
Random seed for reproducibility. |
required |
Returns:
| Type | Description |
|---|---|
float
|
|
float
|
|
float
|
|
delta_r2_ci(r2_a, r2_b, *, b_resamples, alpha, seed)
¶
(lo, point, hi) CI for mean(r2_a) - mean(r2_b).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
r2_a
|
Sequence[float]
|
Per-case \(R^2\) samples for backend A. |
required |
r2_b
|
Sequence[float]
|
Per-case \(R^2\) samples for backend B. |
required |
b_resamples
|
int
|
Number of bootstrap resamples. |
required |
alpha
|
float
|
Significance level; the CI covers |
required |
seed
|
int
|
Random seed for reproducibility. |
required |
Returns:
| Type | Description |
|---|---|
float
|
|
float
|
|
float
|
when either sample has fewer than 2 values. |
tost_equivalence(a, b, *, margin, alpha)
¶
TOST: are mean(a), mean(b) equivalent within \(\pm\)margin?
Two one-sided Welch t-tests; equivalent iff BOTH reject at \(\alpha\) (the diff's \((1-2\alpha)\) CI lies inside [-margin, +margin]). 'Saturated' is then a positive equivalence claim, not an unthresholded r2>0.999.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
a
|
Sequence[float]
|
First sample. |
required |
b
|
Sequence[float]
|
Second sample. |
required |
margin
|
float
|
Equivalence half-width \(\pm\)margin. |
required |
alpha
|
float
|
Significance level for each one-sided test. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
An |
EquivalenceVerdict
|
class: |
EquivalenceVerdict
|
one-sided p-values, and the equivalence verdict. |
tost_paired(deltas, *, margin, alpha)
¶
One-sample TOST on PAIRED differences: is mean(deltas) within \(\pm\)margin?
Use for paired comparisons — e.g. per-case \(\Delta r^2\) between two backends measured on the same case. An unpaired two-sample test would wrongly inflate the standard error with case-to-case spread and mask a real equivalence; the paired test's variance is the spread of the per-case differences, which is tiny when the backends agree case-by-case. Equivalent iff the \((1-2\alpha)\) CI of mean(deltas) lies inside [-margin, +margin].
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
deltas
|
Sequence[float]
|
Per-case paired differences. |
required |
margin
|
float
|
Equivalence half-width \(\pm\)margin. |
required |
alpha
|
float
|
Significance level for each one-sided test. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
An |
EquivalenceVerdict
|
class: |
EquivalenceVerdict
|
one-sided p-values, and the equivalence verdict. |
winner_stability(scores, *, b, seed)
¶
Fraction of bootstrap resamples in which each backend is the winner.
Resamples the per-case score vectors (paired across backends by case index) and records the argmax-mean each round. A low max stability ⇒ no robust winner — which the report must say plainly rather than crowning noise.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scores
|
Mapping[str, Sequence[float]]
|
Per-backend sequences of per-case scores, paired by case index. |
required |
b
|
int
|
Number of bootstrap resamples. |
required |
seed
|
int
|
Random seed for reproducibility. |
required |
Returns:
| Type | Description |
|---|---|
dict[str, float]
|
A dict mapping each backend key to its win fraction across the |
dict[str, float]
|
resamples (values sum to 1.0). |
bh_correct(pvalues, *, q)
¶
Benjamini-Hochberg FDR procedure.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pvalues
|
Sequence[float]
|
Raw p-values to correct. |
required |
q
|
float
|
Target false-discovery rate. |
required |
Returns:
| Type | Description |
|---|---|
list[bool]
|
|
list[float]
|
BH-adjusted q-values, both in the same order as |
coverage_test(pulls, *, nominal=0.6827, alpha=0.025, min_pulls=20, equivalence_margin=0.03)
¶
Test if parameter pulls indicate well-calibrated uncertainties.
A parameter pull is \((\theta_{\mathrm{est}} - \theta_{\mathrm{true}}) / \sigma_{\mathrm{est}}\). In well-calibrated fits, pulls \(\sim \mathcal{N}(0, 1)\), so \(\approx 68.27\%\) should fall within \([-1, 1]\).
The gating verdict (passed) uses a practical-equivalence rule: the
Clopper–Pearson CI of the empirical coverage must lie entirely within
[nominal - equivalence_margin, nominal + equivalence_margin]. This is
the standard CI-inclusion TOST for a one-sample proportion — it avoids the
classic statistical-vs-practical-significance trap at large n where a
point-null binomial test rejects a negligible (< 2 pp) deviation.
The strict point-null binomial_p is always computed and retained as an
honest diagnostic (it appears in CalibrationStat.binomial_p and in the
audit wire evidence string) so the dashboard can still report "coverage
0.668 vs nominal 0.6827, strict p < 0.001" — only the binary gate changes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pulls
|
Sequence[float]
|
Sequence of parameter pulls. |
required |
nominal
|
float
|
Target coverage (default 68.27%, the 1-\(\sigma\) interval). |
0.6827
|
alpha
|
float
|
Significance level for the Clopper-Pearson CI (default 0.025, giving a 97.5% CI used for the equivalence inclusion test). |
0.025
|
min_pulls
|
int
|
Minimum pulls to avoid skipping (default 20). |
20
|
equivalence_margin
|
float
|
Half-width of the pre-registered equivalence band
around |
0.03
|
Returns:
| Type | Description |
|---|---|
CalibrationStat
|
CalibrationStat with coverage point estimate, Clopper-Pearson CI, |
CalibrationStat
|
binomial p-value (honest strict diagnostic), KS test vs N(0,1), |
CalibrationStat
|
equivalence_margin, and pass/skip verdict. |
CalibrationStat
|
If n < min_pulls, skipped=True and passed=False (insufficient data). |
geomean_speedup_test(per_case_speedups, *, b=2000, seed=20260612, alpha=0.025)
¶
Test if subject is significantly faster than baseline using geomean speedup.
The test computes a bootstrap CI for the geometric mean speedup over cases, then validates against \(1\times\) with a sign test and Wilcoxon signed-rank test. Pass criterion: CI lower bound > 1.0 (excludes_one=True).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
per_case_speedups
|
Sequence[float]
|
Sequence of speedup ratios (baseline_ms / subject_ms), one per case. NaN and non-positive values are filtered. |
required |
b
|
int
|
Bootstrap resamples (default 2000). |
2000
|
seed
|
int
|
Random seed for reproducibility (default 20260612). |
20260612
|
alpha
|
float
|
Significance level for CI (default 0.025). |
0.025
|
Returns:
| Type | Description |
|---|---|
SpeedStat
|
SpeedStat with geomean, CI, exclusion verdict, p-values, and pass/skip. |
SpeedStat
|
Empty or <2 valid inputs → skipped=True, passed=False. |
oracles.nested
¶
Nested-model selection statistics for model adequacy V&V.
SelectionStats
pydantic-model
¶
Bases: BaseModel
Selection statistics for nested model comparison.
Computed via LRT, F-test, and information criteria (AIC/BIC). All statistics are "full minus reduced" oriented: negative \(\Delta\)AIC/\(\Delta\)BIC means the full model is preferred.
Show JSON schema:
{
"additionalProperties": false,
"description": "Selection statistics for nested model comparison.\n\nComputed via LRT, F-test, and information criteria (AIC/BIC).\nAll statistics are \"full minus reduced\" oriented: negative $\\Delta$AIC/$\\Delta$BIC\nmeans the full model is preferred.",
"properties": {
"lrt_stat": {
"title": "Lrt Stat",
"type": "number"
},
"lrt_p": {
"title": "Lrt P",
"type": "number"
},
"f_stat": {
"title": "F Stat",
"type": "number"
},
"f_p": {
"title": "F P",
"type": "number"
},
"d_aic": {
"title": "D Aic",
"type": "number"
},
"d_bic": {
"title": "D Bic",
"type": "number"
}
},
"required": [
"lrt_stat",
"lrt_p",
"f_stat",
"f_p",
"d_aic",
"d_bic"
],
"title": "SelectionStats",
"type": "object"
}
Config:
extra:forbid
Fields:
-
lrt_stat(float) -
lrt_p(float) -
f_stat(float) -
f_p(float) -
d_aic(float) -
d_bic(float)
NestedAdequacy
pydantic-model
¶
Bases: BaseModel
Result of a nested-order V&V against a known generative order.
Compares the (true-1) reduced, true, and (true+1) over-fitted models using LRT/F/AIC/BIC to verify that model-selection criteria recover the known order.
Both AIC and BIC verdicts are reported separately so their agreement (or
disagreement) is visible. On the featured real tri-Gaussian, AIC over-selects
the 4-peak model (\(\Delta\)AIC = -1.12, i.e. over_not_preferred_aic = False) while
BIC correctly recovers the true order (\(\Delta\)BIC = +8.46,
over_not_preferred_bic = True). Exposing both criteria makes this known
AIC tendency observable rather than hidden behind a single collapsed flag.
Show JSON schema:
{
"$defs": {
"SelectionStats": {
"additionalProperties": false,
"description": "Selection statistics for nested model comparison.\n\nComputed via LRT, F-test, and information criteria (AIC/BIC).\nAll statistics are \"full minus reduced\" oriented: negative $\\Delta$AIC/$\\Delta$BIC\nmeans the full model is preferred.",
"properties": {
"lrt_stat": {
"title": "Lrt Stat",
"type": "number"
},
"lrt_p": {
"title": "Lrt P",
"type": "number"
},
"f_stat": {
"title": "F Stat",
"type": "number"
},
"f_p": {
"title": "F P",
"type": "number"
},
"d_aic": {
"title": "D Aic",
"type": "number"
},
"d_bic": {
"title": "D Bic",
"type": "number"
}
},
"required": [
"lrt_stat",
"lrt_p",
"f_stat",
"f_p",
"d_aic",
"d_bic"
],
"title": "SelectionStats",
"type": "object"
}
},
"additionalProperties": false,
"description": "Result of a nested-order V&V against a known generative order.\n\nCompares the (true-1) reduced, true, and (true+1) over-fitted models using\nLRT/F/AIC/BIC to verify that model-selection criteria recover the known order.\n\nBoth AIC and BIC verdicts are reported separately so their agreement (or\ndisagreement) is visible. On the featured real tri-Gaussian, AIC over-selects\nthe 4-peak model ($\\Delta$AIC = -1.12, i.e. ``over_not_preferred_aic = False``) while\nBIC correctly recovers the true order ($\\Delta$BIC = +8.46,\n``over_not_preferred_bic = True``). Exposing both criteria makes this known\nAIC tendency observable rather than hidden behind a single collapsed flag.",
"properties": {
"true_order": {
"title": "True Order",
"type": "integer"
},
"reduced_rejected": {
"title": "Reduced Rejected",
"type": "boolean"
},
"over_not_preferred_aic": {
"title": "Over Not Preferred Aic",
"type": "boolean"
},
"over_not_preferred_bic": {
"title": "Over Not Preferred Bic",
"type": "boolean"
},
"selected_order_aic": {
"title": "Selected Order Aic",
"type": "integer"
},
"selected_order_bic": {
"title": "Selected Order Bic",
"type": "integer"
},
"recovered_true_order_aic": {
"title": "Recovered True Order Aic",
"type": "boolean"
},
"recovered_true_order_bic": {
"title": "Recovered True Order Bic",
"type": "boolean"
},
"reduced_vs_true": {
"$ref": "#/$defs/SelectionStats"
},
"true_vs_over": {
"$ref": "#/$defs/SelectionStats"
}
},
"required": [
"true_order",
"reduced_rejected",
"over_not_preferred_aic",
"over_not_preferred_bic",
"selected_order_aic",
"selected_order_bic",
"recovered_true_order_aic",
"recovered_true_order_bic",
"reduced_vs_true",
"true_vs_over"
],
"title": "NestedAdequacy",
"type": "object"
}
Config:
extra:forbid
Fields:
-
true_order(int) -
reduced_rejected(bool) -
over_not_preferred_aic(bool) -
over_not_preferred_bic(bool) -
selected_order_aic(int) -
selected_order_bic(int) -
recovered_true_order_aic(bool) -
recovered_true_order_bic(bool) -
reduced_vs_true(SelectionStats) -
true_vs_over(SelectionStats)
reduced_rejected
pydantic-field
¶
LRT p-value < alpha; the reduced model is significantly worse than the true.
over_not_preferred_aic
pydantic-field
¶
true_vs_over.d_aic > 0 — the over-model is not preferred by AIC.
over_not_preferred_bic
pydantic-field
¶
BIC of the over-model > BIC of the true model — BIC does not prefer the over-model.
selected_order_aic
pydantic-field
¶
Order with the lowest absolute Gaussian-MLE AIC across the three orders.
selected_order_bic
pydantic-field
¶
Order with the lowest absolute Gaussian-MLE BIC across the three orders.
recovered_true_order_aic
pydantic-field
¶
True iff the criterion's argmin equals true_order.
Requires ALL of: reduced_rejected (LRT rejects the reduced model),
over_not_preferred_aic (AIC does not prefer the over-model), AND
selected_order_aic == true_order (AIC argmin lands on the true order).
All three conditions must hold; a subset is insufficient.
recovered_true_order_bic
pydantic-field
¶
True iff the criterion's argmin equals true_order.
Requires ALL of: reduced_rejected (LRT rejects the reduced model),
over_not_preferred_bic (BIC does not prefer the over-model), AND
selected_order_bic == true_order (BIC argmin lands on the true order).
All three conditions must hold; a subset is insufficient.
selection_stats(rss_reduced, rss_full, k_reduced, k_full, n)
¶
Compute nested-model selection statistics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rss_reduced
|
float
|
Residual sum of squares for the reduced model. |
required |
rss_full
|
float
|
Residual sum of squares for the full model. |
required |
k_reduced
|
int
|
Number of parameters in reduced model. |
required |
k_full
|
int
|
Number of parameters in full model. |
required |
n
|
int
|
Number of observations. |
required |
Returns:
| Type | Description |
|---|---|
SelectionStats
|
Object containing LRT, F-test, and information-criterion statistics. |
SelectionStats
|
All statistics are "full minus reduced" oriented: negative \(\Delta\)AIC/\(\Delta\)BIC |
SelectionStats
|
means the full model is preferred. |
nested_adequacy(true_order, fit_order, n, alpha=0.01)
¶
Verify nested-model selection criteria recover true generative order.
Fits three model orders (true_order-1, true_order, true_order+1) via the supplied callback and checks that the reduced model is rejected and the over-model is not preferred under BOTH AIC and BIC, thus recovering the true order.
AIC and BIC verdicts are reported separately. AIC tends to over-select on finite samples (it is not order-consistent); BIC is consistent. On the featured real tri-Gaussian, \(\Delta\)AIC = -1.12 (AIC prefers the 4-peak model) while \(\Delta\)BIC = +8.46 (BIC correctly recovers the true 3-peak order).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
true_order
|
int
|
Known generative model order (m*). |
required |
fit_order
|
Callable[[int], tuple[float, int]]
|
Callable |
required |
n
|
int
|
Number of observations. |
required |
alpha
|
float
|
Significance threshold for the LRT reduced-model rejection test. |
0.01
|
Returns:
| Type | Description |
|---|---|
NestedAdequacy
|
Pydantic model with per-criterion selection verdicts and recovery |
NestedAdequacy
|
flags. |
oracles.stability
¶
Reps-ladder stability study — variance-vs-N data for measurement-uncertainty reporting.
spc-bench stability <dir> aggregates a ladder of benchmark runs executed at
increasing --reps budgets (reps-1/ … reps-100/, each holding one
contract-valid results.json) into per-backend convergence data: how the
suite-level headline numbers (geomean speedup, median speedup, median per-case
runtime) move as the timing-repetition budget grows, and at which budget the
measurement stops moving.
This module is a CI artifact schema, NOT the frozen wire contract — the
:class:StabilityStudy model deliberately lives here and not in
oracles.bench_contract so the web/OpenAPI surface stays untouched. The consumer
is the benchmark:deep:merge GitLab job (.gitlab/55-deep-bench.yml) and,
downstream, measurement-uncertainty reporting.
Conventions: pydantic-first models, match/case dispatch, no per-call maps.
CONVERGENCE_RTOL = 0.02
module-attribute
¶
Relative tolerance for the "measurement converged" verdict: the smallest reps budget where EVERY backend's geomean speedup sits within this band of its reference-budget (highest-reps, normally N=100) value.
LadderRunFile
pydantic-model
¶
Bases: BaseModel
One discovered reps-N/results.json entry in a ladder directory.
Show JSON schema:
{
"additionalProperties": false,
"description": "One discovered ``reps-N/results.json`` entry in a ladder directory.",
"properties": {
"reps": {
"exclusiveMinimum": 0,
"title": "Reps",
"type": "integer"
},
"results_path": {
"format": "path",
"title": "Results Path",
"type": "string"
}
},
"required": [
"reps",
"results_path"
],
"title": "LadderRunFile",
"type": "object"
}
Config:
extra:forbid
Fields:
-
reps(int) -
results_path(Path)
LadderLayout
pydantic-model
¶
Bases: BaseModel
Validated discovery result for a reps-ladder directory.
runs requires at least one entry: pointing the stability command at a
directory with no reps-N subdirectories is a caller error and surfaces
as a pydantic ValidationError at this chokepoint rather than as an
empty study downstream.
Show JSON schema:
{
"$defs": {
"LadderRunFile": {
"additionalProperties": false,
"description": "One discovered ``reps-N/results.json`` entry in a ladder directory.",
"properties": {
"reps": {
"exclusiveMinimum": 0,
"title": "Reps",
"type": "integer"
},
"results_path": {
"format": "path",
"title": "Results Path",
"type": "string"
}
},
"required": [
"reps",
"results_path"
],
"title": "LadderRunFile",
"type": "object"
}
},
"additionalProperties": false,
"description": "Validated discovery result for a reps-ladder directory.\n\n``runs`` requires at least one entry: pointing the stability command at a\ndirectory with no ``reps-N`` subdirectories is a caller error and surfaces\nas a pydantic ``ValidationError`` at this chokepoint rather than as an\nempty study downstream.",
"properties": {
"root": {
"format": "path",
"title": "Root",
"type": "string"
},
"runs": {
"items": {
"$ref": "#/$defs/LadderRunFile"
},
"minItems": 1,
"title": "Runs",
"type": "array"
}
},
"required": [
"root",
"runs"
],
"title": "LadderLayout",
"type": "object"
}
Config:
extra:forbid
Fields:
-
root(Path) -
runs(list[LadderRunFile])
BackendHeadline
pydantic-model
¶
Bases: BaseModel
Suite-level headline numbers for one backend in one run.
Show JSON schema:
{
"additionalProperties": false,
"description": "Suite-level headline numbers for one backend in one run.",
"properties": {
"geomean_speedup": {
"title": "Geomean Speedup",
"type": "number"
},
"median_speedup": {
"title": "Median Speedup",
"type": "number"
},
"median_ms": {
"title": "Median Ms",
"type": "number"
}
},
"required": [
"geomean_speedup",
"median_speedup",
"median_ms"
],
"title": "BackendHeadline",
"type": "object"
}
Config:
extra:forbid
Fields:
-
geomean_speedup(float) -
median_speedup(float) -
median_ms(float)
StabilityPoint
pydantic-model
¶
Bases: BaseModel
One backend's headline numbers at one reps budget.
The rel_dev_* fields are the relative half-width vs the reference
(highest-reps) run: abs(v_N - v_ref) / v_ref.
Show JSON schema:
{
"additionalProperties": false,
"description": "One backend's headline numbers at one reps budget.\n\nThe ``rel_dev_*`` fields are the relative half-width vs the reference\n(highest-reps) run: ``abs(v_N - v_ref) / v_ref``.",
"properties": {
"reps": {
"title": "Reps",
"type": "integer"
},
"geomean_speedup": {
"title": "Geomean Speedup",
"type": "number"
},
"median_speedup": {
"title": "Median Speedup",
"type": "number"
},
"median_ms": {
"title": "Median Ms",
"type": "number"
},
"rel_dev_geomean": {
"title": "Rel Dev Geomean",
"type": "number"
},
"rel_dev_median_speedup": {
"title": "Rel Dev Median Speedup",
"type": "number"
},
"rel_dev_median_ms": {
"title": "Rel Dev Median Ms",
"type": "number"
}
},
"required": [
"reps",
"geomean_speedup",
"median_speedup",
"median_ms",
"rel_dev_geomean",
"rel_dev_median_speedup",
"rel_dev_median_ms"
],
"title": "StabilityPoint",
"type": "object"
}
Config:
extra:forbid
Fields:
-
reps(int) -
geomean_speedup(float) -
median_speedup(float) -
median_ms(float) -
rel_dev_geomean(float) -
rel_dev_median_speedup(float) -
rel_dev_median_ms(float)
BackendStability
pydantic-model
¶
Bases: BaseModel
Convergence trajectory (ascending reps) for one backend.
Show JSON schema:
{
"$defs": {
"StabilityPoint": {
"additionalProperties": false,
"description": "One backend's headline numbers at one reps budget.\n\nThe ``rel_dev_*`` fields are the relative half-width vs the reference\n(highest-reps) run: ``abs(v_N - v_ref) / v_ref``.",
"properties": {
"reps": {
"title": "Reps",
"type": "integer"
},
"geomean_speedup": {
"title": "Geomean Speedup",
"type": "number"
},
"median_speedup": {
"title": "Median Speedup",
"type": "number"
},
"median_ms": {
"title": "Median Ms",
"type": "number"
},
"rel_dev_geomean": {
"title": "Rel Dev Geomean",
"type": "number"
},
"rel_dev_median_speedup": {
"title": "Rel Dev Median Speedup",
"type": "number"
},
"rel_dev_median_ms": {
"title": "Rel Dev Median Ms",
"type": "number"
}
},
"required": [
"reps",
"geomean_speedup",
"median_speedup",
"median_ms",
"rel_dev_geomean",
"rel_dev_median_speedup",
"rel_dev_median_ms"
],
"title": "StabilityPoint",
"type": "object"
}
},
"additionalProperties": false,
"description": "Convergence trajectory (ascending reps) for one backend.",
"properties": {
"solver_id": {
"title": "Solver Id",
"type": "string"
},
"points": {
"items": {
"$ref": "#/$defs/StabilityPoint"
},
"title": "Points",
"type": "array"
}
},
"required": [
"solver_id",
"points"
],
"title": "BackendStability",
"type": "object"
}
Config:
extra:forbid
Fields:
-
solver_id(str) -
points(list[StabilityPoint])
StabilityStudy
pydantic-model
¶
Bases: BaseModel
The full reps-ladder study — serialized to stability.json by the CLI.
Show JSON schema:
{
"$defs": {
"BackendStability": {
"additionalProperties": false,
"description": "Convergence trajectory (ascending reps) for one backend.",
"properties": {
"solver_id": {
"title": "Solver Id",
"type": "string"
},
"points": {
"items": {
"$ref": "#/$defs/StabilityPoint"
},
"title": "Points",
"type": "array"
}
},
"required": [
"solver_id",
"points"
],
"title": "BackendStability",
"type": "object"
},
"StabilityPoint": {
"additionalProperties": false,
"description": "One backend's headline numbers at one reps budget.\n\nThe ``rel_dev_*`` fields are the relative half-width vs the reference\n(highest-reps) run: ``abs(v_N - v_ref) / v_ref``.",
"properties": {
"reps": {
"title": "Reps",
"type": "integer"
},
"geomean_speedup": {
"title": "Geomean Speedup",
"type": "number"
},
"median_speedup": {
"title": "Median Speedup",
"type": "number"
},
"median_ms": {
"title": "Median Ms",
"type": "number"
},
"rel_dev_geomean": {
"title": "Rel Dev Geomean",
"type": "number"
},
"rel_dev_median_speedup": {
"title": "Rel Dev Median Speedup",
"type": "number"
},
"rel_dev_median_ms": {
"title": "Rel Dev Median Ms",
"type": "number"
}
},
"required": [
"reps",
"geomean_speedup",
"median_speedup",
"median_ms",
"rel_dev_geomean",
"rel_dev_median_speedup",
"rel_dev_median_ms"
],
"title": "StabilityPoint",
"type": "object"
}
},
"additionalProperties": false,
"description": "The full reps-ladder study \u2014 serialized to ``stability.json`` by the CLI.",
"properties": {
"reference_reps": {
"title": "Reference Reps",
"type": "integer"
},
"reps_ladder": {
"items": {
"type": "integer"
},
"title": "Reps Ladder",
"type": "array"
},
"tolerance": {
"default": 0.02,
"title": "Tolerance",
"type": "number"
},
"converged_at_reps": {
"anyOf": [
{
"type": "integer"
},
{
"type": "null"
}
],
"title": "Converged At Reps"
},
"backends": {
"items": {
"$ref": "#/$defs/BackendStability"
},
"minItems": 1,
"title": "Backends",
"type": "array"
}
},
"required": [
"reference_reps",
"reps_ladder",
"converged_at_reps",
"backends"
],
"title": "StabilityStudy",
"type": "object"
}
Config:
extra:forbid
Fields:
-
reference_reps(int) -
reps_ladder(list[int]) -
tolerance(float) -
converged_at_reps(int | None) -
backends(list[BackendStability])
discover_ladder(root)
¶
Discover reps-N subdirectories under root, sorted by ascending reps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
root
|
Path
|
Directory expected to contain |
required |
Returns:
| Type | Description |
|---|---|
LadderLayout
|
A validated :class: |
Raises:
| Type | Description |
|---|---|
ValidationError
|
If root contains no |
backend_headline(report, solver_id)
¶
Extract one backend's suite-level headline numbers from a report.
Aggregates over every suite case where the backend reported (a backend may
legitimately be absent from some cases — e.g. jax skips optfn).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
report
|
BenchReport
|
A contract-valid benchmark report. |
required |
solver_id
|
str
|
The backend to summarize. |
required |
Returns:
| Type | Description |
|---|---|
BackendHeadline | None
|
The headline triple, or |
build_stability_study(root, *, tolerance=CONVERGENCE_RTOL)
¶
Aggregate a reps-ladder directory into a :class:StabilityStudy.
The reference run is the highest-reps rung of the ladder (N=100 in the CI matrix); every other rung is compared against it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
root
|
Path
|
Directory containing |
required |
tolerance
|
float
|
Relative band for the converged-at-N verdict. |
CONVERGENCE_RTOL
|
Returns:
| Type | Description |
|---|---|
StabilityStudy
|
The aggregated study (ready for |
Raises:
| Type | Description |
|---|---|
ValidationError
|
If root has no |
render_markdown(study)
¶
Render the study as a human-readable Markdown table + one-line verdict.
Rows are reps budgets, columns are backends; each cell shows the geomean speedup with its signed \(\Delta\%\) vs the reference (highest-reps) value.
verdict_line(study)
¶
One-line convergence verdict for the Markdown report and the CLI echo.
Model kernels¶
oracles.jax_kernels is a deliberately separate module: the numpy evaluate
bodies in oracles.models sit inside the timed fit loop of the lmfit and
scipy-ls benchmark backends and must never grow a jax branch or import cost,
so every peak-shape formula is re-expressed here in jax.numpy and pinned
against its numpy twin by a dedicated parity test.
oracles.jax_kernels
¶
JAX transcriptions of every numpy peak kernel in :mod:oracles.models.
WHY A SEPARATE MODULE. The numpy evaluate bodies in :mod:oracles.models
sit inside the timed fit loop of the lmfit and scipy-ls benchmark backends, so
they must never grow a jax branch, a dispatch, or an import cost — doing so
would give the timed backends an unfair overhead relative to the jax oracle
they are being compared against. The jax oracle needs the same formulas
expressed in jax.numpy so jax.grad/jax.jit can flow through them.
Two implementations of one formula is a real divergence hazard, so every kernel
here is pinned against its numpy twin by tests/unit/test_jax_kernel_parity.py,
which parametrises over the live registry (a new shape is covered the moment it
is registered).
WHY NOT UNDER oracles/backends/. :mod:oracles.models registers these
callables, and oracles.backends.__init__ imports _base, which imports
oracles.cases, which imports :mod:oracles.models — placing this module in
that package would make the registry's own import cycle back through the
backends layer. This module is a leaf: it imports nothing from oracles.
WHY THE IMPORTS ARE LAZY. jax is an optional extra. Every kernel body
imports jax.numpy locally, so import oracles.models (and therefore the
whole test suite and the lmfit/scipy benchmark path) keeps working in an
environment where jax is absent. The kernels are only called by the jax
backend, which already refuses to construct without jax installed.
Signatures mirror the numpy originals exactly — kernel(x, **param_names) in
canonical registry order — so :func:oracles.backends._jax._kernel can splat a
flat parameter segment by name with no per-shape knowledge.
Note
Branch-masking discipline. Wherever the numpy body relies on np.errstate
to let a NaN/inf appear in an unselected np.where branch, the jax
transcription masks the input to that branch instead (the "double-where"
trick). jnp.where selects the right value either way, but a NaN in the
discarded branch poisons the reverse-mode gradient optimistix computes.
gaussian(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.gaussian.
lorentzian(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.lorentzian.
pseudo_voigt(x, amplitude, center, sigma, fraction)
¶
JAX twin of :func:oracles.models.pseudo_voigt (fraction clipped to [0, 1]).
fano(x, amplitude, center, gamma, q)
¶
JAX twin of :func:oracles.models.fano.
constant(x, c)
¶
JAX twin of :func:oracles.models.constant.
jnp.zeros_like(x) + c rather than jnp.full_like(x, c): the latter
would treat a traced c as a fill value and drop it from the autodiff
graph, so the constant background would show a zero gradient.
linear(x, slope, intercept)
¶
JAX twin of :func:oracles.models.linear.
quadratic(x, amplitude, center, offset)
¶
JAX twin of :func:oracles.models.quadratic.
arctan_step(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.arctan_step.
tanh_step(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.tanh_step.
erfc_step(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.erfc_step.
double_exponential(x, A1, lam1, A2, lam2)
¶
JAX twin of :func:oracles.models.double_exponential (A1/A2 mirror it).
true_voigt(x, amplitude, center, sigma, gamma)
¶
JAX twin of :func:oracles.models.true_voigt.
jax.scipy.special.wofz is the same Faddeeva function scipy.special.wofz
provides (agreement to \(\approx 10^{-15}\)), and jax.grad flows through it,
so the true Voigt needs no reformulation — only a complex-dtype x64
context, which :class:~oracles.backends._jax.JaxBackend enables at import.
skewed_gaussian(x, amplitude, center, sigma, gamma)
¶
JAX twin of :func:oracles.models.skewed_gaussian.
exp_gaussian(x, amplitude, center, sigma, gamma)
¶
JAX twin of :func:oracles.models.exp_gaussian (overflow-free EMG).
Identical regime split to the numpy body — \(z \ge 0\) uses the erfcx form,
\(z < 0\) uses \(\exp(\mathrm{arg\_exp})\,\mathrm{erfc}(z)\) with arg_exp < 0 —
but jax.scipy.special has no erfcx, so :func:_erfcx supplies it, and
both branches take masked inputs (numpy instead suppresses the resulting
warnings with np.errstate).
doniach_sunjic(x, amplitude, center, sigma, gamma)
¶
JAX twin of :func:oracles.models.doniach_sunjic.
log_normal(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.log_normal (zero for x <= 0).
pearson7(x, amplitude, center, sigma, m)
¶
JAX twin of :func:oracles.models.pearson7.
split_gaussian(x, amplitude, center, sigma_l, sigma_r)
¶
JAX twin of :func:oracles.models.split_gaussian.
moffat(x, amplitude, center, sigma, beta)
¶
JAX twin of :func:oracles.models.moffat.
students_t(x, amplitude, center, sigma, nu)
¶
JAX twin of :func:oracles.models.students_t.
split_pearson7(x, amplitude, center, sigma_l, sigma_r, m_l, m_r)
¶
JAX twin of :func:oracles.models.split_pearson7.
breit_wigner(x, amplitude, center, sigma, q)
¶
JAX twin of :func:oracles.models.breit_wigner.
asym_ir(x, amplitude, center, sigma, k)
¶
JAX twin of :func:oracles.models.asym_ir (sigmoid exponent capped at 50).
harmonic_ir(x, amplitude, center, sigma)
¶
JAX twin of :func:oracles.models.harmonic_ir.
tauc(x, amplitude, e_gap, exponent)
¶
JAX twin of :func:oracles.models.tauc (zero below the gap).
cauchy_dispersion(x, a, b, c)
¶
JAX twin of :func:oracles.models.cauchy_dispersion (zero for x <= 0).
kww(x, amplitude, tau, beta)
¶
JAX twin of :func:oracles.models.kww (zero for x < 0).
Three-way rather than the numpy body's two-way mask. x = 0 is a real grid
point for this shape (the catalogue generates KWW on \([0, 10]\)) and there the
stretched power \((x/\tau)^\beta\) is exactly \(0^\beta\): the value is fine
(\(\to 0\), so the kernel returns amplitude), but its derivative
\(\beta\,b^{\beta-1}\) diverges, and reverse-mode AD multiplies that \(\infty\)
by the \(\partial b/\partial\tau = -x/\tau^2 = 0\) it sees at \(x = 0\), yielding
a NaN \(\partial/\partial\tau\) that would poison every KWW solve. The function
is constant in \(\tau\) at \(x = 0\), so the honest gradient is zero; splitting
x == 0 out of the power gives exactly that, with the identical value.
saturating_exponential(x, amplitude, rate)
¶
JAX twin of :func:oracles.models.saturating_exponential.
power_saturation(x, amplitude, rate)
¶
JAX twin of :func:oracles.models.power_saturation.
power_law_offset(x, amplitude, offset, shape)
¶
JAX twin of :func:oracles.models.power_law_offset.
mgh09_rational(x, amplitude, num_lin, den_lin, den_const)
¶
JAX twin of :func:oracles.models.mgh09_rational.
rational_cubic(x, a0, a1, a2, a3, b1, b2, b3)
¶
JAX twin of :func:oracles.models.rational_cubic.
generalised_logistic(x, amplitude, shift, rate, shape)
¶
JAX twin of :func:oracles.models.generalised_logistic.
exp_over_linear(x, rate, lin_const, lin_slope)
¶
JAX twin of :func:oracles.models.exp_over_linear.
Case catalog and model registry¶
oracles.cases is the declarative, pydantic-first benchmark catalog: a case
is a typed graph of components (Gaussian, Lorentzian, Voigt, Fano,
background, edge, decay — a discriminated union keyed by model), organized
into categories such as easy, complex, scaling, lineshapes, and
optfn. oracles.models is the model registry itself — one PeakModel
record per shape bundling the numpy formula, the lmfit oracle callable, the
spectrafit ModelType name, canonical parameter names, and the jax twin from
oracles.jax_kernels. See
Adding a new benchmark model
for the registration workflow these two modules drive.
oracles.cases
¶
Declarative, pydantic-first benchmark catalog — typed component graphs.
A case is a graph of components (like spectrafit's FitGraph): each component
is a typed spec (Gaussian, Lorentzian, Voigt, Fano, background, edge, decay) with
its own validated fields — no untyped param dicts. The components form a discriminated
union keyed by model; .to_params() produces the name→value dict only at the
Rust-kernel boundary. This exercises the full model-kernel space, not just Gaussians.
Layers:
- typed :class:Component specs + :class:CaseSpec — serializable case data.
- :class:CaseFamily — declarative generators that expand into many concrete specs.
- :class:BenchCase — the materialized spec (numpy x/y + truth/guess
components), built by :func:materialize.
Category counts: single source of truth in :data:CATEGORY_COUNTS.
Edge-case design philosophy
The edge category encodes hard regimes real LM solvers disagree on as
data, not code: a first-class difficulty axis (overlap fraction, width
ratio, SNR, amplitude decades) drives the component geometry, so a new hard
case is a tuned knob in a recipe function, not a new code path. See
_EDGE_REGIMES / _edge_case for the five regimes this expands into.
Diversity-driven generation
Every case is a distinct (model-set x qualitative condition); each family's
count is derived from its condition grid (see FAMILIES /
:data:CATEGORY_REGISTRY), so a category can never silently grow into Nx
repeats of the same shape. The condition tag is carried onto
:class:CaseSpec and the anti-padding test asserts uniqueness.
CATEGORY_REGISTRY = {(c.id): c for c in (CategoryDef(id='easy', label='Easy', prefix='EZ', hue='var(--ok)', count=20), CategoryDef(id='complex', label='Complex', prefix='CX', hue='var(--accent)', count=35), CategoryDef(id='reality', label='Reality-like', prefix='RL', hue='var(--warn)', count=16), CategoryDef(id='optfn', label='Optimization fns', prefix='OF', hue='var(--c-guess)', count=20, baseline_comparable=False), CategoryDef(id='scaling', label='Large-N scaling', prefix='SC', hue='var(--c-jax)', count=8), CategoryDef(id='edge', label='Edge / ill-conditioned', prefix='ED', hue='var(--bad)', count=20), CategoryDef(id='lineshapes', label='Asymmetric lineshapes', prefix='LS', hue='var(--c-lmfit)', count=27), CategoryDef(id='fixed', label='Fixed-param', prefix='FX', hue='var(--c-spectrafit)', count=4), CategoryDef(id='tied', label='Tied/shared-param', prefix='TI', hue='var(--c-jax)', count=4), CategoryDef(id='robust', label='Robust / contaminated', prefix='RB', hue='var(--warn)', count=6, baseline_comparable=False))}
module-attribute
¶
Single source of truth: one CategoryDef per category. Adding the next category
is a record here, not edits across four dicts. (The genuine 2-D example lives in
the featured MultiDim payload, not a suite category — see engine._multidim.)
CATEGORY_COUNTS = {cid: (c.count) for cid, c in (CATEGORY_REGISTRY.items())}
module-attribute
¶
Backward-compatible projection — engine.py / synth.py import these names
unchanged. Derive, never duplicate: a new CATEGORY_REGISTRY record propagates
to all four projections below.
CATEGORY_LABELSPREFIXCATEGORY_HUE
NON_COMPARABLE_CATEGORIES = frozenset(cid for cid, c in (CATEGORY_REGISTRY.items()) if not c.baseline_comparable)
module-attribute
¶
Categories excluded from the baseline-parity axes (see
CategoryDef.baseline_comparable); engine and reports import this
instead of hard-coding "optfn".
SOLVER_META = [SolverMeta(id='spectrafit', label='spectrafit', color='var(--c-spectrafit)', soft='var(--c-spectrafit-soft)'), SolverMeta(id='lmfit', label='lmfit', color='var(--c-lmfit)', soft='var(--c-lmfit-soft)'), SolverMeta(id='jax', label='jax', color='var(--c-jax)', soft='var(--c-jax-soft)'), SolverMeta(id='scipy-ls-lm', label='scipy-ls-lm', color='var(--c-scipy-ls-lm, var(--c-lmfit))', soft='var(--c-scipy-ls-lm-soft, var(--c-lmfit-soft))'), SolverMeta(id='scipy-ls-trf', label='scipy-ls-trf', color='var(--c-scipy-ls-trf, var(--c-lmfit))', soft='var(--c-scipy-ls-trf-soft, var(--c-lmfit-soft))'), SolverMeta(id='scipy-ls-dogbox', label='scipy-ls-dogbox', color='var(--c-scipy-ls-dogbox, var(--c-lmfit))', soft='var(--c-scipy-ls-dogbox-soft, var(--c-lmfit-soft))')]
module-attribute
¶
Catalog presentation metadata — the single home for solver styling. engine.py and synth.py import this rather than redefining it (silent drift between two copies is worse than a shared import).
RECOVERABLE_PARAMS = ('amplitude', 'center', 'sigma', 'gamma')
module-attribute
¶
Shape-defining params scored for recovery (slope/intercept/offset/c/
fraction/q/lam are fit but not part of the error metric).
CategoryDef
pydantic-model
¶
Bases: BaseModel
One validated record per suite category — the single source of category truth.
Replaces the four parallel dicts (counts/labels/prefix/hue) keyed by the same
strings: a missing key in one of those silently became a runtime KeyError in
engine._CATEGORY_META. The derived dicts below project this registry, so the
public names stay backward-compatible while drift is structurally impossible.
Show JSON schema:
{
"additionalProperties": false,
"description": "One validated record per suite category \u2014 the single source of category truth.\n\nReplaces the four parallel dicts (counts/labels/prefix/hue) keyed by the same\nstrings: a missing key in one of those silently became a runtime ``KeyError`` in\n``engine._CATEGORY_META``. The derived dicts below project this registry, so the\npublic names stay backward-compatible while drift is structurally impossible.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"label": {
"title": "Label",
"type": "string"
},
"prefix": {
"title": "Prefix",
"type": "string"
},
"hue": {
"title": "Hue",
"type": "string"
},
"count": {
"minimum": 0,
"title": "Count",
"type": "integer"
},
"baseline_comparable": {
"default": true,
"title": "Baseline Comparable",
"type": "boolean"
}
},
"required": [
"id",
"label",
"prefix",
"hue",
"count"
],
"title": "CategoryDef",
"type": "object"
}
Config:
extra:forbid
Fields:
-
id(str) -
label(str) -
prefix(str) -
hue(str) -
count(int) -
baseline_comparable(bool)
baseline_comparable = True
pydantic-field
¶
Whether the category counts toward the baseline-parity axes.
The headline \(|\Delta r^2|\), the suite-regression flag and the win-rate all
compare spectrafit against the least-squares oracle. False marks the
categories where spectrafit deliberately solves a different problem from
the baselines — multimodal optfn (global vs local search) and
robust (an M-estimator on contaminated data vs plain least squares) —
so a delta there is not a regression signal. Recovery-vs-truth still
scores them.
GaussianSpec
pydantic-model
¶
Bases: _Component
Gaussian peak.
Show JSON schema:
{
"additionalProperties": false,
"description": "Gaussian peak.",
"properties": {
"model": {
"const": "gaussian",
"default": "gaussian",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "GaussianSpec",
"type": "object"
}
Fields:
-
model(Literal['gaussian']) -
amplitude(float) -
center(float) -
sigma(float)
LorentzianSpec
pydantic-model
¶
Bases: _Component
Lorentzian peak.
Show JSON schema:
{
"additionalProperties": false,
"description": "Lorentzian peak.",
"properties": {
"model": {
"const": "lorentzian",
"default": "lorentzian",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "LorentzianSpec",
"type": "object"
}
Fields:
-
model(Literal['lorentzian']) -
amplitude(float) -
center(float) -
sigma(float)
PseudoVoigtSpec
pydantic-model
¶
Bases: _Component
Pseudo-Voigt / Voigt peak (Voigt is the spectrafit alias, same formula).
Show JSON schema:
{
"additionalProperties": false,
"description": "Pseudo-Voigt / Voigt peak (Voigt is the spectrafit alias, same formula).",
"properties": {
"model": {
"default": "pseudo_voigt",
"enum": [
"pseudo_voigt",
"voigt"
],
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"fraction": {
"title": "Fraction",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"fraction"
],
"title": "PseudoVoigtSpec",
"type": "object"
}
Fields:
-
model(Literal['pseudo_voigt', 'voigt']) -
amplitude(float) -
center(float) -
sigma(float) -
fraction(float)
FanoSpec
pydantic-model
¶
Bases: _Component
Fano resonance.
Show JSON schema:
{
"additionalProperties": false,
"description": "Fano resonance.",
"properties": {
"model": {
"const": "fano",
"default": "fano",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"gamma": {
"title": "Gamma",
"type": "number"
},
"q": {
"title": "Q",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"gamma",
"q"
],
"title": "FanoSpec",
"type": "object"
}
Fields:
-
model(Literal['fano']) -
amplitude(float) -
center(float) -
gamma(float) -
q(float)
ConstantSpec
pydantic-model
¶
Bases: _Component
Constant background.
Show JSON schema:
Fields:
-
model(Literal['constant']) -
c(float)
LinearSpec
pydantic-model
¶
Bases: _Component
Linear background.
Show JSON schema:
{
"additionalProperties": false,
"description": "Linear background.",
"properties": {
"model": {
"const": "linear",
"default": "linear",
"title": "Model",
"type": "string"
},
"slope": {
"title": "Slope",
"type": "number"
},
"intercept": {
"title": "Intercept",
"type": "number"
}
},
"required": [
"slope",
"intercept"
],
"title": "LinearSpec",
"type": "object"
}
Fields:
-
model(Literal['linear']) -
slope(float) -
intercept(float)
QuadraticSpec
pydantic-model
¶
Bases: _Component
Quadratic background / bowl.
Show JSON schema:
{
"additionalProperties": false,
"description": "Quadratic background / bowl.",
"properties": {
"model": {
"const": "quadratic",
"default": "quadratic",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"offset": {
"title": "Offset",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"offset"
],
"title": "QuadraticSpec",
"type": "object"
}
Fields:
-
model(Literal['quadratic']) -
amplitude(float) -
center(float) -
offset(float)
StepSpec
pydantic-model
¶
Bases: _Component
Edge/step background (arctan / tanh / erfc).
Show JSON schema:
{
"additionalProperties": false,
"description": "Edge/step background (arctan / tanh / erfc).",
"properties": {
"model": {
"enum": [
"arctan_step",
"tanh_step",
"erfc_step"
],
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"model",
"amplitude",
"center",
"sigma"
],
"title": "StepSpec",
"type": "object"
}
Fields:
-
model(Literal['arctan_step', 'tanh_step', 'erfc_step']) -
amplitude(float) -
center(float) -
sigma(float)
DecaySpec
pydantic-model
¶
Bases: _Component
Bi-exponential decay.
Show JSON schema:
{
"additionalProperties": false,
"description": "Bi-exponential decay.",
"properties": {
"model": {
"const": "double_exponential",
"default": "double_exponential",
"title": "Model",
"type": "string"
},
"A1": {
"title": "A1",
"type": "number"
},
"lam1": {
"title": "Lam1",
"type": "number"
},
"A2": {
"title": "A2",
"type": "number"
},
"lam2": {
"title": "Lam2",
"type": "number"
}
},
"required": [
"A1",
"lam1",
"A2",
"lam2"
],
"title": "DecaySpec",
"type": "object"
}
Fields:
-
model(Literal['double_exponential']) -
A1(float) -
lam1(float) -
A2(float) -
lam2(float)
AsymPeakSpec
pydantic-model
¶
Bases: _Component
Asymmetric peak: true-Voigt, skewed Gaussian, EMG, or Doniach-Sunjic.
All four share (amplitude, center, sigma, gamma). Gamma is Lorentzian HWHM,
skew, decay-rate, or asymmetry respectively.
Show JSON schema:
{
"additionalProperties": false,
"description": "Asymmetric peak: true-Voigt, skewed Gaussian, EMG, or Doniach-Sunjic.\n\nAll four share ``(amplitude, center, sigma, gamma)``. Gamma is Lorentzian HWHM,\nskew, decay-rate, or asymmetry respectively.",
"properties": {
"model": {
"enum": [
"true_voigt",
"skewed_gaussian",
"exp_gaussian",
"doniach_sunjic"
],
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"gamma": {
"title": "Gamma",
"type": "number"
}
},
"required": [
"model",
"amplitude",
"center",
"sigma",
"gamma"
],
"title": "AsymPeakSpec",
"type": "object"
}
Fields:
-
model(Literal['true_voigt', 'skewed_gaussian', 'exp_gaussian', 'doniach_sunjic']) -
amplitude(float) -
center(float) -
sigma(float) -
gamma(float)
LogNormalSpec
pydantic-model
¶
Bases: _Component
Log-normal peak (positive-axis, chromatography/particle-size style).
Show JSON schema:
{
"additionalProperties": false,
"description": "Log-normal peak (positive-axis, chromatography/particle-size style).",
"properties": {
"model": {
"const": "log_normal",
"default": "log_normal",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "LogNormalSpec",
"type": "object"
}
Fields:
-
model(Literal['log_normal']) -
amplitude(float) -
center(float) -
sigma(float)
Pearson7Spec
pydantic-model
¶
Bases: _Component
Pearson VII peak with tunable tail weight m.
Show JSON schema:
{
"additionalProperties": false,
"description": "Pearson VII peak with tunable tail weight ``m``.",
"properties": {
"model": {
"const": "pearson7",
"default": "pearson7",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"m": {
"title": "M",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"m"
],
"title": "Pearson7Spec",
"type": "object"
}
Fields:
-
model(Literal['pearson7']) -
amplitude(float) -
center(float) -
sigma(float) -
m(float)
SplitGaussianSpec
pydantic-model
¶
Bases: _Component
Split Gaussian with independent widths on each side.
Show JSON schema:
{
"additionalProperties": false,
"description": "Split Gaussian with independent widths on each side.",
"properties": {
"model": {
"const": "split_gaussian",
"default": "split_gaussian",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma_l": {
"title": "Sigma L",
"type": "number"
},
"sigma_r": {
"title": "Sigma R",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma_l",
"sigma_r"
],
"title": "SplitGaussianSpec",
"type": "object"
}
Fields:
-
model(Literal['split_gaussian']) -
amplitude(float) -
center(float) -
sigma_l(float) -
sigma_r(float)
MoffatSpec
pydantic-model
¶
Bases: _Component
Moffat peak (\(\beta\) tail weight).
Show JSON schema:
{
"additionalProperties": false,
"description": "Moffat peak ($\\beta$ tail weight).",
"properties": {
"model": {
"const": "moffat",
"default": "moffat",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"beta": {
"title": "Beta",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"beta"
],
"title": "MoffatSpec",
"type": "object"
}
Fields:
-
model(Literal['moffat']) -
amplitude(float) -
center(float) -
sigma(float) -
beta(float)
StudentsTSpec
pydantic-model
¶
Bases: _Component
Student's-t peak (\(\nu\) degrees of freedom).
Show JSON schema:
{
"additionalProperties": false,
"description": "Student's-t peak ($\\nu$ degrees of freedom).",
"properties": {
"model": {
"const": "students_t",
"default": "students_t",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"nu": {
"title": "Nu",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"nu"
],
"title": "StudentsTSpec",
"type": "object"
}
Fields:
-
model(Literal['students_t']) -
amplitude(float) -
center(float) -
sigma(float) -
nu(float)
SplitPearson7Spec
pydantic-model
¶
Bases: _Component
Split Pearson VII (split width + exponent each side).
Show JSON schema:
{
"additionalProperties": false,
"description": "Split Pearson VII (split width + exponent each side).",
"properties": {
"model": {
"const": "split_pearson7",
"default": "split_pearson7",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma_l": {
"title": "Sigma L",
"type": "number"
},
"sigma_r": {
"title": "Sigma R",
"type": "number"
},
"m_l": {
"title": "M L",
"type": "number"
},
"m_r": {
"title": "M R",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma_l",
"sigma_r",
"m_l",
"m_r"
],
"title": "SplitPearson7Spec",
"type": "object"
}
Fields:
-
model(Literal['split_pearson7']) -
amplitude(float) -
center(float) -
sigma_l(float) -
sigma_r(float) -
m_l(float) -
m_r(float)
BreitWignerSpec
pydantic-model
¶
Bases: _Component
Breit-Wigner-Fano resonance.
Show JSON schema:
{
"additionalProperties": false,
"description": "Breit-Wigner-Fano resonance.",
"properties": {
"model": {
"const": "breit_wigner",
"default": "breit_wigner",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"q": {
"title": "Q",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"q"
],
"title": "BreitWignerSpec",
"type": "object"
}
Fields:
-
model(Literal['breit_wigner']) -
amplitude(float) -
center(float) -
sigma(float) -
q(float)
AsymIrSpec
pydantic-model
¶
Bases: _Component
Asymmetric IR band (Gaussian \(\times\) logistic sigmoid).
Show JSON schema:
{
"additionalProperties": false,
"description": "Asymmetric IR band (Gaussian $\\times$ logistic sigmoid).",
"properties": {
"model": {
"const": "asym_ir",
"default": "asym_ir",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"k": {
"title": "K",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"k"
],
"title": "AsymIrSpec",
"type": "object"
}
Fields:
-
model(Literal['asym_ir']) -
amplitude(float) -
center(float) -
sigma(float) -
k(float)
HarmonicIrSpec
pydantic-model
¶
Bases: _Component
Driven damped harmonic-oscillator IR absorption.
Show JSON schema:
{
"additionalProperties": false,
"description": "Driven damped harmonic-oscillator IR absorption.",
"properties": {
"model": {
"const": "harmonic_ir",
"default": "harmonic_ir",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "HarmonicIrSpec",
"type": "object"
}
Fields:
-
model(Literal['harmonic_ir']) -
amplitude(float) -
center(float) -
sigma(float)
TaucSpec
pydantic-model
¶
Bases: _Component
Tauc optical band-gap edge (power-law absorption above e_gap).
Show JSON schema:
{
"additionalProperties": false,
"description": "Tauc optical band-gap edge (power-law absorption above ``e_gap``).",
"properties": {
"model": {
"const": "tauc",
"default": "tauc",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"e_gap": {
"title": "E Gap",
"type": "number"
},
"exponent": {
"title": "Exponent",
"type": "number"
}
},
"required": [
"amplitude",
"e_gap",
"exponent"
],
"title": "TaucSpec",
"type": "object"
}
Fields:
-
model(Literal['tauc']) -
amplitude(float) -
e_gap(float) -
exponent(float)
CauchyDispersionSpec
pydantic-model
¶
Bases: _Component
Cauchy refractive-index dispersion \(a + b/x^2 + c/x^4\) (positive-x).
Show JSON schema:
{
"additionalProperties": false,
"description": "Cauchy refractive-index dispersion $a + b/x^2 + c/x^4$ (positive-x).",
"properties": {
"model": {
"const": "cauchy_dispersion",
"default": "cauchy_dispersion",
"title": "Model",
"type": "string"
},
"a": {
"title": "A",
"type": "number"
},
"b": {
"title": "B",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
}
},
"required": [
"a",
"b",
"c"
],
"title": "CauchyDispersionSpec",
"type": "object"
}
Fields:
-
model(Literal['cauchy_dispersion']) -
a(float) -
b(float) -
c(float)
KwwSpec
pydantic-model
¶
Bases: _Component
Kohlrausch–Williams–Watts stretched exponential (positive-x relaxation).
Show JSON schema:
{
"additionalProperties": false,
"description": "Kohlrausch\u2013Williams\u2013Watts stretched exponential (positive-x relaxation).",
"properties": {
"model": {
"const": "kww",
"default": "kww",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"tau": {
"title": "Tau",
"type": "number"
},
"beta": {
"title": "Beta",
"type": "number"
}
},
"required": [
"amplitude",
"tau",
"beta"
],
"title": "KwwSpec",
"type": "object"
}
Fields:
-
model(Literal['kww']) -
amplitude(float) -
tau(float) -
beta(float)
SaturatingExponentialSpec
pydantic-model
¶
Bases: _Component
Saturating exponential (NIST BoxBOD, positive-x).
\(\mathrm{amplitude}\cdot(1-\exp(-\mathrm{rate}\cdot x))\).
Show JSON schema:
{
"additionalProperties": false,
"description": "Saturating exponential (NIST BoxBOD, positive-x).\n\n$\\mathrm{amplitude}\\cdot(1-\\exp(-\\mathrm{rate}\\cdot x))$.",
"properties": {
"model": {
"const": "saturating_exponential",
"default": "saturating_exponential",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"rate": {
"title": "Rate",
"type": "number"
}
},
"required": [
"amplitude",
"rate"
],
"title": "SaturatingExponentialSpec",
"type": "object"
}
Fields:
-
model(Literal['saturating_exponential']) -
amplitude(float) -
rate(float)
PowerSaturationSpec
pydantic-model
¶
Bases: _Component
Power-law saturation (NIST Misra1b, positive-x).
\(\mathrm{amplitude}\cdot(1-(1+\mathrm{rate}\cdot x/2)^{-2})\).
Show JSON schema:
{
"additionalProperties": false,
"description": "Power-law saturation (NIST Misra1b, positive-x).\n\n$\\mathrm{amplitude}\\cdot(1-(1+\\mathrm{rate}\\cdot x/2)^{-2})$.",
"properties": {
"model": {
"const": "power_saturation",
"default": "power_saturation",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"rate": {
"title": "Rate",
"type": "number"
}
},
"required": [
"amplitude",
"rate"
],
"title": "PowerSaturationSpec",
"type": "object"
}
Fields:
-
model(Literal['power_saturation']) -
amplitude(float) -
rate(float)
PowerLawOffsetSpec
pydantic-model
¶
Bases: _Component
Power-law with offset (Bennett5-like, positive-x).
Requires offset + x > 0 on the grid — kept safe by a positive x-range plus an
offset comfortably larger than |x_min| (mirrors the parity-test guard).
Show JSON schema:
{
"additionalProperties": false,
"description": "Power-law with offset (Bennett5-like, positive-x).\n\nRequires ``offset + x > 0`` on the grid \u2014 kept safe by a positive x-range plus an\n``offset`` comfortably larger than ``|x_min|`` (mirrors the parity-test guard).",
"properties": {
"model": {
"const": "power_law_offset",
"default": "power_law_offset",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"offset": {
"title": "Offset",
"type": "number"
},
"shape": {
"title": "Shape",
"type": "number"
}
},
"required": [
"amplitude",
"offset",
"shape"
],
"title": "PowerLawOffsetSpec",
"type": "object"
}
Fields:
-
model(Literal['power_law_offset']) -
amplitude(float) -
offset(float) -
shape(float)
Mgh09RationalSpec
pydantic-model
¶
Bases: _Component
Kowalik–Osborne rational function (NIST StRD MGH09, positive-x).
\(\mathrm{amplitude}\cdot(x^2+\mathrm{num\_lin}\cdot x)/(x^2+\mathrm{den\_lin}\cdot x+\mathrm{den\_const})\); the denominator stays positive when \(\mathrm{den\_lin}^2-4\cdot\mathrm{den\_const} < 0\) (mirrors the parity-test param values).
Show JSON schema:
{
"additionalProperties": false,
"description": "Kowalik\u2013Osborne rational function (NIST StRD MGH09, positive-x).\n\n$\\mathrm{amplitude}\\cdot(x^2+\\mathrm{num\\_lin}\\cdot x)/(x^2+\\mathrm{den\\_lin}\\cdot x+\\mathrm{den\\_const})$; the denominator stays\npositive when $\\mathrm{den\\_lin}^2-4\\cdot\\mathrm{den\\_const} < 0$ (mirrors the parity-test param values).",
"properties": {
"model": {
"const": "mgh09_rational",
"default": "mgh09_rational",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"num_lin": {
"title": "Num Lin",
"type": "number"
},
"den_lin": {
"title": "Den Lin",
"type": "number"
},
"den_const": {
"title": "Den Const",
"type": "number"
}
},
"required": [
"amplitude",
"num_lin",
"den_lin",
"den_const"
],
"title": "Mgh09RationalSpec",
"type": "object"
}
Fields:
-
model(Literal['mgh09_rational']) -
amplitude(float) -
num_lin(float) -
den_lin(float) -
den_const(float)
RationalCubicSpec
pydantic-model
¶
Bases: _Component
Rational cubic with unit denominator constant (NIST StRD benchmarks).
\((a_0+a_1x+a_2x^2+a_3x^3)/(1+b_1x+b_2x^2+b_3x^3)\); coefficients are kept near the parity-test values (a0=1.2, a1=0.3, a2=0.15, a3=0.02, b1=0.05, b2=0.02, b3=0.001) so the denominator stays positive across the grid. Named after NIST StRD Kirby2, Hahn1, and Thurber, whose fits share this rational-cubic form.
Show JSON schema:
{
"additionalProperties": false,
"description": "Rational cubic with unit denominator constant (NIST StRD benchmarks).\n\n$(a_0+a_1x+a_2x^2+a_3x^3)/(1+b_1x+b_2x^2+b_3x^3)$; coefficients are kept near the\nparity-test values (a0=1.2, a1=0.3, a2=0.15, a3=0.02, b1=0.05, b2=0.02, b3=0.001)\nso the denominator stays positive across the grid. Named after NIST StRD\nKirby2, Hahn1, and Thurber, whose fits share this rational-cubic form.",
"properties": {
"model": {
"const": "rational_cubic",
"default": "rational_cubic",
"title": "Model",
"type": "string"
},
"a0": {
"title": "A0",
"type": "number"
},
"a1": {
"title": "A1",
"type": "number"
},
"a2": {
"title": "A2",
"type": "number"
},
"a3": {
"title": "A3",
"type": "number"
},
"b1": {
"title": "B1",
"type": "number"
},
"b2": {
"title": "B2",
"type": "number"
},
"b3": {
"title": "B3",
"type": "number"
}
},
"required": [
"a0",
"a1",
"a2",
"a3",
"b1",
"b2",
"b3"
],
"title": "RationalCubicSpec",
"type": "object"
}
Fields:
-
model(Literal['rational_cubic']) -
a0(float) -
a1(float) -
a2(float) -
a3(float) -
b1(float) -
b2(float) -
b3(float)
GeneralisedLogisticSpec
pydantic-model
¶
Bases: _Component
Generalised logistic (Richards) curve (NIST StRD Rat42/Rat43).
\(\mathrm{amplitude}/(1+\exp(\mathrm{shift}-\mathrm{rate}\cdot x))^{1/\mathrm{shape}}\); unrestricted domain, no positive-x requirement.
Show JSON schema:
{
"additionalProperties": false,
"description": "Generalised logistic (Richards) curve (NIST StRD Rat42/Rat43).\n\n$\\mathrm{amplitude}/(1+\\exp(\\mathrm{shift}-\\mathrm{rate}\\cdot x))^{1/\\mathrm{shape}}$;\nunrestricted domain, no positive-x requirement.",
"properties": {
"model": {
"const": "generalised_logistic",
"default": "generalised_logistic",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"shift": {
"title": "Shift",
"type": "number"
},
"rate": {
"title": "Rate",
"type": "number"
},
"shape": {
"title": "Shape",
"type": "number"
}
},
"required": [
"amplitude",
"shift",
"rate",
"shape"
],
"title": "GeneralisedLogisticSpec",
"type": "object"
}
Fields:
-
model(Literal['generalised_logistic']) -
amplitude(float) -
shift(float) -
rate(float) -
shape(float)
ExpOverLinearSpec
pydantic-model
¶
Bases: _Component
Exponential decay over a line (NIST StRD Chwirut1/Chwirut2).
\(e^{-\mathrm{rate}\cdot x}/(\mathrm{lin\_const}+\mathrm{lin\_slope}\cdot x)\); kept
away from the parity-test mgh09_rational's den_const/den_lin names (shared
global param table) and from the denominator's root by holding lin_const
comfortably larger than \(\lvert \mathrm{lin\_slope}\cdot x_{\min}\rvert\) on the grid.
Show JSON schema:
{
"additionalProperties": false,
"description": "Exponential decay over a line (NIST StRD Chwirut1/Chwirut2).\n\n$e^{-\\mathrm{rate}\\cdot x}/(\\mathrm{lin\\_const}+\\mathrm{lin\\_slope}\\cdot x)$; kept\naway from the parity-test mgh09_rational's ``den_const``/``den_lin`` names (shared\nglobal param table) and from the denominator's root by holding ``lin_const``\ncomfortably larger than $\\lvert \\mathrm{lin\\_slope}\\cdot x_{\\min}\\rvert$ on the grid.",
"properties": {
"model": {
"const": "exp_over_linear",
"default": "exp_over_linear",
"title": "Model",
"type": "string"
},
"rate": {
"title": "Rate",
"type": "number"
},
"lin_const": {
"title": "Lin Const",
"type": "number"
},
"lin_slope": {
"title": "Lin Slope",
"type": "number"
}
},
"required": [
"rate",
"lin_const",
"lin_slope"
],
"title": "ExpOverLinearSpec",
"type": "object"
}
Fields:
-
model(Literal['exp_over_linear']) -
rate(float) -
lin_const(float) -
lin_slope(float)
OutlierSpec
pydantic-model
¶
Bases: BaseModel
Gross-outlier contamination applied by :func:materialize after the noise.
round(fraction * n_points) distinct points (chosen from the case's own
id-seeded RNG, so the contamination is reproducible) receive an additive
spike of magnitude × max|y_clean| with a random sign. The spike is far
outside the Gaussian noise floor by construction, which is what makes the
case a test of the loss function rather than of the noise model.
Show JSON schema:
{
"additionalProperties": false,
"description": "Gross-outlier contamination applied by :func:`materialize` after the noise.\n\n``round(fraction * n_points)`` distinct points (chosen from the case's own\nid-seeded RNG, so the contamination is reproducible) receive an additive\nspike of ``magnitude \u00d7 max|y_clean|`` with a random sign. The spike is far\noutside the Gaussian noise floor by construction, which is what makes the\ncase a test of the *loss function* rather than of the noise model.",
"properties": {
"fraction": {
"exclusiveMinimum": 0.0,
"maximum": 0.5,
"title": "Fraction",
"type": "number"
},
"magnitude": {
"exclusiveMinimum": 0.0,
"title": "Magnitude",
"type": "number"
}
},
"required": [
"fraction",
"magnitude"
],
"title": "OutlierSpec",
"type": "object"
}
Config:
extra:forbid
Fields:
-
fraction(float) -
magnitude(float)
CaseSpec
pydantic-model
¶
Bases: BaseModel
A fully concrete, serializable benchmark-case declaration.
Show JSON schema:
{
"$defs": {
"AsymIrSpec": {
"additionalProperties": false,
"description": "Asymmetric IR band (Gaussian $\\times$ logistic sigmoid).",
"properties": {
"model": {
"const": "asym_ir",
"default": "asym_ir",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"k": {
"title": "K",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"k"
],
"title": "AsymIrSpec",
"type": "object"
},
"AsymPeakSpec": {
"additionalProperties": false,
"description": "Asymmetric peak: true-Voigt, skewed Gaussian, EMG, or Doniach-Sunjic.\n\nAll four share ``(amplitude, center, sigma, gamma)``. Gamma is Lorentzian HWHM,\nskew, decay-rate, or asymmetry respectively.",
"properties": {
"model": {
"enum": [
"true_voigt",
"skewed_gaussian",
"exp_gaussian",
"doniach_sunjic"
],
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"gamma": {
"title": "Gamma",
"type": "number"
}
},
"required": [
"model",
"amplitude",
"center",
"sigma",
"gamma"
],
"title": "AsymPeakSpec",
"type": "object"
},
"BreitWignerSpec": {
"additionalProperties": false,
"description": "Breit-Wigner-Fano resonance.",
"properties": {
"model": {
"const": "breit_wigner",
"default": "breit_wigner",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"q": {
"title": "Q",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"q"
],
"title": "BreitWignerSpec",
"type": "object"
},
"CauchyDispersionSpec": {
"additionalProperties": false,
"description": "Cauchy refractive-index dispersion $a + b/x^2 + c/x^4$ (positive-x).",
"properties": {
"model": {
"const": "cauchy_dispersion",
"default": "cauchy_dispersion",
"title": "Model",
"type": "string"
},
"a": {
"title": "A",
"type": "number"
},
"b": {
"title": "B",
"type": "number"
},
"c": {
"title": "C",
"type": "number"
}
},
"required": [
"a",
"b",
"c"
],
"title": "CauchyDispersionSpec",
"type": "object"
},
"ConstantSpec": {
"additionalProperties": false,
"description": "Constant background.",
"properties": {
"model": {
"const": "constant",
"default": "constant",
"title": "Model",
"type": "string"
},
"c": {
"title": "C",
"type": "number"
}
},
"required": [
"c"
],
"title": "ConstantSpec",
"type": "object"
},
"DecaySpec": {
"additionalProperties": false,
"description": "Bi-exponential decay.",
"properties": {
"model": {
"const": "double_exponential",
"default": "double_exponential",
"title": "Model",
"type": "string"
},
"A1": {
"title": "A1",
"type": "number"
},
"lam1": {
"title": "Lam1",
"type": "number"
},
"A2": {
"title": "A2",
"type": "number"
},
"lam2": {
"title": "Lam2",
"type": "number"
}
},
"required": [
"A1",
"lam1",
"A2",
"lam2"
],
"title": "DecaySpec",
"type": "object"
},
"ExpOverLinearSpec": {
"additionalProperties": false,
"description": "Exponential decay over a line (NIST StRD Chwirut1/Chwirut2).\n\n$e^{-\\mathrm{rate}\\cdot x}/(\\mathrm{lin\\_const}+\\mathrm{lin\\_slope}\\cdot x)$; kept\naway from the parity-test mgh09_rational's ``den_const``/``den_lin`` names (shared\nglobal param table) and from the denominator's root by holding ``lin_const``\ncomfortably larger than $\\lvert \\mathrm{lin\\_slope}\\cdot x_{\\min}\\rvert$ on the grid.",
"properties": {
"model": {
"const": "exp_over_linear",
"default": "exp_over_linear",
"title": "Model",
"type": "string"
},
"rate": {
"title": "Rate",
"type": "number"
},
"lin_const": {
"title": "Lin Const",
"type": "number"
},
"lin_slope": {
"title": "Lin Slope",
"type": "number"
}
},
"required": [
"rate",
"lin_const",
"lin_slope"
],
"title": "ExpOverLinearSpec",
"type": "object"
},
"FanoSpec": {
"additionalProperties": false,
"description": "Fano resonance.",
"properties": {
"model": {
"const": "fano",
"default": "fano",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"gamma": {
"title": "Gamma",
"type": "number"
},
"q": {
"title": "Q",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"gamma",
"q"
],
"title": "FanoSpec",
"type": "object"
},
"GaussianSpec": {
"additionalProperties": false,
"description": "Gaussian peak.",
"properties": {
"model": {
"const": "gaussian",
"default": "gaussian",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "GaussianSpec",
"type": "object"
},
"GeneralisedLogisticSpec": {
"additionalProperties": false,
"description": "Generalised logistic (Richards) curve (NIST StRD Rat42/Rat43).\n\n$\\mathrm{amplitude}/(1+\\exp(\\mathrm{shift}-\\mathrm{rate}\\cdot x))^{1/\\mathrm{shape}}$;\nunrestricted domain, no positive-x requirement.",
"properties": {
"model": {
"const": "generalised_logistic",
"default": "generalised_logistic",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"shift": {
"title": "Shift",
"type": "number"
},
"rate": {
"title": "Rate",
"type": "number"
},
"shape": {
"title": "Shape",
"type": "number"
}
},
"required": [
"amplitude",
"shift",
"rate",
"shape"
],
"title": "GeneralisedLogisticSpec",
"type": "object"
},
"HarmonicIrSpec": {
"additionalProperties": false,
"description": "Driven damped harmonic-oscillator IR absorption.",
"properties": {
"model": {
"const": "harmonic_ir",
"default": "harmonic_ir",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "HarmonicIrSpec",
"type": "object"
},
"KwwSpec": {
"additionalProperties": false,
"description": "Kohlrausch\u2013Williams\u2013Watts stretched exponential (positive-x relaxation).",
"properties": {
"model": {
"const": "kww",
"default": "kww",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"tau": {
"title": "Tau",
"type": "number"
},
"beta": {
"title": "Beta",
"type": "number"
}
},
"required": [
"amplitude",
"tau",
"beta"
],
"title": "KwwSpec",
"type": "object"
},
"LinearSpec": {
"additionalProperties": false,
"description": "Linear background.",
"properties": {
"model": {
"const": "linear",
"default": "linear",
"title": "Model",
"type": "string"
},
"slope": {
"title": "Slope",
"type": "number"
},
"intercept": {
"title": "Intercept",
"type": "number"
}
},
"required": [
"slope",
"intercept"
],
"title": "LinearSpec",
"type": "object"
},
"LogNormalSpec": {
"additionalProperties": false,
"description": "Log-normal peak (positive-axis, chromatography/particle-size style).",
"properties": {
"model": {
"const": "log_normal",
"default": "log_normal",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "LogNormalSpec",
"type": "object"
},
"LorentzianSpec": {
"additionalProperties": false,
"description": "Lorentzian peak.",
"properties": {
"model": {
"const": "lorentzian",
"default": "lorentzian",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma"
],
"title": "LorentzianSpec",
"type": "object"
},
"Mgh09RationalSpec": {
"additionalProperties": false,
"description": "Kowalik\u2013Osborne rational function (NIST StRD MGH09, positive-x).\n\n$\\mathrm{amplitude}\\cdot(x^2+\\mathrm{num\\_lin}\\cdot x)/(x^2+\\mathrm{den\\_lin}\\cdot x+\\mathrm{den\\_const})$; the denominator stays\npositive when $\\mathrm{den\\_lin}^2-4\\cdot\\mathrm{den\\_const} < 0$ (mirrors the parity-test param values).",
"properties": {
"model": {
"const": "mgh09_rational",
"default": "mgh09_rational",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"num_lin": {
"title": "Num Lin",
"type": "number"
},
"den_lin": {
"title": "Den Lin",
"type": "number"
},
"den_const": {
"title": "Den Const",
"type": "number"
}
},
"required": [
"amplitude",
"num_lin",
"den_lin",
"den_const"
],
"title": "Mgh09RationalSpec",
"type": "object"
},
"MoffatSpec": {
"additionalProperties": false,
"description": "Moffat peak ($\\beta$ tail weight).",
"properties": {
"model": {
"const": "moffat",
"default": "moffat",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"beta": {
"title": "Beta",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"beta"
],
"title": "MoffatSpec",
"type": "object"
},
"OutlierSpec": {
"additionalProperties": false,
"description": "Gross-outlier contamination applied by :func:`materialize` after the noise.\n\n``round(fraction * n_points)`` distinct points (chosen from the case's own\nid-seeded RNG, so the contamination is reproducible) receive an additive\nspike of ``magnitude \u00d7 max|y_clean|`` with a random sign. The spike is far\noutside the Gaussian noise floor by construction, which is what makes the\ncase a test of the *loss function* rather than of the noise model.",
"properties": {
"fraction": {
"exclusiveMinimum": 0.0,
"maximum": 0.5,
"title": "Fraction",
"type": "number"
},
"magnitude": {
"exclusiveMinimum": 0.0,
"title": "Magnitude",
"type": "number"
}
},
"required": [
"fraction",
"magnitude"
],
"title": "OutlierSpec",
"type": "object"
},
"Pearson7Spec": {
"additionalProperties": false,
"description": "Pearson VII peak with tunable tail weight ``m``.",
"properties": {
"model": {
"const": "pearson7",
"default": "pearson7",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"m": {
"title": "M",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"m"
],
"title": "Pearson7Spec",
"type": "object"
},
"PowerLawOffsetSpec": {
"additionalProperties": false,
"description": "Power-law with offset (Bennett5-like, positive-x).\n\nRequires ``offset + x > 0`` on the grid \u2014 kept safe by a positive x-range plus an\n``offset`` comfortably larger than ``|x_min|`` (mirrors the parity-test guard).",
"properties": {
"model": {
"const": "power_law_offset",
"default": "power_law_offset",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"offset": {
"title": "Offset",
"type": "number"
},
"shape": {
"title": "Shape",
"type": "number"
}
},
"required": [
"amplitude",
"offset",
"shape"
],
"title": "PowerLawOffsetSpec",
"type": "object"
},
"PowerSaturationSpec": {
"additionalProperties": false,
"description": "Power-law saturation (NIST Misra1b, positive-x).\n\n$\\mathrm{amplitude}\\cdot(1-(1+\\mathrm{rate}\\cdot x/2)^{-2})$.",
"properties": {
"model": {
"const": "power_saturation",
"default": "power_saturation",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"rate": {
"title": "Rate",
"type": "number"
}
},
"required": [
"amplitude",
"rate"
],
"title": "PowerSaturationSpec",
"type": "object"
},
"PseudoVoigtSpec": {
"additionalProperties": false,
"description": "Pseudo-Voigt / Voigt peak (Voigt is the spectrafit alias, same formula).",
"properties": {
"model": {
"default": "pseudo_voigt",
"enum": [
"pseudo_voigt",
"voigt"
],
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"fraction": {
"title": "Fraction",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"fraction"
],
"title": "PseudoVoigtSpec",
"type": "object"
},
"QuadraticSpec": {
"additionalProperties": false,
"description": "Quadratic background / bowl.",
"properties": {
"model": {
"const": "quadratic",
"default": "quadratic",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"offset": {
"title": "Offset",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"offset"
],
"title": "QuadraticSpec",
"type": "object"
},
"RationalCubicSpec": {
"additionalProperties": false,
"description": "Rational cubic with unit denominator constant (NIST StRD benchmarks).\n\n$(a_0+a_1x+a_2x^2+a_3x^3)/(1+b_1x+b_2x^2+b_3x^3)$; coefficients are kept near the\nparity-test values (a0=1.2, a1=0.3, a2=0.15, a3=0.02, b1=0.05, b2=0.02, b3=0.001)\nso the denominator stays positive across the grid. Named after NIST StRD\nKirby2, Hahn1, and Thurber, whose fits share this rational-cubic form.",
"properties": {
"model": {
"const": "rational_cubic",
"default": "rational_cubic",
"title": "Model",
"type": "string"
},
"a0": {
"title": "A0",
"type": "number"
},
"a1": {
"title": "A1",
"type": "number"
},
"a2": {
"title": "A2",
"type": "number"
},
"a3": {
"title": "A3",
"type": "number"
},
"b1": {
"title": "B1",
"type": "number"
},
"b2": {
"title": "B2",
"type": "number"
},
"b3": {
"title": "B3",
"type": "number"
}
},
"required": [
"a0",
"a1",
"a2",
"a3",
"b1",
"b2",
"b3"
],
"title": "RationalCubicSpec",
"type": "object"
},
"SaturatingExponentialSpec": {
"additionalProperties": false,
"description": "Saturating exponential (NIST BoxBOD, positive-x).\n\n$\\mathrm{amplitude}\\cdot(1-\\exp(-\\mathrm{rate}\\cdot x))$.",
"properties": {
"model": {
"const": "saturating_exponential",
"default": "saturating_exponential",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"rate": {
"title": "Rate",
"type": "number"
}
},
"required": [
"amplitude",
"rate"
],
"title": "SaturatingExponentialSpec",
"type": "object"
},
"SplitGaussianSpec": {
"additionalProperties": false,
"description": "Split Gaussian with independent widths on each side.",
"properties": {
"model": {
"const": "split_gaussian",
"default": "split_gaussian",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma_l": {
"title": "Sigma L",
"type": "number"
},
"sigma_r": {
"title": "Sigma R",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma_l",
"sigma_r"
],
"title": "SplitGaussianSpec",
"type": "object"
},
"SplitPearson7Spec": {
"additionalProperties": false,
"description": "Split Pearson VII (split width + exponent each side).",
"properties": {
"model": {
"const": "split_pearson7",
"default": "split_pearson7",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma_l": {
"title": "Sigma L",
"type": "number"
},
"sigma_r": {
"title": "Sigma R",
"type": "number"
},
"m_l": {
"title": "M L",
"type": "number"
},
"m_r": {
"title": "M R",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma_l",
"sigma_r",
"m_l",
"m_r"
],
"title": "SplitPearson7Spec",
"type": "object"
},
"StepSpec": {
"additionalProperties": false,
"description": "Edge/step background (arctan / tanh / erfc).",
"properties": {
"model": {
"enum": [
"arctan_step",
"tanh_step",
"erfc_step"
],
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
}
},
"required": [
"model",
"amplitude",
"center",
"sigma"
],
"title": "StepSpec",
"type": "object"
},
"StudentsTSpec": {
"additionalProperties": false,
"description": "Student's-t peak ($\\nu$ degrees of freedom).",
"properties": {
"model": {
"const": "students_t",
"default": "students_t",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"center": {
"title": "Center",
"type": "number"
},
"sigma": {
"title": "Sigma",
"type": "number"
},
"nu": {
"title": "Nu",
"type": "number"
}
},
"required": [
"amplitude",
"center",
"sigma",
"nu"
],
"title": "StudentsTSpec",
"type": "object"
},
"TaucSpec": {
"additionalProperties": false,
"description": "Tauc optical band-gap edge (power-law absorption above ``e_gap``).",
"properties": {
"model": {
"const": "tauc",
"default": "tauc",
"title": "Model",
"type": "string"
},
"amplitude": {
"title": "Amplitude",
"type": "number"
},
"e_gap": {
"title": "E Gap",
"type": "number"
},
"exponent": {
"title": "Exponent",
"type": "number"
}
},
"required": [
"amplitude",
"e_gap",
"exponent"
],
"title": "TaucSpec",
"type": "object"
}
},
"additionalProperties": false,
"description": "A fully concrete, serializable benchmark-case declaration.",
"properties": {
"id": {
"title": "Id",
"type": "string"
},
"name": {
"title": "Name",
"type": "string"
},
"category": {
"title": "Category",
"type": "string"
},
"difficulty": {
"maximum": 1.0,
"minimum": 0.0,
"title": "Difficulty",
"type": "number"
},
"components": {
"items": {
"discriminator": {
"mapping": {
"arctan_step": "#/$defs/StepSpec",
"asym_ir": "#/$defs/AsymIrSpec",
"breit_wigner": "#/$defs/BreitWignerSpec",
"cauchy_dispersion": "#/$defs/CauchyDispersionSpec",
"constant": "#/$defs/ConstantSpec",
"doniach_sunjic": "#/$defs/AsymPeakSpec",
"double_exponential": "#/$defs/DecaySpec",
"erfc_step": "#/$defs/StepSpec",
"exp_gaussian": "#/$defs/AsymPeakSpec",
"exp_over_linear": "#/$defs/ExpOverLinearSpec",
"fano": "#/$defs/FanoSpec",
"gaussian": "#/$defs/GaussianSpec",
"generalised_logistic": "#/$defs/GeneralisedLogisticSpec",
"harmonic_ir": "#/$defs/HarmonicIrSpec",
"kww": "#/$defs/KwwSpec",
"linear": "#/$defs/LinearSpec",
"log_normal": "#/$defs/LogNormalSpec",
"lorentzian": "#/$defs/LorentzianSpec",
"mgh09_rational": "#/$defs/Mgh09RationalSpec",
"moffat": "#/$defs/MoffatSpec",
"pearson7": "#/$defs/Pearson7Spec",
"power_law_offset": "#/$defs/PowerLawOffsetSpec",
"power_saturation": "#/$defs/PowerSaturationSpec",
"pseudo_voigt": "#/$defs/PseudoVoigtSpec",
"quadratic": "#/$defs/QuadraticSpec",
"rational_cubic": "#/$defs/RationalCubicSpec",
"saturating_exponential": "#/$defs/SaturatingExponentialSpec",
"skewed_gaussian": "#/$defs/AsymPeakSpec",
"split_gaussian": "#/$defs/SplitGaussianSpec",
"split_pearson7": "#/$defs/SplitPearson7Spec",
"students_t": "#/$defs/StudentsTSpec",
"tanh_step": "#/$defs/StepSpec",
"tauc": "#/$defs/TaucSpec",
"true_voigt": "#/$defs/AsymPeakSpec",
"voigt": "#/$defs/PseudoVoigtSpec"
},
"propertyName": "model"
},
"oneOf": [
{
"$ref": "#/$defs/GaussianSpec"
},
{
"$ref": "#/$defs/LorentzianSpec"
},
{
"$ref": "#/$defs/PseudoVoigtSpec"
},
{
"$ref": "#/$defs/FanoSpec"
},
{
"$ref": "#/$defs/ConstantSpec"
},
{
"$ref": "#/$defs/LinearSpec"
},
{
"$ref": "#/$defs/QuadraticSpec"
},
{
"$ref": "#/$defs/StepSpec"
},
{
"$ref": "#/$defs/DecaySpec"
},
{
"$ref": "#/$defs/AsymPeakSpec"
},
{
"$ref": "#/$defs/LogNormalSpec"
},
{
"$ref": "#/$defs/Pearson7Spec"
},
{
"$ref": "#/$defs/SplitGaussianSpec"
},
{
"$ref": "#/$defs/MoffatSpec"
},
{
"$ref": "#/$defs/StudentsTSpec"
},
{
"$ref": "#/$defs/SplitPearson7Spec"
},
{
"$ref": "#/$defs/BreitWignerSpec"
},
{
"$ref": "#/$defs/AsymIrSpec"
},
{
"$ref": "#/$defs/HarmonicIrSpec"
},
{
"$ref": "#/$defs/TaucSpec"
},
{
"$ref": "#/$defs/CauchyDispersionSpec"
},
{
"$ref": "#/$defs/KwwSpec"
},
{
"$ref": "#/$defs/SaturatingExponentialSpec"
},
{
"$ref": "#/$defs/PowerSaturationSpec"
},
{
"$ref": "#/$defs/PowerLawOffsetSpec"
},
{
"$ref": "#/$defs/Mgh09RationalSpec"
},
{
"$ref": "#/$defs/RationalCubicSpec"
},
{
"$ref": "#/$defs/GeneralisedLogisticSpec"
},
{
"$ref": "#/$defs/ExpOverLinearSpec"
}
]
},
"title": "Components",
"type": "array"
},
"x_min": {
"title": "X Min",
"type": "number"
},
"x_max": {
"title": "X Max",
"type": "number"
},
"n_points": {
"title": "N Points",
"type": "integer"
},
"noise": {
"title": "Noise",
"type": "number"
},
"guess_scale": {
"default": 0.1,
"title": "Guess Scale",
"type": "number"
},
"solver_hint": {
"default": "lm",
"title": "Solver Hint",
"type": "string"
},
"recover": {
"default": true,
"title": "Recover",
"type": "boolean"
},
"featured": {
"default": false,
"title": "Featured",
"type": "boolean"
},
"landscape": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Landscape"
},
"condition": {
"anyOf": [
{
"type": "string"
},
{
"type": "null"
}
],
"default": null,
"title": "Condition"
},
"fixed_params": {
"additionalProperties": {
"items": {
"type": "string"
},
"type": "array"
},
"title": "Fixed Params",
"type": "object"
},
"expr_edges": {
"items": {
"additionalProperties": {
"type": "string"
},
"type": "object"
},
"title": "Expr Edges",
"type": "array"
},
"outliers": {
"anyOf": [
{
"$ref": "#/$defs/OutlierSpec"
},
{
"type": "null"
}
],
"default": null
}
},
"required": [
"id",
"name",
"category",
"difficulty",
"components",
"x_min",
"x_max",
"n_points",
"noise"
],
"title": "CaseSpec",
"type": "object"
}
Config:
extra:forbid
Fields:
-
id(str) -
name(str) -
category(str) -
difficulty(float) -
components(list[Component]) -
x_min(float) -
x_max(float) -
n_points(int) -
noise(float) -
guess_scale(float) -
solver_hint(str) -
recover(bool) -
featured(bool) -
landscape(str | None) -
condition(str | None) -
fixed_params(dict[str, list[str]]) -
expr_edges(list[dict[str, str]]) -
outliers(OutlierSpec | None)
condition = None
pydantic-field
¶
Qualitative scenario tag (e.g. "gaussian/noisy"). The anti-padding invariant
is: no two cases in a category share (frozenset(model keys), condition).
fixed_params
pydantic-field
¶
Per-node fixed parameter names: {"p0": ["center"]} means the center of
node p0 is held fixed at its guess value during the fit. Empty for all existing
cases (backward-compatible default).
outliers = None
pydantic-field
¶
Gross-outlier recipe (robust category only); None = clean data.
BenchCase
pydantic-model
¶
Bases: BaseModel
A materialized case: spec + generated data + truth/guess components.
Config:
extra:forbidfrozen:Truearbitrary_types_allowed:True
Fields:
-
spec(CaseSpec) -
x(Array) -
y(Array) -
comp_true(list[Component]) -
comp_guess(list[Component])
id
property
¶
Case id.
name
property
¶
Human-readable case name.
category
property
¶
Category id.
difficulty
property
¶
Difficulty in [0, 1].
solver_hint
property
¶
Solver key passed to spectrafit.
recover
property
¶
Whether parameter recovery is meaningful for this case.
featured
property
¶
Whether this is the featured case.
n_components
property
¶
Number of graph components.
n_peaks
property
¶
Alias for component count (the featured case is all-Gaussian peaks).
true_params
property
¶
Truth parameters keyed by dotted p{i}.param (graph order).
CaseFamily
pydantic-model
¶
Bases: BaseModel
Declarative generator: expands into count concrete :class:CaseSpecs.
Config:
extra:forbidarbitrary_types_allowed:True
Fields:
-
category(str) -
count(int) -
compose(Callable[[Random, int], tuple[list[Component], dict[str, Any]]]) -
x_range(tuple[float, float]) -
n_points(int) -
noise(float) -
difficulty(tuple[float, float]) -
guess_scale(float) -
solver_hint(str) -
name_template(str) -
points_scale_per_block(bool)
expand(rng)
¶
Deterministically expand into concrete case specs.
curve(x, comps)
¶
Evaluate and sum a list of typed components over the grid x.
materialize(spec)
¶
Generate the data + truth/guess components for spec (deterministic by id).
build_specs(seed=20260603)
¶
Expand all families into the concrete spec list (deterministic).
build_catalog(seed=20260603)
¶
Build the full deterministic catalog (materialized).
featured_case(catalog)
¶
Return the designated featured case (the reality tri-Gaussian).
oracles.models
¶
Model registry — the extensible catalogue of fittable shapes.
Each shape is registered once as a :class:PeakModel record that bundles
everything the engine needs: the numpy formula (data generation + scoring), the
lmfit shape callable (oracle), the spectrafit ModelType name (subject), the
canonical parameter names, and the shape's jax twin kernel from
:mod:oracles.jax_kernels (which is what makes :attr:PeakModel.jax_supported
true — the flag is derived, never stored). Adding a model is a one-record
registration — backends never hardcode per-model maps.
Conventions (docs/reference/models/index.md): amplitude is the peak value at center (not
area), sigma is the standard-deviation width, and the pseudo-Voigt mixing
weight is always fraction.
This module holds behaviour (formulas), so the registry records are Pydantic
models with Callable fields rather than plain dataclasses — keeping the whole
engine pydantic-first while still carrying code.
Note
:data:SHAPE_BOUNDS gives finite bounds for the long-tail shape
parameters an oracle's LM-family search can otherwise drive into a model
function's overflow/NaN region (e.g. pearson7 m -> 0+ makes
2^(1/m) overflow to NaN and aborts the solver). The truth-side
generators in cases.py draw from much narrower ranges; these
envelopes are ~10x wider so an oracle stays independent (it can still
disagree with spectrafit) but cannot reach a numerically degenerate
corner. Registered once here so the lmfit and scipy-ls oracle backends
read the same table instead of hand-syncing two private copies (the
CX-033 NaN cascade that originally motivated this table is the precedent
for keeping it registry-backed, not duplicated).
SHAPE_BOUNDS = {'m': (1.05, 50.0), 'm_l': (1.05, 50.0), 'm_r': (1.05, 50.0), 'beta': (0.1, 50.0), 'nu': (0.5, 200.0), 'q': (-100.0, 100.0), 'k': (-50.0, 50.0)}
module-attribute
¶
Keyed by shape-parameter name. See the module docstring's Note for the rationale.
- pearson7:
m - split_pearson7:
m_l/m_r - moffat:
beta - students_t:
nu - fano/breit_wigner:
q - asym_ir:
k
PeakModel
pydantic-model
¶
Bases: BaseModel
One registered fittable shape: formula + per-backend adapters + metadata.
Config:
extra:forbidfrozen:Truearbitrary_types_allowed:True
Fields:
-
key(str) -
spectrafit_type(str) -
param_names(tuple[str, ...]) -
evaluate(Callable[..., Array]) -
formula_latex(str) -
jax_evaluate(Callable[..., Any] | None) -
extra_defaults(dict[str, float])
Validators:
-
_spectrafit_type_is_known_member
key
pydantic-field
¶
Registry key, e.g. "gaussian" (also the catalog/case model field).
spectrafit_type
pydantic-field
¶
Name of the spectrafit ModelType enum member, e.g. "GAUSSIAN".
param_names
pydantic-field
¶
Canonical per-peak parameter names in order.
evaluate
pydantic-field
¶
evaluate(x, **params) -> y numpy formula for one peak.
formula_latex = ''
pydantic-field
¶
LaTeX formula string for this model shape (compact, uses the model reference's param names). Empty string for shapes without a simple closed-form (landscapes, etc.).
jax_evaluate = None
pydantic-field
¶
jax_evaluate(x, **params) -> y jax twin of :attr:evaluate, or None.
Registered from :mod:oracles.jax_kernels, whose bodies import jax
lazily — so carrying a kernel here costs nothing when the optional jax
extra is absent. Parity with :attr:evaluate is enforced by
tests/unit/test_jax_kernel_parity.py, which parametrises over this field.
extra_defaults = {}
pydantic-field
¶
Defaults for non-amplitude/center/sigma params (e.g. fraction).
jax_supported
property
¶
Whether the jax oracle implements this shape.
DERIVED from :attr:jax_evaluate rather than stored, so the flag the
jax backend gates on cannot drift out of sync with whether a kernel
actually exists (a stored bool could claim either direction wrongly).
A plain property, not a computed_field: :class:PeakModel is never
serialized (it carries Callable fields), and keeping it off the
schema means adding it changes no wire contract.
one(x, params)
¶
Evaluate a single peak from a param dict.
sum(x, peaks)
¶
Sum this model over a list of peak-parameter dicts (zeros if empty).
gaussian(x, amplitude, center, sigma)
¶
Gaussian peak: amplitude at center, std-dev sigma.
lorentzian(x, amplitude, center, sigma)
¶
Lorentzian peak normalized to amplitude at center (HWHM sigma).
pseudo_voigt(x, amplitude, center, sigma, fraction)
¶
Pseudo-Voigt: fraction-weighted Lorentzian + Gaussian mix (peak amplitude).
\(\mathrm{fraction}\cdot\text{Lorentzian} + (1-\mathrm{fraction})\cdot\text{Gaussian}\).
Formulas are inlined (not calls to :func:lorentzian / :func:gaussian)
so this hot-path body has no coupling to sibling function bodies.
fano(x, amplitude, center, gamma, q)
¶
Fano resonance \(A\cdot(q+\varepsilon)^2/(1+\varepsilon^2)\).
\(\varepsilon=(x-\mathrm{center})/\gamma\).
constant(x, c)
¶
Constant background c.
linear(x, slope, intercept)
¶
Linear background \(\mathrm{slope}\cdot x + \mathrm{intercept}\).
quadratic(x, amplitude, center, offset)
¶
Quadratic bowl \(A\cdot(x-\mathrm{center})^2 + \mathrm{offset}\).
arctan_step(x, amplitude, center, sigma)
¶
Arctan edge (rising).
\(A\cdot(\tfrac{1}{2} + \tfrac{1}{\pi}\arctan((x-\mathrm{center})/\mathrm{sigma}))\).
tanh_step(x, amplitude, center, sigma)
¶
Tanh edge (rising).
\((A/2)\cdot(1 + \tanh((x-\mathrm{center})/\mathrm{sigma}))\).
erfc_step(x, amplitude, center, sigma)
¶
Erfc edge (falling).
\((A/2)\cdot\mathrm{erfc}((x-\mathrm{center})/(\mathrm{sigma}\sqrt{2}))\).
double_exponential(x, A1, lam1, A2, lam2)
¶
Bi-exponential decay \(A_1 e^{-\lambda_1 x} + A_2 e^{-\lambda_2 x}\).
true_voigt(x, amplitude, center, sigma, gamma)
¶
True Voigt (Gaussian \(\otimes\) Lorentzian) via the Faddeeva fn.
Peak height amplitude.
Note
This formula (and its siblings in this section — skewed_gaussian,
emg, doniach) must stay IDENTICAL to the Rust kernels in
crates/spectrafit-models/src/{voigt_true,skewed_gaussian,emg,doniach}.rs
so numpy↔Rust kernel parity holds.
NOTE: the Rust kernel uses the Hui–Armstrong–Wray Faddeeva approximation
(~1e-6 accuracy) while the numpy fallback uses scipy.special.wofz, so
wheel-vs-numpy parity here is ~1e-4 — see test_kernel_parity.py.
skewed_gaussian(x, amplitude, center, sigma, gamma)
¶
Skewed Gaussian (\(\gamma\) = skew).
\(A\exp(-\tfrac12((x-c)/\sigma)^2)\cdot(1+\mathrm{erf}(\gamma(x-c)/(\sigma\sqrt2)))\).
exp_gaussian(x, amplitude, center, sigma, gamma)
¶
Exponentially-modified Gaussian (asymmetric tail); non-finite → 0 (Rust parity).
Numerically stable, overflow-free, and exact — no clamp. The naive form
\(\exp(\mathrm{arg\_exp})\cdot\mathrm{erfc}(z)\) overflows to \(\mathrm{inf}\cdot 0\) → NaN once arg_exp > 709
(e.g. gamma*sigma > 37). Using the algebraic identity
\(\mathrm{arg\_exp} - z^2 = -(x-\mathrm{center})^2/(2\sigma^2)\) we split on the sign of z:
- \(z \ge 0\): \(A\cdot(\gamma/2)\cdot\exp(-(x-\mathrm{center})^2/(2\sigma^2))\cdot\mathrm{erfcx}(z)\) — both factors are bounded (\(\mathrm{erfcx}(z) \in (0,1]\), Gaussian \(\le 1\)), so no overflow.
- \(z < 0\): \(A\cdot(\gamma/2)\cdot\exp(\mathrm{arg\_exp})\cdot\mathrm{erfc}(z)\) — here
arg_exp < 0soexpis safe, and \(\mathrm{erfc}(z) \in (1,2)\).
The branches are continuous at z = 0. scipy.special.erfcx is
machine-precision; the Rust kernel uses the identical split with a Cody
erfcx port, so numpy↔Rust parity holds to ~1e-9 even in the extreme tail.
doniach_sunjic(x, amplitude, center, sigma, gamma)
¶
Doniach–Šunjić lineshape (\(\gamma\) = asym), \(u=(x-c)/\sigma\).
\(A\cos[\pi\gamma/2+(1-\gamma)\arctan(u)]/(1+u^2)^{(1-\gamma)/2}\).
log_normal(x, amplitude, center, sigma)
¶
Log-normal peak, zero for x<=0.
\(A\exp(-(\ln(x/\mathrm{center}))^2/(2\sigma^2))\) for x>0 (else 0).
Numerically identical to the Rust log_normal kernel (the parity oracle):
amplitude is the peak height at x=center>0, sigma the log-space width.
pearson7(x, amplitude, center, sigma, m)
¶
Pearson VII; \(\sigma\) = HWHM, m→1 Lorentzian, m→\(\infty\) Gaussian.
\(A/[1+((x-c)/\sigma)^2\cdot(2^{1/m}-1)]^m\).
Numerically identical to the Rust pearson7 kernel (the parity oracle).
split_gaussian(x, amplitude, center, sigma_l, sigma_r)
¶
Split (asymmetric) Gaussian.
Width sigma_l for x<center, sigma_r for \(x \ge\) center. Covers
catalog #6 (asymmetric split-\(\sigma\) Gaussian) and #10 (bi-Gaussian) — the
same shape. Numerically identical to the Rust split_gaussian kernel.
moffat(x, amplitude, center, sigma, beta)
¶
Moffat profile; parity oracle for the Rust moffat kernel.
\(A/(((x-c)/\sigma)^2+1)^\beta\).
students_t(x, amplitude, center, sigma, nu)
¶
Student's-t lineshape; oracle for the Rust students_t kernel.
\(A/(1+((x-c)/\sigma)^2/\nu)^{(\nu+1)/2}\).
split_pearson7(x, amplitude, center, sigma_l, sigma_r, m_l, m_r)
¶
Split Pearson VII; oracle for the Rust kernel.
Split width and exponent, one side each of center.
breit_wigner(x, amplitude, center, sigma, q)
¶
Breit-Wigner-Fano lineshape, \(g=\sigma/2\); oracle for the Rust kernel.
\(A\cdot(qg+(x-c))^2/(g^2+(x-c)^2)\).
asym_ir(x, amplitude, center, sigma, k)
¶
Asymmetric IR band \(A\cdot G\cdot\text{sigmoid}\).
Sigmoid exponent clamped to match the Rust kernel.
harmonic_ir(x, amplitude, center, sigma)
¶
Harmonic-oscillator IR lineshape; oracle for the Rust harmonic_ir kernel.
\(A/((c^2-x^2)^2+(\sigma x)^2)\).
tauc(x, amplitude, e_gap, exponent)
¶
Tauc band-gap edge; oracle for the Rust kernel.
\(A\cdot(x-e_{\mathrm{gap}})^p\) for x>e_gap (else 0). Heaviside cut-off
at the gap keeps the fractional power real; numerically identical
to the Rust tauc kernel (np.where(x>e_gap, A*(x-e_gap)**p, 0)). Param order
(amplitude, e_gap, exponent) is identical on both sides — verified against
crates/spectrafit-models/src/tauc.rs::param_names during the C2 migration.
cauchy_dispersion(x, a, b, c)
¶
Cauchy dispersion; oracle for the Rust kernel.
\(n(x)=a+b/x^2+c/x^4\) for x>0 (else 0).
kww(x, amplitude, tau, beta)
¶
KWW stretched exponential; oracle for the Rust kernel.
\(A\exp(-(x/\tau)^\beta)\) for \(x \ge 0\) (else 0). The base \(x/\tau\) is
masked to a safe 1.0 where x<0 so the fractional power
never produces a NaN before np.where selects the 0 branch.
saturating_exponential(x, amplitude, rate)
¶
Saturating exponential (BoxBOD model).
\(\mathrm{amplitude}\cdot(1-\exp(-\mathrm{rate}\cdot x))\). Rises monotonically
from 0 toward amplitude with characteristic rate rate.
Numerically identical to the Rust kernel SaturatingExponential.
power_saturation(x, amplitude, rate)
¶
Power-law saturation (Misra1b model).
\(\mathrm{amplitude}\cdot(1-(1+\mathrm{rate}\cdot x/2)^{-2})\). Rises
monotonically from 0 toward amplitude with characteristic rate rate.
Numerically identical to the Rust kernel PowerSaturation.
power_law_offset(x, amplitude, offset, shape)
¶
Power-law with offset (Bennett5 model).
\(\mathrm{amplitude}\cdot(\mathrm{offset}+x)^{-1/\mathrm{shape}}\).
Numerically identical to the Rust kernel PowerLawOffset. The caller must
ensure offset + x > 0 for all data points; negative or zero arguments yield
nan (matching the Rust domain guard).
mgh09_rational(x, amplitude, num_lin, den_lin, den_const)
¶
Kowalik–Osborne rational function (NIST StRD MGH09 model).
Numerically identical to the Rust kernel Mgh09Rational. The denominator
must be non-zero; at the MGH09 certified parameters the discriminant
\(\mathrm{den\_lin}^2 - 4\cdot\mathrm{den\_const} < 0\), ensuring D > 0 for all x.
Param mapping to NIST b-parameters
amplitude = b1, num_lin = b2, den_lin = b3, den_const = b4
rational_cubic(x, a0, a1, a2, a3, b1, b2, b3)
¶
Rational cubic over cubic with the denominator constant pinned at 1.
Numerically identical to the Rust kernel RationalCubic. A lower-order
rational is this form with the unused coefficients at zero, so NIST StRD
Kirby2 (quadratic/quadratic) fixes \(a_3\) and \(b_3\) while Hahn1 and Thurber
(cubic/cubic) vary all seven.
The denominator constant is pinned rather than fitted: scaling numerator and denominator together leaves the curve unchanged, so fitting it would hand the solver an exact rank deficiency.
generalised_logistic(x, amplitude, shift, rate, shape)
¶
Generalised logistic (Richards) curve.
Numerically identical to the Rust kernel GeneralisedLogistic. NIST StRD
Rat43 uses all four parameters; Rat42 is this curve with \(\text{shape} = 1\).
exp_over_linear(x, rate, lin_const, lin_slope)
¶
Exponential decay over a line.
Numerically identical to the Rust kernel ExpOverLinear. This is the NIST
StRD Chwirut model, shared by Chwirut1 and Chwirut2.
register_model(model)
¶
Register model under its key (idempotent overwrite) and return it.
get_model(key)
¶
Look up a registered model by key.
wheel_parity_pairs()
¶
(wheel_key, model) pairs for the 29 MIGRATE-classified kernels.
The numpy evaluate bodies ARE the timing-fair oracle implementations
(lmfit / scipy-ls introspect and call them inside their timed fit loops —
they must never pay wheel/JSON overhead). Parity with the Rust kernels is
enforced by tests/unit/oracles/test_wheel_eval.py via _wheel_eval,
NOT by routing the hot path through the wheel.
voigt is a frozen copy of pseudo_voigt on the Python side, so it
maps to the pseudo_voigt wheel key here (the dedicated voigt Rust
kernel is cross-checked separately in the parity test).