Skip to content

NIST StRD nonlinear regression benchmark

Per-dataset agreement between each backend's fitted result and NIST's certified values, across all 22 NIST Statistical Reference Datasets (StRD) nonlinear regression problems this repository benchmarks against. See NIST StRD validation methodology for how this benchmark is run and what "agreement" means as a validation method.

Run conditions

  • Start point: START2 (Start 2)
  • Tolerances: ftol=1e-12, xtol=1e-12, gtol=1e-12
  • Agreement measure: log-relative error (significant figures of agreement, capped at 15) of the least accurately recovered parameter against NIST's certified values; the sigma column applies the same measure to spectrafit-core's fitted standard errors against NIST's certified standard deviations

Results by dataset

Sorted by difficulty (Higher, then Average, then Lower), then by dataset name. Agreement values are significant figures (higher is better), capped at 15; an em-dash marks a dataset with no certified reference to compare against. Each dataset name links to NIST's own certified-value page for that problem.

Dataset Difficulty Params Points spectrafit-core lmfit scipy-lm scipy-trf sigma
Bennett5 Higher 3 154 8.527 7.134 6.356 6.252 7.044
BoxBOD Higher 2 6 7.151 6.877 7.172 7.376 7.102
Eckerle4 Higher 3 35 9.591 6.871 9.567 9.115 —
MGH09 (Kowalik and Osborne, 1978) Higher 4 11 6.498 5.709 6.200 5.784 6.613
Rat42 Higher 3 9 8.519 6.171 8.018 8.001 8.138
Rat43 Higher 4 15 7.005 4.630 6.681 6.610 7.272
Thurber Higher 7 37 6.496 4.939 6.707 7.292 5.849
Gauss3 Average 8 250 8.874 6.822 8.721 8.875 8.638
Hahn1 Average 7 236 7.023 6.949 2.167 2.166 7.795
Kirby2 Average 5 151 8.476 6.708 4.952 4.955 8.816
Lanczos1 Average 6 24 8.747 10.563 10.553 10.490 0.599
Lanczos2 Average 6 24 7.978 7.421 7.650 6.007 7.932
MGH17 Average 5 33 8.515 5.762 6.754 6.683 8.068
Roszman1 Average 4 25 8.162 6.075 6.582 6.541 8.665
Chwirut1 Lower 3 214 7.348 6.188 7.481 7.420 7.702
Chwirut2 Lower 3 54 7.181 6.646 7.170 7.203 7.599
DanWood Lower 2 6 10.594 7.176 9.199 9.518 —
Gauss1 Lower 8 250 10.582 7.122 8.088 8.075 10.540
Gauss2 Lower 8 250 9.917 6.845 8.952 8.986 9.910
Lanczos3 Lower 6 24 7.881 5.924 5.871 5.005 7.763
Misra1a Lower 2 14 10.298 8.029 7.878 7.775 10.011
Misra1b Lower 2 14 10.302 7.876 7.286 7.260 10.253

Summary

Minimum agreement across all 22 datasets, and how many of the evaluated datasets each backend clears 4 significant figures of agreement on (the "evaluated" denominator excludes datasets with no certified reference for that column — see the sigma row).

Backend Minimum agreement Clears 4 sig figs
spectrafit-core 6.496 22/22
lmfit 4.630 22/22
scipy-lm 2.167 21/22
scipy-trf 2.166 21/22
sigma 0.599 19/20

Dataset difficulty breakdown: 8 Lower / 7 Average / 7 Higher.