NIST StRD nonlinear regression benchmark¶
Per-dataset agreement between each backend's fitted result and NIST's certified values, across all 22 NIST Statistical Reference Datasets (StRD) nonlinear regression problems this repository benchmarks against. See NIST StRD validation methodology for how this benchmark is run and what "agreement" means as a validation method.
Run conditions¶
- Start point:
START2(Start 2) - Tolerances:
ftol=1e-12,xtol=1e-12,gtol=1e-12 - Agreement measure: log-relative error (significant figures of agreement, capped at 15) of the least accurately recovered parameter against NIST's certified values; the sigma column applies the same measure to spectrafit-core's fitted standard errors against NIST's certified standard deviations
Results by dataset¶
Sorted by difficulty (Higher, then Average, then Lower), then by dataset name. Agreement values are significant figures (higher is better), capped at 15; an em-dash marks a dataset with no certified reference to compare against. Each dataset name links to NIST's own certified-value page for that problem.
| Dataset | Difficulty | Params | Points | spectrafit-core | lmfit | scipy-lm | scipy-trf | sigma |
|---|---|---|---|---|---|---|---|---|
| Bennett5 | Higher | 3 | 154 | 8.527 | 7.134 | 6.356 | 6.252 | 7.044 |
| BoxBOD | Higher | 2 | 6 | 7.151 | 6.877 | 7.172 | 7.376 | 7.102 |
| Eckerle4 | Higher | 3 | 35 | 9.591 | 6.871 | 9.567 | 9.115 | — |
| MGH09 (Kowalik and Osborne, 1978) | Higher | 4 | 11 | 6.498 | 5.709 | 6.200 | 5.784 | 6.613 |
| Rat42 | Higher | 3 | 9 | 8.519 | 6.171 | 8.018 | 8.001 | 8.138 |
| Rat43 | Higher | 4 | 15 | 7.005 | 4.630 | 6.681 | 6.610 | 7.272 |
| Thurber | Higher | 7 | 37 | 6.496 | 4.939 | 6.707 | 7.292 | 5.849 |
| Gauss3 | Average | 8 | 250 | 8.874 | 6.822 | 8.721 | 8.875 | 8.638 |
| Hahn1 | Average | 7 | 236 | 7.023 | 6.949 | 2.167 | 2.166 | 7.795 |
| Kirby2 | Average | 5 | 151 | 8.476 | 6.708 | 4.952 | 4.955 | 8.816 |
| Lanczos1 | Average | 6 | 24 | 8.747 | 10.563 | 10.553 | 10.490 | 0.599 |
| Lanczos2 | Average | 6 | 24 | 7.978 | 7.421 | 7.650 | 6.007 | 7.932 |
| MGH17 | Average | 5 | 33 | 8.515 | 5.762 | 6.754 | 6.683 | 8.068 |
| Roszman1 | Average | 4 | 25 | 8.162 | 6.075 | 6.582 | 6.541 | 8.665 |
| Chwirut1 | Lower | 3 | 214 | 7.348 | 6.188 | 7.481 | 7.420 | 7.702 |
| Chwirut2 | Lower | 3 | 54 | 7.181 | 6.646 | 7.170 | 7.203 | 7.599 |
| DanWood | Lower | 2 | 6 | 10.594 | 7.176 | 9.199 | 9.518 | — |
| Gauss1 | Lower | 8 | 250 | 10.582 | 7.122 | 8.088 | 8.075 | 10.540 |
| Gauss2 | Lower | 8 | 250 | 9.917 | 6.845 | 8.952 | 8.986 | 9.910 |
| Lanczos3 | Lower | 6 | 24 | 7.881 | 5.924 | 5.871 | 5.005 | 7.763 |
| Misra1a | Lower | 2 | 14 | 10.298 | 8.029 | 7.878 | 7.775 | 10.011 |
| Misra1b | Lower | 2 | 14 | 10.302 | 7.876 | 7.286 | 7.260 | 10.253 |
Summary¶
Minimum agreement across all 22 datasets, and how many of the evaluated datasets each backend clears 4 significant figures of agreement on (the "evaluated" denominator excludes datasets with no certified reference for that column — see the sigma row).
| Backend | Minimum agreement | Clears 4 sig figs |
|---|---|---|
| spectrafit-core | 6.496 | 22/22 |
| lmfit | 4.630 | 22/22 |
| scipy-lm | 2.167 | 21/22 |
| scipy-trf | 2.166 | 21/22 |
| sigma | 0.599 | 19/20 |
Dataset difficulty breakdown: 8 Lower / 7 Average / 7 Higher.