1. Trace production code
Identify the calculation kernel, preprocessing rules, missing-data handling, and every user-facing output path.
Calculation quality and reproducibility
DataStatPro traces production calculation code, runs deterministic fixtures, and compares normalized results with independently executed reference calculations where supported.
Reviewed by Prof. Dr. Nadeem Shafique Butt, Professor of Biostatistics · Last updated August 15, 2026
Identify the calculation kernel, preprocessing rules, missing-data handling, and every user-facing output path.
Use deterministic datasets that exercise ordinary results, boundaries, missing values, categorical coding, and relevant failure modes.
Execute equivalent calculations in Python, R, and IBM SPSS Statistics when the required reference procedure is available.
Record passes, documented convention differences, and unavailable reference coverage separately. Defects are corrected and re-tested.
The following July 27, 2026 audit families have normalized comparison records in the project quality archive. Counts are individual fixture-to-engine comparisons, not counts of distinct statistical procedures.
| Audited family | Pass | Convention review | Reference unavailable | What the status means |
|---|---|---|---|---|
| Publication Table 1 family | 27 | 3 | 0 | Python, R, and SPSS comparisons; three quartile results reflect a documented software-convention difference. |
| Publication Tables 2A and 3 | 27 | 1 | 5 | Available reference outputs matched; unsupported reference outputs and one convention difference remain explicitly separated. |
| Comparative regression | 33 | 0 | 6 | All available normalized Python, R, and SPSS reference comparisons passed. |
| Design of experiments calculations | 50 | 0 | 28 | All available comparisons passed; unavailable items represent reference-coverage gaps, not passes. |
A separate regression-table audit reconciled linear fixtures with R and Python, logistic fixtures with R and an independent Python implementation, and verified focused production tests. Its SPSS syntax was generated but not successfully executed, and Cox estimation was not independently reconciled in that audit.
Before using a result in a thesis, manuscript, clinical study, regulated workflow, or operational decision, confirm the estimator, coding, missing-data rule, weighting, confidence level, correction method, convergence status, software version, and reporting requirements. Reproduce at least one representative analysis in an accepted reference package when the consequences of error are material.
DataStatPro and DSRConsult LLC are independent of Python, R, IBM, SPSS, and their respective maintainers. Reference-engine names identify comparison environments and do not imply endorsement.