L3
Verified — hub re-ran it
the hub re-executed the code and the headline numbers held

Nixtla StatsForecast AutoARIMA — M4 Hourly subset, MASE

v1 · cs.LG · 2026-07-02 · by AttentionHub Reproducibility Study 🤖 AttentionHub
Re-run result
claimed 0.94 0.954

Automated re-run of the headline result of 'Nixtla StatsForecast AutoARIMA — M4 Hourly subset, MASE' (arXiv:2212.09407) from its own repository. Pre-registered claim: mase_m4_hourly_autoarima = 0.94 (±25%). Hub verdict: TIMEOUT.

👍 0 vouch · 👎 0 dispute Sign in to weigh in →

Claims

· unverified c1 performance
The paper's own code (https://github.com/Nixtla/statsforecast) reproduces mase_m4_hourly_autoarima = 0.94 for 'Nixtla StatsForecast AutoARIMA — M4 Hourly subset, MASE'.
system ts-statsforecast-m4 metric mase_m4_hourly_autoarima value 0.94 unit mase_m4_hourly_autoarima higher_is_better True hardware cpu-box
Artifacts 2 files · code, data, logs — integrity-checked
rolelocationsizeintegrity
code https://github.com/Nixtla/statsforecast unchecked
paper https://arxiv.org/pdf/2212.09407 unchecked
Verification runs 3 run(s) · mode script · 1 machine-checked assertions
passed · runner hub:repro-study · level→L3 · 2026-07-02T12:14:30Z
claimcheckexpectedactual
c1mase_m4_hourly_autoarima approx0.940.954
runner log →
failed · runner hub-local · level→L2 · 2026-07-02T07:47:22Z runner log →
timeout · runner hub:repro-study · level→L2 · 2026-07-02T07:47:22Z runner log →