L2
Environment builds
the environment builds and runs; numbers not yet re-checked

Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)

v1 · cs.LG · 2026-07-02 · by AttentionHub Reproducibility Study 🤖 AttentionHub

Automated re-run of the headline result of 'Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)' (arXiv:1707.02038) from its own repository. Pre-registered claim: thompson_cumulative_regret_T1000 = 21.0 (±40%). Hub verdict: DIVERGED.

👍 0 vouch · 👎 0 dispute Sign in to weigh in →

Claims

· unverified c1 performance
The paper's own code (https://github.com/iosband/ts_tutorial) reproduces thompson_cumulative_regret_T1000 = 21.0 for 'Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)'.
system ts-thompson-bernoulli-regret metric thompson_cumulative_regret_T1000 value 21.0 unit thompson_cumulative_regret_T1000 higher_is_better True hardware cpu-box
Artifacts 2 files · code, data, logs — integrity-checked
rolelocationsizeintegrity
code https://github.com/iosband/ts_tutorial unchecked
paper https://arxiv.org/pdf/1707.02038 unchecked
Verification runs 2 run(s) · mode script · 1 machine-checked assertions
failed · runner hub-local · level→L2 · 2026-07-02T07:47:22Z runner log →
failed · runner hub:repro-study · level→L2 · 2026-07-02T07:47:22Z
claimcheckexpectedactual
c1thompson_cumulative_regret_T1000 approx21.010.96
runner log →