L1
Integrity-checked
claims carry evidence but haven't been re-run (attested)
A Simple but Tough-to-Beat Baseline for Sentence Embeddings (SIF, Arora et al., ICLR 2017) — STS Pearson correlation
Automated re-run of the headline result of 'A Simple but Tough-to-Beat Baseline for Sentence Embeddings (SIF, Arora et al., ICLR 2017) — STS Pearson correlation' (arXiv:1611.01462) from its own repository. Pre-registered claim: sts_pearson = 0.717 (±10%). Hub verdict: RUN_FAILED.
Claims
· unverified
c1
performance
The paper's own code (https://github.com/PrincetonML/SIF) reproduces sts_pearson = 0.717 for 'A Simple but Tough-to-Beat Baseline for Sentence Embeddings (SIF, Arora et al., ICLR 2017) — STS Pearson correlation'.
system nlp-sif-sts-correlation
metric sts_pearson
value 0.717
unit sts_pearson
higher_is_better True
hardware cpu-box
Artifacts 2 files · code, data, logs — integrity-checked
| role | location | size | integrity |
|---|---|---|---|
| code | https://github.com/PrincetonML/SIF | — | unchecked |
| paper | https://arxiv.org/pdf/1611.01462 | — | unchecked |
Verification runs 3 run(s) · mode script · 1 machine-checked assertions
failed · runner hub:repro-study · level→L1 · 2026-07-02T11:39:48Z
runner log →
failed · runner hub-local · level→L2 · 2026-07-02T07:47:21Z
runner log →
failed · runner hub:repro-study · level→L1 · 2026-07-02T07:47:21Z
runner log →