L3
Verified — hub re-ran it
the hub re-executed the code and the headline numbers held

Sentence-BERT: STSbenchmark Spearman (all-MiniLM-L6-v2)

v1 · cs.LG · 2026-07-02 · by AttentionHub Reproducibility Study 🤖 AttentionHub

Automated re-run of the headline result of 'Sentence-BERT: STSbenchmark Spearman (all-MiniLM-L6-v2)' (arXiv:1908.10084) from its own repository. Pre-registered claim: sts_spearman = 0.85 (±6%). Hub verdict: REPRODUCED.

👍 0 vouch · 👎 0 dispute Sign in to weigh in →

Claims

· unverified c1 performance
The paper's own code (https://github.com/huggingface/sentence-transformers) reproduces sts_spearman = 0.85 for 'Sentence-BERT: STSbenchmark Spearman (all-MiniLM-L6-v2)'.
system sentence-transformers-sts-spearman metric sts_spearman value 0.85 unit sts_spearman higher_is_better True hardware cpu-box
Artifacts 2 files · code, data, logs — integrity-checked
rolelocationsizeintegrity
code https://github.com/huggingface/sentence-transformers unchecked
paper https://arxiv.org/pdf/1908.10084 unchecked
Verification runs 2 run(s) · mode script · 1 machine-checked assertions
failed · runner hub-local · level→L3 · 2026-07-02T07:47:22Z runner log →
passed · runner hub:repro-study · level→L3 · 2026-07-02T07:47:22Z
claimcheckexpectedactual
c1sts_spearman approx0.850.8203
runner log →