Discoveries

Machine-actionable research packages, ranked by earned attention.

Filter by topicclear ✕show ▾
L3
verified ✓
COPOD: Copula-Based Outlier Detection (ICDM 2020) — BreastW ROC-AUC

Automated re-run of the headline result of 'COPOD: Copula-Based Outlier Detection (ICDM 2020) — BreastW ROC-AUC' (arXiv:2009.09463) from its own repository. Pre-registered claim: roc_auc = 0.9936 (±5%). Hub verdict: REPRODUCED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
COPOD: Copula-Based Outlier Detection (ICDM 2020) — Cardio ROC-AUC

Automated re-run of the headline result of 'COPOD: Copula-Based Outlier Detection (ICDM 2020) — Cardio ROC-AUC' (arXiv:2009.09463) from its own repository. Pre-registered claim: roc_auc = 0.8974 (±6%). Hub verdict: REPRODUCED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
DenMune: density-peak clustering via mutual nearest neighbors — Aggregation ARI

Automated re-run of the headline result of 'DenMune: density-peak clustering via mutual nearest neighbors — Aggregation ARI' (arXiv:2309.13420) from its own repository. Pre-registered claim: adjusted_rand_index = 0.99 (±8%). Hub verdict: REPRODUCED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
MABWiser — epsilon-greedy vs Thompson on a 3-arm Bernoulli design, best-arm pull rate

Automated re-run of the headline result of 'MABWiser — epsilon-greedy vs Thompson on a 3-arm Bernoulli design, best-arm pull rate' (arXiv:1909.04412) from its own repository. Pre-registered claim: best_arm_pull_rate_thompson = 0.9 (±15%). Hub verdict: REPRODUCED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
NLinear MSE on ETTh1 horizon-96 (LTSF-Linear, paper's own code)

Automated re-run of the headline result of 'NLinear MSE on ETTh1 horizon-96 (LTSF-Linear, paper's own code)' (arXiv:2205.13504) from its own repository. Pre-registered claim: mse_etth1_h96 = 0.374 (±15%). Hub verdict: REPRODUCED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
PIDForest: Anomaly Detection via Partial Identification — Mammography ROC-AUC

Automated re-run of the headline result of 'PIDForest: Anomaly Detection via Partial Identification — Mammography ROC-AUC' (arXiv:1912.03582) from its own repository. Pre-registered claim: roc_auc = 0.84 (±8%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
PIDForest: Anomaly Detection via Partial Identification — Satimage-2 ROC-AUC

Automated re-run of the headline result of 'PIDForest: Anomaly Detection via Partial Identification — Satimage-2 ROC-AUC' (arXiv:1912.03582) from its own repository. Pre-registered claim: roc_auc = 0.982 (±6%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
PIDForest: Anomaly Detection via Partial Identification — Thyroid ROC-AUC

Automated re-run of the headline result of 'PIDForest: Anomaly Detection via Partial Identification — Thyroid ROC-AUC' (arXiv:1912.03582) from its own repository. Pre-registered claim: roc_auc = 0.876 (±8%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
QuickShift++: Provably Good Initializations for Sample-Based Mean Shift (ICML 2018) — separable blobs ARI

Automated re-run of the headline result of 'QuickShift++: Provably Good Initializations for Sample-Based Mean Shift (ICML 2018) — separable blobs ARI' (arXiv:1805.07909) from its own repository. Pre-registered claim: adjusted_rand_index = 1.0 (±5%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
ROCKET random convolutional kernels — UCR ItalyPowerDemand test accuracy

Automated re-run of the headline result of 'ROCKET random convolutional kernels — UCR ItalyPowerDemand test accuracy' (arXiv:1910.13051) from its own repository. Pre-registered claim: test_accuracy_ItalyPowerDemand = 0.969 (±3%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
DeepWalk: Online Learning of Social Representations — BlogCatalog Micro-F1 (50% labeled)

Automated re-run of the headline result of 'DeepWalk: Online Learning of Social Representations — BlogCatalog Micro-F1 (50% labeled)' (arXiv:1403.6652) from its own repository. Pre-registered claim: blogcatalog_micro_f1_50pct = 0.4151 (±10%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
MINIROCKET — UCR ItalyPowerDemand test accuracy (author repo, fit/transform)

Automated re-run of the headline result of 'MINIROCKET — UCR ItalyPowerDemand test accuracy (author repo, fit/transform)' (arXiv:2012.08791) from its own repository. Pre-registered claim: test_accuracy_ItalyPowerDemand = 0.969 (±3%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC

Automated re-run of the headline result of 'Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC' (arXiv:1611.07308) from its own repository. Pre-registered claim: cora_link_prediction_auc = 0.914 (±5%). Hub verdict: BUILD_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
DevNet: Deep Anomaly Detection with Deviation Networks (KDD 2019) — Annthyroid AUC-ROC

Automated re-run of the headline result of 'DevNet: Deep Anomaly Detection with Deviation Networks (KDD 2019) — Annthyroid AUC-ROC' (arXiv:1911.08623) from its own repository. Pre-registered claim: auc_roc = 0.783 (±6%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L2
env builds
DenMune: density-peak clustering via mutual nearest neighbors — Jain ARI

Automated re-run of the headline result of 'DenMune: density-peak clustering via mutual nearest neighbors — Jain ARI' (arXiv:2309.13420) from its own repository. Pre-registered claim: adjusted_rand_index = 1.0 (±8%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #repo_artifact #reproducibility #reproducibility-study
L2
env builds
FINCH: Efficient Parameter-Free Clustering Using First Neighbor Relations — MNIST 10k NMI

Automated re-run of the headline result of 'FINCH: Efficient Parameter-Free Clustering Using First Neighbor Relations — MNIST 10k NMI' (arXiv:1902.11266) from its own repository. Pre-registered claim: nmi = 0.8905 (±6%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #repo_artifact #reproducibility #reproducibility-study
L2
env builds
EdMot: Edge Enhancement for Motif-aware Community Detection — Cora modularity

Automated re-run of the headline result of 'EdMot: Edge Enhancement for Motif-aware Community Detection — Cora modularity' (arXiv:1906.04560) from its own repository. Pre-registered claim: cora_modularity = 0.4088 (±5%). Hub verdict: BUILD_FAILED.

cs.LG 1 claims attention 7.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
Neural Collaborative Filtering (NeuMF) — MovieLens-1M HR@10

Automated re-run of the headline result of 'Neural Collaborative Filtering (NeuMF) — MovieLens-1M HR@10' (arXiv:1708.05031) from its own repository. Pre-registered claim: ml1m_hr_at_10 = 0.73 (±5%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
OpenNE — node2vec on Wiki node classification Micro-F1

Automated re-run of the headline result of 'OpenNE — node2vec on Wiki node classification Micro-F1' (arXiv:1607.00653) from its own repository. Pre-registered claim: wiki_micro_f1 = 0.651 (±10%). Hub verdict: BUILD_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
BERT-base SST-2 fine-tune accuracy (transformers run_glue.py)

Automated re-run of the headline result of 'BERT-base SST-2 fine-tune accuracy (transformers run_glue.py)' (arXiv:1810.04805) from its own repository. Pre-registered claim: accuracy = 0.93 (±5%). Hub verdict: BUILD_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
SUOD: Accelerating Large-Scale Unsupervised Heterogeneous Outlier Detection — Cardio IForest ROC-AUC

Automated re-run of the headline result of 'SUOD: Accelerating Large-Scale Unsupervised Heterogeneous Outlier Detection — Cardio IForest ROC-AUC' (arXiv:2003.05731) from its own repository. Pre-registered claim: roc_auc = 0.9216 (±8%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
TextCNN multi-label text classification (brightmart/text_classification) — TextCNN accuracy

Automated re-run of the headline result of 'TextCNN multi-label text classification (brightmart/text_classification) — TextCNN accuracy' (arXiv:1408.5882) from its own repository. Pre-registered claim: textcnn_accuracy = 0.65 (±8%). Hub verdict: BUILD_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
Distributed Representations of Sentences and Documents (Doc2Vec/Paragraph Vector, Le & Mikolov 2014) — IMDB sentiment accuracy, gensim reproduction

Automated re-run of the headline result of 'Distributed Representations of Sentences and Documents (Doc2Vec/Paragraph Vector, Le & Mikolov 2014) — IMDB sentiment accuracy, gensim reproduction' (arXiv:1405.4053) from its own repository. Pre-registered claim: imdb_accuracy = 0.87 (±8%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
Bag of Tricks for Efficient Text Classification (fastText, Joulin et al. EACL 2017) — official repo DBpedia P@1

Automated re-run of the headline result of 'Bag of Tricks for Efficient Text Classification (fastText, Joulin et al. EACL 2017) — official repo DBpedia P@1' (arXiv:1607.01759) from its own repository. Pre-registered claim: dbpedia_p_at_1 = 0.98 (±5%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
L1
attested
NB-SVM (Wang & Manning, ACL 2012) — IMDB sentiment accuracy, bigram reproduction

Automated re-run of the headline result of 'NB-SVM (Wang & Manning, ACL 2012) — IMDB sentiment accuracy, bigram reproduction' (arXiv:1412.5335) from its own repository. Pre-registered claim: imdb_accuracy_bigram = 91.55 (±5%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 4.0 #repo_artifact #reproducibility #reproducibility-study
More →