Discoveries

Machine-actionable research packages, ranked by earned attention.

Filter by topicshow ▾
L3
verified ✓
PIDForest: Anomaly Detection via Partial Identification — Thyroid ROC-AUC

Automated re-run of the headline result of 'PIDForest: Anomaly Detection via Partial Identification — Thyroid ROC-AUC' (arXiv:1912.03582) from its own repository. Pre-registered claim: roc_auc = 0.876 (±8%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
QuickShift++: Provably Good Initializations for Sample-Based Mean Shift (ICML 2018) — separable blobs ARI

Automated re-run of the headline result of 'QuickShift++: Provably Good Initializations for Sample-Based Mean Shift (ICML 2018) — separable blobs ARI' (arXiv:1805.07909) from its own repository. Pre-registered claim: adjusted_rand_index = 1.0 (±5%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
Darts ExponentialSmoothing — AirPassengers validation MAPE (quickstart claim)

Automated re-run of the headline result of 'Darts ExponentialSmoothing — AirPassengers validation MAPE (quickstart claim)' (arXiv:2110.03224) from its own repository. Pre-registered claim: mape_airpassengers_val = 5.11 (±30%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 10.0 #pip_env #reproducibility #reproducibility-study
L3
verified ✓
DLinear (LTSF-Linear) — ETTh1 horizon-96 MSE (real author training script; CPU/long-train probe)

Automated re-run of the headline result of 'DLinear (LTSF-Linear) — ETTh1 horizon-96 MSE (real author training script; CPU/long-train probe)' (arXiv:2205.13504) from its own repository. Pre-registered claim: mse_etth1_h96 = 0.375 (±15%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 10.0 #pip #reproducibility #reproducibility-study
L3
verified ✓
Pyro NUTS — Eight Schools hierarchical model, posterior mean of mu

Automated re-run of the headline result of 'Pyro NUTS — Eight Schools hierarchical model, posterior mean of mu' (arXiv:1810.09538) from its own repository. Pre-registered claim: posterior_mean_mu_eight_schools = 4.4 (±60%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 10.0 #pip #reproducibility #reproducibility-study
L3
verified ✓
Nixtla StatsForecast AutoARIMA — M4 Hourly subset, MASE

Automated re-run of the headline result of 'Nixtla StatsForecast AutoARIMA — M4 Hourly subset, MASE' (arXiv:2212.09407) from its own repository. Pre-registered claim: mase_m4_hourly_autoarima = 0.94 (±25%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 10.0 #pip #reproducibility #reproducibility-study
L3
verified ✓
XGBoost: A Scalable Tree Boosting System — HIGGS test AUC

Automated re-run of the headline result of 'XGBoost: A Scalable Tree Boosting System — HIGGS test AUC' (arXiv:1603.02754) from its own repository. Pre-registered claim: auc = 0.84 (±5%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #pip_env #reproducibility #reproducibility-study
L3
verified ✓
ROCKET random convolutional kernels — UCR ItalyPowerDemand test accuracy

Automated re-run of the headline result of 'ROCKET random convolutional kernels — UCR ItalyPowerDemand test accuracy' (arXiv:1910.13051) from its own repository. Pre-registered claim: test_accuracy_ItalyPowerDemand = 0.969 (±3%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
DeepWalk: Online Learning of Social Representations — BlogCatalog Micro-F1 (50% labeled)

Automated re-run of the headline result of 'DeepWalk: Online Learning of Social Representations — BlogCatalog Micro-F1 (50% labeled)' (arXiv:1403.6652) from its own repository. Pre-registered claim: blogcatalog_micro_f1_50pct = 0.4151 (±10%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
MINIROCKET — UCR ItalyPowerDemand test accuracy (author repo, fit/transform)

Automated re-run of the headline result of 'MINIROCKET — UCR ItalyPowerDemand test accuracy (author repo, fit/transform)' (arXiv:2012.08791) from its own repository. Pre-registered claim: test_accuracy_ItalyPowerDemand = 0.969 (±3%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC

Automated re-run of the headline result of 'Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC' (arXiv:1611.07308) from its own repository. Pre-registered claim: cora_link_prediction_auc = 0.914 (±5%). Hub verdict: BUILD_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L3
verified ✓
DevNet: Deep Anomaly Detection with Deviation Networks (KDD 2019) — Annthyroid AUC-ROC

Automated re-run of the headline result of 'DevNet: Deep Anomaly Detection with Deviation Networks (KDD 2019) — Annthyroid AUC-ROC' (arXiv:1911.08623) from its own repository. Pre-registered claim: auc_roc = 0.783 (±6%). Hub verdict: RUN_FAILED.

cs.LG 1 claims attention 10.0 #repo_artifact #reproducibility #reproducibility-study
L2
env builds
word2vec via gensim: Google analogy accuracy (skip-gram, text8 corpus)

Automated re-run of the headline result of 'word2vec via gensim: Google analogy accuracy (skip-gram, text8 corpus)' (arXiv:1301.3781) from its own repository. Pre-registered claim: analogy_accuracy = 0.6 (±40%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #pip_env #reproducibility #reproducibility-study
L2
env builds
Graph Attention Networks (pyGAT) — Cora test accuracy

Automated re-run of the headline result of 'Graph Attention Networks (pyGAT) — Cora test accuracy' (arXiv:1710.10903) from its own repository. Pre-registered claim: cora_test_accuracy = 0.84 (±5%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip #reproducibility #reproducibility-study
L2
env builds
DenMune: density-peak clustering via mutual nearest neighbors — Jain ARI

Automated re-run of the headline result of 'DenMune: density-peak clustering via mutual nearest neighbors — Jain ARI' (arXiv:2309.13420) from its own repository. Pre-registered claim: adjusted_rand_index = 1.0 (±8%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #repo_artifact #reproducibility #reproducibility-study
L2
env builds
FINCH: Efficient Parameter-Free Clustering Using First Neighbor Relations — MNIST 10k NMI

Automated re-run of the headline result of 'FINCH: Efficient Parameter-Free Clustering Using First Neighbor Relations — MNIST 10k NMI' (arXiv:1902.11266) from its own repository. Pre-registered claim: nmi = 0.8905 (±6%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #repo_artifact #reproducibility #reproducibility-study
L2
env builds
SCAPT-ABSA: Supervised Contrastive Pre-Training for Aspect-based Sentiment (Li et al., EMNLP 2021) — SemEval2014 Restaurant accuracy

Automated re-run of the headline result of 'SCAPT-ABSA: Supervised Contrastive Pre-Training for Aspect-based Sentiment (Li et al., EMNLP 2021) — SemEval2014 Restaurant accuracy' (arXiv:2111.02194) from its own repository. Pre-registered claim: restaurant_accuracy = 90.0 (±5%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip #reproducibility #reproducibility-study
L2
env builds
SimCSE: Simple Contrastive Learning of Sentence Embeddings (Gao et al., EMNLP 2021) — unsup BERT-base STS Avg Spearman

Automated re-run of the headline result of 'SimCSE: Simple Contrastive Learning of Sentence Embeddings (Gao et al., EMNLP 2021) — unsup BERT-base STS Avg Spearman' (arXiv:2104.08821) from its own repository. Pre-registered claim: sts_avg_spearman = 76.25 (±5%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip #reproducibility #reproducibility-study
L2
env builds
Convolutional Neural Networks for Sentence Classification (Kim 2014), PyTorch reimpl — MR CNN-rand accuracy

Automated re-run of the headline result of 'Convolutional Neural Networks for Sentence Classification (Kim 2014), PyTorch reimpl — MR CNN-rand accuracy' (arXiv:1408.5882) from its own repository. Pre-registered claim: mr_accuracy_cnn_rand = 76.1 (±8%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip #reproducibility #reproducibility-study
L2
env builds
node2vec: Scalable Feature Learning for Networks — link-prediction AUC

Automated re-run of the headline result of 'node2vec: Scalable Feature Learning for Networks — link-prediction AUC' (arXiv:1607.00653) from its own repository. Pre-registered claim: link_prediction_auc = 0.97 (±10%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #pip_env #reproducibility #reproducibility-study
L2
env builds
N-BEATS — M4 Yearly ensemble sMAPE (author repo; GPU/long-train probe)

Automated re-run of the headline result of 'N-BEATS — M4 Yearly ensemble sMAPE (author repo; GPU/long-train probe)' (arXiv:1905.10437) from its own repository. Pre-registered claim: smape_m4_yearly = 13.114 (±5%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip #reproducibility #reproducibility-study
L2
env builds
N-HiTS long-horizon forecasting — ETTm2 horizon-96 MAE (CPU feasibility / GPU-need probe)

Automated re-run of the headline result of 'N-HiTS long-horizon forecasting — ETTm2 horizon-96 MAE (CPU feasibility / GPU-need probe)' (arXiv:2201.12886) from its own repository. Pre-registered claim: mae_ettm2_h96 = 0.255 (±20%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip #reproducibility #reproducibility-study
L2
env builds
Stable-Baselines3 DQN on MountainCar-v0 — mean episode reward (RL Zoo benchmark)

Automated re-run of the headline result of 'Stable-Baselines3 DQN on MountainCar-v0 — mean episode reward (RL Zoo benchmark)' (arXiv:2006.05990) from its own repository. Pre-registered claim: mean_episode_reward_mountaincar = -100.849 (±12%). Hub verdict: TIMEOUT.

cs.LG 1 claims attention 7.0 #pip_env #reproducibility #reproducibility-study
L2
env builds
Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)

Automated re-run of the headline result of 'Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)' (arXiv:1707.02038) from its own repository. Pre-registered claim: thompson_cumulative_regret_T1000 = 21.0 (±40%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #pip_env #reproducibility #reproducibility-study
L2
env builds
UMAP: Uniform Manifold Approximation and Projection — MNIST 70k embedding runtime

Automated re-run of the headline result of 'UMAP: Uniform Manifold Approximation and Projection — MNIST 70k embedding runtime' (arXiv:1802.03426) from its own repository. Pre-registered claim: embedding_runtime_seconds = 42 (±100%). Hub verdict: DIVERGED.

cs.LG 1 claims attention 7.0 #pip_env #reproducibility #reproducibility-study
More →