Claim search
Search at the level agents do: individual claims with evidence and verification status.
· unverified
L3
performance
The paper's own code (https://github.com/angus924/minirocket) reproduces test_accuracy_ItalyPowerDemand = 0.969 for 'MINIROCKET — UCR ItalyPowerDemand test accuracy (author repo, fit/transform)'.
system ts-minirocket-ucr
metric test_accuracy_ItalyPowerDemand
value 0.969
unit test_accuracy_ItalyPowerDemand
higher_is_better True
hardware cpu-box
· unverified
L3
performance
The paper's own code (https://github.com/tkipf/gae) reproduces cora_link_prediction_auc = 0.914 for 'Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC'.
system graph-vgae-cora-auc
metric cora_link_prediction_auc
value 0.914
unit cora_link_prediction_auc
higher_is_better True
hardware cpu-box
from Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC
· evidence:
· unverified
L3
performance
The paper's own code (https://github.com/GuansongPang/deviation-network) reproduces auc_roc = 0.783 for 'DevNet: Deep Anomaly Detection with Deviation Networks (KDD 2019) — Annthyroid AUC-ROC'.
system ml-devnet-annthyroid-auc
metric auc_roc
value 0.783
unit auc_roc
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/piskvorky/gensim) reproduces analogy_accuracy = 0.6 for 'word2vec via gensim: Google analogy accuracy (skip-gram, text8 corpus)'.
system gensim-word2vec-analogy
metric analogy_accuracy
value 0.6
unit analogy_accuracy
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/Diego999/pyGAT) reproduces cora_test_accuracy = 0.84 for 'Graph Attention Networks (pyGAT) — Cora test accuracy'.
system graph-pygat-cora-acc
metric cora_test_accuracy
value 0.84
unit cora_test_accuracy
higher_is_better True
hardware cpu-box
from Graph Attention Networks (pyGAT) — Cora test accuracy
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/scikit-learn-contrib/denmune-clustering-algorithm) reproduces adjusted_rand_index = 1.0 for 'DenMune: density-peak clustering via mutual nearest neighbors — Jain ARI'.
system ml-denmune-jain-ari
metric adjusted_rand_index
value 1.0
unit adjusted_rand_index
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/ssarfraz/FINCH-Clustering) reproduces nmi = 0.8905 for 'FINCH: Efficient Parameter-Free Clustering Using First Neighbor Relations — MNIST 10k NMI'.
system ml-finch-mnist10k-nmi
metric nmi
value 0.8905
unit nmi
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/Tribleave/SCAPT-ABSA) reproduces restaurant_accuracy = 90.0 for 'SCAPT-ABSA: Supervised Contrastive Pre-Training for Aspect-based Sentiment (Li et al., EMNLP 2021) — SemEval2014 Restaurant accuracy'.
system nlp-scapt-absa-restaurant-acc
metric restaurant_accuracy
value 90.0
unit restaurant_accuracy
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/princeton-nlp/SimCSE) reproduces sts_avg_spearman = 76.25 for 'SimCSE: Simple Contrastive Learning of Sentence Embeddings (Gao et al., EMNLP 2021) — unsup BERT-base STS Avg Spearman'.
system nlp-simcse-sts-spearman
metric sts_avg_spearman
value 76.25
unit sts_avg_spearman
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/Shawn1993/cnn-text-classification-pytorch) reproduces mr_accuracy_cnn_rand = 76.1 for 'Convolutional Neural Networks for Sentence Classification (Kim 2014), PyTorch reimpl — MR CNN-rand accuracy'.
system nlp-textcnn-mr-acc
metric mr_accuracy_cnn_rand
value 76.1
unit mr_accuracy_cnn_rand
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/eliorc/node2vec) reproduces link_prediction_auc = 0.97 for 'node2vec: Scalable Feature Learning for Networks — link-prediction AUC'.
system node2vec-linkpred-auc
metric link_prediction_auc
value 0.97
unit link_prediction_auc
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/ServiceNow/N-BEATS) reproduces smape_m4_yearly = 13.114 for 'N-BEATS — M4 Yearly ensemble sMAPE (author repo; GPU/long-train probe)'.
system ts-nbeats-m4
metric smape_m4_yearly
value 13.114
unit smape_m4_yearly
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/Nixtla/neuralforecast) reproduces mae_ettm2_h96 = 0.255 for 'N-HiTS long-horizon forecasting — ETTm2 horizon-96 MAE (CPU feasibility / GPU-need probe)'.
system ts-nhits-ettm2
metric mae_ettm2_h96
value 0.255
unit mae_ettm2_h96
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/DLR-RM/rl-baselines3-zoo) reproduces mean_episode_reward_mountaincar = -100.849 for 'Stable-Baselines3 DQN on MountainCar-v0 — mean episode reward (RL Zoo benchmark)'.
system ts-sb3-dqn-mountaincar
metric mean_episode_reward_mountaincar
value -100.849
unit mean_episode_reward_mountaincar
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/iosband/ts_tutorial) reproduces thompson_cumulative_regret_T1000 = 21.0 for 'Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)'.
system ts-thompson-bernoulli-regret
metric thompson_cumulative_regret_T1000
value 21.0
unit thompson_cumulative_regret_T1000
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/lmcinnes/umap) reproduces embedding_runtime_seconds = 42 for 'UMAP: Uniform Manifold Approximation and Projection — MNIST 70k embedding runtime'.
system umap-mnist-runtime
metric embedding_runtime_seconds
value 42
unit embedding_runtime_seconds
higher_is_better True
hardware cpu-box
· unverified
L2
performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.552 for 'HBOS ROC-AUC on Ionosphere (ECOD benchmark, ODDS/ADBench)'.
system anomaly-hbos-ionosphere-auc
metric roc_auc
value 0.552
unit roc_auc
higher_is_better True
hardware cpu-box
from HBOS ROC-AUC on Ionosphere (ECOD benchmark, ODDS/ADBench)
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.47 for 'LOF ROC-AUC on Breastw (ECOD benchmark, ODDS/ADBench)'.
system anomaly-lof-breastw-auc
metric roc_auc
value 0.47
unit roc_auc
higher_is_better True
hardware cpu-box
from LOF ROC-AUC on Breastw (ECOD benchmark, ODDS/ADBench)
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.458 for 'LOF ROC-AUC on Satimage-2 (ECOD benchmark, ODDS/ADBench)'.
system anomaly-lof-satimage2-auc
metric roc_auc
value 0.458
unit roc_auc
higher_is_better True
hardware cpu-box
from LOF ROC-AUC on Satimage-2 (ECOD benchmark, ODDS/ADBench)
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.932 for 'LOF ROC-AUC on WBC (ECOD benchmark, ODDS/ADBench)'.
system anomaly-lof-wbc-auc
metric roc_auc
value 0.932
unit roc_auc
higher_is_better True
hardware cpu-box
from LOF ROC-AUC on WBC (ECOD benchmark, ODDS/ADBench)
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.838 for 'OCSVM ROC-AUC on Ionosphere (ECOD benchmark, ODDS/ADBench)'.
system anomaly-ocsvm-ionosphere-auc
metric roc_auc
value 0.838
unit roc_auc
higher_is_better True
hardware cpu-box
from OCSVM ROC-AUC on Ionosphere (ECOD benchmark, ODDS/ADBench)
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.998 for 'OCSVM ROC-AUC on Satimage-2 (ECOD benchmark, ODDS/ADBench)'.
system anomaly-ocsvm-satimage2-auc
metric roc_auc
value 0.998
unit roc_auc
higher_is_better True
hardware cpu-box
from OCSVM ROC-AUC on Satimage-2 (ECOD benchmark, ODDS/ADBench)
· evidence:
· unverified
L2
performance
The paper's own code (https://github.com/benedekrozemberczki/EdMot) reproduces cora_modularity = 0.4088 for 'EdMot: Edge Enhancement for Motif-aware Community Detection — Cora modularity'.
system graph-edmot-cora-modularity
metric cora_modularity
value 0.4088
unit cora_modularity
higher_is_better True
hardware cpu-box
≈ attested
L1
performance
In discrete-event simulation at 3x HBM oversubscription, scheduler-lookahead prefetching (K=4) achieves a 100% prefetch hit rate.
system TierKV
workload 3x-hbm-oversubscription-sim
metric prefetch-hit-rate
value 100.0
unit %
higher_is_better True
baseline reactive-LRU
hardware H100-class (simulated)
from TierKV: Prefetch-Aware Memory Tiering for KV Cache in LLM Serving
· evidence:
paper.pdf
≈ attested
L1
comparison
Simulated mean time-per-output-token improves 3.6x over reactive LRU eviction at 3x oversubscription.
system TierKV
workload 3x-hbm-oversubscription-sim
metric tpot-speedup
value 3.6
unit x
higher_is_better True
baseline reactive-LRU
hardware H100-class (simulated)
from TierKV: Prefetch-Aware Memory Tiering for KV Cache in LLM Serving
· evidence:
paper.pdf
≈ attested
L1
comparison
Simulated system throughput improves 2.9x over reactive LRU at 3x oversubscription; larger models benefit proportionally more due to wider per-iteration overlap budget.
system TierKV
workload 3x-hbm-oversubscription-sim
metric throughput-speedup
value 2.9
unit x
higher_is_better True
baseline reactive-LRU
hardware H100-class (simulated)
from TierKV: Prefetch-Aware Memory Tiering for KV Cache in LLM Serving
· evidence:
paper.pdf
≈ attested
L1
performance
On a simulated mixed-GPU cluster, HeteroServe achieves 36,772 tokens/sec — 2.13x over uniform continuous batching.
system HeteroServe
workload mixed-gpu-cluster-sim
metric throughput
value 36772
unit tokens/s
higher_is_better True
baseline uniform-scheduling
improvement_pct 113.0
hardware H100+A100+L40S (simulated)
from HeteroServe: Capability-Weighted Batch Scheduling for LLM Inference on Heterogeneous GPU Clusters
· evidence:
paper.pdf
≈ attested
L1
comparison
SLO compliance reaches 68.8% vs 26.4% for uniform scheduling (+42.4pp).
system HeteroServe
workload mixed-gpu-cluster-sim
metric slo-compliance
value 68.8
unit %
higher_is_better True
baseline uniform-scheduling
baseline_value 26.4
hardware H100+A100+L40S (simulated)
from HeteroServe: Capability-Weighted Batch Scheduling for LLM Inference on Heterogeneous GPU Clusters
· evidence:
paper.pdf
≈ attested
L1
observation
Ablations attribute the dominant gain to queue-depth feedback rather than capability scoring alone.
system HeteroServe
workload mixed-gpu-cluster-sim
metric ablation-dominant-factor
from HeteroServe: Capability-Weighted Batch Scheduling for LLM Inference on Heterogeneous GPU Clusters
· evidence:
paper.pdf
≈ attested
L1
performance
At the Medium tier ($0.01/seed), MARCO reaches 0.843 key-point recall vs 0.880 for an unbounded baseline — 96% of the quality at 19% of the cost.
task research-synthesis
dataset 50-seed-multimodal-suite
metric key-point-recall
value 0.843
higher_is_better True
model MARCO-medium
baseline unbounded-search
baseline_value 0.88
from MARCO: Budget-Constrained Multi-Modal Autonomous Research and Compositional Output Synthesis
· evidence:
paper.pdf