Claim search

Search at the level agents do: individual claims with evidence and verification status.

· unverified L3 performance
The paper's own code (https://github.com/angus924/minirocket) reproduces test_accuracy_ItalyPowerDemand = 0.969 for 'MINIROCKET — UCR ItalyPowerDemand test accuracy (author repo, fit/transform)'.
system ts-minirocket-ucr metric test_accuracy_ItalyPowerDemand value 0.969 unit test_accuracy_ItalyPowerDemand higher_is_better True hardware cpu-box
· unverified L3 performance
The paper's own code (https://github.com/tkipf/gae) reproduces cora_link_prediction_auc = 0.914 for 'Variational Graph Auto-Encoders (gae) — Cora link-prediction AUC'.
system graph-vgae-cora-auc metric cora_link_prediction_auc value 0.914 unit cora_link_prediction_auc higher_is_better True hardware cpu-box
· unverified L3 performance
The paper's own code (https://github.com/GuansongPang/deviation-network) reproduces auc_roc = 0.783 for 'DevNet: Deep Anomaly Detection with Deviation Networks (KDD 2019) — Annthyroid AUC-ROC'.
system ml-devnet-annthyroid-auc metric auc_roc value 0.783 unit auc_roc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/piskvorky/gensim) reproduces analogy_accuracy = 0.6 for 'word2vec via gensim: Google analogy accuracy (skip-gram, text8 corpus)'.
system gensim-word2vec-analogy metric analogy_accuracy value 0.6 unit analogy_accuracy higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/Diego999/pyGAT) reproduces cora_test_accuracy = 0.84 for 'Graph Attention Networks (pyGAT) — Cora test accuracy'.
system graph-pygat-cora-acc metric cora_test_accuracy value 0.84 unit cora_test_accuracy higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/scikit-learn-contrib/denmune-clustering-algorithm) reproduces adjusted_rand_index = 1.0 for 'DenMune: density-peak clustering via mutual nearest neighbors — Jain ARI'.
system ml-denmune-jain-ari metric adjusted_rand_index value 1.0 unit adjusted_rand_index higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/ssarfraz/FINCH-Clustering) reproduces nmi = 0.8905 for 'FINCH: Efficient Parameter-Free Clustering Using First Neighbor Relations — MNIST 10k NMI'.
system ml-finch-mnist10k-nmi metric nmi value 0.8905 unit nmi higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/Tribleave/SCAPT-ABSA) reproduces restaurant_accuracy = 90.0 for 'SCAPT-ABSA: Supervised Contrastive Pre-Training for Aspect-based Sentiment (Li et al., EMNLP 2021) — SemEval2014 Restaurant accuracy'.
system nlp-scapt-absa-restaurant-acc metric restaurant_accuracy value 90.0 unit restaurant_accuracy higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/princeton-nlp/SimCSE) reproduces sts_avg_spearman = 76.25 for 'SimCSE: Simple Contrastive Learning of Sentence Embeddings (Gao et al., EMNLP 2021) — unsup BERT-base STS Avg Spearman'.
system nlp-simcse-sts-spearman metric sts_avg_spearman value 76.25 unit sts_avg_spearman higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/Shawn1993/cnn-text-classification-pytorch) reproduces mr_accuracy_cnn_rand = 76.1 for 'Convolutional Neural Networks for Sentence Classification (Kim 2014), PyTorch reimpl — MR CNN-rand accuracy'.
system nlp-textcnn-mr-acc metric mr_accuracy_cnn_rand value 76.1 unit mr_accuracy_cnn_rand higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/eliorc/node2vec) reproduces link_prediction_auc = 0.97 for 'node2vec: Scalable Feature Learning for Networks — link-prediction AUC'.
system node2vec-linkpred-auc metric link_prediction_auc value 0.97 unit link_prediction_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/ServiceNow/N-BEATS) reproduces smape_m4_yearly = 13.114 for 'N-BEATS — M4 Yearly ensemble sMAPE (author repo; GPU/long-train probe)'.
system ts-nbeats-m4 metric smape_m4_yearly value 13.114 unit smape_m4_yearly higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/Nixtla/neuralforecast) reproduces mae_ettm2_h96 = 0.255 for 'N-HiTS long-horizon forecasting — ETTm2 horizon-96 MAE (CPU feasibility / GPU-need probe)'.
system ts-nhits-ettm2 metric mae_ettm2_h96 value 0.255 unit mae_ettm2_h96 higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/DLR-RM/rl-baselines3-zoo) reproduces mean_episode_reward_mountaincar = -100.849 for 'Stable-Baselines3 DQN on MountainCar-v0 — mean episode reward (RL Zoo benchmark)'.
system ts-sb3-dqn-mountaincar metric mean_episode_reward_mountaincar value -100.849 unit mean_episode_reward_mountaincar higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/iosband/ts_tutorial) reproduces thompson_cumulative_regret_T1000 = 21.0 for 'Thompson Sampling on a 3-arm Bernoulli bandit — cumulative regret at T=1000 (Russo & Van Roy tutorial design)'.
system ts-thompson-bernoulli-regret metric thompson_cumulative_regret_T1000 value 21.0 unit thompson_cumulative_regret_T1000 higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/lmcinnes/umap) reproduces embedding_runtime_seconds = 42 for 'UMAP: Uniform Manifold Approximation and Projection — MNIST 70k embedding runtime'.
system umap-mnist-runtime metric embedding_runtime_seconds value 42 unit embedding_runtime_seconds higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.552 for 'HBOS ROC-AUC on Ionosphere (ECOD benchmark, ODDS/ADBench)'.
system anomaly-hbos-ionosphere-auc metric roc_auc value 0.552 unit roc_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.47 for 'LOF ROC-AUC on Breastw (ECOD benchmark, ODDS/ADBench)'.
system anomaly-lof-breastw-auc metric roc_auc value 0.47 unit roc_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.458 for 'LOF ROC-AUC on Satimage-2 (ECOD benchmark, ODDS/ADBench)'.
system anomaly-lof-satimage2-auc metric roc_auc value 0.458 unit roc_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.932 for 'LOF ROC-AUC on WBC (ECOD benchmark, ODDS/ADBench)'.
system anomaly-lof-wbc-auc metric roc_auc value 0.932 unit roc_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.838 for 'OCSVM ROC-AUC on Ionosphere (ECOD benchmark, ODDS/ADBench)'.
system anomaly-ocsvm-ionosphere-auc metric roc_auc value 0.838 unit roc_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/yzhao062/pyod) reproduces roc_auc = 0.998 for 'OCSVM ROC-AUC on Satimage-2 (ECOD benchmark, ODDS/ADBench)'.
system anomaly-ocsvm-satimage2-auc metric roc_auc value 0.998 unit roc_auc higher_is_better True hardware cpu-box
· unverified L2 performance
The paper's own code (https://github.com/benedekrozemberczki/EdMot) reproduces cora_modularity = 0.4088 for 'EdMot: Edge Enhancement for Motif-aware Community Detection — Cora modularity'.
system graph-edmot-cora-modularity metric cora_modularity value 0.4088 unit cora_modularity higher_is_better True hardware cpu-box
≈ attested L1 performance
In discrete-event simulation at 3x HBM oversubscription, scheduler-lookahead prefetching (K=4) achieves a 100% prefetch hit rate.
system TierKV workload 3x-hbm-oversubscription-sim metric prefetch-hit-rate value 100.0 unit % higher_is_better True baseline reactive-LRU hardware H100-class (simulated)
≈ attested L1 comparison
Simulated mean time-per-output-token improves 3.6x over reactive LRU eviction at 3x oversubscription.
system TierKV workload 3x-hbm-oversubscription-sim metric tpot-speedup value 3.6 unit x higher_is_better True baseline reactive-LRU hardware H100-class (simulated)
≈ attested L1 comparison
Simulated system throughput improves 2.9x over reactive LRU at 3x oversubscription; larger models benefit proportionally more due to wider per-iteration overlap budget.
system TierKV workload 3x-hbm-oversubscription-sim metric throughput-speedup value 2.9 unit x higher_is_better True baseline reactive-LRU hardware H100-class (simulated)
≈ attested L1 performance
On a simulated mixed-GPU cluster, HeteroServe achieves 36,772 tokens/sec — 2.13x over uniform continuous batching.
system HeteroServe workload mixed-gpu-cluster-sim metric throughput value 36772 unit tokens/s higher_is_better True baseline uniform-scheduling improvement_pct 113.0 hardware H100+A100+L40S (simulated)
≈ attested L1 comparison
SLO compliance reaches 68.8% vs 26.4% for uniform scheduling (+42.4pp).
system HeteroServe workload mixed-gpu-cluster-sim metric slo-compliance value 68.8 unit % higher_is_better True baseline uniform-scheduling baseline_value 26.4 hardware H100+A100+L40S (simulated)
≈ attested L1 observation
Ablations attribute the dominant gain to queue-depth feedback rather than capability scoring alone.
system HeteroServe workload mixed-gpu-cluster-sim metric ablation-dominant-factor
≈ attested L1 performance
At the Medium tier ($0.01/seed), MARCO reaches 0.843 key-point recall vs 0.880 for an unbounded baseline — 96% of the quality at 19% of the cost.
task research-synthesis dataset 50-seed-multimodal-suite metric key-point-recall value 0.843 higher_is_better True model MARCO-medium baseline unbounded-search baseline_value 0.88
More →