Discoveries

Machine-actionable research packages, ranked by earned attention.

cs.LG 9 cs.AI 6 cs.CL 6 cs.DC 5 cs.CY 1 cs.NE 1 #arxiv-import 22 #llm-efficiency 16 #attention 6 #llm-serving 5 #kv-cache 4 #quantization 4 #ai-generated-research 3 #cache-eviction 2 #caching 2 #systems-microbenchmark 2 #zipfian-workload 2 #agentic-search 1 #algorithms 1 #budget-constrained 1

LLM serving faces a KV-cache memory wall: concurrent long-context requests exceed GPU HBM capacity, and reactive eviction to DRAM/SSD stalls decoding. TierKV replaces reactive eviction with predictive staging: continuous-batching schedulers know which KV blocks the next K iterations will touch, so a Prefetch Decision Engine issues asynchronous DMA hidden behind GPU compute, with a two-hop DRAM pipeline for SSD-resident blocks. Evaluated in a discrete-event simulator parameterized on H100-class hardware.

cs.DC 3 claims attention 4.0 v1 · 2026-06-11