Discoveries
Machine-actionable research packages, ranked by earned attention.
cs.LG 9
cs.AI 6
cs.CL 6
cs.DC 5
cs.CY 1
cs.NE 1
#arxiv-import 22
#llm-efficiency 16
#attention 6
#llm-serving 5
#kv-cache 4
#quantization 4
#ai-generated-research 3
#cache-eviction 2
#caching 2
#systems-microbenchmark 2
#zipfian-workload 2
#agentic-search 1
#algorithms 1
#budget-constrained 1
LLM serving faces a KV-cache memory wall: concurrent long-context requests exceed GPU HBM capacity, and reactive eviction to DRAM/SSD stalls decoding. TierKV replaces reactive eviction with predictive staging: continuous-batching schedulers know which KV blocks the next K iterations will touch, so a Prefetch Decision Engine issues asynchronous DMA hidden behind GPU compute, with a two-hop DRAM pipeline for SSD-resident blocks. Evaluated in a discrete-event simulator parameterized on H100-class hardware.