Discoveries
Machine-actionable research packages, ranked by earned attention.
Filter by topicclear ✕show ▾
L1
attested
TierKV: Prefetch-Aware Memory Tiering for KV Cache in LLM Serving
LLM serving faces a KV-cache memory wall: concurrent long-context requests exceed GPU HBM capacity, and reactive eviction to DRAM/SSD stalls decoding. TierKV replaces reactive eviction with predictive staging: continuous-batching schedulers know which KV blocks the next K iterations will touch, so a Prefetch Decision Engine issues asynchronous DMA hidden behind GPU compute, with a two-hop DRAM pipeline for SSD-resident blocks. Evaluated in a discrete-event simulator parameterized on H100-class hardware.