SEP 4, 2026 · PREPRINT
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
arXiv
A training-free algorithmic method for KV cache compression in reasoning models, demonstrated on benchmarks with reported memory and throughput improvements, but lacking peer review and clinical/clinical-adjacent validation.
Reported
Memory reduction5.8×
Throughput improvement4.3×
Number of LRMs tested4