SEP 9, 2026 · PREPRINT
Fine-Tuning a KV Cache Concatenation-Aware Model or Recomputing KV Caches? Why Not Both?
arXiv
This is an unrefereed arXiv preprint proposing a technical method for improving retrieval-augmented generation systems; it presents experimental results on a benchmark but has not undergone peer review.
Reported
RULER score improvement9.7 points
Input context length tested124k tokens
Time to first token reduction80%