CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
Gyuwan Kim, Cheoneum Park, Tao Yang
CoinRAG optimizes the Pareto frontier for long-context RAG by reusing fine-grained, query-relevant KV caches to reduce latency and improve answer quality.