Retrieval-Augmented Generation (RAG) has shown significant improvements in
various natural language processing tasks by integrating the strengths of large
language models (LLMs) and external knowledge databases. However, RAG
introduces long sequence generation and leads to high computation and memory
costs. We propose Thoth, a novel multilevel dynamic caching system tailored for
RAG. Our analysis benchmarks current RAG systems, pinpointing the performance
bottleneck (i.e., long sequence due to knowledge injection) and optimization
opportunities (i.e., caching knowledge's intermediate states). Based on these
insights, we design Thoth, which organizes the intermediate states of retrieved
knowledge in a knowledge tree and caches them in the GPU and host memory
hierarchy. Thoth proposes a replacement policy that is aware of LLM inference
characteristics and RAG retrieval patterns. It also dynamically overlaps the
retrieval and inference steps to minimize the end-to-end latency. We implement
Thoth and evaluate it on vLLM, a state-of-the-art LLM inference system and
Faiss, a state-of-the-art vector database. The experimental results show that
Thoth reduces the time to first token (TTFT) by up to 4x and improves the
throughput by up to 2.1x compared to vLLM integrated with Faiss.

通过集成大型语言模型（LLM）和外部知识数据库，检索增强生成（RAG）在各种自然语言处理任务中展现了显著的改进。然而，RAG 引入了长序列生成，导致了高计算和内存成本。我们提出了一种针对 RAG 量身定制的新型多级动态缓存系统 Thoth，通过组织检索的知识的中间状态，并在 GPU 和主机内存层次结构中缓存它们，以减少时间和资源成本。