The cost problem in Graph-RAG systems
Graph-RAG constructs a knowledge graph from text chunks to improve retrieval in LLM-based question answering — but indexing is expensive. Every text chunk must pass through an LLM to extract entities and relationships. For a modest 3.2 MB dataset, costs run $21. For enterprise scale — a single 5 GB legal case — the bill hits $33,000. As document collections scale to gigabytes and terabytes, indexing cost alone blocks adoption for organizations that need structured reasoning over proprietary data.
How KET-RAG solves it: a practical architecture
KET-RAG makes a key observation: not all text chunks matter equally for knowledge graph construction. A small fraction of core chunks connects the entire graph. By focusing LLM extraction on these structurally important chunks and building a lightweight keyword index over the rest, KET-RAG maintains Graph-RAG's retrieval advantages while cutting indexing costs by over an order of magnitude.
The framework combines two complementary index structures:
- Knowledge graph skeleton — LLM-extracted entity relationships and connections from the top β fraction of chunks (typically 20%), selected by PageRank centrality
- Text-keyword bipartite graph — keywords extracted from all chunks, linked to their source text. Zero LLM cost for this layer. Keywords serve as lightweight proxies for entities, and neighboring text chunks act as proxy ego networks.
During retrieval, KET-RAG searches both layers in parallel. Entity-based reasoning traverses the skeleton. Keyword-based reasoning expands across the full text. The two channels complement each other, producing higher-quality context for generation than either alone.
Concrete cost reductions and performance gains
Cost reduction vs. Microsoft Graph-RAG: KET-RAG achieves comparable or superior retrieval quality while reducing indexing costs by over an order of magnitude. When tuned for generation quality, it improves answer accuracy by up to 32.4% while reducing indexing costs by approximately 20%.
Retrieval performance: On multi-hop reasoning benchmarks, KET-RAG outperforms flat text retrieval. Its modular design lets you use just the skeleton for 20% cost savings while maintaining retrieval quality, or just the keyword layer for zero-LLM-cost reasoning.
Retrieval latency: Query time averages 0.29–0.87 seconds across benchmarks, comparable to full Graph-RAG systems. No latency compromise.
Integration into LLM pipelines: three concrete steps
Step 1: Build the intermediate KNN graph. Link text chunks via lexical similarity (co-occurring keywords) and semantic similarity (embedding cosine distance). This serves as scaffolding to identify core chunks.
Step 2: Extract the skeleton. Apply PageRank to rank chunks by structural importance, then run LLM-based entity and relationship extraction on only the top β fraction (default β=0.8 gives 20% cost reduction). This step consumes a tunable fraction of full Graph-RAG indexing cost.
Step 3: Build the keyword layer. Extract keywords from all text chunks and build a bipartite graph linking keywords to the text where they appear. No LLM calls — just vocabulary extraction and multi-level embeddings. Retrieval expands from seed keywords to neighboring text chunks, mimicking entity ego-network traversal.
Tuning cost vs. quality: The β parameter controls the speed-accuracy tradeoff. Adjust β between 0.2 and 1.0 to balance indexing cost against retrieval quality.
When KET-RAG saves the most money
Large proprietary corpora with multi-hop reasoning: Biomedical research (conference papers, research databases), legal discovery (contracts, case law), patent analysis, technical documentation archives. Indexing costs often exceed query-serving costs; KET-RAG flips that equation.
Constrained budgets with relationship-aware QA: Customer support RAG systems, product documentation Q&A, knowledge bases where you need entity reasoning but can't afford full knowledge graph extraction on terabytes of text. KET-RAG gives you relationship-aware retrieval at a fraction of the cost.
Real-time or batch indexing at scale: Adding new documents incrementally. With KET-RAG's PageRank-based chunk selection, new text is classified as core or peripheral automatically, so incremental indexing stays efficient.
Comparison to alternatives: Text-RAG (classic retrieval) is fast and cheap but misses relationships across documents. Full Graph-RAG captures relationships but costs significantly more to index. KET-RAG sits in the practical middle: relationship-aware on core content, keyword-connected on the periphery.
Modular design: use one layer or both
The framework is modular. Skeleton-RAG (just the graph skeleton) reduces indexing cost while maintaining retrieval quality. Keyword-RAG (just the keyword layer) outperforms flat text retrieval at zero LLM cost. Or combine both for maximum efficiency.
Full implementation, reproducible benchmarks on three QA datasets, and ablation studies are available on GitHub. Evaluated against 13 baselines including Microsoft Graph-RAG, HippoRAG, and LightRAG.