Skip to content
Quantum Fax Machine

RAG Is Simpler Than You Think

RAG Is Simpler Than You Think

Most people, Rafael Pierre writes, seem to over-engineer retrieval for their LLM applications, jumping straight to embeddings, vector databases and reranking pipelines when their users just want to find the right document. Pierre sets out six recipes, from minimal to elaborate: full-text search with BM25, agentic query rewriting, hybrid search that reranks keyword results with embeddings, embedding on the fly for fast-changing data, pre-embedding with hot and cold tiers, and full pre-embedding at scale, each with when to use it and a decision tree for choosing. The post's rule of thumb is that 60% of systems should stop at full-text search plus query rewriting: don't build the 5% solution for a 60% problem.

⌘K

Start typing to search...

Search across content, newsletters, and subscribers