Nairobd
All posts

July 24, 2026 · 6 min read

RAG, explained without the jargon

Mirza Md. Shafi Uddin

Founder, AI & Software Engineer

RAG — Retrieval-Augmented Generation — is one of those terms that sounds more complicated than what it describes. Strip the acronym away and it's this: look something up first, then answer using what you found. That's the entire idea.

Why an LLM needs this at all

A language model on its own only knows what it was trained on, plus whatever's in the current conversation. It doesn't know your return policy, your current pricing, or what changed in your SOP last month. Ask it directly and it will either say so — or, worse, guess confidently and get it wrong. RAG fixes this by handing the model your actual documents at the moment it needs to answer, instead of hoping it memorized something close enough.

What's actually happening under the hood

  • Your documents — policies, pricing sheets, SOPs, past tickets — get broken into chunks and turned into a searchable index.
  • When a question comes in, the system searches that index for the most relevant chunks, not the whole document library.
  • Those chunks get handed to the LLM along with the question, so it answers from what's actually in front of it.

The model isn't “learning” your documents in any permanent sense. It's re-reading the relevant bits fresh, every time, the way you'd hand someone the right page of a manual instead of asking them to recite it from memory.

The part everyone underestimates

The LLM is rarely the weak link in a RAG system. The retrieval step is. If the search step pulls the wrong chunk — an outdated pricing page, a policy that was since revised — the model will answer confidently from the wrong source, and it'll sound just as convincing as if it had the right one. Most of the real engineering work in a knowledge assistant is making sure the right document actually gets found, not making the writing sound better.

When you don't need it

If the answer fits in a paragraph and rarely changes, you don't need RAG — you need a well-written system prompt. RAG earns its complexity when you have real volume of documents, they change often enough that hardcoding them is a maintenance problem, and getting an answer wrong has a real cost. Building it when you don't need it just adds infrastructure without adding accuracy.