we_are_coded.by CODE · The world, decoded
БГ
Concept

RAG (search, then answer)

The BasicsUpdated on 16 August 2026we are coded

Before it answers, a good AI checks your documents first. It doesn't rely on memory alone.

Checked on16 August 2026
In short: RAG (Retrieval-Augmented Generation) is a technique described in 2020 by researchers at Facebook AI Research. It means that before it answers, the model first searches your own documents, and only then writes the answer based on what it found. Not from memory, not by feel. That's why companies put it into their internal systems. There, a mistake costs money.

The model - the program behind the chatbot - learned its language from piles of text on the internet. When you ask it something, it doesn't open a book to check. It guesses based on what it's seen many times before. It sounds confident. Sometimes it's true. Sometimes it's fiction, written in the exact same tone of certainty.

RAG does something simple. Before the model answers, it first searches a defined set of documents - your notes, the company's internal rules, old correspondence - and finds the part that's on topic. Then it writes the answer, standing exactly on that.

Say you ask the work chatbot how many vacation days you have left. A plain model guesses based on what usually gets written in similar conversations, and it can get it wrong. A system with RAG does something else. Before it answers, it opens your actual record in HR, reads the exact number, and only then tells you. It doesn't guess. It reads.

The analogy I like here is a manual within reach. Some people answer from memory - they're well-read, they speak with confidence, but sometimes they quote something from years ago that no longer applies. And some people, before they answer, reach for the current manual on the desk, find the exact line, and only then speak. RAG is the model trained to do the second thing.

Companies put RAG into their internal systems because their documents keep changing. Prices, contracts, internal rules. Retraining the whole model every time is slow and expensive. It's easier to update the documents and let the model keep searching them correctly.

RAG doesn't make the model smarter. It makes it honest - it makes it check before it speaks.

Here's what I like

The truth is, RAG is an admission. An admission that the raw model doesn't know your things - it guesses them, dressed up nicely in confidence. Companies didn't add search on top because it's a luxury. They added it because a confident wrong answer costs money and trust in business. For me, as someone who builds these systems, the choice is between a storyteller and a librarian. The storyteller is more fun to listen to. But when you're asking about your vacation, your contract, your money, you want the librarian.

The visual is generated code art. No third-party images.
Official primary sources
→Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv, 2020)