
RAG (retrieval-augmented generation) is the technique behind nearly every assistant that “answers from our documents”. Instead of training a model on your texts, the system looks up the relevant excerpts and hands them to the model together with the question. The model answers from them. It cuts down invented answers, but it does not remove them.
How it works, in five steps
|
|
|
|
|
RAG or something else
| Approach | Good for | Limit |
| Pasting the text into the chat | One document, once | Does not scale, and you pay for the whole text with each request |
| RAG | Many documents that change, and varied questions | Only as good as the documents and the way they are cut |
| Fine-tuning a model | Teaching a style or a format | Not a good way to teach it facts that change |
What you need to build one
|
|
|
|
| Mind who can see what. If you put public and internal documents in the same index, the assistant may quote the internal ones to anyone. Separate the indexes or control access per user. |
| It can still make things up. If the excerpts do not contain the answer, a badly instructed model fills the gap with what it “knows”. Tell it to say “I did not find it”, and always show which document the answer came from. See AI hallucinations. |
| Start with a few documents and a list of test questions whose answer you know. If it does not get those right, it is not ready for customers. |
|
Want to run an assistant like this on a VPS of your own? See the plans and choose the size. See VPS plans |
RECOMMENDED PRODUCT Web hosting with cPanel Domain and SSL included, daily backups and the panel you already know. from $6.60/mo (3-year plan, with coupon) See plans |
- 0 Users Found This Useful











