Skip to content

RAG chatbots on your own documents

AI & Automation4 min read
RAG chatbots on your own documents

A general language model knows nothing about your price list, your service manuals or your internal policies. Retrieval-augmented generation solves that by searching your own documents first and giving the model only the relevant passages to answer from. Done well, it produces answers with sources. Done badly, it produces confident nonsense faster than before.

How the pipeline actually works

  • Documents are parsed and split into chunks of roughly 300 to 800 tokens, along logical boundaries such as headings.
  • Each chunk is converted into a vector by an embedding model and stored with its metadata.
  • A user question is embedded the same way, and the closest chunks are retrieved.
  • A reranking model reorders the candidates by actual relevance, which usually improves answer quality more than any prompt change.
  • The model receives the top passages plus the question and generates an answer with citations.

The order matters. Most quality problems live in steps one and two, not in the model. If a PDF table is flattened into a wall of numbers during parsing, no amount of prompt engineering will recover the meaning.

Data preparation is the project

Plan for 40 to 60 percent of the budget here. Scanned documents need OCR. Tables need structure-preserving extraction. Outdated versions must be removed, because a model cannot tell that the 2019 price list has been superseded. Every chunk needs metadata: source document, version, date, and access level, so answers can be filtered by who is asking.

A retrieval system is only as honest as the newest document in it and only as safe as the oldest one you forgot to delete.

Measuring quality before customers see it

Build an evaluation set of 80 to 150 real questions with approved answers, collected from support tickets and sales emails. Measure retrieval recall, whether the correct passage was fetched at all, and answer correctness separately. Separating the two tells you where to invest: retrieval problems are fixed with chunking and reranking, generation problems with prompting and model choice.

Set a refusal rule as well. An assistant that answers I could not find this in the documentation, here is the contact for a colleague is far more valuable in a business context than one that guesses. In our projects a well-tuned system lands at 85 to 93 percent correct answers with a refusal rate around 5 percent.

Cost and effort

  • Setup for a document assistant on 500 to 5,000 pages: 6,000 to 14,000 EUR.
  • Embedding the corpus once: usually under 50 EUR, re-embedding after major updates.
  • Running model costs at 300 conversations a month: 40 to 200 EUR.
  • Vector database hosted in the EU: 0 to 90 EUR per month depending on size.
  • Ongoing curation of the document base: two to four hours a month, which someone must own.

Data protection and access control

Personal data in the corpus needs a legal basis and a deletion path. Store the vector index in the EU, sign a data processing agreement with the model provider, and disable training on your inputs. If different roles may see different documents, enforce that at retrieval time by filtering on metadata, never by instructing the model to keep a secret. A model instruction is not an access control.

Where it pays off first

Internal use cases return value fastest, because expectations are realistic and errors are caught by colleagues. Support teams answering repeat product questions, field technicians looking up service procedures, and sales staff searching contract terms are the three strongest starting points. Public customer-facing assistants are worth doing, but only once the internal version has been measured for a quarter. Give every answer a visible source link as well: people trust a system far more when they can open the original page and check it themselves.

Author

Tobias Lang

AI Engineer

Share
Miriam Kraus

Your contact person

Miriam Kraus

I read every request personally and get back to you within one business day.

Write to us right now

Fill out the form and we will contact you

What are you interested in?