A RAG chatbot — one that answers from your own documents instead of making things up — looks cheap to build. A working demo can come together in a weekend. That demo is also why so many of these projects blow their budget: the distance between a weekend prototype and something you'd put in front of customers is most of the actual work, and none of it is visible in the demo.
Rather than quote a number that would be meaningless without knowing your situation, here are the decisions that actually set the cost. Answer these and you can have a grounded budget conversation with anyone.
Data preparation is the biggest cost, and the least glamorous
Teams expect the cost to be in the AI. It's mostly in the data.
A RAG system is only as good as what it retrieves, and real source data is messy: PDFs with inconsistent formatting, scanned documents that need OCR, tables that lose their meaning when flattened to text, duplicated and contradictory content, and no clean boundaries between topics. Turning that into something a retriever can work with — parsing, cleaning, chunking sensibly, handling updates — is where the hours go.
This is also the work that most determines quality. A polished interface over badly-prepared data produces a confident chatbot that's wrong just often enough to be untrustworthy, which is worse than no chatbot at all.
Retrieval quality is where "done" gets redefined
Getting a relevant chunk in front of the model is easy. Getting the right context reliably, across the full range of questions real users ask, is the hard part — and it's iterative.
Naive top-k vector search gets you maybe 70% of the way, and the last 30% is where the budget actually goes: hybrid search, reranking, query rewriting, handling questions that span multiple documents, and knowing when the honest answer is "I don't have that information" rather than a plausible fabrication. Each is real engineering, and you can't tell which you need until you test against actual questions.
Evaluation is not optional, and it costs
Here's the difference between a hobby project and a production system: can you tell whether a change made it better or worse?
Without an evaluation set — real questions with known-good answers — every tweak is guesswork, and you're one prompt change away from silently breaking answers you'd already gotten right. Building that eval harness costs time up front and saves far more later. Skipping it is the most expensive way to save money on this list.
The cost that outlives the build
Unlike a website, a RAG chatbot has a running cost that never stops, and it's the line people forget:
- Model inference — every question spends tokens, and a chatbot that retrieves a lot of context spends a lot of them. Cost scales with usage.
- The vector database — hosting and querying your embeddings is an ongoing service, not a one-off.
- Re-embedding — when your source documents change, the affected content has to be re-processed and re-embedded.
- Monitoring and maintenance — models get deprecated, data drifts, and answer quality needs watching.
A build quote that ignores running costs is telling you half the story. For a high-traffic assistant, the ongoing bill can matter more than the build.
What belongs in version one
The cheapest way to control cost is scope. A defensible first version is usually: one well-prepared data source, solid retrieval, honest "I don't know" behaviour, and an evaluation set — deployed to a limited audience.
What can wait: multiple data sources, live document syncing, conversation memory across sessions, actions beyond answering, and a polished multi-channel interface. Each is real work, and each is far easier to scope once you've watched real people use the simple version and seen what they actually ask.
The cost drivers, ranked
| Driver | Effect on budget | Why |
|---|---|---|
| Data preparation and quality | Largest | Messy sources are most of the real work |
| Retrieval accuracy | Large | The last 30% of relevance is iterative |
| Evaluation harness | Moderate up front, saves more later | Without it, every change is guesswork |
| Ongoing inference + hosting | Recurring, not one-off | Scales with usage forever |
| Integrations and access control | Moderate | Depends on what it must connect to |
| The chat interface | Smaller than assumed | Rarely the thing that blows the budget |
What to bring to a quote conversation
- What the source data is, what format it's in, and how clean it is
- How often that data changes, and whether updates must be live
- The range of questions it needs to answer, and how bad a wrong answer is
- Expected query volume, which drives the running cost
- What it connects to, and who's allowed to see what
With those, an estimate can be grounded. Without an honest read on the data in particular, any number is a guess — the demo lies, and the data tells the truth.
We build production RAG systems with the evaluation and monitoring that keep them trustworthy past the demo. Tell us about your data and what you want it to answer for a scoped estimate.
Related: LLM integration & RAG development · LangChain development services · LangChain vs LlamaIndex