The first question in almost every conversation about AI is: "Which model should we use?" I understand the question. It is still the wrong one. Because the moment a company points an assistant at its own content – manuals, product data, statutes, its own website – the knowledge no longer comes from the model. It comes from the company's own documents. The model does the wording. It is the interchangeable part.
That is easily said. Until now, though, nobody had been able to show it to me.
Why the usual comparisons are no help here
There are excellent benchmarks. LMArena, Artificial Analysis, a dozen leaderboards. They all measure models against fixed, public tasks – mathematics, programming, general knowledge. What they do not measure: how differently do models answer when they have the same documents in front of them that you do? Yet that is precisely the situation you face in your company. So we built it. Four chatbots on one page. The same knowledge base, the same software, the same search across the same text passages. The only difference is the language model: Mistral Large 3, Claude Opus 5, GPT-5.5 and Gemini 3.6 Flash. You ask all four the same question and see side by side what comes out.
Three things caught our attention along the way.
1. On the facts, they are close together
That sounds unspectacular, but it is the actual point. All four draw their answers from the same sources – so the facts largely agree. What differs is everything else: how long the answers are. Whether they are brief and to the point or expansive. Whether the model admits it does not know something, or elegantly papers over the gap. Whether sources are cited. These are questions of style and character – and they have a surprisingly strong bearing on whether an assistant suits a public authority or an online shop. They just have little to do with "good" or "bad".
2. The price difference is not small. It is absurd.
This is where things became uncomfortably concrete for me. The list prices, as of September 2026, per one million tokens:
- Mistral Large 3: €0.44 input, €1.30 output
- Gemini 3.6 Flash: €0.64 input, €3.21 output
- Claude Opus 5: €4.28 input, €21.41 output
- GPT-5.5: €4.28 input, €25.69 output
Between the cheapest and the most expensive model lies a factor of 10 to 20. For an identical task, an identical knowledge base and identical software. And because output tokens cost three to six times as much as input tokens, this hits a chatbot hardest of all. An assistant, after all, mostly produces output.
3. Digital sovereignty used to mean going without
This is the part that surprised me most. For years, the argument went like this: choose a European model for data-protection reasons and you pay for it in quality. Sovereignty as a compromise. In this set-up, the European model is not the compromise. It is the cheapest option – with answers drawn from the same sources as those of the US models. Data protection is no longer a cost factor here. It has become an argument for spending less. I am not saying this holds for every use case. But for an assistant that answers from your own documents, it is a calculation worth running at least once.
What I took away from it
Not: "Use model X."
Rather: do not commit to a model. Commit to an architecture in which you can swap it. Because today's prices are no basis for tomorrow's decisions. The introductory price of Gemini 3.6 Flash expires on 31 December 2026. GPT-5.5 did not exist eighteen months ago. Anyone who ties their knowledge firmly to one provider today is making a decision for the next five years based on numbers that will not hold for six months. So the question is not "Which model?" but: how expensive will it be if we want to switch in two years? If the answer is "one configuration entry", you have built it right.
What the comparison is not
It is not a scientific benchmark. The knowledge base is modest, and it is a snapshot. The speed figures fluctuate too – provider load, time of day, network route and caching all have more of a say than one would like. Anyone trying to turn this into a ranking is over-interpreting it. It is a demonstration, not proof of superiority. But it is the only place I know of where you can try this out for yourself instead of having to take it on faith.
Try it for yourself
The most revealing questions are the unfair ones. Ask something that is guaranteed not to be in the knowledge base, and watch which model openly admits it and which starts improvising. That tells you more about everyday suitability than any benchmark score.
www.dreistein.de/en/rag-comparison
Transparency, because it belongs here: we are a TYPO3 agency and we build assistants like these for a living. The software behind all four chatbots is our own. It is called dAI Pro, and the fact that you can swap the model there is no accident – it is why we built it the way we did. This page came about because I wanted to know for myself, and because I no longer wanted to merely tell customers about it.
What interests me now: if you are facing this decision in your company – is that interchangeability even a criterion for you?