← Back to blog

What does AI actually cost?

Asking a question in ChatGPT feels free. Build AI into a business process and suddenly you are calculating in input tokens, output tokens, context windows and API calls. But what does that actually mean for the cost of an AI solution?

What does AI actually cost?

The question is not what a single token costs

Most AI models charge based on tokens. A token is a small piece of text. A word can consist of a single token, but also of several. Everything you send to a model costs tokens, and the same applies to the answer you get back.

That is why you do not simply pay per question. A short classification of an email costs far less computing power than analysing a complete report and generating an extensive analysis.

The prices per million tokens of models such as GPT and Claude are easy to look up. Still, that says surprisingly little about what your AI application will ultimately cost.

The architecture around it matters at least as much.

Are you sending a complete hundred-page document with every question? Then you are reprocessing enormous amounts of information again and again. Use RAG to retrieve only the relevant passages and the model may only have to process a fraction of it.

The same applies to the choice of model. Using the most powerful model for every task sounds logical, but is often unnecessary. For classifying a message or extracting a few fields, a smaller model can be perfectly sufficient.

Only when a task becomes more complex may a larger model be needed.

At scale the calculation changes

In a proof of concept, token costs barely register. Ten, a hundred or even a thousand requests cost relatively little in many applications.

But an application that becomes part of a daily business process works differently.

When hundreds of documents, emails or transactions are processed by AI every day, small technical choices are repeated thousands of times. A prompt that contains more context than necessary, or a model that is much heavier than needed, then directly affects the structural costs.

Optimising therefore does not only mean looking for a cheaper AI provider. It mainly means making sure you do not use more AI than the process needs.

Tokens are only part of the cost

At the same time, too much attention is sometimes paid to token prices.

A model that is a few euros cheaper makes little difference if employees then have to manually check every outcome anyway. Conversely, a relatively expensive AI call can pay off well if it eliminates twenty minutes of manual work.

With private AI the calculation shifts even further. Instead of paying only per token, you deal with infrastructure and capacity that you keep available yourself.

That is why you cannot look at the cost of AI separately from the process in which you use it.

The relevant question is ultimately not:

What does a million tokens cost?

But:

What does it cost to carry out this task reliably, and what does the same task cost without AI?

Only then do you know whether AI is expensive or surprisingly cheap.

CA
Carola Abbenhuis-Mensink

Marketing Coordinator at Wabber B.V.

Ready to put your data to work?

Schedule a no-obligation 30-minute session. Discover how private AI and tracking systems measurably improve your operation.