The cost of an LLM call is the number of input tokens times the input price plus the number of output tokens times the output price, with providers listing both prices per 1 million tokens. A Java app that sends a long system prompt, the chat history and retrieved documents with every request pays for all of those tokens every time, so most savings come from sending fewer tokens or cheaper ones.
We estimate token costs before choosing a model and again whenever a monthly bill grows faster than the traffic. The same numbers tell us which reduction techniques pay off for our workload.
The following example is a small cost calculator for one request. The record holds a model’s prices, and the method splits the input into cached and uncached tokens, because input that the provider has cached from an earlier request with the same prefix is much cheaper.
public record ModelPrice(String model, double input, double cachedInput, double output) {
public double cost(long inputTokens, long cachedTokens, long outputTokens) {
long uncached = inputTokens - cachedTokens;
return (uncached * input + cachedTokens * cachedInput + outputTokens * output) / 1_000_000;
}
}
ModelPrice sonnet = new ModelPrice("Claude Sonnet 5.5", 2.00, 0.10, 10.00);
double noCache = sonnet.cost(6_000, 0, 400); // 0.01600 dollars
double withCache = sonnet.cost(6_000, 5_000, 400); // 0.00650 dollars
Caching 5,000 of the 6,000 input tokens cuts this request’s cost by about 60%. In the rest of the article, we look at what drives the token count, current prices from three providers, how to count tokens and read real usage in Java, and the techniques that reduce the bill.
1. What Makes Up the Token Count of a Request
A model keeps no memory between calls, so every request carries everything the model needs. For a support chatbot, the input of one request often looks like this, and only the last part is the user’s actual question.
| Part of the request | Example size | Sent again on every request? |
|---|---|---|
| System prompt with rules and product facts | 1,000 to 10,000 tokens | Yes, unchanged |
| Tool definitions | 100 to 500 tokens per tool | Yes, unchanged |
| Retrieved documents (RAG, retrieval-augmented generation) | 1,000 to 5,000 tokens | Yes, different each time |
| Chat history | Grows with every turn | Yes, longer each time |
| User question | 20 to 200 tokens | New each time |
| Answer (output) | 100 to 1,000 tokens | New each time, at a higher price |
The sizes are example ranges for a chat app with a long system prompt, not limits, and every app should measure its own. In most chat apps, the repeated parts of the input cost more than the question and the answer together. That is why caching and trimming the context give the biggest savings.
2. Current Token Prices
Each provider lists three prices per model, for input tokens, cached input tokens and output tokens. The table shows one mid-size and one small model from three providers, in US dollars per 1 million tokens, as listed on their pricing pages on October 9, 2026.
| Model | Input | Cached input | Output |
|---|---|---|---|
| OpenAI GPT-5.6 Terra | $2.00 | $0.20 | $12.00 |
| OpenAI GPT-5.6 Luna | $0.20 | $0.02 | $1.20 |
| Anthropic Claude Sonnet 5.5 | $2.00 | $0.10 | $10.00 |
| Anthropic Claude Haiku 5.5 (prompts up to 100k tokens) | $0.10 | $0.01 | $0.50 |
| Google gemini-3.5-flash | $1.50 | $0.15 | $9.00 |
| Google gemini-3.5-flash-lite | $0.30 | $0.03 | $2.50 |
Output tokens cost five to about eight times the input price for these models. A small model’s prices are 5% to 28% of the mid-size model’s prices from the same provider. Prices change often, so we keep them in configuration and check the OpenAI, Anthropic and Gemini pricing pages before a cost review.
3. Estimating the Monthly Cost in Java
A monthly estimate needs the average tokens per request and the number of requests. We price a support chatbot with a 5,000-token system prompt, 1,000 tokens of history and question, a 400-token answer and 100,000 requests a month, priced for all six models.
static final List<ModelPrice> PRICES = List.of(
new ModelPrice("GPT-5.6 Terra", 2.00, 0.20, 12.00),
new ModelPrice("GPT-5.6 Luna", 0.20, 0.02, 1.20),
new ModelPrice("Claude Sonnet 5.5", 2.00, 0.10, 10.00),
new ModelPrice("Claude Haiku 5.5", 0.10, 0.01, 0.50),
new ModelPrice("gemini-3.5-flash", 1.50, 0.15, 9.00),
new ModelPrice("gemini-3.5-flash-lite", 0.30, 0.03, 2.50));
long input = 6_000, cached = 5_000, output = 400, requestsPerMonth = 100_000;
System.out.printf("%-22s %12s %12s %14s %14s%n", "Model", "Per request", "Monthly", "With caching", "Batch API");
for (ModelPrice p : PRICES) {
double perRequest = p.cost(input, 0, output);
double monthly = perRequest * requestsPerMonth;
double withCache = p.cost(input, cached, output) * requestsPerMonth;
double batch = monthly * 0.5;
System.out.printf("%-22s %12s %12s %14s %14s%n", p.model(),
String.format("$%.5f", perRequest), String.format("$%,.0f", monthly),
String.format("$%,.0f", withCache), String.format("$%,.0f", batch));
}
Model Per request Monthly With caching Batch API
GPT-5.6 Terra $0.01680 $1,680 $780 $840
GPT-5.6 Luna $0.00168 $168 $78 $84
Claude Sonnet 5.5 $0.01600 $1,600 $650 $800
Claude Haiku 5.5 $0.00080 $80 $35 $40
gemini-3.5-flash $0.01260 $1,260 $585 $630
gemini-3.5-flash-lite $0.00280 $280 $145 $140

The model choice is the bigger lever. Caching cuts every model’s cost by about half for this workload, whereas the small model’s bill is 5% to 22% of the mid-size model’s bill from the same provider. So the cheapest setup is a small model with caching, if its answers are good enough for the task. The caching column assumes that every request reads the system prompt from the cache, which holds for a busy endpoint.
4. Counting Tokens Before Sending a Request
To estimate cost before a call, or to check that a prompt fits the context window (the maximum number of tokens a model accepts in one request), we count tokens in our own process, without calling the provider. The JTokkit library (com.knuddels:jtokkit, version 1.1.0) implements the tokenizers of OpenAI models in Java. O200K_BASE is the encoding of current OpenAI models.
Encoding encoding = Encodings.newDefaultEncodingRegistry().getEncoding(EncodingType.O200K_BASE);
String question = "My order 1042 arrived damaged. Can I get a replacement or a refund?";
int questionTokens = encoding.countTokens(question); // 17
Other providers use their own tokenizers, so a local count is an estimate for their models. For exact numbers, Anthropic and Google offer token counting endpoints (count_tokens and countTokens), and OpenAI offers one for its Responses API, but each call is an extra HTTP request.
5. Reading Real Token Usage With Spring AI
Every provider returns the real token counts with each response, and Spring AI exposes them through the same Usage interface for all providers. Logging these numbers per feature is the most reliable way to know where the money goes.
The next example uses Spring AI 2.0.1 with a stand-in ChatModel that returns fixed usage numbers, so it runs without an API key. In a real application, a provider starter creates the ChatModel, and the code that reads the usage stays the same.
// A stand-in ChatModel with fixed usage numbers: 6,000 prompt tokens (5,000 read from the cache), 400 output tokens
ChatModel model = (Prompt prompt) -> new ChatResponse(
List.of(new Generation(new AssistantMessage("You can get a replacement within 30 days."))),
ChatResponseMetadata.builder().usage(new DefaultUsage(6000, 400, 6400, null, 5000L, 0L)).build());
ChatClient chatClient = ChatClient.create(model);
ChatResponse response = chatClient.prompt()
.user("My order 1042 arrived damaged. Can I get a replacement?")
.call()
.chatResponse();
Usage usage = response.getMetadata().getUsage();
System.out.println("Prompt tokens: " + usage.getPromptTokens());
System.out.println("Completion tokens: " + usage.getCompletionTokens());
System.out.println("Total tokens: " + usage.getTotalTokens());
System.out.println("Cache read tokens: " + usage.getCacheReadInputTokens());
ModelPrice sonnet = new ModelPrice("Claude Sonnet 5.5", 2.00, 0.10, 10.00);
double cost = sonnet.cost(usage.getPromptTokens(), usage.getCacheReadInputTokens(), usage.getCompletionTokens());
System.out.printf("Cost of this call: $%.5f%n", cost);
Prompt tokens: 6000
Completion tokens: 400
Total tokens: 6400
Cache read tokens: 5000
Cost of this call: $0.00650
The method getCacheReadInputTokens() returns null when a provider does not report cached tokens, so production code checks for null before using it. Providers also count cached tokens differently in their raw responses. OpenAI and Gemini include cached tokens in the prompt count, whereas Anthropic reports them apart from the input tokens, so we check how our provider’s starter fills Usage before reusing the cost formula. Spring AI also publishes the token counts as the Micrometer metric gen_ai.client.token.usage, tagged with the token type, so a dashboard can show tokens per model and per day without extra code.
6. Ten Ways to Reduce Token Costs
Some techniques cut the number of tokens we send, and others lower the price we pay for each token. The table orders them roughly by how much they save in a typical chat app and how little work they need.
| Technique | How it saves | Watch out for |
|---|---|---|
| 1. Prompt caching | Repeated input, such as the system prompt and tool definitions, costs a tenth or a twentieth of the normal price | Put the fixed parts first; each provider has a minimum cacheable size |
| 2. Smaller model for simple tasks | Classification, extraction and short answers often work on a model whose prices are 5% to 28% of a mid-size model’s | Test answer quality with real examples before switching |
| 3. Batch API for offline work | OpenAI, Anthropic and Google charge 50% less for requests that may take up to 24 hours | Only for work that does not need an immediate answer, such as nightly summaries |
| 4. Limit the output | Set the max tokens option, which caps the length of the answer, and ask for short answers, because output is the most expensive token type | A limit that is too low cuts answers in the middle |
| 5. Trim the chat history | Send only the last few turns, or a short summary of older turns | The model forgets details that were trimmed away |
| 6. Retrieve fewer, smaller chunks | In RAG, send the three best chunks instead of ten long ones | Check that answers still find the right facts |
| 7. Shorter system prompts | Remove repeated rules and examples the model does not need | Re-test behavior after every cut |
| 8. Cache whole answers | Return a stored answer for questions we have already answered | Answers that depend on the user or on fresh data must not be shared |
| 9. Route by difficulty | Send simple requests to a small model and hard ones to a larger model | The routing rule itself must be cheap and reliable |
| 10. Structured output | JSON with only the needed fields is shorter than free text | Validate the JSON against our Java type |
Prompt caching deserves a closer look, because it needs almost no code. The rules differ by provider, and the table lists the main ones from the providers’ caching guides.
| Provider | How caching starts | Minimum prompt size | Cache lifetime |
|---|---|---|---|
| OpenAI | On by default; newer models also allow explicit cache points | 1,024 tokens on GPT-5.6 models | 30 minutes on GPT-5.6 models |
| Anthropic | We mark cacheable blocks with cache_control in the request | 512 to 4,096 tokens, depending on the model | 5 minutes by default, 1 hour optional; each hit restarts the timer |
| Google Gemini | Implicit caching on by default, or explicit cache objects | 4,096 tokens on Gemini 3.x Flash models | Explicit caches live 1 hour by default and are billed per hour of storage |
Writing to the cache can cost more than a normal input token, for example 1.25 times the input price for Anthropic’s 5-minute cache and for GPT-5.6 models. A busy endpoint pays that once and reads the cache many times, so the extra cost is spread over many cheap reads, whereas an endpoint with a few requests per hour may never reuse its cache.
7. Token Cost FAQs
Token bills raise a few questions that the pricing pages answer only in the fine print.
7.1. Why Are Output Tokens More Expensive Than Input Tokens?
The model reads all input tokens in one pass but generates output tokens one at a time, so each output token needs more computing work. For the models in section 2, output costs five to about eight times the input price.
7.2. Do Failed or Cut-Off Requests Cost Money?
Yes, if the provider processed them. An answer cut off by the max tokens limit is billed for every token it produced, and the input tokens are billed in full.
7.3. How Do I Estimate Tokens Without a Tokenizer?
For English text, about 1.3 tokens per word or 4 characters per token is close enough for a first estimate. Code and other languages need more tokens, so we measure with a tokenizer or with the usage numbers before a budget decision.
8. Conclusion
Token cost is input tokens times the input price plus output tokens times the output price, and in chat apps the repeated system prompt, history and documents make up most of the input. A small record and one loop are enough to estimate the monthly bill for any model, and Spring AI’s Usage interface gives the real numbers once the app runs.
Prompt caching and a smaller model are the two biggest savings for most apps, and both need little code. The other techniques pay off once we measure where our tokens go.
9. References
- OpenAI API Pricing
- Anthropic Pricing
- Gemini API Pricing
- OpenAI Prompt Caching
- Anthropic Prompt Caching
- Gemini Context Caching
- Spring AI Observability
Happy Learning !!