LLM Token Costs: Estimate and Reduce Them in Java Apps

Estimate LLM token costs in Java with current OpenAI, Anthropic and Gemini prices, read real token usage with Spring AI, and cut the bill with prompt caching.

Bar chart of the monthly cost of 100,000 support chatbot requests per model, without and with prompt caching. GPT-5.6 Terra 1,680 dollars to 780, Claude Sonnet 5.5 1,600 to 650, gemini-3.5-flash 1,260 to 585, gemini-3.5-flash-lite 280 to 145, GPT-5.6 Luna 168 to 78, Claude Haiku 5.5 80 to 35.

The cost of an LLM call is the number of input tokens times the input price plus the number of output tokens times the output price, with providers listing both prices per 1 million tokens. A Java app that sends a long system prompt, the chat history and retrieved documents with every request pays for all of those tokens every time, so most savings come from sending fewer tokens or cheaper ones.

We estimate token costs before choosing a model and again whenever a monthly bill grows faster than the traffic. The same numbers tell us which reduction techniques pay off for our workload.

The following example is a small cost calculator for one request. The record holds a model’s prices, and the method splits the input into cached and uncached tokens, because input that the provider has cached from an earlier request with the same prefix is much cheaper.

public record ModelPrice(String model, double input, double cachedInput, double output) {

  public double cost(long inputTokens, long cachedTokens, long outputTokens) {
    long uncached = inputTokens - cachedTokens;
    return (uncached * input + cachedTokens * cachedInput + outputTokens * output) / 1_000_000;
  }
}

ModelPrice sonnet = new ModelPrice("Claude Sonnet 5.5", 2.00, 0.10, 10.00);
double noCache = sonnet.cost(6_000, 0, 400);          // 0.01600 dollars
double withCache = sonnet.cost(6_000, 5_000, 400);    // 0.00650 dollars

Caching 5,000 of the 6,000 input tokens cuts this request’s cost by about 60%. In the rest of the article, we look at what drives the token count, current prices from three providers, how to count tokens and read real usage in Java, and the techniques that reduce the bill.

1. What Makes Up the Token Count of a Request

A model keeps no memory between calls, so every request carries everything the model needs. For a support chatbot, the input of one request often looks like this, and only the last part is the user’s actual question.

Part of the requestExample sizeSent again on every request?
System prompt with rules and product facts1,000 to 10,000 tokensYes, unchanged
Tool definitions100 to 500 tokens per toolYes, unchanged
Retrieved documents (RAG, retrieval-augmented generation)1,000 to 5,000 tokensYes, different each time
Chat historyGrows with every turnYes, longer each time
User question20 to 200 tokensNew each time
Answer (output)100 to 1,000 tokensNew each time, at a higher price

The sizes are example ranges for a chat app with a long system prompt, not limits, and every app should measure its own. In most chat apps, the repeated parts of the input cost more than the question and the answer together. That is why caching and trimming the context give the biggest savings.

2. Current Token Prices

Each provider lists three prices per model, for input tokens, cached input tokens and output tokens. The table shows one mid-size and one small model from three providers, in US dollars per 1 million tokens, as listed on their pricing pages on October 9, 2026.

ModelInputCached inputOutput
OpenAI GPT-5.6 Terra$2.00$0.20$12.00
OpenAI GPT-5.6 Luna$0.20$0.02$1.20
Anthropic Claude Sonnet 5.5$2.00$0.10$10.00
Anthropic Claude Haiku 5.5 (prompts up to 100k tokens)$0.10$0.01$0.50
Google gemini-3.5-flash$1.50$0.15$9.00
Google gemini-3.5-flash-lite$0.30$0.03$2.50

Output tokens cost five to about eight times the input price for these models. A small model’s prices are 5% to 28% of the mid-size model’s prices from the same provider. Prices change often, so we keep them in configuration and check the OpenAI, Anthropic and Gemini pricing pages before a cost review.

3. Estimating the Monthly Cost in Java

A monthly estimate needs the average tokens per request and the number of requests. We price a support chatbot with a 5,000-token system prompt, 1,000 tokens of history and question, a 400-token answer and 100,000 requests a month, priced for all six models.

static final List<ModelPrice> PRICES = List.of(
    new ModelPrice("GPT-5.6 Terra", 2.00, 0.20, 12.00),
    new ModelPrice("GPT-5.6 Luna", 0.20, 0.02, 1.20),
    new ModelPrice("Claude Sonnet 5.5", 2.00, 0.10, 10.00),
    new ModelPrice("Claude Haiku 5.5", 0.10, 0.01, 0.50),
    new ModelPrice("gemini-3.5-flash", 1.50, 0.15, 9.00),
    new ModelPrice("gemini-3.5-flash-lite", 0.30, 0.03, 2.50));

long input = 6_000, cached = 5_000, output = 400, requestsPerMonth = 100_000;

System.out.printf("%-22s %12s %12s %14s %14s%n", "Model", "Per request", "Monthly", "With caching", "Batch API");
for (ModelPrice p : PRICES) {
  double perRequest = p.cost(input, 0, output);
  double monthly = perRequest * requestsPerMonth;
  double withCache = p.cost(input, cached, output) * requestsPerMonth;
  double batch = monthly * 0.5;
  System.out.printf("%-22s %12s %12s %14s %14s%n", p.model(),
      String.format("$%.5f", perRequest), String.format("$%,.0f", monthly),
      String.format("$%,.0f", withCache), String.format("$%,.0f", batch));
}
Model                   Per request      Monthly   With caching      Batch API
GPT-5.6 Terra              $0.01680       $1,680           $780           $840
GPT-5.6 Luna               $0.00168         $168            $78            $84
Claude Sonnet 5.5          $0.01600       $1,600           $650           $800
Claude Haiku 5.5           $0.00080          $80            $35            $40
gemini-3.5-flash           $0.01260       $1,260           $585           $630
gemini-3.5-flash-lite      $0.00280         $280           $145           $140
Bar chart of the monthly cost of 100,000 support chatbot requests per model, without and with prompt caching. GPT-5.6 Terra 1,680 dollars to 780, Claude Sonnet 5.5 1,600 to 650, gemini-3.5-flash 1,260 to 585, gemini-3.5-flash-lite 280 to 145, GPT-5.6 Luna 168 to 78, Claude Haiku 5.5 80 to 35.
Caching the system prompt cuts the bill by 48 to 60 percent for every model, and moving to a small model saves even more.

The model choice is the bigger lever. Caching cuts every model’s cost by about half for this workload, whereas the small model’s bill is 5% to 22% of the mid-size model’s bill from the same provider. So the cheapest setup is a small model with caching, if its answers are good enough for the task. The caching column assumes that every request reads the system prompt from the cache, which holds for a busy endpoint.

4. Counting Tokens Before Sending a Request

To estimate cost before a call, or to check that a prompt fits the context window (the maximum number of tokens a model accepts in one request), we count tokens in our own process, without calling the provider. The JTokkit library (com.knuddels:jtokkit, version 1.1.0) implements the tokenizers of OpenAI models in Java. O200K_BASE is the encoding of current OpenAI models.

Encoding encoding = Encodings.newDefaultEncodingRegistry().getEncoding(EncodingType.O200K_BASE);

String question = "My order 1042 arrived damaged. Can I get a replacement or a refund?";
int questionTokens = encoding.countTokens(question);     // 17

Other providers use their own tokenizers, so a local count is an estimate for their models. For exact numbers, Anthropic and Google offer token counting endpoints (count_tokens and countTokens), and OpenAI offers one for its Responses API, but each call is an extra HTTP request.

5. Reading Real Token Usage With Spring AI

Every provider returns the real token counts with each response, and Spring AI exposes them through the same Usage interface for all providers. Logging these numbers per feature is the most reliable way to know where the money goes.

The next example uses Spring AI 2.0.1 with a stand-in ChatModel that returns fixed usage numbers, so it runs without an API key. In a real application, a provider starter creates the ChatModel, and the code that reads the usage stays the same.

// A stand-in ChatModel with fixed usage numbers: 6,000 prompt tokens (5,000 read from the cache), 400 output tokens
ChatModel model = (Prompt prompt) -> new ChatResponse(
    List.of(new Generation(new AssistantMessage("You can get a replacement within 30 days."))),
    ChatResponseMetadata.builder().usage(new DefaultUsage(6000, 400, 6400, null, 5000L, 0L)).build());
ChatClient chatClient = ChatClient.create(model);

ChatResponse response = chatClient.prompt()
    .user("My order 1042 arrived damaged. Can I get a replacement?")
    .call()
    .chatResponse();

Usage usage = response.getMetadata().getUsage();
System.out.println("Prompt tokens: " + usage.getPromptTokens());
System.out.println("Completion tokens: " + usage.getCompletionTokens());
System.out.println("Total tokens: " + usage.getTotalTokens());
System.out.println("Cache read tokens: " + usage.getCacheReadInputTokens());

ModelPrice sonnet = new ModelPrice("Claude Sonnet 5.5", 2.00, 0.10, 10.00);
double cost = sonnet.cost(usage.getPromptTokens(), usage.getCacheReadInputTokens(), usage.getCompletionTokens());
System.out.printf("Cost of this call: $%.5f%n", cost);
Prompt tokens: 6000
Completion tokens: 400
Total tokens: 6400
Cache read tokens: 5000
Cost of this call: $0.00650

The method getCacheReadInputTokens() returns null when a provider does not report cached tokens, so production code checks for null before using it. Providers also count cached tokens differently in their raw responses. OpenAI and Gemini include cached tokens in the prompt count, whereas Anthropic reports them apart from the input tokens, so we check how our provider’s starter fills Usage before reusing the cost formula. Spring AI also publishes the token counts as the Micrometer metric gen_ai.client.token.usage, tagged with the token type, so a dashboard can show tokens per model and per day without extra code.

6. Ten Ways to Reduce Token Costs

Some techniques cut the number of tokens we send, and others lower the price we pay for each token. The table orders them roughly by how much they save in a typical chat app and how little work they need.

TechniqueHow it savesWatch out for
1. Prompt cachingRepeated input, such as the system prompt and tool definitions, costs a tenth or a twentieth of the normal pricePut the fixed parts first; each provider has a minimum cacheable size
2. Smaller model for simple tasksClassification, extraction and short answers often work on a model whose prices are 5% to 28% of a mid-size model’sTest answer quality with real examples before switching
3. Batch API for offline workOpenAI, Anthropic and Google charge 50% less for requests that may take up to 24 hoursOnly for work that does not need an immediate answer, such as nightly summaries
4. Limit the outputSet the max tokens option, which caps the length of the answer, and ask for short answers, because output is the most expensive token typeA limit that is too low cuts answers in the middle
5. Trim the chat historySend only the last few turns, or a short summary of older turnsThe model forgets details that were trimmed away
6. Retrieve fewer, smaller chunksIn RAG, send the three best chunks instead of ten long onesCheck that answers still find the right facts
7. Shorter system promptsRemove repeated rules and examples the model does not needRe-test behavior after every cut
8. Cache whole answersReturn a stored answer for questions we have already answeredAnswers that depend on the user or on fresh data must not be shared
9. Route by difficultySend simple requests to a small model and hard ones to a larger modelThe routing rule itself must be cheap and reliable
10. Structured outputJSON with only the needed fields is shorter than free textValidate the JSON against our Java type

Prompt caching deserves a closer look, because it needs almost no code. The rules differ by provider, and the table lists the main ones from the providers’ caching guides.

ProviderHow caching startsMinimum prompt sizeCache lifetime
OpenAIOn by default; newer models also allow explicit cache points1,024 tokens on GPT-5.6 models30 minutes on GPT-5.6 models
AnthropicWe mark cacheable blocks with cache_control in the request512 to 4,096 tokens, depending on the model5 minutes by default, 1 hour optional; each hit restarts the timer
Google GeminiImplicit caching on by default, or explicit cache objects4,096 tokens on Gemini 3.x Flash modelsExplicit caches live 1 hour by default and are billed per hour of storage

Writing to the cache can cost more than a normal input token, for example 1.25 times the input price for Anthropic’s 5-minute cache and for GPT-5.6 models. A busy endpoint pays that once and reads the cache many times, so the extra cost is spread over many cheap reads, whereas an endpoint with a few requests per hour may never reuse its cache.

7. Token Cost FAQs

Token bills raise a few questions that the pricing pages answer only in the fine print.

7.1. Why Are Output Tokens More Expensive Than Input Tokens?

The model reads all input tokens in one pass but generates output tokens one at a time, so each output token needs more computing work. For the models in section 2, output costs five to about eight times the input price.

7.2. Do Failed or Cut-Off Requests Cost Money?

Yes, if the provider processed them. An answer cut off by the max tokens limit is billed for every token it produced, and the input tokens are billed in full.

7.3. How Do I Estimate Tokens Without a Tokenizer?

For English text, about 1.3 tokens per word or 4 characters per token is close enough for a first estimate. Code and other languages need more tokens, so we measure with a tokenizer or with the usage numbers before a budget decision.

8. Conclusion

Token cost is input tokens times the input price plus output tokens times the output price, and in chat apps the repeated system prompt, history and documents make up most of the input. A small record and one loop are enough to estimate the monthly bill for any model, and Spring AI’s Usage interface gives the real numbers once the app runs.

Prompt caching and a smaller model are the two biggest savings for most apps, and both need little code. The other techniques pay off once we measure where our tokens go.

9. References

Happy Learning !!

Source Code on Github

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.