Spring Boot Virtual Threads: Setup, Load Test and WebFlux

One property, spring.threads.virtual.enabled=true, moves Spring Boot requests, @Async methods and scheduled tasks to virtual threads. A weather API measured against platform threads and WebFlux, plus pinning and monitoring on Java 25.

Spring Boot virtual threads let every web request, @Async method and scheduled task run on a virtual thread, a lightweight thread that the JVM manages instead of the operating system. When a virtual thread waits for a database or a remote service, it releases the OS thread under it, so a blocking app can serve many more requests at the same time. In Spring Boot 4.1 on Java 25, one property turns it on:

spring.threads.virtual.enabled=true

We use virtual threads in Spring Boot apps that spend most of each request waiting, such as a REST API that calls slow services or a database.

The following example is a blocking controller of a weather API that calls two slow services, and with the property set, Tomcat runs it on a virtual thread.

@GetMapping("/weather/{city}")
public WeatherReport weather(@PathVariable String city) {
  Forecast forecast = client.forecast(city);         // waits 250 ms, the OS thread is free meanwhile
  AirQuality airQuality = client.airQuality(city);   // waits 250 ms
  return WeatherReport.of(forecast, airQuality);
}
// Thread.currentThread()  = VirtualThread[#35,tomcat-handler-0]/runnable@ForkJoinPool-1-worker-1
// isVirtual()             = true

Notice that the controller has no reactive types, and the thread name starts with VirtualThread.

With a weather API as the running example in this post, we cover what virtual threads are, when they help, how to enable and verify them in Spring Boot, how they compare with platform threads and Spring WebFlux under load, and how to find pinned threads with jcmd and JDK Flight Recorder.

1. What Are Virtual Threads?

A platform thread is the classic java.lang.Thread, a thin wrapper around one operating system thread. OS threads are expensive, so servers keep a fixed pool of them. Tomcat’s pool has 200 threads by default, and a request that waits for a slow service holds one pool thread the whole time.

Virtual threads, final since Java 21 (JEP 444), are also Thread objects, but the JVM schedules them on a few carrier threads (the platform threads of an internal ForkJoinPool, with one worker per CPU core). When a virtual thread blocks on I/O, the JVM unmounts it, which moves its stack to the heap so the carrier can pick up another virtual thread. When the answer arrives, the JVM mounts the virtual thread again, on any free carrier.

Two panels for 1,000 concurrent requests: on the left, Tomcat platform threads exec-1 to exec-200 are each blocked for 500 ms and requests 201 to 1,000 wait in a queue, giving at most 400 requests per second; on the right, 1,000 virtual threads are parked with their stacks on the heap and only the running ones are mounted on two carrier threads, worker-1 and worker-2; a note gives the thread dump of the weather API under 500 users as 506 virtual threads on 2 carrier threads
A waiting platform thread holds an OS thread. A waiting virtual thread holds only some heap memory.

Both kinds of thread run the same Java code, and the difference is in what a blocked thread costs.

Platform threadVirtual thread
Scheduled byThe operating systemThe JVM, on carrier threads
Created withnew Thread(), Thread.ofPlatform()Thread.ofVirtual(), Executors.newVirtualThreadPerTaskExecutor()
Cost of a blocked threadOne OS threadA stack on the heap
How manyHundreds to a few thousandMillions
Pooled?Yes, because platform threads are expensiveNever; one new thread per task
Thread.currentThread().isVirtual()falsetrue

Virtual threads do not run code faster; they let more tasks wait at the same time. A request that takes 500 ms on a platform thread still takes 500 ms on a virtual thread.

1.1. When Do Virtual Threads Help?

Virtual threads help when the number of requests waiting at the same time is larger than the thread pool. For example, a travel site whose search calls several airline APIs spends most of each request waiting. Our load test in section 2.6 shows both cases.

WorkloadVirtual threads help?Why
REST API calling slow services or a database, many concurrent usersYesWaiting requests no longer hold pool threads
Fewer concurrent requests than server.tomcat.threads.maxNoThe pool never runs out; same numbers in our test
CPU-bound work (image resizing, report calculation)NoThe CPU is the limit, not the threads
An app already written with WebFluxNo needThe event loop already does not block
Blocking code inside synchronized on Java 21 to 23RiskyThe thread is pinned to its carrier (section 4.2); fixed in Java 24

2. Spring Boot Virtual Threads Example

The following example is a weather API, where GET /weather/{city} calls two slow services (a forecast service and an air quality service) and combines their answers. The complete project uses Spring Boot 4.1.1 (Spring Framework 7.0.9, Tomcat 11.0.24), Java 25 and JUnit 6.1.3, and checks the results with 15 tests (mvn test).

2.1. The Weather API and Its Slow Services

A small stub server plays both downstream services and answers every call after a fixed delay, which we set in the stub’s own properties file.

server.port=9090
stub.delay=250ms

The stub is a WebFlux app, so it can hold thousands of waiting calls without becoming the bottleneck in a load test. The operator delayElement() sends the answer later without blocking a thread.

@GetMapping("/forecast/{city}")
public Mono<Forecast> forecast(@PathVariable String city) {
  return Mono.just(new Forecast(city, 18, "cloudy"))
      .delayElement(delay);   // 250 ms
}
mvn spring-boot:run -Dspring-boot.run.main-class=com.howtodoinjava.virtualthreads.stub.StubApplication
curl localhost:9090/forecast/london     # {"city":"london","temperature":18,"sky":"cloudy"}

The weather API needs the Spring MVC starter and the RestClient starter, and the project also has the WebFlux starters for the comparison in section 3.

<dependency>
  <groupId>org.springframework.boot</groupId>
  <artifactId>spring-boot-starter-webmvc</artifactId>
</dependency>
<dependency>
  <groupId>org.springframework.boot</groupId>
  <artifactId>spring-boot-starter-restclient</artifactId>
</dependency>

2.2. Turning On Virtual Threads

Virtual threads need Java 21 or later, and the Spring Boot docs strongly recommend Java 24 or later because of the pinning fix in section 4.2. Then we set one property.

spring.threads.virtual.enabled=true

# Kept the same for both runs of the load test (ignored with virtual threads)
server.tomcat.threads.max=200

# RestClient on the JDK HttpClient (WebFlux puts Reactor Netty on the classpath too)
spring.http.clients.imperative.factory=jdk

weather.stub-url=http://localhost:9090
weather.air-quality-limit=50

The spring.threads.virtual.enabled property switches four components of our Spring Boot 4.1.1 app to virtual threads.

Componentvirtual.enabled=falsevirtual.enabled=true
Tomcat request threadsPool http-nio-8080-exec-N, max 200New virtual thread tomcat-handler-N per request
applicationTaskExecutor (@Async)ThreadPoolTaskExecutorSimpleAsyncTaskExecutor, virtual thread task-N
taskScheduler (@Scheduled)ThreadPoolTaskSchedulerSimpleAsyncTaskScheduler
JDK HttpClient behind RestClientThe JDK client’s default executorVirtual threads named httpclient-N

With virtual threads on, pool properties such as server.tomcat.threads.max and spring.task.execution.pool.max-size have no effect, because there is no pool left to size. The HttpClient row applies only when RestClient uses the JDK client, which is why we set spring.http.clients.imperative.factory=jdk.

2.3. A Blocking Controller With RestClient

The client is ordinary blocking code. The method body() waits until the stub answers, and the virtual thread unmounts during that wait.

public Forecast forecast(String city) {
  return restClient.get()
      .uri("/forecast/{city}", city)
      .retrieve()
      .body(Forecast.class);           // waits 250 ms
}

The controller calls both services one after the other. WeatherReport.of() also records the current thread, so we can see it in the response.

@GetMapping("/weather/{city}")
public WeatherReport weather(@PathVariable String city) {
  Forecast forecast = client.forecast(city);
  AirQuality airQuality = client.airQuality(city);
  return WeatherReport.of(forecast, airQuality);
}

// WeatherReport.of(): thread = Thread.currentThread().toString(), virtual = isVirtual()
mvn spring-boot:run
curl localhost:8080/weather/london
{"city":"london","temperature":18,"sky":"cloudy","aqi":42,
 "thread":"VirtualThread[#35,tomcat-handler-0]/runnable@ForkJoinPool-1-worker-1","virtual":true}

We read the thread name as virtual thread number 35, named tomcat-handler-0 by Tomcat, currently mounted on carrier ForkJoinPool-1-worker-1.

2.4. Checking That a Request Runs on a Virtual Thread

A test makes the check repeatable. We start the app on a random port, call the endpoint and read the virtual flag that Thread.currentThread().isVirtual() set on the server.

@SpringBootTest(classes = WeatherApplication.class, webEnvironment = WebEnvironment.RANDOM_PORT)
class VirtualThreadsTest {

  @Test
  void requestRunsOnVirtualThread() {
    WeatherReport report = api().get().uri("/weather/london").retrieve().body(WeatherReport.class);

    assertThat(report.virtual()).isTrue();
    assertThat(report.thread()).contains(",tomcat-handler-");
  }

  @Test
  void asyncMethodRunsOnVirtualThread() throws Exception {
    Thread thread = auditService.audit(report).get();   // an @Async method that returns its thread
    assertThat(thread.isVirtual()).isTrue();
  }
}

The same tests with spring.threads.virtual.enabled=false show the platform threads.

virtual.enabled=true
  Request thread: VirtualThread[#75,tomcat-handler-0]/runnable@ForkJoinPool-1-worker-1
  @Async thread:  VirtualThread[#126,task-1]/terminated
  HttpClient executor thread: VirtualThread[#125,httpclient-42]/terminated

virtual.enabled=false
  Request thread: Thread[#50,http-nio-auto-1-exec-1,5,main]
  @Async thread:  Thread[#65,task-1,5,main]

The name http-nio-auto-1 is what Tomcat uses for a random port, whereas on port 8080 the name is http-nio-8080-exec-1.

2.5. Calling Both Services in Parallel

The two calls do not depend on each other, so we can run them at the same time. For example, the weather page shows the forecast and the air quality side by side, and neither call needs the other’s result. For a short task, we create a new virtual thread instead of borrowing one from a pool. The executor from Executors.newVirtualThreadPerTaskExecutor() starts one virtual thread per submit(), and closing the executor in try-with-resources waits for both tasks.

@GetMapping("/weather/{city}/parallel")
public WeatherReport weatherParallel(@PathVariable String city) throws Exception {
  try (ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor()) {
    Future<Forecast> forecast = executor.submit(() -> client.forecast(city));
    Future<AirQuality> airQuality = executor.submit(() -> client.airQuality(city));
    return WeatherReport.of(forecast.get(), airQuality.get());
  }
}

Each wait parks only its own virtual thread, so the two waits overlap.

Timeline of one request with two 250 ms service calls: a platform thread http-nio-exec-1 is blocked for the full 500 ms; a virtual thread tomcat-handler-0 runs for a few milliseconds, is parked and unmounted during each wait, while carrier worker-1 runs other virtual threads VT 2, VT 7, VT 31, VT 5, VT 12 and VT 40; with the parallel endpoint two new virtual threads wait for forecast and air quality at the same time and the response arrives after about 250 ms
A virtual thread does not make one call faster, but it frees the carrier during every wait. Two virtual threads let the waits overlap.
sequential 555 ms, parallel 287 ms

Java 25 also has StructuredTaskScope for this pattern, but structured concurrency is still a preview API (JEP 505).

2.6. Load Testing Platform vs Virtual Threads

We sent the same load to two copies of the weather API, one with spring.threads.virtual.enabled=false and one with true, both with server.tomcat.threads.max=200. The load tool is hey, a small HTTP load generator, where -c is the number of concurrent users and -n the total number of requests.

go install github.com/rakyll/hey@latest

hey -n 2000 -c 200 http://localhost:8080/weather/london     # warm-up, result ignored
hey -n 10000 -c 1000 http://localhost:8080/weather/london
Summary:
  Total:        14.0603 secs
  Requests/sec: 711.2200

Latency distribution:
  50%% in 1.2192 secs
  99%% in 2.4161 secs

Status code distribution:
  [200] 10000 responses

Every request waits 2 x 250 ms, so with 200 threads, the platform version can finish at most 200 / 0.5 s = 400 requests per second. During the load tests, we raised weather.air-quality-limit to 10,000, so the semaphore from section 5.2 did not limit the results.

The numbers are measurements on a small sandbox with 2 CPUs shared with other jobs, and hey, the stub and the app run on the same machine. Each value is the median of three runs.

LoadPlatform threadsVirtual threads
100 users, 3,000 requestsRequests/s188180
Median latency0.52 s0.53 s
1,000 users, 10,000 requestsRequests/s374 (342 to 382)506 (425 to 1,154)
Median latency2.63 s1.74 s
99th percentile2.94 s4.34 s
CPU time per request0.58 to 0.72 ms0.40 to 0.47 ms

With 100 users, both versions are the same, because 100 waiting requests fit in 200 threads. With 1,000 users, the platform version stays under its 400 req/s limit, and requests queue for a thread. The virtual thread version passed that limit in every run, but its numbers varied, because our two shared CPUs, not the threads, became the limit (we read the CPU time of each process from /proc/PID/stat before and after a run).

The 99th percentile of the virtual thread version also moved a lot between runs, from 2.42 s to 4.88 s, so we do not read much into it.

3. Virtual Threads vs Spring WebFlux

Spring WebFlux solves the same problem in another way. WebFlux never blocks a thread. Each call returns a Mono (a value that arrives later), and a few Netty event loop threads run all the callbacks. Virtual threads keep blocking code and make the blocking cheap.

3.1. The Same Endpoint in WebFlux

The WebFlux version uses WebClient. The two sequential calls become a flatMap() chain, and the parallel version uses Mono.zip().

@GetMapping("/weather/{city}")
public Mono<WeatherReport> weather(@PathVariable String city) {
  return forecast(city)
      .flatMap(forecast -> airQuality(city)
          .map(airQuality -> WeatherReport.of(forecast, airQuality)));
}

@GetMapping("/weather/{city}/parallel")
public Mono<WeatherReport> weatherParallel(@PathVariable String city) {
  return Mono.zip(forecast(city), airQuality(city), WeatherReport::of);
}

private Mono<Forecast> forecast(String city) {
  return webClient.get().uri("/forecast/{city}", city).retrieve().bodyToMono(Forecast.class);
}
{"city":"london","temperature":18,"sky":"cloudy","aqi":42,
 "thread":"Thread[#30,reactor-http-epoll-4,5,main]","virtual":false}

Netty builds the report on an event loop thread, not on a virtual thread. The WebFlux test measured sequential 531 ms, parallel 277 ms, the same pattern as in section 2.5.

3.2. Performance Comparison

We ran the same hey commands against the WebFlux app on the same sandbox. With 1,000 users, WebFlux reached 883 requests per second (699 to 899) with a median latency of 1.08 s and a 99th percentile of 1.35 s.

Bar chart of requests per second, median of three runs: with 100 users platform threads 188, virtual threads 180, WebFlux 188, median latency about 0.52 s for all; with 1,000 users platform threads 374 with median latency 2.63 s, virtual threads 506 with 1.74 s, WebFlux 883 with 1.08 s; a dashed line marks the platform limit of 400 requests per second; a note says virtual threads ranged from 425 to 1,154 requests per second between runs
Virtual threads remove the 400 req/s limit of the thread pool. On our two shared CPUs, WebFlux used the least CPU per request and went furthest.

WebFlux used 0.28 to 0.33 ms of CPU time per request in our runs, less than both MVC versions, and on a machine where the CPU is the limit, less CPU per request means more requests per second. The waits alone would allow 1,000 users / 0.5 s = 2,000 requests per second, but on our sandbox the CPU ran out first for every version.

3.3. Which One Should We Choose?

For a new blocking-style Spring MVC app, we enable virtual threads first and measure. WebFlux is worth its learning curve when the app streams data with back-pressure (a slow client slows down the producer), or when CPU per request matters most. For example, a stock price feed that pushes updates to thousands of browsers fits WebFlux.

QuestionVirtual threads (Spring MVC)Spring WebFlux
Code stylePlain blocking code, normal stack tracesMono/Flux chains, operators to learn
Existing code and librariesJDBC, JPA, RestClient work as they areNeeds reactive drivers (R2DBC, WebClient)
DebuggingThread dumps and debuggers show the requestStack traces show the event loop, not the request
Streaming, back-pressure, server-sent eventsPossible, but not the strengthBuilt in
CPU per request in our test0.40 to 0.47 ms0.28 to 0.33 ms
Migration costOne propertyRewrite of the I/O code

4. Debugging and Monitoring Virtual Threads

Virtual threads are real Thread objects, so the JDK tools see them. Two tools cover most questions. A thread dump shows what each request is waiting for, and JDK Flight Recorder (JFR) records pinning.

4.1. Thread Dumps With jcmd

The classic jstack thread dump lists only platform threads. The command jcmd Thread.dump_to_file lists virtual threads too, in plain text or JSON.

jcmd <pid> Thread.dump_to_file -format=json threads.json
jcmd <pid> Thread.dump_to_file threads.txt

We took a dump while hey sent 500 concurrent users, and each request was a virtual thread parked in RestClient, waiting for the stub.

#2069 "tomcat-handler-525" virtual WAITING 2026-10-03T17:39:01.179774508Z
    at java.base/java.lang.VirtualThread.park(VirtualThread.java:745)
    - parking to wait for <java.util.concurrent.CompletableFuture$Signaller@10183f17>
    ...
    at java.base/java.util.concurrent.CompletableFuture.get(CompletableFuture.java:2093)
    at org.springframework.http.client.JdkClientHttpRequest.executeInternal(JdkClientHttpRequest.java:126)
    ...

The JSON file groups threads by container. In our dump the root container held 506 tomcat-handler and 42 httpclient virtual threads, and the ForkJoinPool-1 container held 2 carrier threads, one per CPU.

4.2. Pinning After JDK 24

A virtual thread is pinned when it cannot unmount while it blocks, so it keeps its carrier busy. With only one carrier per CPU, a few pinned threads can stall every other request. On Java 21 to 23, blocking inside a synchronized block or method pinned the thread, but JEP 491 in JDK 24 removed that case, so on Java 25 synchronized is fine again.

Some cases still pin. We ran each case on a virtual thread that blocks for 50 ms and recorded the JFR events (PinningTest in the project).

Virtual thread blocks while …Pinned on Java 25?pinnedReason in JFR
Inside synchronized (lock)NoNo event
Holding a ReentrantLockNoNo event
Inside a Java callback called from native code (FFM qsort upcall)YesNative or VM frame on stack
Inside a static initializer (static { } block)YesVM call to PinningTest$SlowInit.<clinit> on stack

The FFM (Foreign Function and Memory) API lets Java call C functions, and in an upcall, the C function calls back into Java. Few Spring apps do either in their own code, but a library with native code can.

static class SlowInit {
  static final String VALUE;

  static {
    pause();          // blocks 50 ms: jdk.VirtualThreadPinned, the carrier is held
    VALUE = "ready";
  }
}

The old -Djdk.tracePinnedThreads option that many tutorials show has no effect since JDK 24, so we use JFR instead.

4.3. Recording Pinned Threads With JFR

JFR writes a jdk.VirtualThreadPinned event when a virtual thread stays pinned longer than 20 ms. The event is on by default, so we only start a recording and print the events.

java -XX:StartFlightRecording=filename=recording.jfr -jar target/spring-boot-virtual-threads-1.0.0.jar
jfr print --events jdk.VirtualThreadPinned recording.jfr
jdk.VirtualThreadPinned {
  startTime = 23:14:27.143 (2026-10-03)
  duration = 67.6 ms
  blockingOperation = "LockSupport.park"
  pinnedReason = "Native or VM frame on stack"
  carrierThread = "ForkJoinPool-1-worker-2" (javaThreadId = 79)
  eventThread = "" (javaThreadId = 644, virtual)
  stackTrace = [ ... ]
}

The stackTrace shows the code that blocked, and pinnedReason says why the thread could not unmount. The weather API itself recorded no jdk.VirtualThreadPinned event while hey sent it 5,000 requests from 500 users. JFR has three more virtual thread events, namely jdk.VirtualThreadStart and jdk.VirtualThreadEnd (off by default, because they are frequent) and jdk.VirtualThreadSubmitFailed (on).

5. Spring Boot Virtual Threads FAQs

5.1. Should We Pool Virtual Threads?

No. Thread pools exist because platform threads are expensive, whereas a virtual thread is cheap, so each task gets a new one. Both versions compile, but only the second one creates a new virtual thread for each task.

// Wrong: a pool of virtual threads
ExecutorService pool = Executors.newFixedThreadPool(50, Thread.ofVirtual().factory());

// Right: one new virtual thread per task
try (ExecutorService executor = Executors.newVirtualThreadPerTaskExecutor()) {
  executor.submit(() -> client.forecast(city));
}

The same applies to Spring beans. Any Executor bean of our own switches off the virtual applicationTaskExecutor that Spring Boot creates. In our test, a ThreadPoolTaskExecutor bean with the prefix report- did not even run the @Async method on it. Spring found two task executors (ours and the scheduler bean, which is also a TaskExecutor), could not pick one, and fell back to a new platform thread. We let Spring Boot create the executor instead.

@Async thread with a custom pool: Thread[#652,SimpleAsyncTaskExecutor-1,5,main]

5.2. How Do We Limit Calls to a Slow Downstream Service?

A thread pool used to limit concurrency as a side effect, because 50 threads meant at most 50 calls, whereas with virtual threads, 1,000 requests make 1,000 calls at once. When a downstream service allows only a few calls at a time, we use a Semaphore, not a pool. A semaphore hands out a fixed number of permits, and a thread without a permit waits.

private final Semaphore airQualityPermits = new Semaphore(properties.airQualityLimit());

public AirQuality airQuality(String city) {
  airQualityPermits.acquireUninterruptibly();   // a virtual thread waits here cheaply
  try {
    return restClient.get().uri("/air-quality/{city}", city).retrieve().body(AirQuality.class);
  } finally {
    airQualityPermits.release();
  }
}
30 requests, max air quality calls at the same time: 5

A database connection pool already limits the concurrent queries, so with virtual threads, the pool size, not the thread count, decides how many queries run at once.

5.3. Is ThreadLocal Safe With Virtual Threads?

Yes, a ThreadLocal works the same way, and Spring still uses ThreadLocal for the security context and transactions. The problem is caching, because code that keeps an expensive object in a ThreadLocal to reuse it across tasks gets no reuse, because each task has a new thread. The fix is a shared, thread-safe object instead of a per-thread cache.

// Creates a new SimpleDateFormat for every request on virtual threads
static final ThreadLocal<SimpleDateFormat> FORMAT = ThreadLocal.withInitial(SimpleDateFormat::new);

// Thread-safe and shared by all threads
static final DateTimeFormatter FORMATTER = DateTimeFormatter.ISO_LOCAL_DATE;

Large values in thread locals also multiply the memory per request, and there can be thousands of requests at once.

5.4. Does server.tomcat.threads.max Still Matter?

No. With virtual threads, Tomcat starts a new virtual thread for every request and ignores server.tomcat.threads.max, and our thread dump in section 4.1 showed 506 request threads while the property was 200. The limits that still apply are the connection settings.

PropertyDefaultMeaning with virtual threads
server.tomcat.threads.max200Ignored
server.tomcat.max-connections8192Open connections Tomcat accepts; the real cap on concurrent requests
server.tomcat.accept-count100Queue for connections above max-connections

5.5. Why Does the Application Stop Right After Startup?

Virtual threads are daemon threads, threads that do not keep the JVM running, and the JVM exits when only daemon threads are left. An app without a web server, for example one that only runs @Scheduled jobs, can stop right after startup. For that case, we set one more property, which keeps the JVM running.

spring.main.keep-alive=true

6. Conclusion

The property spring.threads.virtual.enabled=true moves Tomcat requests, @Async methods, scheduled tasks and the JDK HttpClient to virtual threads, and our tests confirmed the switch with Thread.currentThread().isVirtual(). Virtual threads remove the thread pool limit for I/O-bound apps without changing the code, but they do not make a single request faster or add CPU.

On Java 25, synchronized no longer pins, whereas native callbacks and class initializers still pin, and JFR shows those cases. We create a new virtual thread per task, limit downstream calls with a Semaphore, and keep WebFlux for streaming or for apps where CPU per request is the main cost.

7. References

Happy Learning !!

Source Code on Github

Leave a Comment

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.