Hibernate Batch Insert and Update: batch_size Explained

Hibernate sends inserts, updates and deletes in JDBC batches once hibernate.jdbc.batch_size is set. Measured on Hibernate 7.4: why IDENTITY ids stop batching, how flush() and clear() keep memory flat, and when to use order_inserts, StatelessSession or a JPQL bulk statement.

Jakarta_EE

Batch processing in Hibernate means sending many INSERT, UPDATE or DELETE statements to the database in groups, one JDBC call per group, instead of one call per row. Each call is a round trip (our code sends the request to the database and waits for the answer), so a Hibernate batch insert of 10,000 rows can take 200 calls instead of 10,000.

We use batch processing for imports and nightly jobs that write thousands of rows, for example loading a day of readings from home temperature sensors. We turn it on with the hibernate.jdbc.batch_size setting and keep memory flat with flush() and clear(). When a job does not need tracked entities, we switch to a StatelessSession or a JPQL bulk statement.

The following example turns on JDBC batching and imports 10,000 sensor readings in chunks of 50.

HibernatePersistenceConfiguration config = new HibernatePersistenceConfiguration("sensor-import")
    .property("hibernate.jdbc.batch_size", 50)
    .property("hibernate.order_inserts", true)
    .property("hibernate.order_updates", true);
emf.runInTransaction(em -> {
  for (int i = 0; i < readings.size(); i++) {
    em.persist(readings.get(i));
    if ((i + 1) % 50 == 0) {
      em.flush();        // executeBatch(50): insert into SensorReading ...
      em.clear();        // the persistence context is empty again
    }
  }
});                      // 10,000 inserts in 200 round trips (10,000 without batch_size)

Notice that the flush interval matches hibernate.jdbc.batch_size, so each flush() sends one full batch of 50 inserts.

Next, we look at how JDBC batching works and why IDENTITY ids disable it. After that, we batch updates and deletes, and compare batching with JPQL bulk statements and StatelessSession.

1. How Does JDBC Batching Work in Hibernate?

Hibernate does not send SQL at persist(). It keeps every new or changed entity in the persistence context, a map of the objects one EntityManager tracks, and writes the SQL at flush time (the step that turns pending changes into SQL statements), in most cases at commit. Without batching, each of those statements is a separate executeUpdate() call and waits for the database to answer. With hibernate.jdbc.batch_size set, Hibernate collects up to that many statements for the same table with JDBC addBatch() and sends them with one executeBatch() call.

Two sequence diagrams for importing 10,000 readings: without hibernate.jdbc.batch_size, Hibernate sends insert reading 1, 2, 3 up to 10,000 as separate calls, 10,000 round trips; with batch_size = 50 it sends executeBatch with readings 1-50, 51-100 up to 9,951-10,000, 200 round trips; a note says both send 10,000 INSERT statements and call the sequence 201 times, and IDENTITY ids need 10,000 round trips even with batch_size = 50
The database still runs 10,000 INSERT statements. Batching cuts the number of round trips, which is where most of the time goes.

Batching is a JDBC feature, so it does not change the SQL. The same statements reach the database, only in fewer calls. Four settings control it.

SettingDefault in Hibernate 7.4What it does
hibernate.jdbc.batch_sizeNot set: no batchingMaximum statements per executeBatch() call
hibernate.order_insertsfalseGroups INSERT statements by table so batches stay full
hibernate.order_updatesfalseSorts UPDATE statements by table and id
Session.setJdbcBatchSize(n)Global valueOverrides the batch size for one session

Batching is different from a bulk statement, because a batch still has one statement per row, whereas a JPQL update … where … is one statement for all rows and skips the entities in memory. The example uses both.

2. Hibernate Batch Insert Example

The following example imports temperature readings from home sensors (“Kitchen”, “Garage”) into an in-memory H2 database with Hibernate 7.4.11 and Java 25. To count what reaches the database, the project wraps the H2 DataSource with datasource-proxy, which reports every execute() and executeBatch() call.

2.1. The Sensor Reading Model

A reading has a sensor name, a time as LocalDateTime and a temperature. Its id comes from a database sequence, which matters for batching, as section 2.3 shows.

@Entity
public class SensorReading {

  @Id
  @GeneratedValue(strategy = GenerationType.SEQUENCE)
  @SequenceGenerator(sequenceName = "reading_seq", allocationSize = 50)
  private Long id;

  private String sensorName;
  private LocalDateTime measuredAt;
  private double temperature;
}

Two more entities show special cases. LegacyReading has the same columns but an auto-increment id, and Sensor lists its readings with a @OneToMany on the sensorName column and saves new ones by cascade.

// LegacyReading
@Id
@GeneratedValue(strategy = GenerationType.IDENTITY)
private Long id;

// Sensor
@Version
private int version;

@OneToMany(cascade = CascadeType.PERSIST)
@JoinColumn(name = "sensorName", referencedColumnName = "name", insertable = false, updatable = false)
private List<SensorReading> readings = new ArrayList<>();

The counting listener is registered on the proxy, and we pass the proxy to HibernatePersistenceConfiguration as the DataSource.

DataSource dataSource = ProxyDataSourceBuilder.create(h2DataSource)
    .listener(counter)          // afterQuery() is called once per execute() or executeBatch()
    .build();

HibernatePersistenceConfiguration config = new HibernatePersistenceConfiguration("sensor-import")
    .managedClasses(Sensor.class, SensorReading.class, LegacyReading.class)
    .property("hibernate.connection.datasource", dataSource)
    .property("hibernate.jdbc.batch_size", 50);

2.2. Turning On Batch Inserts

We persist five readings in a loop, once without and once with hibernate.jdbc.batch_size = 50. The persist() code is the same in both runs.

emf.runInTransaction(em -> readings.forEach(em::persist));
execute  select next value for reading_seq
execute  select next value for reading_seq
execute  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
execute  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
execute  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
execute  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
execute  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
execute          select next value for reading_seq
execute          select next value for reading_seq
executeBatch(5)  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)

Without hibernate.jdbc.batch_size, Hibernate 7.4 sends every INSERT on its own. For 10,000 readings that is 10,000 round trips, whereas a batch size of 50 gives 200 full batches of 50.

To turn batching on for one job only, we set the size on the session and leave the global setting alone.

em.unwrap(Session.class).setJdbcBatchSize(50);   // 10,000 inserts in 200 batches, no global setting

2.3. Why IDENTITY Ids Disable Insert Batching

With @GeneratedValue IDENTITY, the database assigns the id while it inserts the row. Hibernate must know the id after persist() to put the entity in the persistence context, so it runs each INSERT at once and reads the generated key.

execute  insert into LegacyReading (measuredAt,sensorName,temperature,id) values (?,?,?,default)
execute  insert into LegacyReading (measuredAt,sensorName,temperature,id) values (?,?,?,default)
execute  insert into LegacyReading (measuredAt,sensorName,temperature,id) values (?,?,?,default)
...                                                                    // 10,000 calls, 0 batches

For batch inserts, map the id with GenerationType.SEQUENCE. With a sequence, Hibernate gets the id before the INSERT, so the statements can wait in the batch. The allocationSize = 50 uses Hibernate’s pooled optimizer: one select next value reserves 50 ids, and Hibernate assigns them in memory.

Id strategyCalls to get ids for 10,000 rowsINSERT round trips (batch_size 50)
SEQUENCE, allocationSize = 50201 select next value for reading_seq200
IDENTITYNone, the INSERT returns the id10,000

The pooled optimizer called the sequence twice for the first block and once per 50 ids after that, which gives 201 calls. MySQL has no sequences, so on MySQL the common choices are a sequence table or ids generated in Java, such as UUIDs.

2.4. Keeping Memory Flat With flush() and clear()

Batching reduces round trips but not memory. Every persisted reading stays in the persistence context until the transaction ends, together with a copy of its state for dirty checking, the comparison that finds changed entities at flush time. The method flush() writes the pending statements, and clear() detaches all entities so the garbage collector can free them.

Line chart of managed entities while 100,000 readings are persisted: with persist() only the count rises in a straight line to 100,000; with flush() and clear() every 50 it stays at the bottom, never above 50; a zoom on the first 200 readings shows a sawtooth from 0 to 50; a note gives the sandbox heap after GC as 46 MB without clear() and 27 MB with it, both including the H2 rows
Without clear(), the persistence context grows with every row. With flush() and clear() every 50 rows, it never holds more than 50.
emf.runInTransaction(em -> {
  for (int i = 0; i < readings.size(); i++) {
    em.persist(readings.get(i));
    if ((i + 1) % 50 == 0) {                   // same number as hibernate.jdbc.batch_size
      em.flush();                              // executeBatch(50)
      em.clear();                              // managed entities: 50 -> 0
    }
  }
});
Loop for 100,000 readingsManaged entities at the end of the loopUsed heap after GC (sandbox)
persist() only100,00046 MB
persist() + flush() + clear() every 50At most 5027 MB

The heap numbers come from one run in our sandbox and include the H2 in-memory rows, which are in the same JVM. Use the batch size as the flush interval, so each flush() sends full batches. After clear(), the readings are detached entities, and later changes to them are not saved.

2.5. Mixing Entity Types With order_inserts

A JDBC batch holds one SQL string, so it can contain rows for only one table. When the flush order alternates between tables, every switch closes the open batch. For example, a smart-home app registers two new sensors with two readings each.

for (String name : List.of("Kitchen", "Garage")) {
  Sensor sensor = new Sensor(name, "Home");
  readings(name, 2).forEach(sensor::addReading);
  em.persist(sensor);
}
executeBatch(1)  insert into Sensor (location,name,version,id) values (?,?,?,?)
executeBatch(2)  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
executeBatch(1)  insert into Sensor (location,name,version,id) values (?,?,?,?)
executeBatch(2)  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)
executeBatch(2)  insert into Sensor (location,name,version,id) values (?,?,?,?)
executeBatch(4)  insert into SensorReading (measuredAt,sensorName,temperature,id) values (?,?,?,?)

With 100 sensors and 3 readings each, the difference is 200 round trips against 8.

Two rows of statement boxes: with order_inserts false the flush sends Sensor 1, three readings, Sensor 2, three readings and so on, giving batches of 1 and 3 and 400 inserts in 200 round trips; with order_inserts true it sends Sensor 1-50, Sensor 51-100, then readings in groups of 50, giving batches of 50 and 400 inserts in 8 round trips; a note says updates behave the same with order_updates, 200 against 8
order_inserts sorts the pending inserts by table before the flush, so every batch is full.

Turn on hibernate.order_inserts whenever a job saves more than one entity type, for example a parent with cascade on its children. Sorting takes some CPU time, so we measure the job before and after turning it on.

2.6. Batch Updates and Deletes

Updates and deletes use the same batch size. For example, the Kitchen sensor turns out to read 0.5 degrees too high, so we load its 1,000 readings and correct each temperature. Hibernate’s dirty checking writes the changes at commit.

emf.runInTransaction(em -> {
  List<SensorReading> kitchen = em.createQuery(
          "from SensorReading where sensorName = :name", SensorReading.class)
      .setParameter("name", "Kitchen")
      .getResultList();
  kitchen.forEach(r -> r.setTemperature(r.getTemperature() - 0.5));
});   // executeBatch(50) update SensorReading set measuredAt=?,sensorName=?,temperature=? where id=?  (x 20)

Deleting works the same way. A call to em.remove() in a loop queues one DELETE per reading, and the flush sends them in batches.

kitchen.forEach(em::remove);   // executeBatch(50) delete from SensorReading where id=?  (x 20)
Operation on 1,000 readingsStatementsRound trips
Change each reading, no batch_size1,000 UPDATE1,000
Change each reading, batch_size = 501,000 UPDATE20
em.remove() each reading, batch_size = 501,000 DELETE20

Updates of two entity types can alternate like the inserts in section 2.5. We moved 100 sensors to a new location and changed their 3 readings each, so the UPDATE statements alternate between Sensor and SensorReading. The setting hibernate.order_updates = true brought that from 200 round trips to 8, but only after one more change. A query inside the loop flushes the pending changes before it runs (flush mode AUTO), so each flush held only one sensor and its readings.

em.setFlushMode(FlushModeType.COMMIT);     // queries in the loop no longer flush
for (String name : names) {
  Sensor sensor = em.createQuery("from Sensor where name = :name", Sensor.class)
      .setParameter("name", name)
      .getSingleResult();
  sensor.setLocation("Office");            // update Sensor set location=?,name=?,version=? where id=? and version=?
  em.createQuery("from SensorReading where sensorName = :name", SensorReading.class)
      .setParameter("name", name)
      .getResultList()
      .forEach(r -> r.setTemperature(r.getTemperature() - 0.5));
}
// AUTO:   200 round trips, with or without order_updates
// COMMIT: 200 round trips without order_updates, 8 with it

Sensor has a @Version column, and its updates still went out in batches of 50. With COMMIT, a query filters on the values in the database, not on the changes waiting in memory. For example, after setLocation(“Office”), a count of sensors in the “Office” returned 0. We use it only when the loop does not query what it changed.

2.7. Bulk Updates and Deletes With JPQL

When the change is the same for every row, a JPQL bulk statement is faster than any batch, because it sends one statement in one round trip and loads no entities into memory. The method executeUpdate() returns the number of rows changed.

int updated = em.createQuery(
        "update SensorReading r set r.temperature = r.temperature + :offset where r.sensorName = :name")
    .setParameter("offset", -0.5)
    .setParameter("name", "Kitchen")
    .executeUpdate();     // updated = 1000
// update SensorReading sr1_0 set temperature=(sr1_0.temperature+cast(? as float(53))) where sr1_0.sensorName=?

int deleted = em.createQuery("delete from SensorReading r where r.measuredAt < :before")
    .setParameter("before", LocalDateTime.of(2027, 10, 1, 0, 0))
    .executeUpdate();     // deleted = 1300, all readings of all sensors
// delete from SensorReading sr1_0 where sr1_0.measuredAt<?

A bulk statement goes straight to the database and does not change entities that are already loaded. We call refresh() to read the new value, or run the bulk statement before loading anything.

SensorReading first = em.createQuery("from SensorReading order by measuredAt", SensorReading.class)
    .setMaxResults(1)
    .getSingleResult();                        // first.getTemperature() = 20.0
em.createQuery("update SensorReading r set r.temperature = r.temperature - 0.5").executeUpdate();
double before = first.getTemperature();        // still 20.0
em.refresh(first);
double after = first.getTemperature();         // 19.5

2.8. Bulk Inserts With StatelessSession

A StatelessSession is Hibernate’s command-style API for bulk work. It has no persistence context, so insert(), update() and delete() run their SQL at once or add it to the JDBC batch. Nothing is cached, and there is no dirty checking. We get one from the SessionFactory behind our EntityManagerFactory.

emf.unwrap(SessionFactory.class).inStatelessTransaction(ss -> {
  ss.setJdbcBatchSize(50);                    // required in Hibernate 7
  readings.forEach(ss::insert);
});                                           // 10,000 inserts in 201 round trips

emf.unwrap(SessionFactory.class).inStatelessTransaction(ss ->
    ss.insertMultiple(readings));             // 10,000 inserts in 201 round trips

In Hibernate 7, hibernate.jdbc.batch_size has no effect on a StatelessSession. Without setJdbcBatchSize(), our loop made 10,000 separate INSERT calls even with the global setting at 50. The new insertMultiple() batches the list on its own, and updateMultiple() and deleteMultiple() do the same for changes.

The missing persistence context also removes features we get with an EntityManager.

ss.insert(sensorWithTwoReadings);         // insert into Sensor ... only: readings = 0, no cascade
SensorReading a = ss.get(SensorReading.class, id);
SensorReading b = ss.get(SensorReading.class, id);   // a == b is false: no cache
a.setTemperature(30.0);                   // no SQL: the row keeps 20.0 until ss.update(a)

2.9. Comparing the Approaches

Each approach writes the same 10,000 readings, and the number of JDBC calls differs by a factor of 50. The time column comes from a separate run with 100,000 readings against H2 over TCP on localhost, best of three, so it includes a short network hop. The timings are sandbox measurements, so single values changed by up to a factor of 2.5 between runs, but the order stayed the same.

Approach (10,000 readings)INSERT statementsINSERT round tripsManaged entities at the endTime for 100,000 (sandbox)
persist(), no batch_size10,00010,00010,00012.4 s
persist(), IDENTITY id, batch_size 5010,00010,00010,00012.0 s
persist(), batch_size 5010,00020010,0001.3 s
persist() + flush()/clear() every 5010,000200At most 501.1 s
StatelessSession.insert(), global batch_size only10,00010,000None13.2 s
StatelessSession.insert() + setJdbcBatchSize(50)10,000201None1.1 s
StatelessSession.insertMultiple()10,000201None1.1 s

All rows with a SEQUENCE id also made 201 sequence calls. For a regular import, persist() with flush() and clear() keeps cascades and dirty checking and was as fast as the stateless options in the sandbox timings. A StatelessSession fits plain row copying, and a JPQL bulk statement fits a change that one WHERE clause can describe.

The complete project on GitHub prints the JDBC calls for each case and checks these numbers with 30 JUnit tests (mvn test, mvn -q compile exec:java).

3. Hibernate Batch Processing FAQs

3.1. Why Is Hibernate Not Batching My Inserts?

The setting alone is not enough, because several causes split or disable the batches. Each cause shows up as execute calls or as batches of 1 to 3 statements.

CauseWhat we sawFix
hibernate.jdbc.batch_size not set10,000 calls for 10,000 rowsSet it to 20 to 50
IDENTITY idsNo executeBatch() for insertsUse SEQUENCE with an allocationSize
Two entity types saved in turnBatches of 1 and 3hibernate.order_inserts = true
A query inside the loopEach query flushed a small batchRun queries before the loop, or flush mode COMMIT
StatelessSession in Hibernate 710,000 calls despite the settingsetJdbcBatchSize() or insertMultiple()

3.2. How Can We Check That Batching Works?

We set the org.hibernate.orm.jdbc.batch logger to DEBUG level, because Hibernate 7 logs each batch there, so in production we do not need a proxy.

INFO  org.hibernate.orm.jdbc.batch - HHH100501: Automatic JDBC statement batching enabled (maximum batch size 50)
DEBUG org.hibernate.orm.jdbc.batch - Adding to JDBC batch (1 / 50) - [com.howtodoinjava.hibernate.batch.SensorReading#INSERT]
...
DEBUG org.hibernate.orm.jdbc.batch - Executing JDBC batch (5 / 50) - [com.howtodoinjava.hibernate.batch.SensorReading#INSERT]

The org.hibernate.engine.jdbc.batch.internal.BatchingBatch logger from Hibernate 5 tutorials no longer exists; the class is gone from Hibernate 7.4.

3.3. What Batch Size Should We Use?

We use a batch size between 10 and 50. Going from 1 to 50 removes 98% of the round trips, whereas going from 50 to 500 removes only a few more and makes each batch and each persistence context chunk larger. We use the same number for batch_size and for the flush()/clear() interval, and the same allocationSize on the sequence.

3.4. How Do We Set the Batch Size in Spring Boot?

We add the Hibernate settings under spring.jpa.properties, because Spring Boot passes every *spring.jpa.properties.** entry to Hibernate with the prefix removed. In application.properties, the three batching settings get that prefix.

spring.jpa.properties.hibernate.jdbc.batch_size=50
spring.jpa.properties.hibernate.order_inserts=true
spring.jpa.properties.hibernate.order_updates=true

3.5. What Changed About Batching in Hibernate 7?

Old tutorials show APIs and settings that no longer exist or behave differently in Hibernate 7.4.

TopicBeforeHibernate 7.4
StatelessSession and batch_sizeUsed the global settingIgnores it; call setJdbcBatchSize() or use insertMultiple()
hibernate.jdbc.batch_versioned_dataSetting to batch @Version entitiesRemoved; versioned updates are batched
Session.save()Common in batch loopsRemoved; use persist()
Batch loggerBatchingBatchorg.hibernate.orm.jdbc.batch

The StatelessSession batching change came with Hibernate 7.0.

4. Conclusion

JDBC batching sends the same statements in fewer round trips. We set hibernate.jdbc.batch_size and use SEQUENCE ids, and we turn on order_inserts and order_updates when a job touches more than one table.

For large imports, we add flush() and clear() every batch to keep memory flat. For plain row copying, a StatelessSession with an explicit batch size works as well, and one JPQL bulk statement replaces row-by-row changes when a WHERE clause can describe them.

5. References

Happy Learning !!

Source Code on Github

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.