Jobs, Steps & Chunks
Die Anatomie eines Batch-Jobs: ein Job aus Steps und das chunk-orientierte Modell, das in transaktionalen Blöcken liest, verarbeitet und schreibt.
Deutsche Übersetzung in Arbeit
Diese Lektion ist noch nicht ins Deutsche übersetzt und wird daher auf Englisch angezeigt. Der Rest der Seite ist vollständig lokalisiert.
Auf dieser Seite
A Spring Batch Job is the unit you launch. Inside, it's a sequence of Steps, and the workhorse step is chunk-oriented: it reads items one at a time, processes each, and writes them out in fixed-size chunks - each chunk wrapped in its own transaction. That structure is the whole reason batch scales, so let's build the mental model precisely.
Job → Steps → Chunks
A Job is a named, restartable process. It's composed of one or more Steps run in order (or branching on success/failure). Each Step does one phase of work. For BookVault's nightly report, the Job might be:
- Step 1 - read active loans, keep the overdue ones, write report rows to a staging table.
- Step 2 - read staged rows grouped by member, email each member their summary.
Within a chunk-oriented Step, work flows as read → process → write, batched:
@Bean
Job overdueReportJob(JobRepository repo, Step reportStep) {
return new JobBuilder("overdueReportJob", repo)
.start(reportStep) // step 1; .next(emailStep) would chain step 2
.build();
}
@Bean
Step reportStep(JobRepository repo, PlatformTransactionManager tx,
ItemReader<Loan> reader, ItemProcessor<Loan, ReportRow> processor,
ItemWriter<ReportRow> writer) {
return new StepBuilder("reportStep", repo)
.<Loan, ReportRow>chunk(100, tx) // 100 items per transaction
.reader(reader)
.processor(processor)
.writer(writer)
.build();
}The chunk is a transaction boundary
The magic number chunk(100, tx) means: read 100 items (one at a time), process each, then hand the whole
batch of 100 to the writer once, and commit. Then the next 100. This is the key idea:
- Bounded memory - only 100 items are in flight, never the whole 2-million-row dataset.
- Incremental durability - each chunk commits independently. If the job dies at record 1,900,050, the 1,900,000 records in already-committed chunks are safe.
- Efficient writes - the writer receives a
Listof 100, so it can do one batchedINSERTinstead of 100 round-trips.
read read read ... (×100) → process each → writer.write(List<100>) → COMMIT
read read read ... (×100) → process each → writer.write(List<100>) → COMMIT
...Choosing chunk size
Chunk size trades throughput against memory and rollback cost. Too small (say 1) means a transaction commit per item - slow. Too large (say 100,000) means huge memory use and, on failure, a big chunk to roll back and reprocess. A few hundred to a few thousand is typical; tune with your real data and row size.
JobInstance, JobExecution, StepExecution
Spring Batch records every run in the JobRepository with three concepts worth knowing:
- JobInstance - a logical run identified by its job parameters (e.g.
reportDate=2026-07-22). The same parameters = the same instance. - JobExecution - one attempt at a JobInstance. If Monday's run fails and you restart it, that's a second JobExecution of the same JobInstance.
- StepExecution - one attempt at a Step, tracking read/write/skip counts and the commit position.
This is what makes restart work: the JobRepository knows a given JobInstance's last position, so a restart resumes instead of starting fresh (next lesson).
The Job is the whole book you set out to read; each Step is a chapter you finish before starting the next. The chunk is a page you fully read before slipping in a bookmark - you never re-read pages behind the bookmark, and if you're interrupted, the bookmark shows exactly where to resume. The JobInstance is 'reading this edition, starting tonight'; a JobExecution is one evening's sitting; if you fall asleep and pick it up tomorrow, same book, second sitting, resumed from the bookmark.
In a Step configured with chunk(50), your ItemWriter does a batched database insert. Over 500 total items, how many transactions commit, and how many times is the writer's write() method called? What does this buy you if the job crashes at item 220?
What does chunk(100) control in a chunk-oriented step?
Key takeaways
- A Job is a restartable process made of ordered Steps; the chunk-oriented step is the workhorse.
- A chunk-oriented step flows read → process → write, batched: read/process per item, write and commit per chunk.
- The chunk size is the transaction boundary - it bounds memory, gives incremental durability, and enables batched writes.
- Tune chunk size between throughput (larger) and memory/rollback cost (smaller); hundreds to low thousands is typical.
- JobInstance (identified by parameters), JobExecution (one attempt), and StepExecution (per-step progress) are how the JobRepository tracks runs and enables restart.