Loslegen
Javaneer
Zurück zur Stufe
Modul 10·Spring Batch

Jobs, Steps & Chunks

Die Anatomie eines Batch-Jobs: ein Job aus Steps und das chunk-orientierte Modell, das in transaktionalen Blöcken liest, verarbeitet und schreibt.

16 Min. LesezeitExperte

Deutsche Übersetzung in Arbeit

Diese Lektion ist noch nicht ins Deutsche übersetzt und wird daher auf Englisch angezeigt. Der Rest der Seite ist vollständig lokalisiert.

Auf dieser Seite

A Spring Batch Job is the unit you launch. Inside, it's a sequence of Steps, and the workhorse step is chunk-oriented: it reads items one at a time, processes each, and writes them out in fixed-size chunks - each chunk wrapped in its own transaction. That structure is the whole reason batch scales, so let's build the mental model precisely.

Job → Steps → Chunks

A Job is a named, restartable process. It's composed of one or more Steps run in order (or branching on success/failure). Each Step does one phase of work. For BookVault's nightly report, the Job might be:

  • Step 1 - read active loans, keep the overdue ones, write report rows to a staging table.
  • Step 2 - read staged rows grouped by member, email each member their summary.

Within a chunk-oriented Step, work flows as read → process → write, batched:

@Bean
Job overdueReportJob(JobRepository repo, Step reportStep) {
    return new JobBuilder("overdueReportJob", repo)
        .start(reportStep)          // step 1; .next(emailStep) would chain step 2
        .build();
}

@Bean
Step reportStep(JobRepository repo, PlatformTransactionManager tx,
                ItemReader<Loan> reader, ItemProcessor<Loan, ReportRow> processor,
                ItemWriter<ReportRow> writer) {
    return new StepBuilder("reportStep", repo)
        .<Loan, ReportRow>chunk(100, tx)   // 100 items per transaction
        .reader(reader)
        .processor(processor)
        .writer(writer)
        .build();
}

The chunk is a transaction boundary

The magic number chunk(100, tx) means: read 100 items (one at a time), process each, then hand the whole batch of 100 to the writer once, and commit. Then the next 100. This is the key idea:

  • Bounded memory - only 100 items are in flight, never the whole 2-million-row dataset.
  • Incremental durability - each chunk commits independently. If the job dies at record 1,900,050, the 1,900,000 records in already-committed chunks are safe.
  • Efficient writes - the writer receives a List of 100, so it can do one batched INSERT instead of 100 round-trips.
read read read ... (×100) → process each → writer.write(List<100>) → COMMIT
read read read ... (×100) → process each → writer.write(List<100>) → COMMIT
...

Choosing chunk size

Chunk size trades throughput against memory and rollback cost. Too small (say 1) means a transaction commit per item - slow. Too large (say 100,000) means huge memory use and, on failure, a big chunk to roll back and reprocess. A few hundred to a few thousand is typical; tune with your real data and row size.

JobInstance, JobExecution, StepExecution

Spring Batch records every run in the JobRepository with three concepts worth knowing:

  • JobInstance - a logical run identified by its job parameters (e.g. reportDate=2026-07-22). The same parameters = the same instance.
  • JobExecution - one attempt at a JobInstance. If Monday's run fails and you restart it, that's a second JobExecution of the same JobInstance.
  • StepExecution - one attempt at a Step, tracking read/write/skip counts and the commit position.

This is what makes restart work: the JobRepository knows a given JobInstance's last position, so a restart resumes instead of starting fresh (next lesson).

A book split into chapters and page bookmarks

The Job is the whole book you set out to read; each Step is a chapter you finish before starting the next. The chunk is a page you fully read before slipping in a bookmark - you never re-read pages behind the bookmark, and if you're interrupted, the bookmark shows exactly where to resume. The JobInstance is 'reading this edition, starting tonight'; a JobExecution is one evening's sitting; if you fall asleep and pick it up tomorrow, same book, second sitting, resumed from the bookmark.

Where does the transaction commit?

In a Step configured with chunk(50), your ItemWriter does a batched database insert. Over 500 total items, how many transactions commit, and how many times is the writer's write() method called? What does this buy you if the job crashes at item 220?

What does chunk(100) control in a chunk-oriented step?

Key takeaways

  • A Job is a restartable process made of ordered Steps; the chunk-oriented step is the workhorse.
  • A chunk-oriented step flows read → process → write, batched: read/process per item, write and commit per chunk.
  • The chunk size is the transaction boundary - it bounds memory, gives incremental durability, and enables batched writes.
  • Tune chunk size between throughput (larger) and memory/rollback cost (smaller); hundreds to low thousands is typical.
  • JobInstance (identified by parameters), JobExecution (one attempt), and StepExecution (per-step progress) are how the JobRepository tracks runs and enables restart.
War diese Lektion hilfreich?
Diese Seite auf GitHub bearbeiten