Start Learning
Javaneer
Back to stage
Module 10·Spring Batch

Why Batch Processing

When you need to process huge volumes on a schedule - reports, migrations, billing - and why a robust batch framework beats a naive loop.

13 min readIntermediate
On this page

Some work doesn't fit a web request. A user clicks a button and waits a second or two - that's the online world. But a nightly overdue-loan report for BookVault has to scan every active loan, find the ones past their due date, group them by member, and email a summary. That might touch millions of rows and run for minutes. No one is sitting there watching a spinner. This is batch processing: large volumes, no interactive user, run on a schedule.

What makes work "batch"

Three signatures tell you a job belongs in a batch, not an HTTP handler:

  • Volume - it processes a whole dataset (every loan, every invoice, every user), not one item.
  • No user waiting - it runs on a timer or is kicked off by an operator, and completes minutes or hours later.
  • Must be reliable - if it dies halfway through 2 million records, you cannot simply "click again"; it has to resume, skip bad rows, and never double-charge or double-email.

Reports, data migrations, monthly billing, ETL (extract-transform-load), reindexing a search engine - all batch.

The naive loop, and why it breaks

Your first instinct is a plain loop:

List<Loan> loans = loanRepository.findAllActive();   // load 2 million rows into memory 💥
for (Loan loan : loans) {
    if (loan.isOverdue()) {
        reportRow.add(summarize(loan));
    }
}
emailService.send(buildReport(reportRow));

This looks fine on 200 rows and collapses on 2 million. It has four fatal problems:

  1. Memory - findAllActive() pulls every row into a List, exhausting the heap.
  2. No transactions - one giant commit at the end (or none), so a crash loses everything or corrupts state.
  3. No restart - if it dies at row 1,900,000, the next run starts again from zero.
  4. No fault tolerance - one malformed loan throws, and the whole run aborts.

'It works on my test data'

A loop over a small list is the classic trap. Batch problems only show up at production volume - out-of-memory on the millionth row, or a single bad record aborting an eight-hour job at hour seven. Batch frameworks exist precisely because these failures are invisible until they cost you.

What a batch framework gives you

Spring Batch solves exactly these four problems so you don't reinvent them:

  • Chunk-oriented processing streams records in small transactional batches (say 100 at a time) instead of loading everything - bounded memory.
  • A JobRepository persists progress in the database, so a failed job restarts where it stopped.
  • Skip and retry policies let a bad record be logged and skipped, or a transient failure retried, without killing the run.
  • Scaling hooks - multi-threaded steps, partitioning - for when one thread is too slow.

You describe what to read, process, and write; the framework handles the transactions, the bookmarking, and the failure recovery.

A payroll clerk vs. a mailroom conveyor

The naive loop is one clerk trying to process the entire month's payroll in a single sitting - if they faint at employee 900, the whole stack scatters and they start over tomorrow. Spring Batch is a conveyor belt with a tally sheet: envelopes move past in trays of 100, each completed tray is checked off, and if the power cuts, the tally sheet shows exactly which tray to resume from. Bad envelope? It's set aside in a reject bin and the belt keeps moving.

Spot the batch job

Of these four features in BookVault, which belong in a batch job rather than a web request, and why? (a) Borrowing a book, (b) a nightly email of overdue loans, (c) importing a 500 MB CSV of a partner library's catalog, (d) showing a member their current loans.

Why does a plain for-loop over findAllActive() fail as a production batch job?

Key takeaways

  • Batch work is high-volume, runs without an interactive user, and must survive mid-run failures - reports, migrations, billing, ETL.
  • A naive loop over the whole dataset fails at scale: out of memory, all-or-nothing commits, no restart, no fault tolerance.
  • Spring Batch streams records in small transactional chunks for bounded memory.
  • A JobRepository persists progress so a failed job restarts where it stopped, and skip/retry policies handle bad records.
  • You describe what to read/process/write; the framework owns transactions, bookmarking, and recovery.
Was this lesson helpful?
Edit this page on GitHub