Start Learning
Javaneer
Back to stage
Module 10·Spring Batch

Readers, Processors & Writers

The three pluggable pieces of a chunk step - ItemReader, ItemProcessor, ItemWriter - and the built-ins for files, databases, and queues.

15 min readAdvanced
On this page

A chunk-oriented step has exactly three pluggable parts: an ItemReader that produces items one at a time, an optional ItemProcessor that transforms or filters each item, and an ItemWriter that persists a whole chunk at once. Spring Batch ships production-grade implementations for files, databases, and queues, so most of the time you configure rather than write them.

ItemReader: one item at a time

An ItemReader<T> returns the next item on each read() call, or null when the source is exhausted - that null is how the step knows to stop:

public interface ItemReader<T> {
    T read();   // next item, or null when done
}

You rarely implement this yourself. For BookVault's overdue report, a JdbcCursorItemReader (or the paging variant) streams active loans straight from the database without loading them all:

@Bean
JdbcPagingItemReader<Loan> loanReader(DataSource ds) {
    return new JdbcPagingItemReaderBuilder<Loan>()
        .name("loanReader")
        .dataSource(ds)
        .selectClause("SELECT id, book_id, member_id, due_date")
        .fromClause("FROM loans")
        .whereClause("WHERE returned_at IS NULL")
        .sortKeys(Map.of("id", Order.ASCENDING))   // stable order = safe restart
        .pageSize(100)
        .rowMapper(new LoanRowMapper())
        .build();
}

Built-ins cover most sources: FlatFileItemReader (CSV/fixed-width), JdbcPagingItemReader / JdbcCursorItemReader, JpaPagingItemReader, StaxEventItemReader (XML), and Kafka readers.

Paging vs. cursor readers

A cursor reader holds one long-lived connection and streams rows - simple, but the connection is pinned for the whole step. A paging reader issues repeated 'give me the next 100 ordered rows' queries, so it needs a stable sort key or rows can be skipped or duplicated across pages. For restartable, multi-threaded jobs, paging with a sort key is usually the safer choice.

ItemProcessor: transform or filter

The ItemProcessor<I, O> takes one input item and returns the output to be written - or null to filter it out of this run entirely:

public class OverdueProcessor implements ItemProcessor<Loan, ReportRow> {
    public ReportRow process(Loan loan) {
        if (!loan.isOverdue()) {
            return null;              // not overdue → drop it, never written
        }
        long daysLate = DAYS.between(loan.dueDate(), LocalDate.now());
        return new ReportRow(loan.memberId(), loan.bookId(), daysLate);
    }
}

Returning null is the idiomatic way to skip items that don't belong in the output - no exception, no special case. The input type (Loan) and output type (ReportRow) can differ, which is where transformation happens. The processor is optional; a step that only moves data can omit it.

ItemWriter: a chunk at a time

The ItemWriter<T> receives the whole chunk as a list and writes it in one go - the batched write that makes the chunk model efficient:

public interface ItemWriter<T> {
    void write(Chunk<? extends T> items);   // the whole chunk, e.g. 100 rows
}

Because it gets a list, a JdbcBatchItemWriter issues a single batched INSERT for all 100 rows rather than 100 separate statements. Built-ins mirror the readers: FlatFileItemWriter, JdbcBatchItemWriter, JpaItemWriter, and Kafka writers. You can also compose several writers with a CompositeItemWriter.

An assembly line: picker, inspector, packer

The reader is the picker pulling one part off the shelf at a time. The processor is the inspector who reshapes each part and tosses defects in the bin (returns null). The writer is the packer at the end who waits until a full tray of 100 has accumulated, then boxes and ships the whole tray in one motion. Three specialists, each doing one job, and only the packer works in batches.

Filter and transform in one processor

BookVault's import step reads raw catalog rows. You must drop any row missing an ISBN, and convert the rest from RawBook to CatalogEntry. Sketch the ItemProcessor. What does returning null accomplish, and what happens to a row you return normally?

What does an ItemProcessor returning null do?

Key takeaways

  • A chunk step has three parts: ItemReader (one item per read()), ItemProcessor (transform/filter each), ItemWriter (a whole chunk at once).
  • The reader returns null to signal end-of-input; use built-ins like JdbcPagingItemReader or FlatFileItemReader instead of writing your own.
  • Paging readers need a stable sort key for safe restart and multi-threading; cursor readers pin one connection.
  • The processor's input and output types can differ (transform); returning null filters the item out of the run.
  • The writer receives the chunk as a List, enabling one batched insert per chunk instead of per-item writes.
Was this lesson helpful?
Edit this page on GitHub