Reader, Processor & Writer
Die drei austauschbaren Teile eines Chunk-Steps - ItemReader, ItemProcessor, ItemWriter - und die Bausteine für Dateien, Datenbanken und Queues.
Deutsche Übersetzung in Arbeit
Diese Lektion ist noch nicht ins Deutsche übersetzt und wird daher auf Englisch angezeigt. Der Rest der Seite ist vollständig lokalisiert.
Auf dieser Seite
A chunk-oriented step has exactly three pluggable parts: an ItemReader that produces items one at a time, an optional ItemProcessor that transforms or filters each item, and an ItemWriter that persists a whole chunk at once. Spring Batch ships production-grade implementations for files, databases, and queues, so most of the time you configure rather than write them.
ItemReader: one item at a time
An ItemReader<T> returns the next item on each read() call, or null when the source is exhausted - that
null is how the step knows to stop:
public interface ItemReader<T> {
T read(); // next item, or null when done
}You rarely implement this yourself. For BookVault's overdue report, a JdbcCursorItemReader (or the paging
variant) streams active loans straight from the database without loading them all:
@Bean
JdbcPagingItemReader<Loan> loanReader(DataSource ds) {
return new JdbcPagingItemReaderBuilder<Loan>()
.name("loanReader")
.dataSource(ds)
.selectClause("SELECT id, book_id, member_id, due_date")
.fromClause("FROM loans")
.whereClause("WHERE returned_at IS NULL")
.sortKeys(Map.of("id", Order.ASCENDING)) // stable order = safe restart
.pageSize(100)
.rowMapper(new LoanRowMapper())
.build();
}Built-ins cover most sources: FlatFileItemReader (CSV/fixed-width), JdbcPagingItemReader /
JdbcCursorItemReader, JpaPagingItemReader, StaxEventItemReader (XML), and Kafka readers.
Paging vs. cursor readers
A cursor reader holds one long-lived connection and streams rows - simple, but the connection is pinned for the whole step. A paging reader issues repeated 'give me the next 100 ordered rows' queries, so it needs a stable sort key or rows can be skipped or duplicated across pages. For restartable, multi-threaded jobs, paging with a sort key is usually the safer choice.
ItemProcessor: transform or filter
The ItemProcessor<I, O> takes one input item and returns the output to be written - or null to filter it
out of this run entirely:
public class OverdueProcessor implements ItemProcessor<Loan, ReportRow> {
public ReportRow process(Loan loan) {
if (!loan.isOverdue()) {
return null; // not overdue → drop it, never written
}
long daysLate = DAYS.between(loan.dueDate(), LocalDate.now());
return new ReportRow(loan.memberId(), loan.bookId(), daysLate);
}
}Returning null is the idiomatic way to skip items that don't belong in the output - no exception, no special
case. The input type (Loan) and output type (ReportRow) can differ, which is where transformation happens.
The processor is optional; a step that only moves data can omit it.
ItemWriter: a chunk at a time
The ItemWriter<T> receives the whole chunk as a list and writes it in one go - the batched write that
makes the chunk model efficient:
public interface ItemWriter<T> {
void write(Chunk<? extends T> items); // the whole chunk, e.g. 100 rows
}Because it gets a list, a JdbcBatchItemWriter issues a single batched INSERT for all 100 rows rather than
100 separate statements. Built-ins mirror the readers: FlatFileItemWriter, JdbcBatchItemWriter,
JpaItemWriter, and Kafka writers. You can also compose several writers with a CompositeItemWriter.
The reader is the picker pulling one part off the shelf at a time. The processor is the inspector who reshapes each part and tosses defects in the bin (returns null). The writer is the packer at the end who waits until a full tray of 100 has accumulated, then boxes and ships the whole tray in one motion. Three specialists, each doing one job, and only the packer works in batches.
BookVault's import step reads raw catalog rows. You must drop any row missing an ISBN, and convert the rest
from RawBook to CatalogEntry. Sketch the ItemProcessor. What does returning null accomplish, and what
happens to a row you return normally?
What does an ItemProcessor returning null do?
Key takeaways
- A chunk step has three parts: ItemReader (one item per read()), ItemProcessor (transform/filter each), ItemWriter (a whole chunk at once).
- The reader returns null to signal end-of-input; use built-ins like JdbcPagingItemReader or FlatFileItemReader instead of writing your own.
- Paging readers need a stable sort key for safe restart and multi-threading; cursor readers pin one connection.
- The processor's input and output types can differ (transform); returning null filters the item out of the run.
- The writer receives the chunk as a List, enabling one batched insert per chunk instead of per-item writes.