Neustartfähigkeit & Fehlertoleranz
Das JobRepository merkt sich den Fortschritt, sodass ein fehlgeschlagener Job dort neu startet, wo er stoppte; Skip- und Retry-Policies behandeln fehlerhafte Datensätze, ohne den Lauf abzubrechen.
Deutsche Übersetzung in Arbeit
Diese Lektion ist noch nicht ins Deutsche übersetzt und wird daher auf Englisch angezeigt. Der Rest der Seite ist vollständig lokalisiert.
Auf dieser Seite
The whole point of a batch framework is surviving failure. Two mechanisms deliver that: restartability - the JobRepository remembers how far a job got, so a failed run resumes instead of starting over - and fault tolerance - skip and retry policies let a single bad record be tolerated without aborting a multi-hour job.
Restart: the JobRepository remembers
Every chunk commit updates the JobRepository (a set of Spring Batch tables in your database) with the step's position. So if BookVault's import dies at record 1,900,000, the repository already recorded the 1,899,900 successfully committed. Re-launching the same JobInstance (same job parameters) resumes from where it stopped:
JobParameters params = new JobParametersBuilder()
.addString("reportDate", "2026-07-22") // identifies the JobInstance
.toJobParameters();
jobLauncher.run(overdueReportJob, params); // first run fails at chunk 19,001
// ... fix the problem, launch again with the SAME params ...
jobLauncher.run(overdueReportJob, params); // resumes at chunk 19,001, not chunk 1The rule that makes this work: job parameters identify the instance. Same parameters = same JobInstance =
resume. Different parameters (e.g. a new reportDate) = a brand-new instance that starts fresh. This is why
a completed instance can't be re-run with identical parameters - Spring Batch refuses, because it already
finished.
Design steps to be restartable
Restart only helps if your step can safely re-run its last incomplete chunk. That means a stable sort order in the reader (so 'resume from item 200' means the same rows every time) and idempotent writes where possible. A step that emails members is dangerous to blindly restart - you could double-send. Stage results to a table first, then email in a separate, carefully-designed step.
Skip: tolerate bad records
By default, one exception aborts the step. A skip policy says "these exceptions on a single item are survivable - log it and move on":
@Bean
Step importStep(JobRepository repo, PlatformTransactionManager tx,
ItemReader<RawBook> reader, ItemWriter<CatalogEntry> writer) {
return new StepBuilder("importStep", repo)
.<RawBook, CatalogEntry>chunk(100, tx)
.reader(reader).writer(writer)
.faultTolerant()
.skip(FlatFileParseException.class) // a malformed CSV line...
.skipLimit(50) // ...tolerate up to 50, then fail the job
.build();
}Now a malformed row is recorded as skipped and the run continues - up to skipLimit, after which Spring
Batch concludes the input is too broken and fails the job. Skips are counted in the StepExecution so you can
report "processed 1,000,000, skipped 37".
Retry: survive transient failures
Some failures aren't the record's fault - a deadlock, a brief network blip to a downstream service. Retry re-attempts the operation before giving up:
.faultTolerant()
.retry(TransientDataAccessException.class) // e.g. a lock timeout
.retryLimit(3) // try up to 3 times
.skip(TransientDataAccessException.class) // if still failing, skip it
.skipLimit(10)Retry and skip compose: retry a flaky operation a few times, and if it still fails, skip that item rather than abort. The distinction matters - retry is for transient errors that might succeed on a second try; skip is for permanent bad data that never will.
A cashier scanning a huge order hits two kinds of trouble. A smudged barcode on one item is a bad record - no amount of re-scanning fixes it, so they set the item aside (skip) and keep going, but if half the cart won't scan they call a manager (skip limit hit). A momentarily jammed scanner is transient - they wipe it and try the same item again (retry), and only set it aside if it still won't read after a few tries. Meanwhile the register tape (JobRepository) records every item rung up, so if the power flickers, they resume from the last recorded item rather than re-ringing the whole cart.
BookVault's import step calls an external ISBN-validation service per book. Two failures occur: (1) the service returns a 503 for a few seconds during a deploy, and (2) one book has a genuinely malformed ISBN the service rejects with 400. Which failure wants retry, which wants skip, and why not use the same policy for both?
What is the difference between a retry policy and a skip policy in Spring Batch?
Key takeaways
- The JobRepository persists each chunk's progress, so re-launching the same JobInstance resumes from the last commit rather than restarting.
- Job parameters identify the JobInstance: same parameters resume; different parameters start a fresh instance; a completed instance won't re-run with identical parameters.
- Restartable steps need a stable reader sort order and, ideally, idempotent writes; stage results before side-effecting steps like emailing.
- A skip policy tolerates permanently-bad records (log and continue) up to a skip limit, then fails the job.
- A retry policy re-attempts transient failures a few times; combine retry (transient) with skip (permanent) so each failure gets the right response.