Scheduling & Scaling
Trigger jobs on a schedule, and scale a heavy job with multi-threaded steps, partitioning, and remote chunking when one thread isn't enough.
On this page
A batch job needs two more things to be production-ready: something to launch it on a schedule (the "nightly" in "nightly report"), and, when one thread can't finish in the window, a way to scale the heavy step across threads or machines. Spring Batch stays deliberately unopinionated about scheduling and gives you several scaling strategies to reach for in order.
Launching on a schedule
Spring Batch runs a job through a JobLauncher; when to call it is left to you. The simplest trigger is
Spring's own @Scheduled:
@Component
class OverdueReportScheduler {
private final JobLauncher launcher;
private final Job overdueReportJob;
@Scheduled(cron = "0 0 2 * * *") // 02:00 every day
void runNightly() throws Exception {
JobParameters params = new JobParametersBuilder()
.addString("reportDate", LocalDate.now().toString()) // unique per run
.toJobParameters();
launcher.run(overdueReportJob, params);
}
}Note the unique parameter (reportDate): each night is a new JobInstance, so it runs fresh rather than
being rejected as "already completed". For anything beyond a single instance, teams use an external scheduler -
cron, Kubernetes CronJobs, Quartz, or an enterprise scheduler - to invoke the launcher, which gives you
retries, alerting, and a dashboard the framework itself doesn't provide.
Don't schedule from inside a scaled-out app
If your service runs as three replicas, a naive @Scheduled fires the job three times at 02:00. Guard against this with a leader-election lock (e.g. ShedLock), or - cleaner - run batch jobs from a dedicated launcher process or a Kubernetes CronJob that starts a single job pod. Batch and always-on web serving often belong in separate deployables.
Scaling, in order of reach
When a single-threaded step is too slow, escalate through these strategies - each more powerful and more complex than the last:
1. Multi-threaded step - process chunks on a pool of threads within one JVM. A one-line change, but your reader/writer must be thread-safe and you lose strict ordering:
return new StepBuilder("reportStep", repo)
.<Loan, ReportRow>chunk(100, tx)
.reader(reader).processor(processor).writer(writer)
.taskExecutor(new SimpleAsyncTaskExecutor()) // chunks run concurrently
.build();2. Partitioning - split the input into ranges (e.g. loan IDs 1-1M, 1M-2M, ...) and run a copy of the step per partition, each with its own reader/writer over its slice. Partitions can run on local threads or be distributed to remote workers. This is the go-to for large, cleanly-divisible datasets.
3. Remote chunking - a manager reads items and sends chunks over a message queue to worker JVMs that process and write. Useful when processing is the bottleneck and reading is cheap, but it adds messaging infrastructure.
Reach for scaling last
Most jobs finish comfortably single-threaded - correct chunk sizing and an indexed reader query solve more 'slow job' problems than parallelism does. Scaling adds thread-safety hazards, ordering loss, and (for partitioning/remote chunking) real infrastructure. Profile first; parallelize the step that's actually the bottleneck, not every step.
Scheduling is the alarm clock that says 'start at 2 a.m.' Single-threaded is one clerk working the whole stack - fine for most days. A multi-threaded step is several clerks at one office sharing the same in-tray (fast, but they can tread on each other, so the in-tray must be shareable). Partitioning is splitting the work by postcode and giving each branch office its own stack - clean division, near-linear speedup. Remote chunking is one office reading the mail and couriering batches to other offices to process - powerful when the processing, not the reading, is what's slow.
BookVault's overdue report now covers 40 million loans and misses its 6 a.m. deadline. The reader query is fast and indexed; the slow part is a per-loan fine calculation. Loans have sequential numeric IDs. Which scaling strategy fits best, and why not just multi-thread the single step?
Which scaling strategy runs a copy of the step over an isolated slice of the input, each with its own reader/writer?
Key takeaways
- Spring Batch launches jobs via a JobLauncher; scheduling is up to you - @Scheduled for simple cases, an external scheduler (cron, K8s CronJob, Quartz) for production.
- Give each scheduled run a unique job parameter so it's a new JobInstance rather than a rejected duplicate.
- Guard scheduled launches against multi-replica double-firing with leader election or a dedicated launcher process.
- Scale in order of reach: multi-threaded step (one JVM, thread-safe reader/writer), then partitioning (isolated slices), then remote chunking (distribute processing over a queue).
- Reach for parallelism last - correct chunk sizing and indexed queries fix most slow jobs; parallelize only the proven bottleneck.