Start Learning
Javaneer
Back to stage
Module 11·GraphQL with Spring

The N+1 Problem & Batching

GraphQL's per-field resolution makes N+1 queries dangerously easy; @BatchMapping and DataLoader coalesce them back into one efficient fetch.

15 min readAdvanced
On this page

Per-field resolution, the source of GraphQL's flexibility, is also the source of its most notorious performance trap: the N+1 problem. A single innocent-looking client query can fan out into hundreds of database queries on the server. This lesson shows exactly how it happens and the two Spring tools - @BatchMapping and DataLoader - that collapse the fan-out back into one efficient fetch.

How one query becomes N+1

Recall that a nested field resolver runs once per parent object. Now a client asks for a list of books with each author:

query { books { title author { name } } }

The engine resolves it like this:

books()              → 1 query: SELECT * FROM books           (returns 100 books)
author(book #1)      → 1 query: SELECT * FROM authors WHERE id = ?
author(book #2)      → 1 query: SELECT * FROM authors WHERE id = ?
...
author(book #100)    → 1 query: SELECT * FROM authors WHERE id = ?

That's 1 + 100 = 101 queries for one client request - the "N+1". With nesting (books → authors → their other books) it compounds into thousands. The client wrote one clean query; the server quietly melted.

It's invisible in development

With 5 books in your test database, N+1 is 6 queries - fast, unnoticed. In production with 500 books and a loaded database, it's 501 queries and a timeout. Like most batch and data problems, N+1 hides until real volume exposes it. Watch your SQL logs (or use Hibernate's statistics) when a resolver returns a list.

Fix 1: @BatchMapping

Spring's simplest fix: instead of a resolver that fetches one author for one book, write a @BatchMapping that fetches all authors for all books in the chunk, in one call. Spring collects the parent books, invokes your method once with the whole list, and you return a map from book to author:

    @BatchMapping                                   // resolves Book.author for ALL books at once
    Map<Book, Author> author(List<Book> books) {
        Set<Long> authorIds = books.stream()
            .map(Book::authorId).collect(toSet());
        Map<Long, Author> byId = authorService.findAllByIds(authorIds)   // ONE query
            .stream().collect(toMap(Author::id, a -> a));
        return books.stream()
            .collect(toMap(b -> b, b -> byId.get(b.authorId())));
    }

Now the query is 2 queries total: one for the books, one for all their authors (WHERE id IN (...)). The method name author still maps to Book.author - you've only changed how it's fetched, from per-item to batched. This is the go-to fix in Spring for GraphQL.

Fix 2: DataLoader

DataLoader is the more general mechanism (from the GraphQL spec) that @BatchMapping builds on. It defers each field request, collects all the keys requested during one tick of query execution, then calls a single batch-load function - and it caches within the request so the same author fetched twice is loaded once. You register a batch loader and reference it in a resolver:

    @SchemaMapping
    CompletableFuture<Author> author(Book book, DataLoader<Long, Author> loader) {
        return loader.load(book.authorId());   // deferred; batched behind the scenes
    }

Use @BatchMapping for the common case (it's terser); reach for an explicit DataLoader when you need its per-request caching or want to share one loader across several fields.

A runner making one trip to the stockroom

The N+1 resolver is a shop assistant who, handed a list of 100 orders, walks to the stockroom and back once per order - 100 exhausting round-trips. @BatchMapping and DataLoader turn them into a smart runner who collects all 100 slips, walks to the stockroom once, grabs everything in a single armful (SELECT ... WHERE id IN (...)), and hands each order its item. Same result, one trip instead of a hundred - and DataLoader even remembers items already fetched so it never grabs the same one twice.

Diagnose and fix

A BookVault GraphQL query { members { name currentLoans { book { title } } } } returns 200 members and is timing out. The SQL log shows 400+ queries. Explain the query count, then describe how you'd fix it with batching.

Why does a GraphQL query for a list of books with their authors cause an N+1 problem?

Key takeaways

  • Per-field resolution runs a nested resolver once per parent, so a list query becomes 1 + N database queries - the N+1 problem.
  • N+1 is invisible on tiny test data and causes timeouts at production volume; watch SQL logs when a resolver returns a list.
  • @BatchMapping replaces a per-item resolver with one that receives the whole list of parents and returns a Map, fetching all children in one query.
  • DataLoader is the general spec mechanism @BatchMapping builds on: it defers and batches keys per execution tick and caches within the request.
  • Any nested resolver that returns data for a list of parents is a latent N+1 - batch it with @BatchMapping or a DataLoader.
Was this lesson helpful?
Edit this page on GitHub