Incremental Updates for Derived Datasets Jump to heading
Applying an .osc file to a local OSM database is the easy part, and for most pipelines it is not the part that costs anything. The expense sits downstream: the vector tiles cut from that database, the search index built over its place names, the analytics table aggregating its features by administrative area. Each of those took hours to build, each of them is now wrong in a handful of places, and rebuilding all of them every minute is not an option anybody can afford.
This topic is about the translation step between those two facts — turning a change file into the smallest set of downstream work that restores correctness. The general shape is always the same: read the diff, determine which derived artifacts the changed features participate in, invalidate or recompute exactly those, and record what you did well enough to prove the derived dataset corresponds to a specific upstream state. What varies, and varies enormously, is how hard that second step is for a given derived dataset.
It assumes the sequence model from Replication Sequence Numbers & State Tracking and the apply mechanics from Applying .osc Change Files with Osmium. Everything below starts from a successfully applied diff and asks what happens next.
The Problem This Topic Solves Jump to heading
A minutely diff touching four hundred features is a trivial database update and a potentially enormous downstream one. Four hundred features scattered across a continent can invalidate tiles at every zoom level they appear in, which at zoom fourteen alone might be a few thousand tiles and across the whole pyramid considerably more. One of those features might be an administrative boundary whose edit changes which country a hundred thousand other features are counted in. Another might be a deleted node that three ways referenced, whose removal changes the geometry of each of them.
The naive responses both fail. Rebuilding everything on each diff is correct and impossibly expensive. Updating only the features named in the diff is cheap and wrong, because derived datasets contain relationships the diff does not mention — the tile that contains a feature, the relation whose geometry depends on a member, the aggregate whose total includes a row. The work of this topic is finding the middle: a derivation of affected downstream units that is neither the whole dataset nor merely the changed features.
Prerequisites Jump to heading
Deriving the Affected Set Jump to heading
Every derived dataset needs a function from changed features to affected units, and writing that function honestly is the substance of the work.
For tiles, the unit is a tile coordinate and the function is geometric: the bounding box of the changed feature’s old and new geometry, expanded to cover every tile it intersects at every zoom the layer is rendered at. Both geometries matter — a feature that moved invalidates where it was as well as where it is — which means the pipeline needs the previous geometry, and a diff that only carries the new version does not supply it. That is the single most common source of stale tiles: the old location is never invalidated because nothing remembered it.
For a search index, the unit is a document and the function is usually identity-based, which makes it the easy case. A changed place gets reindexed; a deleted one gets removed. The complication is that search documents typically denormalise context — a street’s document may carry its city and country names — so an edit to a city name invalidates every document that embedded it, and identity alone will not find those.
For analytics aggregates, the unit is a group and the function is whatever the grouping key is. If features are counted by administrative area, a changed feature invalidates its area’s total, and a changed boundary invalidates the totals of every area whose extent moved plus the neighbours it borrowed from. Aggregates are where incremental update most often goes subtly wrong, because a wrong total looks exactly like a right one.
The general rule is that the affected set must be derived from both the old and the new state of each changed feature. Anything computed from the new state alone will leave the consequences of the old state in place, and those consequences are invisible: a tile showing a building that has been demolished, a count including a feature that moved away.
Handling Deletions and Identity Churn Jump to heading
Deletions are the case incremental pipelines handle worst, for a structural reason: a deleted feature carries almost no information. An .osc <delete> block gives you a type, an identifier and a version, and frequently nothing else. If the derived dataset needs the feature’s geometry to compute the affected set — as tiles do — that geometry has to come from somewhere, and the only somewhere is the local database as it stood before the diff was applied.
This forces an ordering constraint that is easy to get backwards. Compute the affected set before applying the diff, not after. Once the delete has landed, the geometry is gone and the tiles covering it can no longer be identified. Pipelines that apply first and derive afterwards work fine for creates and modifications and silently fail for deletions, which is a failure mode that takes months to notice because deletions are relatively rare and the resulting staleness is localised.
Identity churn is the related problem covered in OSM Feature Identity & ID Stability. A feature can be deleted and recreated with a new identifier while representing the same real-world thing, or a way can be split into two ways where one keeps the original identifier and one does not. A derived dataset keyed on the OSM identifier sees a delete and a create, which is the correct behaviour for an index and the wrong behaviour for anything tracking a thing over time.
Recording What the Derived State Corresponds To Jump to heading
A derived dataset that has been incrementally updated needs to record which upstream sequence it reflects, for the same reason the primary database does: without it, nobody can tell whether it is current, and after any failure nobody can tell where to resume.
The subtlety is that the derived dataset’s sequence is not the primary database’s sequence. They diverge whenever an incremental update fails, is skipped, or is still running, and treating them as one number is how a tile cache ends up being reported as current while serving data from an hour ago. Each derived dataset owns its own marker, advanced only when its update for that sequence has fully succeeded.
Recording it also makes the rebuild decision expressible. If the tile cache is at sequence 6,231,004 and the database is at 6,231,890, the gap is 886 diffs — and at some gap size, catching up incrementally costs more than starting over.
Deciding When to Rebuild Instead Jump to heading
Incremental update is an optimisation, and like every optimisation it has a region where it stops paying. Three signals say you have left it.
The gap is large. Catching up on a thousand diffs sequentially takes a thousand round trips and a thousand affected-set computations, most of which touch overlapping units. A rebuild processes the current state once. The crossover point is dataset-specific and worth measuring rather than guessing, but it exists, and it is usually lower than people expect.
The derivation logic changed. If the tile schema, the index mapping or the aggregate definition has been edited, incremental update propagates new changes through new logic while leaving old data computed under the old logic. The result is a dataset that is internally inconsistent in a way no diff will ever repair.
Correctness is in doubt. After any incident where the affected set may have been computed wrongly — a crash mid-update, a bug fixed in the derivation — the cheap remedy is a rebuild, because the alternative is proving a negative about data you cannot enumerate.
A useful discipline is to rebuild on a schedule regardless: a weekly or monthly full rebuild bounds how long any undetected incremental error can persist, and it keeps the rebuild path exercised so it works when you need it urgently. A rebuild path that has not been run in six months is not a fallback.
Validation and Error Handling Jump to heading
| Check | What it catches | Response |
|---|---|---|
| Derived sequence versus primary sequence | A derived dataset silently falling behind | Alert on gap; rebuild beyond a threshold |
| Affected set size distribution | A derivation returning nothing, or everything | Investigate before applying the update |
| Old-geometry availability on delete | Deriving after applying rather than before | Reorder the loop; rebuild the affected area |
| Sample re-derivation against a rebuild | Incremental and full results diverging | Rebuild, then fix the derivation logic |
| Update duration trend | Incremental cost approaching rebuild cost | Reassess the crossover point |
| Per-dataset failure isolation | One derived dataset failing the whole loop | Advance the others; retry the failed one |
Performance and Scale Jump to heading
The cost of an incremental update decomposes into deriving the affected set and doing the work on it, and which dominates varies by dataset. For tiles, the derivation is cheap — a bounding-box-to-tile computation — and the work is expensive, because each invalidated tile must be re-cut from the database. For aggregates, the derivation may involve a spatial join that costs more than recomputing the aggregate would.
Two practical levers matter more than micro-optimisation. The first is batching across diffs: processing ten minutely diffs together and deduplicating the affected set is dramatically cheaper than processing them one at a time, because the same tiles and the same aggregates are invalidated repeatedly by consecutive edits in the same area. The cost is freshness, and an hour of freshness is a reasonable trade for most derived datasets even when the primary database is minutely.
The second is deferring the expensive half. An invalidation can be recorded without being serviced: mark the tile dirty and let it be re-cut on the next request, or on a background worker draining the dirty list by priority. This turns a synchronous cost into an asynchronous one and lets popularity decide what is actually worth recomputing, which for tile pyramids is a very large saving because most tiles are never requested.
Failure Modes and Gotchas Jump to heading
Deriving after applying. Covered above, and worth repeating because it survives code review: it works perfectly for the ninety-odd percent of changes that are creates and modifications.
Assuming the diff is complete. A change file contains the features that changed, not the features whose derived representation changed. Relations whose members moved, and features whose containing boundary shifted, are affected without appearing.
One sequence marker for everything. Derived datasets fail independently and therefore must be tracked independently, or the first failure makes every subsequent freshness claim untrue.
Dirty lists that grow unboundedly. Deferring work is only a saving if the deferred work is eventually done or explicitly expired. A dirty tile list that grows faster than it drains is a rebuild that has not admitted it yet.
No rebuild path. Every incremental pipeline eventually needs to start over, and the ones that cannot are the ones that stay subtly wrong for years.
Integration Points Jump to heading
Incremental derived updates sit between replication and every consumer. Upstream they depend on the sequence discipline in Replication Sequence Numbers & State Tracking and the apply step in Applying .osc Change Files with Osmium. Downstream they feed OSM Vector Tiles & Rendering Pipelines, where the dirty-tile list is the interface, and the analytics shapes in Modelling OSM for Analytics Warehouses.
They also interact with quality gating: an incremental update that lands a bad batch is harder to identify than a bad rebuild, because only part of the dataset changed. The thresholds discussed in Continuous QA for OSM Pipelines apply to the update, not just the whole dataset.
Guides in This Topic Jump to heading
- Propagating OSM Diffs Into a GeoParquet Lake — applying changes to immutable files, where update means rewriting partitions.
- Updating a Search Index from OSM Diffs — the identity-keyed case, and what denormalised context does to it.
- Computing a Dirty Tile List from an .osc File — the geometric derivation, including the old-geometry problem.
Frequently Asked Questions Jump to heading
Should derived datasets update synchronously with the diff apply?
Usually not. Coupling them means the slowest derived dataset sets the pace of the whole replication loop, and a failure in a tile cutter stops the primary database from advancing. The better shape is that the apply loop records what changed and advances its own state, and each derived dataset consumes that record at its own rate with its own marker. The cost is that freshness varies between datasets, which is true anyway and better made explicit.
How do you verify an incrementally updated dataset is actually correct?
By rebuilding a sample and comparing. Pick a bounded area — a city, a tile range, an aggregate group — rebuild it from the current database, and diff it against the incrementally maintained version. Doing this on a schedule catches derivation bugs that no amount of logging will, because the failure signature of a wrong affected set is data that looks entirely plausible.
What about changes that affect nothing downstream?
Most changes are like this, and filtering them early is the largest available saving. An edit to a tag your tile schema never reads, or to a feature type your index does not contain, produces an empty affected set and should cost nothing. The filter belongs at the start of the derivation, and it is worth measuring what proportion of diffs it eliminates, because on a narrow derived dataset that proportion is often above ninety percent.
Can an incremental update be rolled back?
Only if the derived dataset supports it, which most do not. A tile cache has no history; a search index has no previous version of a document once it is replaced. The practical answer is that rollback for derived data means rebuilding from a primary database that has itself been rolled back, which is why the primary database’s recoverability matters more than the derived dataset’s. Keeping the derived marker separate at least tells you which upstream state to rebuild from.
How often should a full rebuild run even when nothing is wrong?
Often enough that an undetected incremental error cannot persist beyond the interval, and often enough that the rebuild path stays working. For most pipelines that lands between weekly and monthly. The second reason is the one people forget: a rebuild procedure that has not run since the schema changed is not a fallback, and discovering that during an incident is expensive.
Related Jump to heading
- OSM Replication & Diff Sync — the parent section.
- Replication Sequence Numbers & State Tracking — where the markers come from.
- Applying .osc Change Files with Osmium — the apply step this follows.
- Replication Monitoring & Lag Alerting — the signals that show a derived dataset falling behind.
- OSM Vector Tiles & Rendering Pipelines — the largest consumer of incremental invalidation.
- OSM Feature Identity & ID Stability — why delete-and-recreate breaks identity-keyed derivations.
Up one level: OSM Replication & Diff Sync.