Overpass API Query Language Jump to heading
A query that works on a city and times out on a country is not a scaling problem — it is a query that was always doing the wrong work and only got away with it because the data was small. Overpass QL looks like a filter expression language, and that resemblance is what leads people to write node["amenity"] with a continent-sized bounding box and then blame the server. The language is a set algebra over a spatial database, and once you read it that way the cost of every clause becomes visible before you press run.
The Problem This Topic Solves Jump to heading
You need a bounded set of OpenStreetMap features — every pharmacy in a metropolitan area, every cycleway crossing an administrative boundary, every building without an address in one district — and you need it now, without downloading and parsing a regional extract. Overpass is built exactly for that, and used within its design envelope it answers in seconds what a local pipeline would need an hour of setup to reproduce.
The failure scenario when this goes wrong is specific and recognisable. A team prototypes against a small bounding box, the query returns in two seconds, and the job is scheduled nightly against a whole country. The first run takes ninety seconds. The second week it starts returning HTTP 504. By the third week the client’s address is rate limited and unrelated interactive queries from the same office stop working. Nothing about the query changed; the only thing that changed was the number of elements the filters had to touch. This page is about seeing that number before it bites.
Prerequisites Jump to heading
Before writing queries in anger, be comfortable with the node, way and relation data model, because Overpass returns raw elements and a way without its nodes is not geometry. Know the tag conventions in Tag Taxonomy & Key-Value Standards, since every filter you write is an assertion about key-value practice. And read the parent Querying OSM: Overpass, Nominatim & APIs overview for the quota context this page assumes throughout.
Sets, Not Pipelines Jump to heading
Every Overpass statement produces a set of elements and writes it somewhere. By default it writes into the unnamed set _, and by default the next statement reads from _. That is the whole model, and two consequences follow immediately.
The first is that the default set is overwritten, not accumulated. Two consecutive statements do not mean “and also”; the second replaces the first. This is the single most common cause of “my query returns nothing” — a spatial filter that reads the default set after a previous statement already emptied it.
The second is that you can name sets and combine them. Writing ->.a at the end of a statement binds its result to a; writing .a before a filter reads from a; and a union block ( .a; .b; ) produces their combination. Once you name sets deliberately, complex queries become readable and, more importantly, debuggable — you can out an intermediate set to see exactly where the element count collapsed.
[out:json][timeout:90];
// bind the search area once, reuse it twice
area["name"="Kraków"]["admin_level"="8"]->.city;
(
node["amenity"="pharmacy"](area.city)->.pharm;
way["amenity"="pharmacy"](area.city);
);
out center tags;
Two details in that snippet carry real weight. The area statement resolves an administrative relation into an area object once, and both element queries reference it, rather than each re-resolving the boundary. And out center tags asks for a representative point plus tags rather than full geometry, which for a point-of-interest layer is all you need and a fraction of the payload.
Filters, Ordered by What They Cost Jump to heading
Overpass evaluates a query against indexes, and the practical cost model is simple: a spatial filter is cheap and selective; a tag filter is cheap per element but selective only in proportion to how rare the tag is; a regular-expression tag filter cannot use the value index at all. Ordering follows directly.
The rules that follow from that ordering are short:
- Put the geography first, always.
node["amenity"="cafe"](50.0,19.8,50.1,20.1)and the same filters in the other order describe the same set, but a query with no spatial bound at all asks the server to consider every matching element on Earth before anything narrows it. - Prefer an exact value to a key existence test.
["highway"="residential"]touches a small slice;["highway"]touches every road, path, crossing and kerb in the area. - Prefer a union of exact values to a regular expression.
["amenity"~"^(cafe|bar|pub)$"]is a regex;(node["amenity"="cafe"](area.a); node["amenity"="bar"](area.a); node["amenity"="pub"](area.a););is three indexed lookups and is usually faster despite being longer. - Never write a regex on a key unless you genuinely cannot enumerate the keys —
[~"^addr:"~"."]is occasionally the only way to ask “does this have any address tag”, and it is correspondingly expensive. - Filter by element type. If you want buildings, ask for
wayandrelation; asking fornwradds every node in the area to the candidate set for no benefit.
Spatial Filters in Detail Jump to heading
Three spatial forms cover essentially all real use, and they differ in more than syntax.
A bounding box (south,west,north,east) is the cheapest and the most predictable. It is also the only one whose cost you can estimate by eye, which makes it the right choice for scheduled jobs where a surprise is worse than a slightly larger result.
An area filter (area.name) resolves an administrative or other closed relation into a polygon and tests membership properly. It is the right answer when the question is genuinely “inside this city”, but it carries two traps: the area has to exist as a mapped relation, and area ids are derived from the relation id by an offset, so a hand-written id is a common source of empty results. Resolve areas by name and admin level, bind the result, and check the bound set is non-empty before relying on it.
An around filter (around:500,50.06,19.94) selects elements within a radius of a point, or — in its set form (around.set:100) — within a radius of every element of another set. The set form is extraordinarily useful (every bus stop within a hundred metres of a rail station) and extraordinarily expensive, because it evaluates a buffer per element in the driving set. Use it with a driving set you have already made small.
Recursion: Completing Partial Results Jump to heading
A query for way["building"] returns ways whose geometry is a list of node references and nothing else. To get usable geometry you either ask the server to inline it with out geom, or you complete the set yourself with a recursion operator.
>(“recurse down”) adds the nodes of every way in the set, and the members — and their nodes — of every relation.<(“recurse up”) adds the ways and relations that reference the elements in the set. This is how you answer “which routes include this stop”.>>and<<are the transitive forms, following nested relation membership all the way. They are the ones that occasionally return a surprising fraction of a country when a route relation turns out to belong to a superroute that spans it.
Output Modes and What They Cost You Jump to heading
The out statement takes modifiers that change the payload by an order of magnitude, and choosing among them is the cheapest optimisation available.
| Modifier | Returns | Typical use | Cost note |
|---|---|---|---|
out ids; |
Element ids only | Diffing against a cached set | Smallest possible response |
out tags; |
Ids and tags, no geometry | Tag analysis where location is irrelevant | Small |
out body; |
Tags plus node refs for ways | You will resolve geometry yourself | Needs a recursion to be usable |
out center; |
One representative coordinate per element | Point-of-interest layers, markers | Usually the right default |
out geom; |
Inlined coordinates on every element | Rendering shapes directly | Largest; duplicates shared nodes |
out meta; |
Adds version, timestamp, changeset, user | Provenance and audit trails | Adds roughly a third to the payload |
out count; |
Just the number of elements | Estimating a query before running it | Effectively free — use it first |
out count; deserves a special mention: running your query with out count before running it for real tells you exactly how big the answer is, in one cheap request. Making that a habit removes most accidental continent-scale downloads.
Validation and Error-Handling Matrix Jump to heading
| Condition | Root cause | Detection | Remediation |
|---|---|---|---|
| Empty result, no error | A later statement overwrote the default set | out count on the intermediate set is zero |
Bind sets with ->.name and read them explicitly |
| Empty result from an area | Area relation not found, or an id offset used by hand | The area set itself is empty |
Resolve by name and admin level; assert non-empty |
| HTTP 400 with a parse error | Missing semicolon, or a filter after out |
Server returns the offending line | Keep one statement per line; end every statement |
| HTTP 504 gateway timeout | Query exceeded its declared or server-capped timeout | Response arrives at the timeout boundary | Narrow spatially first; only then raise [timeout:] |
| HTTP 429 | Too many concurrent slots from your address | Rate-limit body naming the slot state | Back off, serialise requests, cache successes |
| “runtime error: Query run out of memory” | Candidate set exceeded [maxsize:] |
Explicit runtime error in the body | Reduce the candidate set; do not just raise maxsize |
| Ways with no coordinates | out body without a recursion |
Consumers see node refs, not points | Add >; before out, or switch to out geom |
| Result larger than expected | nwr used where one type was meant |
Element type mix in the response | Name the element types explicitly |
Performance and Scale Jump to heading
Three levers matter, in this order.
Narrow before you filter. Every measurement of Overpass query cost comes back to the size of the candidate set that the expensive filters have to examine. A spatial bound applied first turns a regex filter from a planet scan into a city scan, and the regex then costs nothing worth discussing.
Ask once, not many times. Per-request overhead dominates small queries, and the server’s slot accounting punishes concurrency far more than it punishes size. One query returning ten thousand elements is cheaper for everybody than fifty queries returning two hundred each. When a loop over inputs seems unavoidable, look for a set-based reformulation — an around.set filter, or a union of bounding boxes — that expresses it as one request.
Stop before the ceiling. Both [timeout:n] and [maxsize:n] are ceilings, not allocations; raising them does not make a query faster, it only lets a bad query run longer before failing. If a query needs an unusual ceiling, that is information about the query, and the right response is usually to narrow it — or, past a certain volume, to stop querying and read a file. Handling Overpass Timeouts and Rate Limits builds the client side of that discipline, and Running a Local Overpass Instance for Bulk Queries covers the point at which hosting your own is the honest answer.
Failure Modes and Gotchas Jump to heading
- The default set is a footgun. Reading
_after a statement you forgot writes to it is the cause of most mysterious empty results. Bind everything you will reuse. areaids are derived, not raw. An area is not a relation; its id is offset from the relation id. Resolve areas by tags, never by an id you computed yourself.aroundon a large driving set is quadratic in feel. It evaluates a buffer per driving element. Shrink the driving set first, and check its count.- Regex filters skip the value index.
~is not a convenience with a small cost; it changes which index can be used. Enumerate values when you can. out geomduplicates shared nodes. On a road network where a junction node belongs to five ways, its coordinate appears five times. That duplication is often most of the response size.- Timestamps need
out meta. If you need version or changeset information, ask for it explicitly; the default output has none, and discovering that after building a pipeline is expensive. nwris rarely what you want. It is shorthand for all three element types and quietly triples the candidate set.
Integration Points Jump to heading
The output of a query is JSON or XML that still has to become usable geometry and a typed table. Converting Overpass JSON to a GeoDataFrame covers that conversion, including the center and geom shapes and the relation cases that need care. Downstream, the tag values in a response are raw OSM values and need the same treatment as anything parsed from a file — the normalization stage in Parsing & Tag Normalization Workflows applies unchanged, and running validation before trusting the result is as sensible here as for a local extract.
Upstream, the decision to use Overpass at all belongs to Choosing Between Overpass and a Local Extract, and for scheduled work the answer is often a local file.
Guides in This Topic Jump to heading
- Writing Overpass QL Area and Bounding Box Queries — resolving areas reliably and choosing between area and bounding box spatial filters.
- Handling Overpass Timeouts and Rate Limits — a client with backoff, caching and a size guard that survives a shared server’s throttle.
- Converting Overpass JSON to a GeoDataFrame — turning each output mode into typed rows with valid geometry.
- Running a Local Overpass Instance for Bulk Queries — sizing, importing and keeping a private instance current.
Frequently Asked Questions Jump to heading
Why does my Overpass query return nothing when the data clearly exists?
Almost always because a statement overwrote the default set before the filter you care about read it. Each statement writes into the unnamed set unless you bind it with an arrow, and the next statement reads that same set, so two consecutive queries replace rather than accumulate. Bind every set you intend to reuse to a name, read it explicitly, and use out count on the intermediate sets to find the exact statement where the element count fell to zero.
Should I raise the timeout when a query times out?
Only after narrowing it. The timeout setting is a ceiling, not an allocation, so raising it lets a badly shaped query run longer before failing rather than making it finish. Apply the spatial filter first so the expensive tag filters examine a city rather than a continent, replace regular expressions with unions of exact values, and run out count to see how large the answer really is. If the query still needs an unusual ceiling after that, it is a signal to read a local extract instead.
What is the difference between out center and out geom?
out center returns one representative coordinate for each way and relation, which is all a marker layer or a point-of-interest table needs and is by far the smallest payload. out geom inlines the full coordinate list onto every element, which lets you render real shapes without any client-side assembly but repeats every shared node once per way that references it. On a dense street network that duplication is usually most of the response size.
How do I select everything inside a named city?
Resolve the boundary to an area object by its tags rather than by an id — match on the name and the administrative level, bind the result to a named set, and reference that set from each element query with an area filter. Assert that the bound area set is not empty before using it, because an unmatched name produces an empty area and therefore an empty final result with no error to explain it.
Is a regular expression filter ever the right choice?
Sometimes, but rarely as a first choice. A regular expression on a value cannot use the value index, so the server tests candidates one at a time; a regular expression on a key is worse still. When the set of values is enumerable, a union of exact key-value queries is longer to write and usually faster to run. Keep regular expressions for genuinely open-ended questions, such as whether an element carries any key in a namespace, and always apply a spatial filter first.
Related Jump to heading
- Querying OSM: Overpass, Nominatim & APIs — the parent section with the quota model these queries live inside.
- Choosing Between Overpass and a Local Extract — when to stop querying and read a file instead.
- Node, Way & Relation Data Model — the element model every Overpass response is expressed in.
- Tag Taxonomy & Key-Value Standards — the key-value practice your filters assert.
- Parsing & Tag Normalization Workflows — the normalization every fetched tag still needs.
- Spatial Indexing for OSM Extracts — the local equivalent of the spatial narrowing Overpass does server-side.
Up one level: Querying OSM: Overpass, Nominatim & APIs.