Instanced Geometry, Explained

Many 3D scenes draw the same shape over and over: a hundred columns in a temple, a thousand atoms in a molecule, a field of identical bolts, trees or crates. Drawn naively, each copy costs a separate draw call, and the CPU spends more time telling the GPU to draw than the GPU spends drawing. Instancing is the OpenGL feature that fixes this: you hand the GPU the shape once and say “now draw it 1000 times, here is the little bit of data that differs for each copy.” This page explains how OpenGLContext applies instancing automatically, what it can and cannot batch, and the switches that control it. The indented technical notes point to the code that implements each piece.

Screenshot of the OpenGLContext forest demo with instanced grass and trees

The problem: one draw call per object is expensive

To draw an object, your program sends the GPU a stream of commands: bind this geometry, set this transform, set this colour, now draw. That handshake has a fixed cost paid per object, on the CPU, whether the object fills the screen or is a single pixel. Draw ten objects and it is nothing. Draw ten thousand and the CPU is still setting up object number 9,999 while the GPU sits idle, waiting for work. This is the draw-call bottleneck, and it is the wall large scenes hit first.

In OpenGLContext the per-object cost lives in Shape.Render and the surrounding loop (matrix upload, material upload, id assignment, the draw), on the order of tens of microseconds per object. A scene with tens of thousands of objects can spend a whole frame on setup before a single triangle is shaded. Instancing collapses that loop into one call.

The core idea: draw the shape once, vary a little data per copy

A normal draw call is glDrawElements: “draw this mesh once.” The instanced version is glDrawElementsInstanced: “draw this mesh N times.” One command, N copies. The GPU runs the vertex shader for every vertex of every copy, but your CPU issued a single call.

For that to be useful, each copy needs to end up somewhere different, with maybe a different colour. So alongside the shared mesh you give the GPU a second, small array — one entry per instance — holding the data that differs. OpenGL calls this a per-instance vertex attribute, and you mark it with glVertexAttribDivisor(location, 1). The divisor is the whole trick:

So the shape's vertices are shared by all copies, while the per-instance array supplies each copy's own transform, id and material.

OpenGLContext packs three things per instance, at fixed attribute locations that the shaders (pbr.vert, vrml97_lighting.vert, shadow_depth.vert) agree on: locations 5–8 a mat4 model-view matrix (a matrix occupies four consecutive vec4 locations), location 9 a uint object id for picking, location 10 a uint material index. The shaders multiply by the per-instance matrix instead of a uniform, and are gated by an instancingEnabled uniform so the ordinary non-instanced path is byte-for-byte unchanged. See passes/instancing.py (pack_instance_buffer, draw_instanced_mesh).

How shapes are grouped into one draw

Instancing only helps when copies genuinely share geometry, so the first job is to decide which shapes may ride in the same instanced draw. OpenGLContext does this every frame by computing an instance key for each shape and bucketing shapes with equal keys. A bucket with enough members becomes one instanced draw; everything else falls back to the ordinary per-shape path, unchanged.

Two shapes share a key — and therefore one draw — when they agree on:

Two ways instances are found

The grouping is fed from two directions that converge on the same batcher:

Content collapse: matching identical-but-separate meshes

Node identity (“is this literally the same object in memory?”) catches the explicit cases cheaply. To catch the opportunistic ones, a geometry can supply a content key — a short signature of its shape — so two distinct nodes with the same signature land in one bucket. A Sphere's signature is just ('Sphere', radius, tessellation); a general mesh hashes its vertex arrays.

The keys live in passes/instancing.py: geometry_content_key (PBR pass, keys on content + texture set + pass signature, lets material factors vary) and geometry_content_instance_key (VRML97 lit pass, which binds one material per group so it also splits on material identity). A geometry opts in by implementing instanceContentKey(); without one, a mesh's vertex arrays are hashed once and cached on the node (_geometry_content_id). Grouping itself is build_instance_groups; the minimum bucket size is OPENGLCONTEXT_INSTANCE_MIN (below it, per-shape drawing is cheaper than instancing's fixed setup). Opportunistic collapse can be turned off with OPENGLCONTEXT_INSTANCE_COLLAPSE=0, which falls back to node-identity grouping only.

Per-instance materials without per-instance draws

A field of “same sphere, different colour” should still be one draw. It is: the group's distinct materials are packed into a small array on the GPU, and each instance carries an index (that location-10 attribute) selecting its own entry. The shader reads materials[index]. So colour, metalness, roughness and the other numeric factors vary per instance for free.

The array is a std140 uniform block (MaterialBlock in pbr.frag), sized to the guaranteed 16 KB UBO minimum — 90 materials — though desktop drivers report much more (~370 at 64 KB). A group with more distinct materials than fit is split into a few instanced draws (“chunking”), still a huge win over per-shape. Only the PBR pass has the material array; the VRML97 lit pass binds one material per group, so it splits colour variants into separate groups instead.

Picking works, per instance

OpenGLContext lets you click objects to select them. It does this by drawing each object's id into a hidden buffer and reading back the id under the cursor. Instancing must not break that — you should be able to click one atom out of a thousand in a single instanced draw. It works because the id is just another per-instance value (location 9): every copy writes its own stable id, so the hidden buffer ends up with a thousand distinct ids exactly as if they had been drawn separately.

Ids come from _objectIdFor(path), stable per scene path across frames. A shape marked pickable=False is handled by masking the id attachment (glColorMaski) so it writes no id and picks read through it — the water-surface / gizmo case. Masking is per-draw state, so a mix of pickable and non-pickable instances is split into two draws, the non-pickable one drawn with the id attachment masked.

What instances — and what deliberately does not

A geometry becomes instanceable by exposing its mesh in the shared attribute layout (an instanceGPU(mode) method); the passes pick it up automatically. Today that covers:

Some geometry is deliberately left out, because instancing it would cost more than it saves:

The optimisations, and how they work

Reusing GPU objects instead of rebuilding them

The first time a group is drawn, OpenGLContext builds the GPU objects it needs — the VAO that wires up the mesh and per-instance arrays, and the per-instance buffer — then caches them on the mesh. Each subsequent frame re-uploads only the instance data into the existing buffer; the material array's buffer is reused the same way. There is no per-frame create/destroy of GPU objects.

See _build_instance_vao / draw_instanced_mesh (the VAO + instance vbo.VBO live on the mesh's _MeshGPU, reclaimed with it) and _bind_material_array (one persistent UBO on the pass, orphaned and re-uploaded per group). The instance matrices are eye-space (they fold in the camera), so the buffer is re-uploaded each frame even for a static scene; keeping it static would mean uploading model-space matrices and doing the view transform in the shader. tests/test_instanced_caching_gl.py counts GPU-object creations per frame and requires zero in steady state.

Not drawing what is off-screen

An object outside the camera's view (the frustum) should not be drawn at all. OpenGLContext already tests every object against the frustum before it forms instanced groups, so off-screen copies never reach the instanced draw — they cost no vertex-shader work. (If you point the camera so half a field is off-screen, the instanced draw contains only the visible half.)

The remaining cost is the test itself: checking a very large field one object at a time, every frame, on the CPU. Cluster culling makes that cheaper. It sorts the instances by spatial locality (a Morton, or Z-order, code from each instance's position), groups them into small contiguous clusters each with a combined bounding box, and tests the box. A cluster wholly outside the view is thrown away in one test instead of one-test-per-member. Clusters that straddle the edge fall back to the exact per-object test, so the result is identical, just reached with less work.

Implemented as pure, testable functions in passes/instancing.py (morton_order, build_clusters, cluster_cull) and wired into frustumVisibilityFilter. The cluster box is padded by the largest member's world radius so it can never reject a cluster that has a visible member — correctness is provably identical to the per-object filter (tests/test_instance_cluster_cull_gl.py checks the drawn set matches exactly). Because per-object culling already gives correctness, this is a pure cost optimisation and is opt-in: OPENGLCONTEXT_INSTANCE_CLUSTER_CULL=1, active only above 256 objects.

Switches that control instancing

Instancing is on by default and needs no code changes to benefit from — load a scene with repeated geometry and it batches. These environment variables let you tune or disable it, mostly for debugging and benchmarking:

VariableDefaultEffect
OPENGLCONTEXT_INSTANCING1 (on) Master switch. Set 0 to draw every shape per-object (compare correctness / measure the speed-up).
OPENGLCONTEXT_INSTANCE_COLLAPSE1 (on) Opportunistic content collapse of distinct-but-identical meshes. Set 0 to batch only explicitly-shared geometry (node identity).
OPENGLCONTEXT_INSTANCE_MIN4 Minimum copies before a bucket is worth an instanced draw; smaller buckets draw per-shape (instancing has a fixed per-batch setup cost).
OPENGLCONTEXT_INSTANCE_CLUSTER_CULLoff Turn on cluster culling to lower the per-frame CPU cost of frustum-testing very large static fields. Correctness is unchanged; it only saves work.

Limitations at a glance

Where the code lives

The engine is OpenGLContext/passes/instancing.py (grouping keys, build_instance_groups, capability detection, cluster culling, and the draw_instanced_mesh draw path). Each pass wires it in via instancing_enabled / _instanceable / _instanceKey / _drawInstanceGroup: the PBR pass in passes/pbrpass.py (with the per-instance material array), the VRML97 lit pass in passes/flatcore.py, and the shadow depth pass in passes/shadowmixin.py (so an instanced scene also casts shadows in one draw per light instead of re-drawing every caster). Geometry nodes opt in with instanceGPU(mode) + instanceContentKey() (see scenegraph/box.py, quadrics.py, pbrmesh.py, indexedfaceset.py). The design record, including the deferred hardware-accelerated paths (SSBO material arrays, multi-draw-indirect, bindless textures), is in plans/INSTANCED-GEOMETRY.md.