Focus Stacking: Tack-Sharp from Front to Back

 

 

12 August 2026Technology

 

Anyone who has photographed a wood grain, a blossom or a populated circuit board from close range knows the moment of disillusionment: everything looks fine on the display, but at 100 per cent magnification a narrow strip is sharp and the rest dissolves. This is not operator error, nor a question of lens price — it is physics. The closer the camera moves to the subject, the thinner the zone that can be rendered sharp at one time. Focus stacking is the answer. Behind this article stands Gigapixel GmbH — the specialized large-format stock portal: real gigapixel captures from 100 MP, no AI upscaling. In our production workflow, stacking is a daily tool, not a special case.

 

The problem: depth of field shrinks to millimetres

 

Depth of field is the range in front of and behind the plane of focus that the eye still accepts as sharp. In a landscape photograph this range extends from a few metres to infinity. Up close, the ratio changes dramatically. In his documentation on focus stacking, Kurt Wirz (2023) describes how depth of field at a reproduction ratio of 1:1 — where the subject is projected onto the sensor at life size — amounts to no more than a few millimetres. At higher magnifications it shrinks into the micrometre range.

 

A joinery shot shows what this means in practice. A planed oak board lies at an angle in the frame, the grain running away from the front edge towards the back. The height difference between the frontmost fibre and the rear edge may be three centimetres. For a close-up intended to show the pores of the wood, that is an eternity: focus on the middle and the front and back go soft. The same applies to a blossom whose stamens occupy different planes, or to a component with solder joints at varying heights.

 

The obvious reflex is to stop down. Smaller aperture, more depth of field. That reflex leads into a trap with two floors.

 

Why stopping down does not solve the problem

 

First floor: effective aperture

 

Up close, the aperture you set is not the aperture that actually takes effect. The optics follow a simple, well-known relationship: effective aperture = set aperture × (1 + reproduction ratio). At a ratio of 1:1, a set aperture of f/16 therefore behaves like f/32. Anyone who believes they are stopping down moderately in the macro range is in truth already working in extreme territory — with all the consequences described by the second floor.

 

Second floor: diffraction

 

Light passing through a small opening is diffracted. A point is no longer rendered as a point but as a disc. The smaller the aperture opening, the larger that disc — and the fewer details the sensor can still separate. Beyond the optimum aperture, a sobering rule applies: every further aperture stop halves the resolution that can be rendered. photoscala.de (thoMas, 2012) worked this out for full frame:

Aperture (full frame)Renderable resolutionRatio to f/5.6
f/5.6approx. 60 MP100 %
f/11approx. 15 MPapprox. 25 %
f/22approx. 3.7 MPapprox. 6 %

 

Going from f/5.6 to f/22 leaves roughly six per cent of the originally renderable resolution. At f/22, a 60-megapixel sensor delivers the detail volume of an old compact camera. You buy depth of field and pay in detail resolution — precisely the wrong currency for large-format printing.

 

Wirz (2023) measured this effect: conventionally stopped down to f/16 or f/22, the loss of resolution in the centre of the image exceeds 40 per cent. A series captured at a critical aperture between f/4 and f/8 and subsequently stacked loses only 0 to 4 per cent. That is the real reason for focus stacking: not convenience, but detail preservation. More on the physical background in our article on the diffraction limit in gigapixel photography.

 

How focus stacking works

 

Instead of a single frame with forced depth of field, a series is captured. The camera stays fixed, the subject stays fixed, and the plane of focus travels through the subject in defined steps — from the frontmost to the rearmost edge. Every individual frame is taken at the critical aperture, where the lens reaches its resolution maximum and diffraction has not yet set in.

 

The software then fuses the series. For each image area it assesses which frame shows the highest local edge sharpness, and assembles the final image from those sharp zones. The result is an image that is sharp throughout its entire depth without ever having been stopped down. The methodological foundations for capturing and viewing very large image datasets were described by Kopf et al. (2007); the principles for acquisition and multiresolution display developed there remain the basis of practically every gigapixel workflow today.

 

Two conditions are non-negotiable: the subject must be still, and the step width must match the depth of field. If the steps are too large, bands of unsharpness appear between the planes of focus — an error that cannot be repaired afterwards.

 

Production practice at Gigapixel GmbH

 

The following section describes our own production practice and rests on in-house knowledge of Gigapixel GmbH, not on an external source.

 

In our production, stacking has a second, less frequently discussed effect: it recovers the resolution that diffraction takes from the individual frame. Multiple exposures of the same spot sample the subject with slightly different pixel positions each time — because however still a subject stands, it never stands perfectly still for the sensor grid. Every frame lands minimally offset, and during fusion more genuine image information is therefore available than a single shutter release could deliver. Stacking thus brings real resolution back up to the sensor's nominal value — and, depending on the subject, beyond it.

 

The order of magnitude from our practice: because of diffraction, a 50-megapixel sensor resolves only around 30 megapixels of genuine detail per individual frame at a working aperture of f/8 — one stop below the optimum, the halving rule from the table above. A stack of around eight exposures of the same spot compensates for exactly that: the result again carries roughly 50 megapixels of genuine image information, and depending on the subject somewhat more. The camera industry actively uses the same underlying principle: Hasselblad's Multi-Shot process combines six exposures of a 50-megapixel medium-format sensor, offset by whole and half pixels, into a 200-megapixel file (Hasselblad, 2014) — there with deliberate sensor shifting, in our case via the natural micro-shift between exposures.

 

What matters here is the difference from generative methods: every pixel remains measured. It originates from light that actually fell onto the sensor. This is the opposite of what AI upscaling does. Kim et al. (2025) showed that upscaling methods collapse at extreme magnifications: they invent structures instead of measuring them. For a stock portal whose customers enlarge images to wall size, that is the heart of the matter — covered in more detail in our comparison of genuine captures with AI upscaling.

 

Our gigapixel production runs serially and automatically, with stacking down to a subject height of one centimetre.

 

Where end-to-end sharpness makes the difference

 

Craft and material. A joinery that wants to show its work is not selling the piece of furniture but the surface: pore pattern, plane marks, edge glue lines. These details rarely lie in a single plane.

 

Product and technical photography. A tool, a valve, an assembly should remain legible from the frontmost edge to the rear mount. Only then can the image be cropped later without unsharp zones becoming visible.

 

Macro. With blossoms, insect wings or mineral structures, stacking is quite simply the only way to show the subject sharp in its entirety.

 

Close viewing on walls and ceilings. Anyone viewing a subject from half a metre faces a hard benchmark: according to Ashraf, Chapiro & Mantiuk (2025), the eye resolves approximately 94 ppd — pixels per degree — and up to 120 ppd for individuals. Converted to a viewing distance of 0.5 m, this corresponds to 274 ppi. At one metre it is 137 ppi, at two metres 68 ppi. A subject viewed at arm's length must carry that detail density across its entire area — including its depth. Fundamentals on this in our article What is gigapixel photography?.

 

Industrial inspection. In automated optical inspection, stacking has long been standard. Basler AG (2025) describes a solution in which an FPGA fuses ten exposures of 5.1 megapixels each into one consistently sharp image in 67 milliseconds — fast enough for inline operation. In printed circuit board and BGA inspection, the method makes micro-voids and coplanarity deviations visible, that is, defects that would remain undetected in a single plane of focus.

 

Frequently asked questions

 

How many frames does a stack need?

There is no fixed number, because the quantity required depends directly on the depth of field of the individual frame and the depth extent of the subject. The calculation is simple in essence: divide the depth of the subject by the depth of field of a single frame and add a safety margin so that the sharp zones overlap. With a flat subject such as a slightly angled wooden board, eight to fifteen frames may suffice. In the macro range, where depth of field at 1:1 amounts to only a few millimetres according to Wirz (2023) and falls into the micrometre range at higher magnification, several dozen to more than a hundred frames are normal. What matters is not the total number but the step width: it must be smaller than the depth of field of the individual frame. Otherwise bands of unsharpness appear between the planes of focus, and no fusion algorithm can repair them, because the sharp information at that point was simply never captured. In our serial production we work with around ten exposures per spot — not as a maximum, but as an operating point that provides security against focus drift and local disturbances.

 

Does stacking really increase resolution, or only depth of field?

Both — for two reasons that should be kept apart. Depth of field grows directly through the method: an image sharp throughout its depth is assembled from several planes of focus. The resolution gain has two sources. The first is externally documented: because every individual frame is captured at a critical aperture between f/4 and f/8 rather than at f/16 or f/22, the diffraction penalty is avoided — Wirz (2023) measures more than 40 per cent resolution loss with conventional stopping down against 0 to 4 per cent with stacking. The second source is multiple sampling: even at the critical aperture, diffraction limits what a single frame resolves — a 50-megapixel sensor delivers only around 30 megapixels of genuine detail at f/8. Because the subject never stands perfectly still for the sensor grid, every further exposure samples it with minimally different pixel positions; a stack of around eight exposures thus recovers the full 50 megapixels or so, and depending on the subject more. Hasselblad's Multi-Shot uses the same principle with deliberate sub-pixel shifting: a 50-megapixel sensor, a 200-megapixel file (Hasselblad, 2014). The specific values of our production are in-house knowledge of Gigapixel GmbH — every pixel remains measured in the process.

 

Why not simply use f/22?

Because f/22 destroys the very detail resolution you are trying to gain. Beyond the optimum aperture, every further aperture stop halves the resolution that can be rendered. photoscala.de (thoMas, 2012) quantifies this for full frame: at f/5.6 around 60 megapixels are renderable, at f/11 roughly 15 megapixels, and at f/22 only about 3.7 megapixels. Around six per cent of the original detail volume therefore remains. You trade depth of field for resolution — at a very poor rate. Up close the problem is compounded, because the effective aperture differs from the one you set: the rule is: effective aperture = set aperture × (1 + reproduction ratio), so f/16 already behaves like f/32 at a ratio of 1:1. Anyone who believes they are stopping down moderately is long since in extreme territory. For an image later printed at wall size or examined down to its structure in a zoom viewer, this loss is unacceptable. Stacking at a critical aperture avoids the conflict entirely.

 

Does stacking work with moving subjects?

No. Focus stacking strictly requires that subject and camera remain unchanged between the individual frames. The fusion algorithm compares the same image areas across the entire series and selects the sharpest frame for each area. If the subject moves, the frames show different parts of the subject at the same image position — the algorithm then assembles fragments that do not belong together. This becomes visible as ghost edges, doubled contours or frayed transitions. Even slight movement is enough: a draught on a blossom, vibration in the floor, a swaying capture platform. All classic stacking applications are therefore static — craft pieces, products, specimens, mineral structures, circuit boards. The method is not suited to moving scenes, and there is no sensible way around this. In industrial inspection the problem is solved through speed rather than stillness: Basler AG (2025) describes the fusion of ten exposures in 67 milliseconds, so that the component stands practically still for the duration of the series. This shortcut does not apply to classic subject photography — there, a stable tripod, a vibration-free base and a wind-sheltered setup remain the prerequisite. Anyone working with plants outdoors waits for the calm minute or shields the subject.

 

Conclusion

Depth of field up close cannot be gained by stopping down without losing precisely what matters: detail resolution. Focus stacking resolves the conflict by capturing every frame at a critical aperture and then fusing the sharp zones. For us this is the foundation of every close-up — and the reason our subjects hold up even at half a metre. Gigapixel GmbH — the specialized large-format stock portal: real gigapixel captures from 100 MP, no AI upscaling.

 

Sources

  • Ashraf, M., Chapiro, A. & Mantiuk, R. K. (2025): Resolution limit of the eye — how many pixels can we see? Nature Communications.
  • Kim et al. (2025): Chain-of-Zoom.
  • photoscala.de / thoMas (2012): Wie viele Megapixel verkraftet eine Kamera?
  • Wirz, K. (2023): Focus Stacking. focus-stacking.ch.
  • Hasselblad (2014): H5D-200c MS — multi-shot with sub-pixel sensor shifting, 50 MP sensor to 200 MP file. Manufacturer documentation.
  • Basler AG (2025): Integrated Focus Stacking Solution for AOI Systems.
  • Kopf, J. et al. (2007): Capturing and Viewing Gigapixel Images. ACM ToG.