Conventional bright-field autofocus is not up to the task. It is necessary to rely on specialised methods, such as oblique focusing systems, the use of structured light to assist with depth estimation, or the use of confocal/Z-stack imaging to reconstruct three-dimensional topography; in such cases, ‘autofocus’ has evolved into software-based reconstruction following a full-height scan.


I. Why Are Deep Grooves Considered ‘Optical Traps’?
Imagine a laser beam as the light from a torch, shining straight down into a narrow, deep well. If the mouth of the well is wide, the light can easily reach the bottom; but if the mouth is extremely narrow and the walls are steep, the beam will be blocked by the side walls as soon as it enters, leaving the bottom forever hidden in shadow. This is the first challenge of deep-trench focusing: geometric obstruction. In three-dimensional microelectronic packaging or MEMS devices, the width of a trench is often only a few micrometres, whilst the aspect ratio easily exceeds 10:1; the vertically incident detection light is ‘swallowed up’ by the sidewalls before it even reaches the bottom of the trench.
What is even more problematic is that any light that manages to enter the trench by chance will bounce back and forth between the sidewalls, generating countless chaotic secondary and tertiary reflection signals. Autofocus systems attempt to identify the ‘sharpest’ point amongst these erratic reflections, but the result is akin to attempting to locate one’s position by echoes in a room full of mirrors—range measurements become chaotic, and the focusing algorithm simply breaks down. Consequently, conventional laser triangulation or image contrast-based focusing not only fails to locate the bottom of the groove but frequently locks onto stray light spots halfway down, yielding completely erroneous depth readings.


Since the vertical optical path is not feasible, the first practical solution is an oblique focusing system. By tilting the detection optical path at an angle, the beam can enter obliquely from one side of the groove opening, travelling directly to the bottom or the side wall to be measured. This ingeniously bypasses the vertical obstruction; the focused spot can return along the same path, enabling the system to accurately read the relative height of that point. This is akin to probing the bottom of a well with a long, inclined rod, rather than looking straight down from the well’s mouth.
This is complemented by structured light to assist with depth determination. By projecting known stripe or grid patterns onto the grooved area, the structured light will bend, break or undergo phase shifts due to variations in depth. A high-speed camera captures these distortions, enabling the system to calculate a rough three-dimensional ‘depth map’ in advance. Armed with this map, the autofocus system can rapidly move the lens to a position near the approximate focal plane before fine-tuning the focus. At this stage, focusing is no longer a haphazard process but is guided by precise navigation.


As the aspect ratio of the grooves continues to increase, and complex structures such as inverted cones and concave sidewalls begin to appear, it becomes difficult to reliably capture a single focal plane using oblique light paths or structured light alone. Consequently, a more fundamental approach has emerged: to abandon the traditional concept of ‘autofocus’ altogether and instead employ confocal or Z-axis layer-by-layer imaging to scan the entire height dimension layer by layer.
Confocal microscopy uses a pinhole to receive only the clear signal at the focal point, naturally rejecting out-of-focus and stray light. During imaging, the objective lens moves at high speed along the Z-axis, capturing a single, paper-thin optical section at each step. Every point on the bottom, top and side walls of the groove will exhibit an extremely bright signal at its corresponding Z-height, whilst the remaining heights appear almost entirely black. Once hundreds or even thousands of image layers have been acquired, software algorithms extract the maximum brightness of each pixel and its corresponding Z-height, enabling the synthesis of a ‘sharp throughout’ all-in-focus image, whilst simultaneously generating a precise three-dimensional point cloud. This is precisely what the conclusion refers to—autofocus has evolved into software-based synthesis following full-height scanning. You no longer need to focus the camera on a specific layer; instead, you first ‘swallow’ all the information from the entire space, and then reconstruct everything from the data.
Similarly, layer-by-layer imaging techniques such as white-light interferometry and optical coherence tomography rely on vertical scanning combined with interference fringes to reconstruct the three-dimensional topography inside the grooves. At this stage, the concept of ‘focusing’ has evolved into ‘topography reconstruction’, revealing every minute detail of the steep walls and the entire cross-section of the deep grooves.
You may also be interested in the following information
Let’s help you to find the right solution for your project!