+1 (415) 360-7596

Segmentation that measures: semantic and instance masks in OpenCV 5, cleaned up and converted to square millimetres

Every few months a client asks for "a detector" and what they actually need is a measurement of area. How much of this weld seam is porous. What percentage of the pallet is covered. How many square millimetres of coating are missing. How wide is this crack at its widest point. A bounding box cannot answer any of those questions — it answers "roughly here, roughly this big". The answer lives in a mask.

Segmentation is the part of a vision pipeline most teams skip straight past, because detection demos better and labels cheaper. This tutorial covers the other path: running semantic and instance segmentation models through OpenCV 5, cleaning masks with the morphology and contour tools that have been in the library for twenty years, and converting pixels into square millimetres with an error figure attached.

1. Pick the right kind of segmentation

Three things get called "segmentation" and they have different cost profiles:

  • Semantic segmentation — every pixel gets a class, instances are not separated. Right for coverage, corrosion, vegetation, road surface, "how much of the frame is X". Cheapest to label if you are painting regions.
  • Instance segmentation — per-object masks with identity. Right for counting, per-part measurement, pick-point selection. YOLO-family -seg variants and Mask R-CNN descendants live here.
  • Promptable / zero-shot segmentation — SAM 2 and friends, segment anything given a point or box prompt. Right for labelling assistance and for the "we have no training data yet" phase. Rarely the right thing to deploy on a line at frame rate.

A pattern that works well in consulting engagements: cheap box detector to localise, promptable segmenter offline to generate masks, train a small purpose-built -seg model on those masks, deploy that. You get instance-quality masks without a human painting ten thousand polygons. Our notes on open-vocabulary detection with SAM 2 and on dataset curation cover that loop in more detail.

2. Running a segmentation model through the OpenCV 5 DNN engine

OpenCV 5's rebuilt DNN engine reads ONNX, and a segmentation model is just a model with awkward outputs. Export at a fixed input size — dynamic shapes are the single most common reason an ONNX graph that works in Python fails to load in the C++ runtime.

import cv2
import numpy as np

net = cv2.dnn.readNetFromONNX("seg_model_640.onnx")
net.setPreferableBackend(cv2.dnn.DNN_BACKEND_OPENCV)
net.setPreferableTarget(cv2.dnn.DNN_TARGET_CPU)

INPUT = 640

def preprocess(frame):
    h, w = frame.shape[:2]
    scale = min(INPUT / w, INPUT / h)
    nw, nh = int(round(w * scale)), int(round(h * scale))
    resized = cv2.resize(frame, (nw, nh), interpolation=cv2.INTER_LINEAR)
    canvas = np.full((INPUT, INPUT, 3), 114, dtype=np.uint8)
    dx, dy = (INPUT - nw) // 2, (INPUT - nh) // 2
    canvas[dy:dy + nh, dx:dx + nw] = resized
    blob = cv2.dnn.blobFromImage(canvas, 1 / 255.0, (INPUT, INPUT), swapRB=True, crop=False)
    return blob, scale, dx, dy

Keep scale, dx, dy around. Every mask you produce has to travel back through that letterbox transform to land on the original pixels, and getting this wrong produces masks that are subtly offset — easy to miss in a demo, fatal in a measurement.

Semantic output

A semantic model emits (1, C, H, W) logits. The whole job is an argmax and a resize:

out = net.forward()                      # (1, C, h, w)
logits = out[0]
class_map = np.argmax(logits, axis=0).astype(np.uint8)   # (h, w)

# Confidence of the winning class, for a reject rule later
probs = np.exp(logits - logits.max(axis=0, keepdims=True))
probs /= probs.sum(axis=0, keepdims=True)
conf_map = probs.max(axis=0)

Resize the class map with INTER_NEAREST — never with bilinear, which invents class IDs that do not exist between 3 and 5. If you want smooth boundaries, upsample the per-class probability maps with INTER_LINEAR and argmax afterwards. That costs C times the memory and is worth it when boundary precision drives the measurement.

mask = cv2.resize(class_map, (INPUT, INPUT), interpolation=cv2.INTER_NEAREST)
mask = mask[dy:dy + nh, dx:dx + nw]                       # undo letterbox padding
mask = cv2.resize(mask, (w, h), interpolation=cv2.INTER_NEAREST)

Instance output (prototype masks)

YOLO-style -seg heads output two tensors: detections with an extra block of mask coefficients, and a prototype tensor of shape (1, 32, 160, 160). The instance mask is a linear combination of prototypes, sigmoided, then cropped to the box:

dets, protos = net.forward(net.getUnconnectedOutLayersNames())
protos = protos[0].reshape(32, -1)                        # (32, 160*160)

def instance_mask(coeffs, box, proto_hw=(160, 160), thresh=0.5):
    m = 1.0 / (1.0 + np.exp(-(coeffs @ protos)))          # sigmoid
    m = m.reshape(proto_hw)
    m = cv2.resize(m, (INPUT, INPUT), interpolation=cv2.INTER_LINEAR)
    crop = np.zeros_like(m, dtype=np.uint8)
    x1, y1, x2, y2 = [int(v) for v in box]
    crop[y1:y2, x1:x2] = 1                                # mask is only valid inside the box
    return ((m > thresh) & (crop > 0)).astype(np.uint8)

Two details people lose a day to. First, the box crop is mandatory: prototype masks bleed well outside the detection, and without the crop you get ghost blobs on the other side of the image. Second, thresh=0.5 is a tunable, not a constant — it directly moves the boundary and therefore the measured area. Pick it on validation data against ground-truth area, not by eye.

3. Clean the mask with classical tools

Raw network masks are speckled and have ragged edges. The fix is not a bigger model, it is thirty lines of morphology — and this is where OpenCV earns its place over a pure-PyTorch pipeline.

k3 = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (3, 3))
k7 = cv2.getStructuringElement(cv2.MORPH_ELLIPSE, (7, 7))

clean = cv2.morphologyEx(mask, cv2.MORPH_OPEN, k3)    # drop isolated speckle
clean = cv2.morphologyEx(clean, cv2.MORPH_CLOSE, k7)  # fill pinholes inside regions

Then drop fragments by area with connected components, which is far faster than contour-filtering every blob:

MIN_AREA_PX = 120
n, labels, stats, centroids = cv2.connectedComponentsWithStats(clean, connectivity=8)
keep = np.zeros_like(clean)
for i in range(1, n):
    if stats[i, cv2.CC_STAT_AREA] >= MIN_AREA_PX:
        keep[labels == i] = 1

Set MIN_AREA_PX from physics, not taste: if the smallest defect that matters is 0.5 mm across and your resolution is 0.08 mm/px, that is about 6 px diameter, roughly 30 px of area — so anything below ~30 px is noise by definition, and anything above it that you discard is a missed detection you will have to explain.

For boundary quality, a guided or joint-bilateral refinement against the original image snaps the mask edge to the real intensity edge, which matters when the measurement is a width or an area of something small:

refined = cv2.ximgproc.guidedFilter(guide=cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY),
                                    src=(keep * 255).astype(np.uint8),
                                    radius=8, eps=1e-2)
refined = (refined > 128).astype(np.uint8)

4. Turn pixels into millimetres

This is the step that makes the output a measurement rather than a picture. You need a pixel-to-world scale, and there are only three honest ways to get one:

  1. A calibrated camera at a known, fixed working distance on a flat plane — scale is Z / f metres per pixel. See our camera calibration walkthrough.
  2. A homography to the object plane, computed from four known points or a fiducial in the scene. Warp the mask with cv2.warpPerspective before measuring, because area is not preserved under perspective — a mask in the far half of an oblique frame covers far more real-world area per pixel than one at the near edge.
  3. A depth source (stereo, ToF) giving per-pixel scale, which is the only option for non-planar surfaces.
MM_PER_PX = 0.0824                    # from calibration, on the rectified plane
area_mm2 = int(keep.sum()) * MM_PER_PX ** 2

Shape metrics come from contours, and cv2.minAreaRect is usually what you want for "how long and how wide", not the axis-aligned bounding box:

contours, _ = cv2.findContours(keep, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
for c in contours:
    area_px = cv2.contourArea(c)
    perim = cv2.arcLength(c, True)
    (cx, cy), (rw, rh), angle = cv2.minAreaRect(c)
    length_mm = max(rw, rh) * MM_PER_PX
    width_mm = min(rw, rh) * MM_PER_PX
    circularity = 4 * np.pi * area_px / (perim ** 2 + 1e-9)
    print(f"{area_px * MM_PER_PX ** 2:.2f} mm^2  {length_mm:.2f} x {width_mm:.2f} mm  circ={circularity:.2f}")

Note cv2.contourArea and mask.sum() disagree slightly — the contour is a polygon through pixel centres, the sum counts whole pixels. For anything under a few hundred pixels the difference is several percent. Pick one convention, write it down, and use it in both your validation and your production code.

5. The error budget nobody asks for until the audit

Before quoting an accuracy figure, note where area error comes from. Area is a two-dimensional quantity, so boundary error hurts twice:

  • Boundary uncertainty. A blob of diameter d pixels with ±1 px of boundary uncertainty carries roughly ±4/d relative area error. At d = 20 px that is ±20%. At d = 200 px it is ±2%. If small features must be measured to ±5%, you need more resolution — not a better model.
  • Scale error. A 1% error in MM_PER_PX becomes 2% in area, because it is squared.
  • Threshold sensitivity. Sweep your mask threshold from 0.4 to 0.6 and record how measured area moves. If it moves 15%, your number is a threshold choice dressed up as a measurement.
  • Non-planarity. A surface tilted 10° out of the calibration plane changes effective scale by ~1.5%.

Report these together: "coverage area, ±7% for features above 3 mm, at 0.08 mm/px, mask threshold 0.5." That sentence is what separates a deliverable from a demo.

6. Validate against area, not IoU

Mean IoU is the metric the literature reports and it is the wrong acceptance criterion for a measurement system. Two masks can share an IoU of 0.85 and differ in area by 20% in opposite directions. Validate the quantity the customer will actually read:

  • Build a held-out set with ground-truth areas — machined coupons, printed targets of known size, or carefully painted reference masks.
  • Plot predicted vs. true area and fit a line. The slope exposes systematic bias (usually a threshold or scale problem, and it is fixable with a single calibration constant). The scatter is your real repeatability.
  • Add a Bland–Altman style check: is the error constant, or does it grow with size? Both are fixable, but they are fixed differently.
  • Measure the same static part 30 times. Any spread at all is run-to-run noise from lighting and sensor, and it bounds every claim you can make.

Then freeze the whole thing into a regression test — golden frames in, expected areas within tolerance out — as described in our pipeline regression testing post. A model swap that improves mIoU by 2 points and shifts measured area by 8% is a failed release, and only an area-based gate will catch it.

7. Performance notes

Segmentation costs more than detection, mostly in post-processing rather than inference. Three things that reliably buy back frame time:

  • Do not upsample masks to full resolution if you do not need to. Measure in model resolution and scale the area by (1 / scale)². Resizing a 4K mask per frame is often more expensive than the forward pass.
  • Keep masks as uint8, never float32 or bool arrays, in the hot path. Morphology on uint8 hits SIMD paths; on other dtypes it falls back.
  • Batch the prototype multiply. Stack all instance coefficients into one (N, 32) matrix and do a single @ protos, instead of one matmul per detection.

For INT8 export and NPU targets, the usual caveat applies more strongly to segmentation than detection: quantisation moves boundaries, and boundaries are the measurement. Re-validate area accuracy after quantising, not just mIoU — see our INT8 quantisation notes.

Where this fits

If the deliverable is a number with a unit on it — square millimetres of defect, percentage coverage, crack width, fill level — segmentation plus calibration is the pipeline, and the hard engineering is in the mask cleanup, the scale chain and the validation protocol rather than the model choice.

SentientSight's OpenCV consultants build and audit measurement-grade vision systems: target design, mask post-processing, scale calibration, error budgets and regression gates. If you need an area or dimensional figure that survives a customer audit, get in touch.