Every OpenCV engagement that involves a learned model eventually stalls in the same place. The architecture is fine, the deployment target is chosen, the inference budget is met — and accuracy is stuck at 91% when the contract says 98%. Nine times out of ten the fix is not a better backbone. It is the dataset: near-duplicate frames inflating the validation score, three annotators who disagreed about what "scratch" means, and 40,000 unlabelled images of which maybe 900 are worth a human minute.
This tutorial is about the part of a vision project nobody quotes for: curating and labelling data with OpenCV 5 as the workhorse, and closing an active-learning loop so each labelling round buys measurable accuracy instead of more of the same.
1. Deduplicate before you label, not after
Video-sourced datasets are overwhelmingly redundant. A 30 fps camera watching a conveyor gives you 30 almost-identical images per second, and if those near-duplicates end up split across train and validation your metrics become fiction — the model has effectively seen the validation set.
Two passes, cheap to expensive.
Pass one: perceptual hashing. OpenCV's img_hash module (mainline since 4.x, still there in 5) catches exact and near-exact repeats in milliseconds per image.
import cv2, glob, collections
hasher = cv2.img_hash.PHash_create()
buckets = collections.defaultdict(list)
for path in sorted(glob.glob("raw/**/*.jpg", recursive=True)):
img = cv2.imread(path)
h = hasher.compute(img).tobytes()
buckets[h].append(path)
exact_dupes = [p for group in buckets.values() for p in group[1:]]
print(f"{len(exact_dupes)} exact/near-exact duplicates")
For near duplicates rather than identical hashes, compare Hamming distance between pHash values; anything under about 6 bits out of 64 is the same scene for labelling purposes.
Pass two: embedding-space clustering. Perceptual hashes miss semantic redundancy — same part, same defect, slightly different lighting. Run a small self-supervised backbone (DINOv2-S or a MobileNet feature extractor) through the OpenCV 5 DNN engine, then cluster:
import numpy as np
net = cv2.dnn.readNet("dinov2_small.onnx")
net.setPreferableBackend(cv2.dnn.DNN_BACKEND_OPENCV)
def embed(path):
img = cv2.imread(path)
blob = cv2.dnn.blobFromImage(img, 1/255.0, (224, 224),
mean=(0.485, 0.456, 0.406), swapRB=True)
net.setInput(blob)
v = net.forward().ravel()
return v / (np.linalg.norm(v) + 1e-9)
paths = sorted(glob.glob("raw/**/*.jpg", recursive=True))
E = np.stack([embed(p) for p in paths]) # N x D, L2-normalised
S = E @ E.T # cosine similarity
Greedily keep one representative per cluster above a similarity threshold (0.95 is a sensible starting point, tuned by eyeballing 50 rejected pairs). On real conveyor and CCTV datasets this routinely removes 60–85% of frames with no loss of coverage — and it removes them before you pay for annotation.
Crucially, split into train/val/test by source group, not by frame: by camera, by shift, by production lot, by site. Random per-frame splits on video data are the single most common reason a model that scored 0.97 mAP offline collapses in the factory.
2. Write the labelling guide before the first label
Annotation disagreement is a spec problem wearing a data costume. Before anyone draws a box, write a one-page guide that answers, with reference images:
- Class boundaries. Is a 0.3 mm scratch a defect? At what length does "scuff" become "gouge"?
- Occlusion rule. Label objects visible below 30%? Pick a number.
- Truncation rule. Objects touching the frame edge: box the visible part, or skip?
- Crowds. Individual boxes or a single ignore-region?
- Ambiguity escape hatch. A
uncertainclass annotators can use instead of guessing. You will learn more from those 200 images than from anything else in the set.
Then measure agreement. Have two annotators label the same 200-image control set independently and compute per-class IoU-matched agreement:
def iou(a, b):
ax1, ay1, ax2, ay2 = a; bx1, by1, bx2, by2 = b
ix1, iy1 = max(ax1, bx1), max(ay1, by1)
ix2, iy2 = min(ax2, bx2), min(ay2, by2)
iw, ih = max(0, ix2 - ix1), max(0, iy2 - iy1)
inter = iw * ih
union = (ax2-ax1)*(ay2-ay1) + (bx2-bx1)*(by2-by1) - inter
return inter / union if union > 0 else 0.0
If two humans only agree at 0.78, 0.78 is your model's ceiling and no training run will beat it. Fixing the guide is cheaper than buying GPUs. Re-run the control set whenever a new annotator joins.
3. Pre-label with the model you already have
Once you have a weak first model (even one trained on 500 images, or an open-vocabulary detector — see open-vocabulary detection with YOLO-World and SAM 2), stop labelling from scratch. Generate proposals, let humans correct them. Correction is roughly 3–5× faster than drawing.
def pre_label(path, net, conf_thresh=0.35):
img = cv2.imread(path)
h, w = img.shape[:2]
blob = cv2.dnn.blobFromImage(img, 1/255.0, (640, 640), swapRB=True)
net.setInput(blob)
out = net.forward()[0] # e.g. N x (cx,cy,w,h,score,...)
boxes, scores, cls = [], [], []
for det in out:
score = float(det[4])
if score < conf_thresh:
continue
cx, cy, bw, bh = det[:4]
boxes.append([int((cx-bw/2)*w/640), int((cy-bh/2)*h/640),
int(bw*w/640), int(bh*h/640)])
scores.append(score); cls.append(int(det[5:].argmax()))
keep = cv2.dnn.NMSBoxes(boxes, scores, conf_thresh, 0.5)
return [(boxes[i], scores[i], cls[i]) for i in np.array(keep).ravel()]
Two guardrails. Set the proposal threshold low — a missing box costs an annotator more than a spurious one, because deleting is one click and drawing is twenty. And keep a provenance flag on every annotation (model_proposed vs human_drawn vs human_corrected), because automation bias is real: annotators accept plausible-looking wrong boxes. Audit a sample of untouched model_proposed labels every round.
4. The active-learning loop: choose what to label next
This is where the money is. Given 40,000 unlabelled frames and budget for 1,000 labels, picking well is worth several points of mAP over picking randomly. Combine three signals.
Uncertainty. For detection, low max-class confidence and boxes sitting near the NMS threshold both signal confusion. A serviceable scalar per image:
def uncertainty(dets):
if not dets:
return 0.5 # empty predictions: mildly interesting
# peak near 0.5 confidence = maximally uncertain
return float(np.mean([1.0 - abs(2*s - 1.0) for _, s, _ in dets]))
Disagreement. Run the image twice under different augmentations (or through the FP32 and the INT8 build — see INT8 quantisation and NPU inference) and measure prediction instability. Images where a horizontal flip changes the detection count are exactly the images the model has not learned.
Diversity. Pure uncertainty sampling collapses onto one hard corner of the data. Re-use the embeddings from §1: cluster the top-uncertainty pool and sample across clusters rather than taking the top-N.
def select_batch(paths, unc, E, k=1000, n_clusters=50):
pool = np.argsort(unc)[-4*k:] # 4k most uncertain
Ep = E[pool]
crit = (cv2.TERM_CRITERIA_EPS + cv2.TERM_CRITERIA_MAX_ITER, 50, 1e-4)
_, labels, _ = cv2.kmeans(Ep.astype(np.float32), n_clusters, None,
crit, 5, cv2.KMEANS_PP_CENTERS)
labels = labels.ravel()
chosen, per = [], max(1, k // n_clusters)
for c in range(n_clusters):
idx = pool[labels == c]
idx = idx[np.argsort(unc[idx])[-per:]]
chosen.extend(idx.tolist())
return [paths[i] for i in chosen[:k]]
Always reserve 10–20% of each batch for random sampling. Without it the loop never discovers failure modes the current model is confidently wrong about, and your validation set slowly stops representing production.
5. Freeze a golden test set and never touch it
Before the first loop iteration, carve out a test set — stratified across cameras, shifts, lighting conditions and rare classes — label it with your best annotator, double-check it, and lock it. It never gets extended by active learning (that would bias it toward hard cases) and it never gets used for threshold tuning.
Track, per iteration: labels added, cumulative label cost, mAP on the golden set, and per-class recall for the classes the customer actually cares about. A typical healthy curve rises steeply for two or three rounds, then flattens. The flattening is the signal to stop labelling and change something else — more classes, better imaging (see camera control, HDR and input-quality gates), or a different model. Continuing to buy labels past the knee is the most common way vision budgets get burned.
Wire the golden-set evaluation into CI alongside your pipeline regression tests (golden frames, model-swap gates and drift alarms) so that every dataset version and every model swap produces a comparable number.
6. Version the data like code
Datasets rot quietly. Minimum viable discipline:
- Content-addressed images. Filename = SHA-256 of the bytes. No more "final_v2_FIXED.jpg".
- Annotations in git (COCO JSON or similar), images in object storage, with a manifest mapping dataset version → image hashes → annotation commit.
- Every trained model records its dataset version. When accuracy regresses you need to know whether the data or the code moved.
- Keep the raw frames — not just the curated subset. Class definitions change; re-mining is only possible if you kept the originals.
- Log label provenance and timestamps. When a customer challenges a result, "which human labelled this class, under which guide revision" is a question you want to be able to answer.
7. A realistic first-30-days plan
For a typical industrial inspection engagement:
- Days 1–3. Ingest raw video, dedupe (§1), report how much genuinely distinct data exists. This alone frequently changes the project plan — "you have 400 distinct examples of the defect, not 40,000".
- Days 4–6. Write the labelling guide, run the two-annotator control set, fix the guide, record the agreement ceiling.
- Days 7–12. Label a diverse 800–1,200 image seed set plus the frozen golden test set.
- Days 13–18. Train a baseline, export to ONNX, measure on the golden set, stand up the pre-labelling service.
- Days 19–30. Two active-learning rounds of ~500 labels each, plotting mAP against label count.
At the end of that you do not just have a model — you have a defensible curve that tells the customer what another 2,000 labels are worth. That is the number that actually settles budget arguments.
Where this fits
Model architecture is largely a solved, commoditised choice. Data strategy is not, and it is where the accuracy gap in most stalled computer vision projects actually lives. If your detector has plateaued, the honest first question is not "which backbone next?" but "how much distinct, consistently-labelled data do we really have, and what is the next 1,000 labels worth?"
SentientSight's OpenCV consultants build dataset curation and active-learning pipelines — dedup, annotation specs and QA, pre-labelling services, and golden-set evaluation wired into CI — for teams who need an accuracy number they can defend. If your project is stuck on the data rather than the model, get in touch.