Skip to content
MH
← all work
Four-Corner Keypoints Static-camera bin localization

AP 1.000 — published with its own caveat attached

Detect the four corners of a bin in fixed CCTV imagery and export to ONNX. Its output becomes the crop stage of the fill-level pipeline — the two projects are one delivery chain.

2026 · for client deployments

MMPoseHRNet-w32MSRA heatmapsAlbumentationsONNX Runtime

1.000

Validation AP

epoch 70 @ 1024x576

0.9804

Held-out AP

COCO AP, verified

460

Training set

samples, 79% synthetic

4

Keypoints

graded at strict sigma 0.25

Pipeline

  1. CVAT
    polygons
  2. polygon →
    COCO keypoints
  3. HRNet-w32
    heatmap head
  4. argmax +
    sub-pixel
  5. ONNX
    export
  6. fill-level
    pipeline

The bug that made the model constant

Four legacy configs enabled horizontal flip augmentation. Every keypoint swap field was empty, so the framework produced an identity mapping: the image mirrored and the labels did not. The model converged on predicting fixed coordinates for everything. On the deployed ONNX, the first keypoint x-coordinate varied only between 658 and 694 pixels across completely different scenes — a 36px window. The classic failure is spatial augmentation without a coordinate update, and this repo had shipped one.

Two more transforms that were doing nothing

RandomHalfBody had been carried across four client configs for a year. It needs 11 total and 8 half keypoints to trigger; with four it never fires once. And flip-test at inference compounds the same flip bug — averaging a heatmap with its own mirror mostly smears predictions toward the centre.

Chasing the wrong variable, then finding the right one

The input-size sweep was built on the theory that aspect-ratio distortion was hurting accuracy. Matching 16:9 exactly at 896x504 crashed training outright — HRNet needs both dimensions divisible by 32, and 504 is not. The sweep did eventually produce the answer, just not the expected one: the real issue was bbox-centred crop semantics, a policy borrowed from flexible human pose estimation and applied to rigid corner-sensitive geometry.

Curated augmentation over stacked augmentation

The legacy pipeline sampled thirteen transforms independently, so four or five could fire on one image and produce something resembling no camera on earth. It was replaced by a tree: 30% of samples pass through untouched, 70% receive exactly one of six hand-picked two-transform recipes — blur with JPEG, gamma with JPEG. Pairs that genuinely co-occur in lossy video. ImageCompression was singled out as the most valuable of all, because real CCTV always goes through a codec and training on lossless RGB is a documented domain gap.

What the AP of 1.000 actually means

It is a validation number, and he flagged that both the validation and test evaluators point at the same split. The verified held-out figure is COCO AP 0.9804. The evaluation also uses OKS sigma 0.25 across all four points — strict by design, since typical human-pose sigmas run 0.025 to 0.107, so the model is graded on tight localisation rather than rough placement.

From the archive

“More augmentation is not better augmentation. Augmentation should simulate realistic production variation.”

“The real issue was NOT aspect ratio distortion. The real issue was bbox-centered crop semantics.”

“Store a frozen config file inside every saved run folder.”

Run log

Every run recorded on this project.