AP 1.000 — published with its own caveat attached
Detect the four corners of a bin in fixed CCTV imagery and export to ONNX. Its output becomes the crop stage of the fill-level pipeline — the two projects are one delivery chain.
2026 · for client deployments
1.000
Validation AP
epoch 70 @ 1024x576
0.9804
Held-out AP
COCO AP, verified
460
Training set
samples, 79% synthetic
4
Keypoints
graded at strict sigma 0.25
Pipeline
-
CVAT
polygons -
polygon →
COCO keypoints -
HRNet-w32
heatmap head -
argmax +
sub-pixel -
ONNX
export -
fill-level
pipeline
The bug that made the model constant
Four legacy configs enabled horizontal flip augmentation. Every keypoint swap field was empty, so the framework produced an identity mapping: the image mirrored and the labels did not. The model converged on predicting fixed coordinates for everything. On the deployed ONNX, the first keypoint x-coordinate varied only between 658 and 694 pixels across completely different scenes — a 36px window. The classic failure is spatial augmentation without a coordinate update, and this repo had shipped one.
Two more transforms that were doing nothing
RandomHalfBody had been carried across four client configs for a year. It needs 11 total and 8 half keypoints to trigger; with four it never fires once. And flip-test at inference compounds the same flip bug — averaging a heatmap with its own mirror mostly smears predictions toward the centre.
Chasing the wrong variable, then finding the right one
The input-size sweep was built on the theory that aspect-ratio distortion was hurting accuracy. Matching 16:9 exactly at 896x504 crashed training outright — HRNet needs both dimensions divisible by 32, and 504 is not. The sweep did eventually produce the answer, just not the expected one: the real issue was bbox-centred crop semantics, a policy borrowed from flexible human pose estimation and applied to rigid corner-sensitive geometry.
Curated augmentation over stacked augmentation
The legacy pipeline sampled thirteen transforms independently, so four or five could fire on one image and produce something resembling no camera on earth. It was replaced by a tree: 30% of samples pass through untouched, 70% receive exactly one of six hand-picked two-transform recipes — blur with JPEG, gamma with JPEG. Pairs that genuinely co-occur in lossy video. ImageCompression was singled out as the most valuable of all, because real CCTV always goes through a codec and training on lossless RGB is a documented domain gap.
What the AP of 1.000 actually means
It is a validation number, and he flagged that both the validation and test evaluators point at the same split. The verified held-out figure is COCO AP 0.9804. The evaluation also uses OKS sigma 0.25 across all four points — strict by design, since typical human-pose sigmas run 0.025 to 0.107, so the model is graded on tight localisation rather than rough placement.
From the archive
“More augmentation is not better augmentation. Augmentation should simulate realistic production variation.”
“The real issue was NOT aspect ratio distortion. The real issue was bbox-centered crop semantics.”
“Store a frozen config file inside every saved run folder.”
Run log