For a decade, progress in computer vision has been measured in tenths of a point on MS-COCO. A survey out of Sejong University argues that a meaningful share of those tenths were never real to begin with — they were artifacts of ground truth that was simply wrong.

The paper, “Quality over quantity: a data-centric survey of annotation errors in object detection datasets,” appeared in Springer’s Artificial Intelligence Review (volume 59, article 107) under an open-access CC BY licence. It was received on August 11, 2025, accepted on January 17, 2026, and published online on February 7. TechXplore brought it to a wider audience on September 7, 2026, under the headline “Better AI starts with better data.”

The team — Adnan Hussain, Kaleem Ullah, Muhammad Afaq, Muhammad Munsif, Altaf Hussain and corresponding author Professor Sung Wook Baik, all of the Department of Software at Sejong University in Seoul — did something the field rarely bothers to do. Instead of proposing another detector, they audited the labels.

What They Found

Fifteen inspectors worked the problem in two tiers: ten running a Python screening script to surface suspect annotations, five verifying candidates by hand, iterating until the queue cleared. Between them they went through roughly 47,000 images pulled from seven benchmarks that appear on nearly every model card in the field — MS-COCO, Open Images, PASCAL VOC, LVIS, DOTA, FSOD and Objects365. Each confirmed error was logged with its image ID and published in a supplementary file.

They confirmed 596 errors across six of the datasets. Incorrect labels dominated, at 236 cases — roughly 40 percent of the total. Inconsistent annotations accounted for 111, localization errors where the box fails to align with the object for 100, duplicate annotations for 64, debatable examples for 52, and group errors for 33. In raw counts Open Images and Objects365 fared worst, at 149 and 140 respectively; PASCAL VOC, the oldest and smallest of the group, logged 46. Missed annotations — visible objects never labeled at all — showed up in MS-COCO, DOTA, PASCAL VOC, Objects365 and FSOD.

Those are conservative numbers, drawn from a hand-checked sample. The survey’s wider literature review, spanning 196 references published between 2016 and 2025, turns up figures that dwarf them. When Google researchers rebuilt a slice of Open Images as the More Inclusive Annotations for People dataset, the survey notes, “approximately 30% of individuals in the images were not tagged in the original dataset.” A re-annotation effort on a 6,000-image MS-COCO subset found more than 12 percent of images carried missing or incorrect labels. A separate pseudo-labelling pipeline flagged 52,621 of MS-COCO 2017’s 123,287 images — 42.7 percent — as anomalous enough to warrant re-annotation.

The survey also imposes order on the fixes. Baik’s team sorts existing error-detection work into four families: manual review, which is the most precise and the least scalable; weakly supervised methods that work from image-level tags and class activation maps; semi-supervised approaches built on pseudo-labelling and teacher–student consistency; and fully automatic pipelines that scale but introduce confident false corrections of their own.

Why It Matters

The uncomfortable implication is not that models are worse than advertised. It is that we may not know which models are better.

“Such errors introduce label noise during training, corrupt evaluation metrics, and can lead to misleading conclusions about model performance, especially for long-tailed categories, few-shot settings, or safety-critical applications,” the authors write. That is a claim with hard evidence behind it. COCO-ReM, a 2024 re-annotation of COCO by a team from IIT Roorkee, Georgia Tech and the University of Michigan, refined roughly 860,000 training and 36,000 validation instances. Every one of the fifty detectors they re-evaluated scored higher on the cleaned annotations; the then state-of-the-art OneFormer InternImage-H gained 7.7 AP points. More consequentially, model rankings flipped. Cascade ViTDet-L outranks Mask2Former Swin-L on COCO-2017 and loses to it on COCO-ReM. Query-based architectures gained about 6.7 points on average, region-based ones as little as 2 — meaning the noise was not uniform, it was systematically biased toward one design philosophy.

That is benchmark leaderboards adjudicating architectural debates on the strength of bad boxes. Years of published comparisons, funding decisions and product roadmaps rest on those orderings.

The stakes climb sharply downstream. Detectors trained on these benchmarks are the pretraining backbone for perception stacks in autonomous vehicles, warehouse robotics and medical imaging. A localization error that shifts a bounding box by a few pixels is a rounding error on a leaderboard and a materially different braking distance in a car. The survey is blunt that the field has under-invested here: error identification, the authors write, “receives little attention compared to the creation of algorithms for object detectors and learning strategies.”

What to Watch

The fix the paper argues for is not another architecture. “Addressing these issues requires data-centric methodologies that can detect, validate, and correct annotation noise at scale, rather than focusing solely on new architectures or loss functions,” the authors write. Watch for whether the cleaned variants — COCO-ReM, Sama-COCO, MJ-COCO, Mini6KClean — start appearing alongside COCO-2017 in results tables at CVPR and ICCV, and whether reviewers begin asking for it. Watch, too, for whether benchmark maintainers publish error rates the way dataset cards now publish demographic breakdowns. Baik’s team has handed the field a taxonomy and a logged list of image IDs. The harder question is whether a research culture organized around beating a number will accept that the number itself needs auditing.

“Such errors introduce label noise during training, corrupt evaluation metrics, and can lead to misleading conclusions about model performance, especially for long-tailed categories, few-shot settings, or safety-critical applications.”
— Sung Wook Baik and colleagues, Department of Software, Sejong University, Seoul
47,000
Images manually inspected across seven benchmarks
596
Confirmed annotation errors, 40% incorrect labels
+7.7 AP
Score gain for the SOTA detector on re-annotated COCO
42.7%
MS-COCO 2017 images flagged as anomalous