When the CVPR 2026 Awards Selection Committee climbed the stage at the Colorado Convention Center in Denver on June 5, the field it was about to honor had never been bigger. Out of 16,092 submissions, the conference accepted 4,089 papers — a record-shattering haul for the preeminent gathering in computer vision. From that ocean of research, the IEEE Computer Society and the Computer Vision Foundation singled out a handful of works that, taken together, read like a map of where the field is heading: into three dimensions, into generative modeling, and into systems that act in the physical and virtual world rather than merely describe it.
The headline prize went to a paper with a deceptively playful title — "Efficiently Reconstructing Dynamic Scenes One D4RT at a Time" — and a serious technical claim behind it.
The winners
The CVPR 2026 Best Paper went to D4RT, the work of a team spanning Google DeepMind, University College London, and the University of Oxford, with authors including Chuhan Zhang, Guillaume Le Moing, Skanda Koppula, Andrew Zisserman, Zoubin Ghahramani, Raia Hadsell, and Mehdi S. M. Sajjadi, among others. D4RT is a transformer-based network that reconstructs the geometry and motion of dynamic 4D scenes directly from video. In a single unified architecture, it estimates depth, spatio-temporal correspondence, and full camera parameters, then lets a user independently probe the 3D position of any point in space and time.
The significance is as much about engineering as elegance. Traditional 4D reconstruction stitches together separate models for depth, optical flow, and camera pose, often with costly test-time optimization. D4RT collapses that pipeline into one query interface: encode a video once, then answer depth, point-tracking, and camera-pose questions from the same latent representation, treating moving objects no differently from static ones. The result, its authors report, is a lightweight, highly scalable method enabling "remarkably efficient training and inference" — and, according to independent write-ups of the work, a new state of the art across the 4D benchmarks it was tested on, including scenes where prior methods stumble.
The CVPR 2026 Best Student Paper went to "Native and Compact Structured Latents for 3D Generation," a collaboration among Tsinghua University, Microsoft Research, the University of Science and Technology of China, and Microsoft AI. Its authors — Jianfeng Xiang, Xiaoxue Chen, Sicheng Xu, Ruicheng Wang and colleagues — introduce O-Voxel, a representation designed to capture complex shapes and surface attributes more faithfully than existing approaches. The payoff is sharper, more realistic AI-generated 3D assets, a contribution the committee credited with meaningfully advancing 3D generative modeling.
Three further papers earned honorable mentions, and their subjects are telling. Two received Best Paper honorable mentions: NitroGen, an open foundation model for generalist gaming agents from a team including NVIDIA, Stanford, Caltech, the University of Chicago, and UT Austin, trained on 40,000 hours of gameplay across more than 1,000 games; and SAM 3D: 3Dfy Anything in Images, from Meta Superintelligence Labs, which reconstructs the geometry, texture, and layout of objects from a single image and posted at least a 5:1 win rate in human-preference tests. The Best Student Paper honorable mention went to ChordEdit, a training-free, inversion-free image-editing method from a Chinese university consortium that the authors say achieves true real-time editing.
What the awards signal
Read as a set, the slate points unmistakably in one direction. "From advances in dynamic scene reconstruction to breakthroughs in 3D generative modeling, these works address fundamental challenges in computer vision while opening new possibilities for applications across AI, robotics, and more," said Alexander G. Schwing, an associate professor at the University of Illinois Urbana-Champaign and a CVPR 2026 Program Co-Chair.
Four of the five honored papers are explicitly about 3D or 4D — recovering, generating, or manipulating structure in space and time rather than classifying flat pixels. That is a notable shift for a conference whose history is rooted in 2D recognition. The other throughline is agency: NitroGen learns to act inside thousands of games, and D4RT's per-point queryable geometry is precisely the kind of spatial scaffolding embodied agents and robots need to plan and move. Generative modeling, long associated with images and text, now sits squarely in the geometry stack.
Chen Change Loy, the Tan Lip-Bu Professor in AI at Nanyang Technological University and a fellow Program Co-Chair, framed the stakes plainly. "These papers represent some of the most impactful and forward-looking research presented at CVPR 2026," he said. "The selected papers introduce concepts and practices that undoubtedly will help shape the next generation of computer vision systems and applications."
What to watch
The institutional fingerprints are worth noting too: Google DeepMind, Microsoft, Meta, and NVIDIA all appear among the winners, alongside Oxford, UCL, Tsinghua, Stanford, and others — a reminder that frontier computer-vision research increasingly runs on industrial-scale compute and data, even when the prize-winning ideas are conceptually lean.
The papers were presented at the IEEE Computer Society's TCPAMI meeting on June 6, the day after the awards ceremony, capping a conference that ran June 3–7 and drew well over 10,000 attendees. The next test of these ideas is adoption: whether D4RT's unified query model becomes the default for 4D reconstruction, whether O-Voxel-style latents reshape 3D generation pipelines, and whether agents like NitroGen migrate from games to the messier physical world. CVPR 2027 — set for June 19–26 at the Seattle Convention Center — will be the venue where we find out.
"From advances in dynamic scene reconstruction to breakthroughs in 3D generative modeling, these works address fundamental challenges in computer vision while opening new possibilities for applications across AI, robotics, and more."- Alexander G. Schwing, CVPR 2026 Program Co-Chair, UIUC