Rhoda AI, a Palo Alto robotics startup that spent 18 months building in stealth, has stepped into public view with one of the largest first institutional rounds the robotics sector has seen: a $450 million Series A that values the company at roughly $1.7 billion. Alongside the raise, Rhoda unveiled FutureVision, a robot-intelligence platform built on a technique it calls video-predictive control, and made a deceptively simple pitch about what the money is for: robots that work in the real world, not just in controlled lab demonstrations.

The company is led by cofounder and CEO Jagdeep Singh, a serial deep-tech founder, with a founding team drawn heavily from Stanford's computer-vision and generative-modeling community, including Chief Science Officer Eric Ryan Chan, a Stanford researcher and former generative-model architect at World Labs, and Gordon Wetzstein, the Stanford professor who heads the university's Computational Imaging Lab. The investor list reads like a who's-who of deep-tech capital: Khosla Ventures, Temasek, Capricorn Investment Group, Mayfield, Premji Invest, Prelude Ventures, Leitmotif, Matter Venture Partners, and Xora, with venture pioneer John Doerr also among the backers.

What video-predictive control actually means

Most industrial robots today follow pre-programmed paths and excel only in tightly structured settings. Even the newer wave of AI-driven systems, particularly vision-language-action (VLA) models, tend to stumble when an unexpected object appears, a layout shifts, or a workflow turns irregular. Those failures usually mean a stall or a call for human intervention, which is precisely the bottleneck that keeps automation confined to the most predictable corners of a factory.

Rhoda's bet is architectural. Rather than learning primarily from teleoperated robot demonstrations, FutureVision is pretrained on hundreds of millions of internet videos so the model develops an intuition for how the physical world moves, including motion, physics, and physical interaction, before it ever touches a robot. It is then fine-tuned with comparatively small amounts of robot-specific data. The result, the company says, is that a new task can sometimes be learned with as little as ten hours of teleoperation data, a fraction of what conventional approaches demand.

The runtime behavior is where the "predictive" label earns its keep. Rhoda calls its proprietary architecture a Direct Video Action (DVA) model. The system continuously observes its surroundings, forecasts upcoming states as video, converts those predictions into actions, executes them, and re-observes, closing the loop every few hundred milliseconds. That distinguishes it from open-loop planners that generate a plan once and run it blind. Because the model updates its behavior as conditions change, Rhoda argues it can hold accuracy under the kind of variability that breaks scripted systems.

"We believe the next era of robotics requires models that understand how the world moves – not just what it looks like or how it's described in language," Singh said in announcing the launch. "By learning from internet-scale video and operating in closed loop, our systems are designed to adapt to real-world variability in ways conventional approaches struggle to achieve. The goal is simple: robots that work in the real world, not just controlled lab settings."

Early traction and target markets

Rhoda is pointing FutureVision at manufacturing and logistics first, the markets where high task variability has historically resisted automation and where customers feel labor constraints most acutely. The company says the technology has already been tested in production environments handling constantly changing materials and workflows. In one high-volume manufacturing evaluation, it reported that a robot system completed a component-processing workflow in under two minutes per cycle without human intervention, exceeding the customer's performance targets.

The plan is to operate FutureVision as an intelligence layer for Rhoda's own systems initially, then eventually license it as a foundation model to partners building robotic hardware and software, a strategy that mirrors how large-language-model providers turned a core model into a platform business. The Series A, the company says, will fund continued research and engineering, the expansion of customer pilots and industrial deployments, and growth of a multidisciplinary team spanning generative AI, computer vision, and robotics.

Investors framed the appeal in terms of expanding the universe of automatable work. "In manufacturing, tasks with high variability have historically resisted automation. The real challenge isn't solving it once, it's delivering consistent, reliable output under real-world production conditions," said Jens Wiese, managing partner at Leitmotif and a former Volkswagen Group executive. He added that what impressed his firm was the system's ability to adapt to conditions that typically require human intervention, technology he argued could play a role in re-industrializing mature economies.

Analysis: another bet in the robotics-foundation-model wave

Rhoda lands in the middle of an extraordinary capital surge into "physical AI." A $450 million Series A would have been almost unheard of a few years ago; in 2026 it is one of several nine-figure rounds chasing the thesis that the transformer playbook, which scaled language and image models, can be ported to robotics if the right data and architecture are found. The central disagreement in the field is what that data and architecture should be.

The dominant approach until recently has been the VLA model, which grafts robot control onto large language and vision models, effectively asking a system that reasons in tokens and text to also produce motor actions. Critics argue language is a lossy intermediary for physics. Rhoda's video-predictive framing is a direct challenge to that orthodoxy: if the model can predict the next frame of reality, the reasoning argument goes, it has internalized dynamics that language can only gesture at, and control becomes a matter of acting on those predictions rather than translating instructions. It is a thesis shared in spirit by world-model research elsewhere in the industry, and the heavy Stanford generative-vision pedigree on Rhoda's founding team signals where the company believes the edge lies.

The skeptic's caveats are familiar. Internet video is abundant but messy, and the gap between predicting plausible pixels and executing reliable force-controlled actions is exactly where many demos quietly fail. A single sub-two-minute cycle in one evaluation is encouraging but not yet evidence of fleet-scale reliability across customers and tasks. And a $1.7 billion valuation on a company emerging from stealth bakes in expectations that the model generalizes far beyond its first pilots.

What to watch next

The tells over the coming months will be concrete: named manufacturing or logistics customers moving from pilot to paid production, throughput and uptime figures across more than one task, and whether the promised licensing model attracts hardware partners or stays an in-house intelligence layer. Watch, too, for whether Rhoda publishes any technical detail on the DVA architecture, since the field will judge the video-predictive claim on benchmarks, not press releases. If FutureVision holds accuracy across genuinely variable real-world settings, it strengthens the case that video, not language, is the right substrate for robot foundation models, and it puts pressure on the well-funded VLA camp to prove otherwise.

"We believe the next era of robotics requires models that understand how the world moves – not just what it looks like or how it's described in language."
— Jagdeep Singh, Cofounder and CEO, Rhoda AI
$450M
Series A
$1.7B
Valuation
100s of millions
Training videos