Most of the world's data is trapped inside video, and almost none of it can be searched, questioned, or put to work. TwelveLabs Inc. has spent five years betting that solving that problem is the real frontier of machine intelligence. On July 1, the San Francisco startup gave investors a fresh reason to agree, closing a $100 million Series B to bankroll what it calls "video superintelligence."
The round was co-led by NEA and NAVER Ventures, with participation from Amazon, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital and Red Bull Ventures. The new capital brings TwelveLabs' total funding to more than $207 million and comes as the company pushes beyond its foundation models into a full-stack "agentic" system that combines perception, knowledge and reasoning about video into a single architecture.
The pitch rests on a simple but staggering statistic: video accounts for as much as 90% of the world's data, yet the overwhelming majority of it is opaque, unsearchable and effectively unusable. TwelveLabs wants to turn those idle archives into a "living, searchable system" that grows more valuable with every clip it ingests.
"Five years ago, we made a contrarian bet: the substrate of machine intelligence is recorded reality in motion, not language," said Jae Lee, co-founder and chief executive of TwelveLabs. "Language is downstream of understanding. Video is the data understanding has to answer to. This funding lets us take TwelveLabs from foundation models to a full-stack video cognition system that meets every user, every agent, and every machine that needs to understand the world. The road to Video Superintelligence starts here."
The Technology
TwelveLabs' core argument is that the tools of the large-language-model era were built for text and break down when applied to moving images. Today's LLMs cannot consume a video all at once; they sample a handful of frames, miss everything in between, and start from zero with every new query. Feeding entire libraries into a context window, meanwhile, would demand compute that does not exist at a cost no enterprise could justify.
The company's answer is what it describes as "genuine multimodality" — models born in video rather than language models retrofitted to watch it. Its Marengo 3.0 embedding model, released late last year, converts raw footage into a semantic layer machines can search at scale, understanding sound, speech and motion across time. A companion model, Pegasus 1.5, turns video into structured data: scene boundaries, entities, temporal segments and semantic context that other systems can reason over. Both are distributed through Amazon Bedrock and TwelveLabs' own API.
The Series B funds the leap from those models to an agentic architecture that builds a persistent memory of every video it ingests and reasons across all of it — "intelligence that compounds with every video processed," in the company's framing, rather than a tool that resets with each request. Earlier this month TwelveLabs took its first step up the stack with Rodeo, its first application-layer product.
Its earliest backers frame the bet in even grander terms. "When we first met Jae, he described TwelveLabs as the visual cortex for future AI agents," said YJ Park, general partner at NAVER Ventures, which made TwelveLabs its first-ever investment. "As agents and machines move into roles where they need to perceive and reason about the physical world, video is the modality that matters most."
Why It Matters
Video understanding has quietly become one of the AI industry's most contested frontiers. Text and images have well-funded model families and mature tooling; motion-first, temporally aware video reasoning remains comparatively unsolved — and it is the modality that autonomous agents, robots and physical-world systems will most need to interpret reality.
The enterprise appetite is already concrete. TwelveLabs says it has built deep traction in media and entertainment and is expanding into the public sector, working with governments to apply video intelligence to mission-critical workflows. Advertising, security, sports and automotive round out a customer base sitting on vast, underused footage libraries — archives that today register as storage costs rather than strategic assets.
The competitive backdrop is unavoidable. Google, OpenAI and Meta are all pushing multimodal models that ingest video, and any of them could turn video understanding into a commodity feature. TwelveLabs' counter is depth and ownership of the full stack — perception, knowledge, reasoning and the orchestration binding them. "Models commoditize," Lee said. "The intelligence layer that composes them does not."
The Amazon relationship sharpens that positioning. Beyond investing, AWS is TwelveLabs' preferred cloud provider, and the two have signed a multiyear commitment to optimize the startup's video-inference workloads on AWS Trainium chips, with new TwelveLabs models slated to launch first on AWS. The Korean thread is equally notable: NAVER Ventures and Korea Investment Partners anchor a company that keeps major operations in Seoul alongside San Francisco, a rare bi-continental profile among frontier AI labs.
What to Watch
The money will flow into research and development plus geographic expansion, with new offices opening in New York and London to complement the existing San Francisco and Seoul hubs. The sharper question is whether TwelveLabs can climb the stack fast enough. Rodeo signals ambitions well beyond selling API access to models — and applications are where durable enterprise revenue, and durable defensibility, tend to live.
Watch three things: whether the compounding-memory architecture delivers measurably better results than frame-sampling rivals in production; whether the deepening AWS alliance becomes a genuine moat or a dependency; and whether "video superintelligence" matures into shipping products or remains an aspirational tagline. With more than $207 million now behind it, TwelveLabs has the runway to find out.
"Five years ago, we made a contrarian bet: the substrate of machine intelligence is recorded reality in motion, not language. The road to Video Superintelligence starts here."— Jae Lee, CEO and Co-founder, TwelveLabs