Sakana AI's pitch for its newest model is, in a sense, that it built no model at all. On June 22, 2026, the Tokyo lab released Sakana Fugu and its flagship tier, Fugu Ultra — and the surprise is what it is not. It is not a giant Japanese-language frontier LLM, not an evolutionary merge of open weights, and not a scaled-up monolith. Fugu is an orchestration model: a language model trained to break a task apart, route the pieces to a swappable pool of other frontier LLMs, verify their work, and synthesize a single answer — all behind one OpenAI-compatible API endpoint.

The framing is pure Sakana. Since its 2023 founding by David Ha, transformer co-author Llion Jones, and Ren Ito, the lab has bet that the most powerful AI systems will be "collaborative ecosystems" rather than "isolated monoliths." Fugu is that thesis shipped as a product. "You send a request to one endpoint, and Fugu decides how to handle it: solving it directly when that is enough, or assembling and coordinating a team of expert models when a task calls for more," the company wrote. Crucially, Fugu can call instances of itself recursively, and the underlying agent pool is designed to be "entirely swappable."

What Fugu Ultra actually is

The product ships in two tiers. Plain Fugu trades some quality for low latency and slots into everyday coding and chatbot workloads; it even lets teams opt specific agents out of the pool for privacy or compliance reasons. Fugu Ultra (model ID `fugu-ultra-20260615`) is tuned for maximum answer quality on long, multi-step problems and coordinates a deeper, fixed pool of experts. Sakana says early users leaned on it for AI research, paper reproduction, cybersecurity analysis, and patent investigation. The work builds on two Sakana papers accepted to ICLR 2026 — TRINITY, an evolved coordinator that hands out Thinker/Worker/Verifier roles, and Conductor, an RL-trained system that learns natural-language coordination strategies.

The benchmark numbers are striking, with a caveat. By Sakana's own measurements, a Fugu model tops 10 of 11 published benchmarks, and the orchestrator beats the very models it coordinates. Fugu Ultra posts 73.7 on SWE-Bench Pro (vs. 69.2 for Opus 4.8, 58.6 for GPT-5.5), 50.0 on Humanity's Last Exam, and leads all four coding benchmarks. Sakana also claims Fugu Ultra stands "shoulder-to-shoulder" with Anthropic's Fable 5 and Mythos Preview — notably, two models it cannot include in its pool because they are not publicly accessible. Those scores are self-reported and unaudited, and early testers flagged a gap between the benchmarks and messy real-world use.

The geopolitics built into the product

What sets this launch apart is that Sakana explicitly frames orchestration as a hedge against vendor and geopolitical risk. The announcement pointed directly to recent export controls on Anthropic's Fable and Mythos models as proof that "access can shift or disappear overnight." CEO David Ha made the stakes plain: "Relying on a single company's APIs for critical infrastructure, finance, or governance is a material vulnerability. This risk is no longer a hypothetical possibility, but a reality." If one provider goes dark, Sakana argues, Fugu simply routes around it — what it calls "the realistic, resilient blueprint required for AI sovereignty."

That sovereignty claim is also where the skepticism concentrates. An orchestrator that depends on a rotating cast of foreign frontier models is more resilient than sovereign in any strict sense, and early reaction on Hacker News and X skewed doubtful, with the "is this just a router or a wrapper?" critique dominating. Because Fugu's per-query routing is proprietary and hidden, users can't see which model answered them — a transparency gap that matters for the same compliance-conscious buyers Sakana is courting.

Japan's distinctive lane in the AI race

Fugu underscores how Japan's most visible AI lab has chosen to compete sideways rather than head-on. Sakana has never tried to out-scale OpenAI, Google, or Anthropic on raw pre-training compute; its identity is nature-inspired AI — evolutionary model merging, swarm behavior, the "AI Scientist" — that squeezes new capability out of combining existing models. In an era of escalating GPU costs and tightening export regimes, an architecture that turns the entire global model ecosystem into a resource, rather than a competitor, is a genuinely different strategic posture. It also dovetails with national-sovereignty anxieties that resonate well beyond Tokyo.

What to watch

Three things. First, independent benchmarks: Sakana's self-reported leads need third-party confirmation before the "beats the models it orchestrates" claim can be trusted. Second, the economics and transparency of a black-box router that bills $5/$30 per million input/output tokens while hiding which model did the work. Third, Sakana's promise to fold open models and "Sakana AI's own models" into the pool over time — the moment Fugu starts orchestrating Sakana-trained agents is the moment its no-model-at-all pitch quietly becomes a model company after all.

"Relying on a single company's APIs for critical infrastructure, finance, or governance is a material vulnerability. This risk is no longer a hypothetical possibility, but a reality."
- David Ha, CEO and Co-founder, Sakana AI
73.7
Fugu Ultra score on SWE-Bench Pro (vs 69.2 for Opus 4.8)
10 of 11
Benchmarks where a Fugu model posts the top score
~500
Early users in the Fugu closed beta
$5 / $30
Fugu Ultra price per million input / output tokens