ARK Invest
ARK Invest5d ago
Tech

Will China’s Kimi K3 Win The AI Model Race?| The Brainstorm 141

24 min video4 key momentsWatch original
TL;DR

Kimi K3 matches GPT-4.5 on benchmarks but costs the same per task due to lower token efficiency; the real threat to frontier labs isn't raw performance but margin compression as open-source models force pricing pressure across the inference market.

Key Insights

1

Same average cost per taskKimi K3 lists at $15 per million output tokens versus GPT-4o's $30, but requires twice as many tokens per response—ending up at identical cost per task despite the lower headline price.

2

75% tokens, 80% spending divergenceOn OpenRouter, 75% of tokens consumed are from open-source models, yet 80% of spending flows to closed-source ones—revealing that token volume and dollar spend move in opposite directions.

3

Ambiguity demands intelligenceWell-defined tasks warrant cheap models; ambiguous work needs frontier intelligence to avoid token waste and higher total costs—the scope for premium models depends entirely on task clarity, not raw capability.

4

Token maxing is overCompanies are shifting from token-maxing in early 2026 to measuring actual productivity lifts and ROI, meaning inference cost declines alone won't help if they don't translate to real business outcomes.

5

Infrastructure is the constraintMoonshot's infrastructure crashed from user demand spikes, and weights won't be available on US clouds until next week—proving compute bottlenecks remain the real constraint, not model quality.

Want this for every new video ARK Invest posts? Brevyd summarizes each upload automatically, the morning it drops.

Deep Dive

Kimi K3 Arrives but Pricing Doesn't Match Hype

Frank opens by framing Kimi K3 as the latest open-source model surge—similar to the DeepSeek moment from January 2024. The release triggered a 28% selloff in Moonshot's Chinese competitors, and benchmarks show the 2.8 trillion-parameter model landing between Opus and GPT-4.5. But Frank cuts through the excitement with cost math. Moonshot lists output at $15 per million tokens versus GPT-4.5's $30, a 50% discount that looks compelling. The catch: Kimi's lower efficiency means it needs twice as many tokens per response, erasing the savings. Average cost per task stays identical. Frank also notes the model's sheer size creates higher operating costs across the board, while infrastructure—Moonshot's website crashed under user load and API weights won't hit US clouds until next week—reveals compute is still the bottleneck, not model innovation.

The Frontier Keeps Getting Crowded

Nick shifts the frame from model benchmarks to market reality. He argues the frontier—however you define it—has grown crowded over the last six months. OpenAI and Anthropic still trade leadership on a few percentage points of benchmark gains, but for actual knowledge work, dozens of models are now good enough. That's bad news for frontier model companies still private, Nick reckons, whose stocks would be crushed if public. The real value, he argues, accrues upstream at the infrastructure layer. AMD, Nvidia, and cloud providers don't lose when new open models launch; they win because more models mean more demand for chips and compute. Nick references Brian Armstrong's chart showing tokens rising while costs fall—the dynamic that should terrify model companies. The market is drifting from token-maximization hype toward cost-conscious productivity, and that pressure flows downhill.

Why You Still Pay for the Best Lawyer

Frank pivots to task design, borrowing Brett's metaphor about lawyers passing the bar. Not all lawyers cost the same despite minimum competence thresholds. He applies this to models: for well-defined tasks, use the cheapest model that works. But when ambiguity exists—and Frank argues most real work involves ambiguity—you need frontier-level intelligence to fill gaps. A cheaper model struggling with ambiguity wastes tokens trying to route around it, often costing more than just paying for quality upfront. He cites UiPath's market size as evidence: if all software tasks were discrete and well-defined, RPA would be much bigger. Knowledge work—legal analysis, investment research, even software itself—remains messy. The scope for premium models isn't shrinking; it depends on whether your task can be perfectly specified. Nick agrees but pushes back slightly, noting that productivity improvement is what matters, not theoretical intelligence. If a team feels more productive with AI but sales don't move, something's broken in how the tool's being deployed.

The Real Test: Actual ROI, Not Token Metrics

Nick surfaces the core issue: companies need to measure AI adoption against real business outcomes, not just time-savings or token consumption. The wave to watch, he says, is companies realizing employees are busier but revenue hasn't budged—triggering a reckoning about whether they're using the right models for the right work. Frank agrees, then frames it as a natural evolution. Early electricity adoption was chaotic and dangerous; over time, applications and industries formed around it. AI is at the touching-the-wire phase. Token-maxing was the equivalent of playing with raw power and getting burned. The real narrative, Frank says, isn't that Kimi K3 kills frontier labs—it's that cost declines will continue forcing model companies to rethink margins and business models. Even if individual customers spend more on AI absolute dollars, per-token pricing will keep falling. Companies on the wrong side of that curve burn; companies providing cheaper intelligence at scale thrive.

Takeaways

  • Measure your AI productivity gains against actual business metrics—revenue, throughput, customer retention—not internal efficiency scores or token usage; if employee output jumps but revenue doesn't follow, your deployment is broken.
  • For routine, well-defined work, route to the cheapest capable model; reserve frontier-level inference for ambiguous tasks where intelligence saves tokens and time; build workflows that route intelligently rather than overpaying for blanket coverage.
  • Watch infrastructure-as-a-service and chip makers (Nvidia, AMD) as the safer long-term plays; open-source model releases accelerate compute demand, not kill it, making hardware the real beneficiary of margin compression upstream.
  • Track per-token pricing trends alongside absolute spending on prediction markets like Kalshi; current H100 futures pricing will reflect compute capacity inflection when Blackwell and Vera Rubin scale.

Key moments

1:25Cost Math Erases Kimi's Price Advantage

It takes twice the amount of tokens to respond. So you end up with the same average cost per task across the two models. So this is a competitive open-source model, but it is not a lower cost model.

5:05Open vs. Closed Spending Divergence

If you look at token volume, open models versus closed models, it's about 75% open-source models. But at the same time, if you look at where the dollars are flowing, about 80% is going towards the closed-source models.

12:35Ambiguity Demands Frontier Intelligence

If you have some level of ambiguity in that task, you need intelligence to fill in the gaps. And in which case I think it is beneficial to have the smartest model figuring out how to fill in the gaps.

17:15The Real Fear: Cost Declines, Not Model Competition

There is nuance here. It's not just Kimi K3 came in and everyone's dead. It's more like everything's getting more crowded. Everything's becoming cost competitive. Companies are looking at the bottom line.

You just read one. Brevyd does this for every upload.

Follow ARK Invest and every new video comes back as a summary like this, in your morning briefing. No watching required.