Deep Dive
Kimi K3 Arrives but Pricing Doesn't Match Hype
Frank opens by framing Kimi K3 as the latest open-source model surge—similar to the DeepSeek moment from January 2024. The release triggered a 28% selloff in Moonshot's Chinese competitors, and benchmarks show the 2.8 trillion-parameter model landing between Opus and GPT-4.5. But Frank cuts through the excitement with cost math. Moonshot lists output at $15 per million tokens versus GPT-4.5's $30, a 50% discount that looks compelling. The catch: Kimi's lower efficiency means it needs twice as many tokens per response, erasing the savings. Average cost per task stays identical. Frank also notes the model's sheer size creates higher operating costs across the board, while infrastructure—Moonshot's website crashed under user load and API weights won't hit US clouds until next week—reveals compute is still the bottleneck, not model innovation.
The Frontier Keeps Getting Crowded
Nick shifts the frame from model benchmarks to market reality. He argues the frontier—however you define it—has grown crowded over the last six months. OpenAI and Anthropic still trade leadership on a few percentage points of benchmark gains, but for actual knowledge work, dozens of models are now good enough. That's bad news for frontier model companies still private, Nick reckons, whose stocks would be crushed if public. The real value, he argues, accrues upstream at the infrastructure layer. AMD, Nvidia, and cloud providers don't lose when new open models launch; they win because more models mean more demand for chips and compute. Nick references Brian Armstrong's chart showing tokens rising while costs fall—the dynamic that should terrify model companies. The market is drifting from token-maximization hype toward cost-conscious productivity, and that pressure flows downhill.
Why You Still Pay for the Best Lawyer
Frank pivots to task design, borrowing Brett's metaphor about lawyers passing the bar. Not all lawyers cost the same despite minimum competence thresholds. He applies this to models: for well-defined tasks, use the cheapest model that works. But when ambiguity exists—and Frank argues most real work involves ambiguity—you need frontier-level intelligence to fill gaps. A cheaper model struggling with ambiguity wastes tokens trying to route around it, often costing more than just paying for quality upfront. He cites UiPath's market size as evidence: if all software tasks were discrete and well-defined, RPA would be much bigger. Knowledge work—legal analysis, investment research, even software itself—remains messy. The scope for premium models isn't shrinking; it depends on whether your task can be perfectly specified. Nick agrees but pushes back slightly, noting that productivity improvement is what matters, not theoretical intelligence. If a team feels more productive with AI but sales don't move, something's broken in how the tool's being deployed.
The Real Test: Actual ROI, Not Token Metrics
Nick surfaces the core issue: companies need to measure AI adoption against real business outcomes, not just time-savings or token consumption. The wave to watch, he says, is companies realizing employees are busier but revenue hasn't budged—triggering a reckoning about whether they're using the right models for the right work. Frank agrees, then frames it as a natural evolution. Early electricity adoption was chaotic and dangerous; over time, applications and industries formed around it. AI is at the touching-the-wire phase. Token-maxing was the equivalent of playing with raw power and getting burned. The real narrative, Frank says, isn't that Kimi K3 kills frontier labs—it's that cost declines will continue forcing model companies to rethink margins and business models. Even if individual customers spend more on AI absolute dollars, per-token pricing will keep falling. Companies on the wrong side of that curve burn; companies providing cheaper intelligence at scale thrive.