Deep Dive
How the founding team came together
The conversation opens with co-founders tracing their long interconnected history. Chris met Dario and Jared at 19 when visiting the Bay Area, and later sat beside Dario at Google Brain. Sam joined after Dario showed him enough AI model results to convince him the work was serious. Daniela arrived at OpenAI where she found brilliant but disorganized people doing scaling work — the same team that would eventually become Anthropic's core. Jack recalled interviewing Dario in 2015 about his 'Concrete Problems in AI Safety' paper at Google, then later joined OpenAI believing engineers could contribute to safety research. Tom came from Stripe where he'd introduced his boss Greg to Dario, believing he was the smartest person he knew. What bound them wasn't forced alignment but genuine shared conviction: they'd watched scaling laws work, they believed AI could become very powerful, and they worried about safety implications that few others took seriously at the time.
Breaking through the AI winter mentality
Dario explained that in 2014-2018, AI researchers were psychologically damaged by the previous AI winter and hostile to ambitious claims. They avoided bold visions and grand schemes because ambition itself felt disallowed. The team's insight was that safety actually requires believing AI could be powerful and transformative — you can't worry about existential risks from technology you think will never work. Chris noted this was an extension of broader academic risk-aversion, but physicists like Jared brought a different culture: they're arrogant, constantly ambitious, and comfortable with grand schemes. The 'Concrete Problems' paper served as a political consensus-building exercise rather than definitive framework — it simply made safety-thinking legible and respectable across institutions. By demonstrating that serious researchers took these ideas seriously, it shifted the overton window. Only around 2022 did the industry broadly escape that mentality, but Anthropic's founders saw it coming years earlier when others dismissed scaling as impossible.
Constitutional AI and the RSP's origins
Around 2020-2021, the team developed constitutional AI based on a simple insight: AI systems excel at multiple-choice problems and following instructions. Why not write down a constitution of principles and let models compare their own behavior against it? Jared remembered when this sounded crazy, but the team found that simple things work really well in AI — if you give models clear targets and data to train on, they'll optimize toward them. The Responsible Scaling Policy grew from late 2022 conversations between Dario and Paul Christiano about whether to cap scaling until certain safety problems were solved. Rather than a single cap, they designed thresholds: at each stage, you run specific tests, measure capabilities, and implement increasing safety measures before advancing. Dario pushed for a third party to design it so other companies wouldn't dismiss it as self-serving. Paul Christiano led the external design, and Anthropic then built their own version within weeks. The RSP went through more drafts than any document in company history because it's literally the constitution governing their scaling trajectory.
The RSP as organizational alignment tool
Tom explained the RSP forces unity across the organization: if any part of the company isn't aligned with safety values, the RSP blocks them, forcing everyone to confront trade-offs together rather than at the top. Daniela emphasized that recent rewrites aimed to make it less technocratic and adversarial — the goal is a document everyone in the company can understand and champion, like OKRs they naturally reference. When she talks to senators, she frames it simply: we have systems that make our models hard to steal and safe to deploy, which they respond to as normal business practice. Chris noted the RSP prevents weaponizing the word 'safety' — you can't just claim something needs stopping because of safety without referencing the actual policy. Sam highlighted that building the institutions to implement the RSP proved far harder than the theory suggested; there are enormous gray areas you can't predict until implementation. The solution isn't perfect first drafts but rapid iteration: do three or four passes to get it right, learning what goes wrong as early as possible while stakes are still lower.
Race to the top, not noble failure
The founders rejected the narrative that safety requires noble failure or purity. Jack and Dario argued that if you intentionally fail to prove safety can't coexist with competitiveness, you only ensure the people making decisions are those who don't care about safety at all. Instead, Anthropic pursues a 'race to the top': build great products, prove safety is compatible with winning, and let market pressure and gravitational force pull competitors toward similar practices. Dario explained that success is the best argument — the more successful Anthropic becomes while maintaining safety practices, the more incentive others have to copy what works. They've already seen this happen: three major AI companies adopted RSPs within months, the Frontier Red Team got cloned immediately, and interpretability research is now taken seriously industry-wide. The world needs to see this transition managed successfully at the company level, then the industry level, to prove it's possible to go from 'technology doesn't exist' to 'very powerful technology exists and society manages it well.' Sam added that customers actively prefer safer models, which creates real market pressure independent of policy. This is fundamentally different from restricting innovation; it's expanding what's possible by showing safety is a feature, not a bug.
What the founders are excited about next
Chris sees interpretability as both the key to safe AI and a field about to break consensus. Neural networks contain incredible structure and beauty that we're just beginning to understand — they have an artificial biology inside them comparable to how evolution produces complexity. He's serious that interpretability breakthroughs will connect to understanding mental illness and neuroscience in ways that could earn future Nobel Prizes. Dario highlighted three areas where consensus is breaking: interpretability, AI for biology (following AlphaFold's Nobel-winning success), and using AI to enhance democracy rather than enable authoritarianism. He wants to build a hundred AlphaFolds and prove AI can be a tool for freedom and self-determination. Daniela emphasized that people join Anthropic because they become convinced by progress that both advancing AI and ensuring safety matter — increasingly people are united in both the vision for technology and the responsibility it entails. Sam noted that recent advances are now letting them study what risks might actually emerge from very advanced systems directly, using interpretability and safety mechanisms empirically rather than theoretically. This lets them ground the safety mission in real scientific observation of what can actually go wrong with powerful systems, making the next phase both more rigorous and more achievable.