Skip to content
Video Lenny's Podcast Jul 2026

Lenny's Podcast: Anthropic's first technical PM on building Claude from the inside

This episode of Lenny’s Podcast was published on July 26, 2026. The guest is Dianne Penn, Head of Product for Anthropic’s AI Research and Labs teams. Penn joined Anthropic in 2023 as its first technical product manager, when the entire product team was five engineers. She has worked on every model release from Claude 2 through Fable, and helped build Claude Code, MCP, Skills, computer use, tool use, and reasoning capabilities from inception. Before Anthropic, she worked on Alexa’s AI at Amazon and traded high-yield bonds at JPMorgan Chase.

The episode focuses on the specific decisions — not the abstract principles — that moved Claude from a credible but underdog model to a dominant position in coding and enterprise workflows. Penn explains these decisions from inside the product development process, which distinguishes the conversation from most external analysis of Anthropic’s trajectory.

Key takeaways

  1. The coding pivot was a deliberate bet, not an emergent strength. Penn describes how Anthropic identified coding as the domain where a model could create the most measurable, verifiable value and concentrated product and model development around it. The decision to go deep on one high-value domain rather than competing on generalist benchmarks shaped everything that followed. For PMs building AI products, this is a positioning decision worth examining: which domain is the right one to win before expanding?

  2. Evaluation-driven development is the actual methodology. Penn explains that Anthropic does not ship features based on intuition or demo performance. Every product decision is grounded in an evaluation suite — a structured set of test cases with measurable outputs. Building the evaluation infrastructure before the feature itself is the process, not an afterthought. For product managers overseeing AI features, this reframes what “done” means at the spec stage: a feature is not ready to build until you can describe how you will know whether it works.

  3. Token maxing and the jagged edge. Penn introduces two concepts that explain why AI products often behave unexpectedly in production. Token maxing describes the tendency of users to push models to their context limits in ways that reveal failure modes at scale. The jagged edge refers to the uneven capability profile of large language models — exceptional in some dimensions, surprisingly weak in adjacent ones. Understanding both patterns helps product managers design better constraints, onboarding, and fallback behaviors.

  4. Living in the future as a PM methodology. At Anthropic, product managers are expected to use the tools they are building at capability levels that most users will not reach for months. Penn describes this as “living in the future” — deliberately exploring the edge of what the model can do so that product decisions account for where the capability is going, not just where it is today. This is harder to replicate without access to frontier models, but the underlying principle — staying ahead of your users’ actual usage — applies broadly.

  5. What comes after coding is solved. Penn offers her read on where agentic AI is heading as coding becomes a baseline capability. She expects the next domain to be higher-stakes professional work that requires longer context, stronger reasoning, and trust infrastructure — areas where reliability and verifiability matter more than raw speed.

Worth watching if

You are building AI features and want a practitioner’s account of how product decisions get made at the company that ships the models you are using. Penn’s frame is not prescriptive — she is not telling you how to run your team — but the specifics of how Anthropic approaches evaluation, prioritization, and product-model alignment offer concrete reference points for teams working at smaller scale.