The discussion kicked off with Claire from OpenAI acknowledging the challenge of shipping imperfect products, a stark contrast to her previous experience at Stripe where deep polish was the norm. She emphasized that in the current AI era, urgency is paramount, and empirical evidence from users iterating on products outweighs theoretical perfection. The "toggle," an initial agentic harness for ChatGPT's billion-plus users, exemplifies this: an imperfect solution deemed necessary for immediate functionality and future iteration.
Non then delved into the decision of what to deprecate versus what to keep. He highlighted the importance of a "coherent story" for users, allowing them to follow the product's evolution. If users understand the journey, they are more accepting of changes, even if it means building and discarding a lot of code or product along the way.
The conversation then shifted to the constraints that maintain a high-quality bar despite rapid development. Claire outlined three key elements: first, ensuring the product is "additive for users" by unlocking genuine value; second, meeting an "internal bar" where products are trialed internally for uptake, delight, and novel use cases; and third, aiming for where the models will be in "two to three months," avoiding being overly anchored in the present or too futuristic to be usable. Non added that the biggest constraint is users' "understanding and ability to absorb" new capabilities, noting a "capability overhang" where models can do more than users currently leverage.
They addressed the common belief that B2B customers cannot absorb change. Claire argued that while the pace might feel "beyond breakneck," not shipping frontier advancements risks enterprises being "leapfrogged." She used the example of the agent revolution, where enterprises needed agents fast, even if it broke existing processes, to unlock greater value. This reinforces the truism of building what users *need*, not just what they *say* they need.
A debate emerged around single versus multi-identity agents. Non suggested that successful designs map onto human nature, noting that managing "40 agents is quite a lot." He observed users creating "chief of staff agents" to consolidate management, indicating a natural inclination towards grouping activities. Claire, however, emphasized the practical, tactical considerations like data access, permissions, and segmented memory, concluding that the "use case matters so much in deciding this."
Regarding the necessary skill sets for product builders, both agreed on classic PM traits like user empathy and systems thinking, now applied to a new technological universe. Claire added "a relentlessness to like try it out and keep iterating" and enduring "a lot of pain," while Non highlighted the ability to learn from those feedback loops.
The discussion moved to the broader ecosystem and tools for productivity. Claire described ChatGPT as a platform offering layers of capability, from native functions to third-party plugins, and a "computer use" fallback. This layered approach ensures that if one tool doesn't work, users still have a path to accomplish their tasks. Non underscored the difference between a product that gets "99% there" and one that finishes the job, praising "computer use" for its ability to always get tasks "all the way done," even if it's slower. He likened the ideal user experience to being a "guest in your home," anticipating all needs.
Claire, new to working with research at OpenAI, shared her learning curve. She advised providing specific use cases, clear user goals, and, most importantly, learning to write "evals" to drive the iterative loop of model training.
On the topic of time horizons for planning, Claire maintained her ideal of "two to three months out," cautioning against longer-term predictions due to their inherent inaccuracy, while also stressing the need to build for the future, not just the present.
Closing with classic PM craft versus new table stakes, Non highlighted "onboarding" as significantly more critical now due to the "capability overhang" and the need for users to absorb new offerings. Claire added the increased importance of users understanding "privacy and their data" and ensuring predictable behavior from semi-autonomous agents. She also noted that classic PM skills still apply, but "the importance of testing things empirically and moving faster has only increased." Both agreed that "direct user relationships," or being a "dmable PM," has become increasingly vital, especially for understanding subtle user feedback like an agent having "attitude."
Finally, for predictions for 2027: Claire is "so bullish on voice," citing its natural interface and improved model capabilities. Non predicted the rise of "self-driving" products that can use themselves to solve the "empty input box problem" and provide gentle onboarding. The host, Claire, predicted "hardware that replaces carrying around our laptop," envisioning a "baby bot."