Why Tejal Patwardhan stopped underestimating models – Episode 21
Why Tejal Patwardhan stopped underestimating models – Episode 21


Author: OpenAI – Duration: 00:44:23
Old tests become too easy. Tejal Patwardhan leads OpenAI's Frontier Assessment team, which finds new ways to measure and predict progress as models become more capable. She and host Andrew Mayne discuss the importance of assessments for research, how benchmarks can be broken or manipulated, and what models should then be judged on. Chapters 00:00:24 Growing up with OpenAI 00:03:10 Why reasoning changed everything 00:06:28 What made o1 surprising 00:11:20 Why old benchmarks stopped working 00:14:45 What makes a good benchmark 00:17:35 Why assessments are increasingly difficult 00:22:09 Measuring voice and vision models 00:24:48 Testing models on real science 00:33:23 How OpenAI tracks frontier progress 00:40:47 What AI means for work






