Post by Turing
2,158,158 followers
A model can pass every benchmark you throw at it and still fall apart the moment it's deployed as an agent. Why? Because the real test isn't the eval. It's the messy, multi-step, tool-using work agents actually get asked to do once they're live. Join Fiddler AI for AI Explained, an AMA series with the people building at the edge of agentic AI. Featured speakers: Juhi Parekh, GM of Key Frontier AGI Accounts at Turing. She helps frontier AI labs test and train their models and sees firsthand how they perform once companies put them to work. Juhi previously held product roles at Apple, Amazon, Niantic Spatial, and Samsung Research US. Buddy Brewer, VP of Product Management at Fiddler AI, where he leads product and design. He brings more than 20 years of experience building monitoring and observability products, including serving as Chief Product Officer at Mezmo, GVP and GM at New Relic, product executive roles at Akamai, and co-founding front-end monitoring startup Log-Normal, which was acquired in 2012. They'll get into: -How teams design test problems hard enough to break a model, so they can fix it before it ships -What's actually changed: better reasoning, more reliable tool use, and staying coherent through long, multi-step tasks -Why an agent that aces every benchmark can still fail at basics like generating a document or reading an image correctly Training the model is one problem. Building an agent that holds up in the real world is another. This conversation is about closing that gap. Save your spot for the next AI Explained AMA: https://lnkd.in/g9dB-XZV