Post by Samaya AI
5,835 followers
Query understanding is not the glamorous part of agentic research, but it is often where the answer is won or lost. At Samaya AI, we trained a specialist query understanding model that outperforms Claude Sonnet 4.6 and Gemini Flash, with lower latency. Before an AI agent can reason well, it has to retrieve the right evidence. That is especially hard in financial research. Even a simple query about fetching the right earnings call or filing requires the system to understand crucial attributes such as the company in question and the reporting period. This first step is called query understanding, and it sits on the critical path of nearly every agentic research workflow. We trained a specialist model using a two-stage recipe that combines Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL). The downstream impact was significant: š¹ +1.5 percentage points improvement in rubric satisfaction š¹ 4% fewer tool calls š¹ 23% lower median time to answer generation The big takeaway: specialist models can outperform generalized frontier APIs on critical workflow components, improving quality, latency, and cost at the same time. š Blog in comments. Special thanks to: Thejas Venkatesh, O. Ozan Koyluoglu, Ashwin Paranjape, Yuhao Zhang, Richard Diehl Martinez, Vishank Bhatia, Kyle C., and Paddu Raghavan