Post by Samaya AI
5,833 followers
We are launching FrontierFinance, the world’s hardest benchmark for measuring frontier financial intelligence of agentic systems. Built by finance and AI experts at Samaya AI for realistic and reproducible comparison of frontier AI systems, it has 220 diverse queries and a total of 11,543 expert-crafted rubrics, making it the largest open benchmark of its kind. FrontierFinance differentiates itself on diversity and hardness. Unlike existing benchmarks which largely focus on financial data extraction, FrontierFinance covers a diverse range of use cases essential to an investor’s workflow and are harder to evaluate: from screening to research to analysis to monitoring. The benchmark tests for an agent’s ability to exhaustively find information, perform numerical analysis and most importantly synthesize information using the taste and judgement of professional investors. This makes FrontierFinance hardest among existing finance benchmarks with a ~50% pass rate for the best system currently. Samaya's agentic system outperforms frontier models with 50.8% on the same benchmark, at roughly a 4x lower inference cost than Fable 5 and ~2.7x lower than Opus and GPT. Among frontier models, Claude Fable 5 is the best-performing on FrontierFinance, scoring 49.2%, followed by Claude Opus 4.8 at 45% and GPT 5.5 at 43.5%. The data, code and analysis for FrontierFinance is linked in the comments below. Samaya’s mission is to take us from information to conviction and FrontierFinance takes a large step in that direction. It comes from a larger internal set of almost 5000 Agent examples, and we plan to release subsequent harder benchmarks and technical reports.