Post by Perplexity

1,717,694 followers

We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer. https://lnkd.in/gfCh95F3 Wide-and-deep research requires two capabilities. 1) Agents must search broadly enough to find all qualifying entities 2) Agents must investigate deep enough to support every claim with evidence. WANDR represents these requirements as hierarchical, independently verifiable records. It consists of 500 research tasks that require 170,495 source-backed records across three tiers of difficulty. WANDR tests how well agents discover large sets of entities and verify specific facts about each one. It provides a dense, interpretable eval signal that reveals whether an agent fails, and where. The pipeline also doubles as a semi-automated factory for training data. WANDR is constructed from de-identified production use cases covering people’s day-to-day research tasks like competitive research, due diligence, literature review, market analysis, product comparison, talent sourcing, and more. Instead of grading against a gold solution, WANDR re-fetches every cited page and checks each claim against the underlying evidence. This allows certain tasks to contain time-varying facts. Soft scores give partial credit. Hard scores require every component for a member to be complete and correct. The benchmark tasks and evaluation harness are available at: https://lnkd.in/gvaYBDdG

Post contentPost contentPost contentPost contentPost content