3 ms·Wandr Benchmark: Evaluating Research Agents That Must Search Wide and Deep1 points by tagawa 3mo agotagawa 3mo agoRepo with benchmark tasks, evaluation harness, tech report: https://github.com/perplexityai/wandr https://github.com/perplexityai/wandr