Benchmarking Forecasting Agents with Point-in-Time Search
A 100-question forecasting benchmark made possible by Arise PiT web search.
A 100-question forecasting benchmark made possible by Arise PiT web search.
Built for billion-scale vector search, with fast queries and continuous updates.
Search for agents still returns what it returned for humans: a ranked list of passages. Our paper argues it should return something an agent can work in: a bounded slice of the corpus, with tools to explore it.