TimeSeek: Temporal Reliability of Agentic Forecasters (April 2026)
FreeBenchmark built from 150 regulated prediction markets evaluated at 5 lifecycle checkpoints — models are most competitive early and on high-uncertainty markets; search improves pooled accuracy but degrades 12% of conditions
About TimeSeek: Temporal Reliability of Agentic Forecasters (April 2026)
TimeSeek is a benchmark designed to study how the reliability of agentic LLM forecasters changes over a prediction market's lifecycle. It evaluates 10 frontier models on 150 CFTC-regulated Kalshi binary markets at five temporal checkpoints, with and without web search, resulting in 15,000 forecasts. Key findings reveal that models are most competitive early in a market's life and on high-uncertainty markets, but less competitive near resolution and on strong-consensus markets. Web search improves pooled Brier Skill Score (BSS) for every model overall, yet hurts in 12% of model-checkpoint pairs. Simple two-model ensembles reduce error without surpassing the market overall. These results motivate time-aware evaluation and selective-deference policies.
Key Features
Pros & Cons
- Uses real regulated (CFTC) prediction markets for ecological validity
- Large-scale evaluation with 150 markets and 15,000 forecasts
- Reveals important temporal dynamics often ignored in static benchmarks
- Includes both search-augmented and non-search settings
- Simple ensemble analysis provides practical insights
- Limited to binary (yes/no) prediction markets
- Models still underperform the market overall, especially near resolution
- Web search can degrade performance in some conditions (12% of cases)
- Academic workshop paper; not a production-ready tool
- Requires access to proprietary Kalshi market data (not publicly available)