TimeSeek: Temporal Reliability of Agentic Forecasters (April 2026) logo

TimeSeek: Temporal Reliability of Agentic Forecasters (April 2026)

Free

Benchmark built from 150 regulated prediction markets evaluated at 5 lifecycle checkpoints — models are most competitive early and on high-uncertainty markets; search improves pooled accuracy but degrades 12% of conditions

FreeFree tier
Type
Open Source

About TimeSeek: Temporal Reliability of Agentic Forecasters (April 2026)

TimeSeek is a benchmark designed to study how the reliability of agentic LLM forecasters changes over a prediction market's lifecycle. It evaluates 10 frontier models on 150 CFTC-regulated Kalshi binary markets at five temporal checkpoints, with and without web search, resulting in 15,000 forecasts. Key findings reveal that models are most competitive early in a market's life and on high-uncertainty markets, but less competitive near resolution and on strong-consensus markets. Web search improves pooled Brier Skill Score (BSS) for every model overall, yet hurts in 12% of model-checkpoint pairs. Simple two-model ensembles reduce error without surpassing the market overall. These results motivate time-aware evaluation and selective-deference policies.

Key Features

Evaluates 10 frontier LLM models on 150 CFTC-regulated Kalshi binary markets
Five temporal checkpoints across market lifecycle
15,000 total forecasts with and without web search
Measures Brier Skill Score (BSS) and pooled accuracy
Analyzes two-model ensemble performance
Identifies temporal patterns: models strongest early and on high-uncertainty markets
Reveals non-uniform benefit of web search (improves overall but hurts in 12% of cases)

Pros & Cons

Pros
  • Uses real regulated (CFTC) prediction markets for ecological validity
  • Large-scale evaluation with 150 markets and 15,000 forecasts
  • Reveals important temporal dynamics often ignored in static benchmarks
  • Includes both search-augmented and non-search settings
  • Simple ensemble analysis provides practical insights
Cons
  • Limited to binary (yes/no) prediction markets
  • Models still underperform the market overall, especially near resolution
  • Web search can degrade performance in some conditions (12% of cases)
  • Academic workshop paper; not a production-ready tool
  • Requires access to proprietary Kalshi market data (not publicly available)

Best For

Evaluating reliability of AI forecasters over time in prediction marketsDeveloping time-aware evaluation protocols for agentic LLMsInforming selective-deference policies between AI forecasts and market outcomesResearch on when to trust model predictions vs. human consensusBenchmarking frontier LLMs on temporal forecasting tasks

FAQ

What is TimeSeek?
TimeSeek is a benchmark introduced in a workshop paper for studying how the reliability of agentic LLM forecasters changes over a prediction market's lifecycle.
Which models were evaluated?
The benchmark evaluates 10 frontier LLM models on 150 CFTC-regulated Kalshi binary markets.
How many forecasts were made?
A total of 15,000 forecasts were made, covering five temporal checkpoints, with and without web search.
What are the main findings?
Models are most competitive early in a market's life and on high-uncertainty markets, but less competitive near resolution and on strong-consensus markets. Web search improves overall accuracy but hurts in 12% of model-checkpoint pairs. Simple two-model ensembles reduce error without surpassing the market.
What metric is used?
The primary metric is the Brier Skill Score (BSS), along with pooled accuracy measures.