Will any xAI Grok model score at least 30% on the FrontierMath Exam?

Predicted at2026-02-16 07:09 UTC
Prediction83.7%
Market (at prediction)74.5%
Market (live)

Analysis

Agent 3 (8%) misread the resolution date as Feb 28 instead of June 30, creating a misleading outlier. Excluding it, agents cluster 58-82% with mean ~68%. The market at 74.5% is supported by sibling market structure (40% threshold at 72% implies 30% should be higher) and the recent price jump suggesting new information. Grok 4 Heavy already at ~26% means only modest improvement needed. However, the key risk is the leaderboard requirement - even if a Grok model can score 30%, Epoch AI must evaluate and publish it. My estimate of 72% is close to market price (edge ~2.5%), far below the 5% threshold for a trade. Agent confidence is generally low (0.35-0.65), reflecting genuine uncertainty about xAI's release timeline and Epoch's evaluation schedule.


View on Polymarket

Timestamped via OpenTimestamps · Block 957898

SHA-256: c1f2abdee77384c8733dcefcfe5e8e0000b806f08e1376681355c3dacc9c6a14

Download content.md · Download .ots proof

Verification needs both files. Download them, confirm the text’s hash with sha256sum content.md, then check the proof against it — drag both into opentimestamps.org or run ots verify content.md.ots locally. The .ots proof anchors that SHA-256 into a Bitcoin block, showing this analysis existed before the outcome was known.

This page is for informational and research purposes only. Nothing here constitutes financial advice. Do not make investment decisions based on these predictions.