基本信息

要点解读

这是什么

在2026年世界杯的赛期(39 天)里,六款具备扩展思考和服务器端搜索功能的前沿大语言模型在每场比赛开球前填写七市场预测卡,涵盖104场小组赛、12支小组冠军以及赛前全赛预测。提问时答案尚未出现,评估因此实现了无泄漏的前瞻性,冻结档案保留了4 494条已评分预测。

用在哪里

适用于需要验证大模型在真实、尚未发生事件上的预测能力的研究场景,特别是体育比赛、金融走势或舆论热点等前瞻预测任务;也可以作为构建和发布类似公开基准的参考。

可以推断的

推测:在实时预测任务中,模型通常倾向于押注人气最高的选项,导致对平局或低概率结果的预测偏少。
推测:若在此类任务中加入投票机制,可能难以显著提升整体准确率,因为模型的错误模式高度相似。

来源摘要/节选

Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs – all with extended thinking and native server-side web search – were asked before every kickoff, one match at a time, to fill in a seven-market prediction card for all 104 matches, plus 12 group winners and a pre-tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage-free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker’s favourite – which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under-commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.

来源说明

当前保存的是来源摘要,不代表论文全文。请以原始来源为准。

「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。