EN Submit a tool
Paper

WorldCupArena: A Fine-Grained Evaluation Benchmark for Football Prediction with Large Language Models and Deep Research Agents

Published: Source: HuggingFace Daily Papers (Community Hot Papers)

ShareXFacebookTelegramWhatsApp

Shanghai Jiao Tong University and other institutions released WorldCupArena, a dynamic evaluation benchmark for large language models and deep research agents, first assessing all 104 matches of the 2026 FIFA World Cup across 13 systems. Results show that models with similar outcome accuracy differ significantly in fine-grained predictions such as scores and players; the best system only has a clear advantage over betting markets and fan baselines in score proximity. Code, prompts, predictions, and evaluation scripts are open-sourced.

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category
· X: Elvis Saravia (@omarsar0, DAIR.AI)
· Hacker News Hot (buzzing.cc Chinese translation)
· The Decoder: AI News (RSS)
· X: Rohan Paul (@rohanpaul_ai)
· X: Rohan Paul (@rohanpaul_ai)