Knockout Stage Recap: Upsets Dominate, AI Prediction Divergence Emerges
The World Cup knockout stage has delivered a series of surprises: Canada edged South Africa 1-0, Brazil narrowly beat Japan 2-1, Germany was eliminated by Paraguay after a penalty shootout, and the Netherlands fell to Morocco on penalties. Belgium vs Senegal went to extra time after a 2-2 draw, amplifying the tournament's unpredictability. Against this backdrop, Odaily Planet Daily asked six AI models—ChatGPT, Grok, Qwen, DeepSeek, Gemini, and Claude—identical pre-match questions for each knockout game, then retrospected the actual outcomes to identify which models merely sounded convincing and which truly captured the match trajectory. Preliminary results reveal clear differentiation among the models.


DeepSeek and Gemini: The 'Scriptwriters' of Upset Predictions
The most striking performance came from DeepSeek and Gemini on the Netherlands vs Morocco match. Most models recognized Morocco's threat but ultimately favored the Netherlands. Gemini directly forecast a 1-1 draw in regulation and a Morocco penalty shootout win; the actual match ended 1-1, and Morocco advanced 3-2 on penalties. Not only did it predict the upset, but it also spelled out the entire scenario—regulation draw, penalties, winner. DeepSeek predicted a 1-1 or 0-0 regulation scoreline and leaned toward Morocco's defensive upset. Both models demonstrated the ability to foresee not just the direction but also the match narrative—extra time, penalty drama—while others either hesitated or backed the favorite.

Grok and Qwen: Precision Score Predictors
In South Africa vs Canada, Grok and Qwen both predicted a 1-0 win for Canada; Canada indeed scored only one goal to advance. Brazil vs Japan saw both models forecast a 2-1 win for Brazil, matching the exact result. For Côte d'Ivoire vs Norway, Grok and Qwen called Norway 2-1, which came true. These two models excel in hot-favorite matches: they not only get the winner right but also deliver accurate scorelines. They are less adept at catching blockbuster upsets but shine in predicting whether a strong team will cruise or scrape through—useful for Canada, Brazil, Norway, and France. In essence, they know when a favorite will win comfortably vs. barely.

ChatGPT and Claude: Analytical Depth Without Betting the Farm, and the Collective Failure
ChatGPT did not call the Morocco penalty upset nor consistently predict exact scores like Grok/Qwen. Its strength lies in identifying match resistance early: in Brazil vs Japan, it noted Japan's pressing and discipline would make Brazil uncomfortable; in Côte d'Ivoire vs Norway, it highlighted physicality and wing threat; in England vs DR Congo, it foresaw a tight game with low-block defense. ChatGPT and similar-model Claude serve as 'analytical players'—their reasoning is thorough, direction mostly correct, and they often signal extra-time risk but stop short of picking the upset. For Netherlands vs Morocco, they saw the danger of penalties but still sided with the Dutch.

However, Germany vs Paraguay became a spectacular collective failure. Every model backed Germany, predicting scores like 2-0, 3-0, or 3-1, citing superior squad depth and firepower. None foresaw Paraguay's ability to drag the game into a grinding contest; Germany failed to break the deadlock in regulation or extra time and lost on penalties. This match exposed a common blind spot across models: underestimating the power of defensive disruption and tournament pressure.

In summary, DeepSeek and Gemini lead in upset detection; Grok and Qwen dominate exact score predictions for favorites; ChatGPT and Claude are best for understanding game dynamics but lack decisive edge. Prediction market users can choose their reference model based on whether they seek upsets, precise scorelines, or process insight.


