Methodology: Identical Questions Before Each Knockout Match, Results Compared Post-Match
Starting from the first knockout round, Odaily Planet Daily asked the same set of pre-match questions to different AI models — including ChatGPT, Grok, Qwen (Tongyi Qianwen), DeepSeek, Gemini, and Claude — and then compared their predictions with actual results. Completed knockout matches so far include: Canada 1-0 South Africa, Brazil 2-1 Japan, Germany eliminated by Paraguay after a penalty shootout, Netherlands eliminated by Morocco on penalties, and Belgium's comeback victory after 2-2 draw with Senegal. These games fully demonstrate the unpredictability of knockout football.


Star Performers: DeepSeek and Gemini Wrote the Upset Script in Advance
The most memorable case was Netherlands vs Morocco. Most models acknowledged Morocco would be tough but still favored the Netherlands. DeepSeek and Gemini went further. Gemini predicted a 1-1 draw in regular time and Morocco winning on penalties — exactly what happened (1-1, Morocco 3-2 penalties). DeepSeek also judged a 1-1 or 0-0 draw in regular time, leaning toward a Morocco upset via defensive counter-attacks. This match elevated both models' credibility, especially Gemini, which seemed to have previewed the match script.

Score-Savvy Players: Grok and Qwen Hit Exact Scores in Favorable Matches
Grok and Qwen excelled in matches with a clear favorite. For Canada vs South Africa, they predicted a 1-0 win for Canada, matching the actual narrow victory. Brazil vs Japan: both predicted 2-1, correct. Ivory Coast vs Norway: both predicted Norway 2-1, also correct. These models may not be best at catching upsets, but they accurately judge whether a strong team will cruise or win narrowly, delivering precise scorelines in multiple matches.

Analytical Players: ChatGPT and Claude Explain Risks but Avoid Upset Calls
ChatGPT did not predict the Morocco upset nor consistently nail scorelines, but it consistently highlighted match difficulty. For Brazil vs Japan, it noted Japan's pressing, running, and discipline would trouble Brazil. For England vs DR Congo, it predicted a dull game with DR Congo's low block. ChatGPT is strong at describing match resistance but lacks the decisiveness to call an upset. Claude behaves similarly — thorough analysis, usually correct direction, but ultimately siding with the favorite.

Collective Failure: Germany vs Paraguay — All Models Wrong
If previous rounds showed individual strengths, Germany vs Paraguay was a collective failure. Every AI model — ChatGPT, Grok, Qwen, Gemini, Claude — favored Germany, predicting scores like 2-0, 3-0, or 3-1. Their reasoning was identical: Germany's superior squad, depth, and attack. But Germany failed to break down Paraguay in regular time, failed in extra time, and lost on penalties. The models underestimated Paraguay's ability to drag the game into a grind.

Conclusion: Different Models Suit Different Scenarios
Based on the completed knockout matches, DeepSeek and Gemini are most valuable for upset predictions, daring to script a shock result. Grok and Qwen are reliable for confirming how a favorite will win and hitting exact scores. ChatGPT and Claude provide thorough analysis but tend to stay conservative. Prediction market users should match the AI's style to their strategy: for catching upsets, lean on DeepSeek/Gemini; for confirming favorite outcomes, use Grok/Qwen. No single model is universally best — choose according to your risk appetite and prediction focus.


