As the World Cup moved into the knockout stage, AI forecasting became easier to evaluate in practice. Odaily tested multiple models with roughly the same pre-match prompts before each game, then compared their answers with the actual results afterward. The key question was not which model sounded the most detailed, but which one produced conclusions that were genuinely useful once the match had been played.

That distinction matters because many models can sound convincing in sports prediction. Some focus on squad value, others cite group-stage data, injuries, tactics, and historical matchup patterns. A few go further and offer exact scorelines, extra-time scenarios, or penalty scripts. On the surface, ChatGPT, Grok, Qwen, DeepSeek, Gemini, and Claude can all appear “knowledgeable.” But knockout football quickly exposes the difference between fluent analysis and usable judgment.
The sample of completed knockout matches already includes several high-variance outcomes: Canada edged South Africa 1:0, Brazil survived Japan 2:1, Germany were dragged into a shootout and eliminated by Paraguay, the Netherlands were knocked out by Morocco on penalties, and Belgium vs. Senegal escalated into a 2:2 draw before an extra-time turnaround. These matches made it possible to see where each model was strong, and where each one remained overly dependent on favorite bias.
DeepSeek and Gemini stood out by committing to upset scenarios
The clearest high-value case so far was Netherlands vs. Morocco. Before kickoff, this was an easy fixture to misread. On paper, the Netherlands had the stronger squad and the more complete roster, and many models acknowledged that Morocco would be difficult without actually moving their final call away from the Dutch side. DeepSeek and Gemini differentiated themselves because they did not stop at saying the match would be “tight.” They extended that logic into a full match script.

Gemini was especially impressive. It projected a 1:1 result in regular time and then a Moroccan win in the penalty shootout. That is almost exactly how the match unfolded: the game finished 1:1, and Morocco ultimately advanced by winning the shootout 3:2. This was not just a correct directional call. It captured the mechanism of the upset, including how the match would be dragged into penalties and who would emerge from them.
DeepSeek was also very close. It suggested that regular time would likely end 1:1 or 0:0, with the game potentially extending into extra time or even penalties, while leaning toward Morocco advancing through defensive resilience and counterattacking efficiency. In other words, DeepSeek and Gemini were not merely better at recognizing uncertainty. They were more willing to convert uncertainty into a concrete upset conclusion.
That willingness matters. Many models can identify risk around a favorite, but far fewer will turn that into a forecast that explicitly backs the underdog. In this case, DeepSeek and Gemini gained visibility precisely because they crossed that line and were then validated by the result.
Grok and Qwen looked strongest when forecasting favorites and score ranges
If DeepSeek and Gemini were the standout upset readers, Grok and Qwen looked more like reliable scoreline specialists in matches where the overall direction was relatively clear. Their edge was not necessarily in writing dramatic underdog scripts, but in distinguishing between easy wins and narrow escapes for stronger teams.

South Africa vs. Canada is one of the best examples. Most AI models favored Canada before the game, but the real question was whether Canada would win comfortably or only by a slim margin. Grok projected a 1:0 win for Canada, and Qwen also pointed toward a one-goal victory. The actual result matched that reading: Canada advanced, but only by a single goal rather than by the lopsided margin some might have expected.
Brazil vs. Japan followed a similar pattern. Most models agreed that Brazil were the stronger side, but the key issue was whether Japan could keep the match competitive. Grok and Qwen both forecast 2:1, and the game indeed ended with Brazil winning 2:1. What they got right was not just “Brazil will win.” They also identified that Japan could make life difficult enough for the scoreline to stay close.
The same pattern appeared in Ivory Coast vs. Norway. Norway’s direction of advancement was not especially hard to justify given Erling Haaland’s presence, but Ivory Coast’s physicality and wing pressure suggested the match would not become one-sided. Grok and Qwen both landed on Norway 2:1, which also aligned closely with the outcome. Across these fixtures, their strength was granularity: they read the shape of favorite-led matches with more precision than simply calling the winner.
That gives both models a specific use case. They may not be the best tools for identifying major knockout upsets, but they are useful when the task is to estimate whether a favorite is likely to cruise, struggle, or survive by a narrow margin.

ChatGPT and Claude were better at process analysis than final upset calls
ChatGPT and Claude behaved more like analytical assistants than aggressive forecasters. They did not produce the headline moment that Gemini delivered with Morocco over the Netherlands, nor did they string together multiple exact score hits like Grok and Qwen. Their comparative advantage was different: they were often good at identifying where the resistance in a match would come from.
Brazil vs. Japan illustrates this well. ChatGPT backed Brazil to advance, but it did not describe the game as a routine mismatch. Instead, it emphasized that Japan’s pressing, movement, and discipline could make Brazil uncomfortable, even creating scenarios in which Japan might score first or equalize. In Ivory Coast vs. Norway, it similarly leaned toward Norway while warning that Ivory Coast’s physical duels, wing attacks, and transition game would create complications.
The England vs. DR Congo knockout match showed the same tendency. Rather than projecting an easy England blowout, ChatGPT suggested the match could become slow and muted, with DR Congo using a low block to drag the tempo down. England did advance, but not in an easy fashion. That outcome was broadly consistent with the model’s reading of the match process.

This is why ChatGPT, and to a degree Claude, are more useful for understanding why a game could become awkward, tense, or low quality from the favorite’s perspective. They tend to explain tactical friction reasonably well. The limitation is that they often stop short of making the bolder final leap. In matches where they recognize extra-time or penalty risk, they still frequently end up siding with the traditional powerhouse.
That same pattern appeared in the Netherlands vs. Morocco case. Even when some models could see the danger of a drawn-out match, they still leaned toward the stronger-name team. The analysis was not entirely wrong, but the final forecast lacked conviction.
Germany vs. Paraguay was the clearest collective failure
If the earlier matches highlighted the different specialties of each model, Germany vs. Paraguay showed how they can all fail together. Before the match, virtually every major AI model sided with Germany. ChatGPT, Grok, Qwen, Gemini, and Claude all placed Germany on the advancing side, with most score predictions clustering around 2:0, 3:0, or 3:1.
The logic was also highly uniform. Germany were judged to have the stronger squad on paper, greater depth, and more attacking firepower. Those are all reasonable inputs in a standard pre-match framework. But the result exposed a common blind spot: the models underestimated Paraguay’s ability to drag the game into an ugly, low-fluidity contest in which the stronger team could not easily convert structural advantages into goals.

Germany failed to finish the job in regular time, failed again in extra time, and were ultimately eliminated after a penalty shootout. From a forecasting perspective, this was not just one model getting it wrong. It was a cross-model miss driven by similar assumptions. When many systems rely on squad quality, reputation, and attacking depth in the same way, they become vulnerable to the same type of upset.
This matters because knockout football is exactly where those assumptions are most likely to break. Defensive discipline, pace control, fatigue, game-state management, and penalty variance can all overwhelm paper advantages. Germany vs. Paraguay became the strongest reminder that “better team” logic is not enough if the model cannot price in how the weaker side can weaponize match structure.
Different models now have clearly different use cases
Based on the completed knockout matches so far, a clearer map of model strengths is emerging. DeepSeek and Gemini have delivered the most memorable upside because they were willing to go beyond cautious uncertainty and actually write the upset script. In the Netherlands vs. Morocco match, that difference had real value: they did not just say the game would be competitive, they identified Morocco as a live knockout threat and treated penalties as part of the likely path.
Grok and Qwen, by contrast, look more dependable in favorite-driven matches where scoreline placement matters. Their performance in Canada, Brazil, Norway, and France-related fixtures suggests they can be useful for estimating how cleanly a stronger team might progress. They are not necessarily the best at spotting the biggest bracket-breaking surprise, but they often read the margin correctly.

ChatGPT and Claude fit a different role. They are stronger when the objective is to understand the match itself: where the tactical stress points are, how an underdog could slow the game down, and why a nominally superior team might struggle. For readers who want context and pre-match interpretation, that is valuable. For readers who specifically want upset conviction, it is less sufficient.
The practical takeaway is not that one model simply “knows football best.” It is that different systems are useful for different forecasting tasks. For upset hunting, DeepSeek and Gemini have shown the most willingness to commit. For scoreline calibration in favorite-heavy matches, Grok and Qwen have looked sharper. For reading match friction and process risk, ChatGPT and Claude remain useful, though more conservative.
That division is likely more informative than any single leaderboard. In real-world use, the better question is not which model sounds smartest, but which one is best suited to the decision you are trying to make.

