As the World Cup entered the knockout stage, AI forecasting moved from being an interesting pre-match exercise to something easier to evaluate against real outcomes. In Odaily’s review, the focus was not on which model sounded the most persuasive, wrote the longest tactical explanation, or listed the most variables before kickoff. The real question was narrower and more practical: which model was actually more useful as a reference when trying to anticipate what would happen next.

To test that question, Odaily compared a group of mainstream models under roughly the same prompt conditions before matches and then revisited the predictions after the games ended. The set included ChatGPT, Grok, Qwen, DeepSeek, Gemini, and Claude. Because knockout football tends to amplify variance, this stage offered a stronger stress test than the group phase. A model could not rely only on broad statements such as “the stronger side should advance.” It also had to deal with low-margin matches, the possibility of extra time and penalties, and the tendency for underdogs to reshape the game into something far less predictable.
The completed knockout matches already offered a volatile sample. Canada beat South Africa 1:0, Brazil edged Japan 2:1, Germany were dragged into a penalty shootout and eliminated by Paraguay, the Netherlands were knocked out by Morocco on penalties, and Belgium vs. Senegal turned into a 2:2 game before an extra-time reversal. In that kind of environment, model differences become clearer. Some systems are better at spotting upset paths, some are better at calibrating narrow scorelines, and some are stronger at explaining friction points without fully committing to a contrarian result.
Netherlands vs. Morocco was the clearest separator
The standout case in the review was the Netherlands vs. Morocco knockout match. Before kickoff, this was exactly the kind of tie where many forecasters could sound nuanced and still end up on the wrong side. On paper, the Netherlands held the edge in squad strength and overall depth. Many models acknowledged that Morocco would be difficult to break down, but still leaned toward the Dutch advancing.

What separated the top performers was their willingness to move beyond a generic “this will be close” framing and write out a more specific scenario. Gemini reportedly predicted a 1:1 draw in regular time and a Morocco win on penalties. The actual result followed that script remarkably closely: the match finished 1:1, and Morocco eliminated the Netherlands in the shootout by 3:2. That matters because it was not merely a directional call on the underdog. It was a call on match flow, on how the favorite could be contained, and on how the game might ultimately be decided.
DeepSeek also came close. Its pre-match view was that regulation time would most likely end 1:1 or 0:0, with the contest potentially stretching into extra time or penalties, and with Morocco having a realistic upset path through defense and counterattacking. Compared with models that recognized the risk but still defaulted to the favorite, DeepSeek and Gemini stood out for committing to a concrete upset narrative.
This distinction is important in prediction markets and similar use cases. Many models can articulate uncertainty. Far fewer are willing to translate that uncertainty into an explicit probability-weighted outcome that contradicts the stronger paper favorite. In Odaily’s framing, that is what gave DeepSeek and especially Gemini their highest-profile moment in the knockout-stage review.

Grok and Qwen performed better in favored-team scoreline calls
Not every useful forecast has to come from an upset. In several matches where the likely winner was relatively clear but the margin was not, Grok and Qwen were described as the most practical “scoreline-type” performers. Their strength was not in writing blockbuster underdog scripts. Instead, they were more effective at estimating whether the favorite would cruise or merely scrape through.
Canada vs. South Africa was one example. Before the match, most AI models favored Canada to advance, but there was a split over whether Canada would win comfortably or by a narrow margin. Grok reportedly projected a 1:0 Canada win, and Qwen also leaned toward a one-goal result. That turned out to be the actual outcome. The value in that forecast was not simply saying “Canada will win.” It was correctly identifying that the match would stay tight rather than turning into the sort of comfortable favorite victory that some pre-match narratives implied.
Brazil vs. Japan followed a similar pattern. Most models favored Brazil, but the key issue was whether Japan could keep the match alive deep into the game. Grok and Qwen both forecast a 2:1 result, which matched the final score as Brazil advanced by that margin. Their edge here was in recognizing that Japan would create enough resistance to prevent a one-sided result.

The same applied to Ivory Coast vs. Norway. With Erling Haaland in the picture, Norway were an understandable favorite in directional terms, but that did not automatically mean an easy win. Odaily noted that Grok and Qwen both leaned toward Norway winning 2:1, and the actual scoreline landed inside that exact script. Across these fixtures, the pattern was consistent: these models were especially useful when the task was not “find the upset” but “estimate how difficult the favorite’s path will be.”
ChatGPT and Claude offered stronger process analysis, but were more conservative on outcomes
In Odaily’s breakdown, ChatGPT and Claude were less notable for nailing exact scores and more notable for identifying where a match could become uncomfortable for the favorite. They tended to provide fuller explanations of tactical friction, pressing intensity, defensive structure, physical duels, wing play, transition risk, and the chance that a game could become slower, uglier, or harder than public expectation suggested.
Brazil vs. Japan illustrates that pattern. ChatGPT favored Brazil to advance, but it did not frame the match as a straightforward domination. Instead, it reportedly highlighted Japan’s pressing, work rate, and discipline as factors that could make the game difficult for Brazil, including scenarios in which Japan might score first or level the score. That kind of forecast may not win headlines the same way an upset call does, but it can still be useful if the goal is to understand match resistance rather than just choose a winner.
The same logic applied to Ivory Coast vs. Norway. ChatGPT leaned Norway, yet also emphasized that Ivory Coast’s physicality, wing pressure, and transition threat could keep the game competitive. In England vs. DR Congo, it again avoided the simplistic “England will roll over the opponent” narrative, suggesting instead that the game might become cagey because DR Congo could sit deep and slow the rhythm. England did advance, but not in an especially comfortable manner.

This gives ChatGPT and Claude a distinctive profile. They can often point to the correct trouble spots in advance, helping readers understand why a favorite may underperform relative to public expectation. But when it comes to moving from “this could get messy” to “the underdog may actually go through,” they appear more hesitant than models like Gemini or DeepSeek in the most memorable upset case of the round.
Germany vs. Paraguay exposed a broad failure across models
If earlier matches revealed the strengths of different systems, Germany vs. Paraguay revealed a weakness shared by almost all of them. According to Odaily, every major model leaned toward Germany before the match. ChatGPT, Grok, Qwen, Gemini, and Claude all backed Germany, and many of the score predictions clustered around 2:0, 3:0, or 3:1. The reasoning was also consistent: Germany had the stronger squad on paper, better depth, and greater attacking firepower.
But the match did not obey that hierarchy. Paraguay were able to drag the game into a muddy, low-clarity contest. Germany failed to settle it in regular time, failed again in extra time, and were eventually eliminated in the penalty shootout. The result was more than just one bad call. It highlighted a recurring modeling bias: over-reliance on paper strength and broad reputational priors in situations where knockout football allows the weaker side to compress the game, reduce available space, and increase randomness through game state management.

That matters because it suggests a limit to AI forecasting in tournament settings. A model may be very good at summarizing objective strength differences and historical tendencies, but still underestimate an underdog’s ability to alter match conditions. The Germany-Paraguay result is the clearest example in this review of models failing not because they lacked information, but because they weighted that information in a way that was too favorable to the established power.
The practical takeaway: model selection depends on the use case
After reviewing the completed knockout fixtures, Odaily’s conclusion was less about naming a single best model and more about identifying which models fit which tasks. That is a useful distinction for prediction-oriented users. In practice, one model may be better for upset detection, another for precise scoreline estimation, and another for understanding tactical risk without overcommitting to a contrarian result.
DeepSeek and Gemini emerged as the strongest performers in high-variance fixtures. Their edge was not simply that they recognized a match would be close. It was that they were willing to convert that uncertainty into an explicit upset scenario, including possibilities such as a draw in regulation, extra time, or a penalty shootout. The Netherlands vs. Morocco match, especially Gemini’s reported “1:1 plus Morocco on penalties” forecast, remains the most striking example in the sample described by Odaily.

Grok and Qwen looked more like reliable scoreline operators in favorite-heavy fixtures. They were particularly effective in matches involving Canada, Brazil, Norway, and France, where the challenge was to judge whether the stronger side would dominate or merely survive. Their limitation, at least in this review, was that they still leaned more heavily toward the traditional favorites in matches such as Germany’s and the Netherlands’.
ChatGPT and Claude, meanwhile, appeared better suited to explanatory use. They can be valuable when the reader wants to understand why a match may be tighter than expected, where tactical stress points may emerge, and why extra time or a low-scoring grind is plausible. But in the context of outright upset prediction, they were portrayed as more cautious.
The broader implication is straightforward. Asking which AI “knows football best” is less useful than asking what kind of forecasting job needs to be done. For users focused on identifying upset paths, DeepSeek and Gemini looked stronger in the reviewed sample. For those prioritizing scoreline proximity in favorite-led matches, Grok and Qwen appeared more useful. For those trying to understand the likely shape and resistance points of a game, ChatGPT and Claude remained relevant despite being less decisive on underdog outcomes. Source: Odaily.

