---
title: "Superhuman AI Models Struggled to Predict Football Scores"
publisher: "Stockmark.IT"
author: "Stockmark.IT Website"
published: "2026-04-11T10:26:14+00:00"
modified: "2026-04-11T05:27:57+00:00"
date: 2026-04-11
canonical: "https://stockmark.it/superhuman-ai-models-struggled-to-predict-football-scores/"
category: "AI"
categories: ["AI", "Artificial intelligence", "Financial", "Gambling"]
image: "https://i0.wp.com/stockmark.it/wp-content/uploads/stencil.default-2023-10-27T045855.748.jpg?fit=1200%2C800&quality=89&ssl=1"
format: "news"
language: "en-GB"
---

# Superhuman AI Models Struggled to Predict Football Scores

**Published:** April 11, 2026
**Author:** Stockmark.IT Website
**Categories:** AI, Artificial intelligence, Financial, Gambling
**Featured image:** ![Two football players compete for the ball on a dark pitch; one wears a white kit with blue stripes and the number 16, the other wears an orange and black kit. Only their legs and the football are visible. @ stockmark.it](https://i0.wp.com/stockmark.it/wp-content/uploads/stencil.default-2023-10-27T045855.748.jpg?fit=1200%2C800&quality=89&ssl=1)

---

In a recent examination of artificial intelligence capabilities, eight AI models were put to the test to predict outcomes in a virtual re-creation of the 2023-24 Premier League season. The analysis conducted by General Reasoning, a London-based AI startup, sought to determine whether these advanced models could successfully navigate the complexities of football betting.

The AI models, which included leading technologies from Google, OpenAI, and Anthropic, were provided with a starting bank of £100,000 alongside extensive historical data. This data encompassed past results, line-ups, and public betting odds. The models were tasked with developing strategies to maximise returns while managing risk throughout the season, requiring them to adapt to evolving game outcomes and identify edges in betting markets.

Among the participants, Anthropic’s Claude Opus 4.6 emerged as the best performer, demonstrating an average loss of 11 per cent over three attempts. Google’s Gemini 3.1 pro achieved a notable gain of approximately £33,000 during its most successful outing, yet subsequently faced bankruptcy in another simulation. Only two models, Claude Opus 4.6 and GPT-5.4, managed to avoid financial ruin across the three testing scenarios.

The findings from General Reasoning reveal a concerning trend; all evaluated AI models ultimately incurred losses during the season, with many facing significant ruin. The company’s report highlights a critical observation that existing AI frameworks excel in well-defined tasks with clear objectives. Open-ended, long-term scenarios prove significantly more challenging for these models.

Analysts noted that the ability of these AI systems to maintain coherent decision-making over extended periods was lacking. Many failed to act on their analyses or adjust strategies according to changing circumstances. The sophistication of the strategies employed by the models was found to be considerably lower compared to human decision-making processes, indicating substantial potential for enhancement in future AI developments.

This assessment suggests a vital need to shift evaluations toward more complex environments that challenge long-horizon and sequential decision-making capabilities. As AI technology evolves, its application within unpredictable domains like sports betting may prove to be an ongoing area for research and improvement.

---

**Original URL:** https://stockmark.it/superhuman-ai-models-struggled-to-predict-football-scores/
*Created by [WP Markdown Endpoint](https://wpmarkdownendpoint.com/)*
