In the famous 1997 match, IBM’s Deep Blue defeated Russian grandmaster and former world chess champion Garry Kasparov. Google’s new tournament continues this tradition, now using language models. On 5 August, the company launched a two-day chess tournament on the Kaggle Gaming Arena platform, in which leading artificial intelligence models competed against each other in a live test of their cognitive abilities. The company launched the tournament after Elon Musk claimed that his chatbot Grok demonstrated outstanding thinking abilities.
Viewers could observe the reasoning behind each model’s moves. According to Google, this transparency enables one to assess whether the models are actually thinking through problems or merely imitating the training data.
Kaggle Gaming Arena for testing AI agents
Google held a chess tournament between artificial intelligence language models on the new Kaggle Gaming Arena platform. The company designed this space for testing universal AI agents in real time in a competitive environment. The first tournament was a public stress test designed to evaluate the logic, strategy, and decision-making abilities of the models.
The competition took the form of daily chess matches between six leading language models: ChatGPT, Gemini, Claude, Grok, DeepSeek and Kimi. The tournament featured the following pairs: o4-mini and DeepSeek-R1 from OpenAI, Gemini 2.5 Pro and Claude Opus 4, Kimi K2 Instruct from Moonshot AI and o3 from OpenAI, as well as Grok 4 and Gemini 2.5 Flash.
The winner was OpenAI’s o3 model, which defeated Grok 4 in the final exhibition match. In the game for third place, Gemini 2.5 Pro defeated o4-mini with a score of 3.5:0.5, winning the bronze medal.
Google has announced plans to expand Kaggle Gaming Arena beyond chess by adding new games and challenges. According to the company, such competitions allow for the assessment of AI’s actual capabilities, distinguishing between genuine thinking and imitation, and identifying the strengths of the models.
Google DeepMind co-founder and CEO Demis Hassabis noted that games have always served as an effective platform for testing AI. He recalled the successes of AlphaGo and AlphaZero and confirmed his intention to develop Arena as a platform for evaluating AI thinking in a dynamic environment.
The first format for evaluating AI strategy
Unlike standard tests, this format demonstrates AI strategy by evaluating how models think, adapt and recover under pressure, according to a statement from Google. The company intends to identify differences in reasoning abilities that other tests cannot detect.
The competition format is as follows:
- Each round consists of a series of four matches, with the winners advancing to the next round in a playoff system.
- The two best models compete in the final match for the gold medal.
The matches are broadcast live on YouTube.
The competition is a logical continuation of other Google game tests to test AI reasoning abilities. Previously, the company tested agents in Atari, AlphaGo and AlphaStar games.
‘Applications are ranked using a Bayesian skill rating system that is regularly updated, allowing for a thorough long-term assessment,’ Google reports. The Bayesian system utilises probability to update a player’s skill rating over time, based on their performance relative to other competitors.
Chess in the Esports World Cup 2025 programme
Following the chess trend in esports, the European organisation Team Secret signed a contract with Grandmaster Anish Giri. He became the brand representative at the Esports World Cup 2025, where chess was included as a primary discipline for the first time. Team Secret was the latest esports organisation to sign agreements with chess players ahead of the championship.
In addition to Team Secret, teams such as Team Liquid, Wolves Esports and Team Falcons have signed contracts with top-class chess players. All of them competed for the Esports World Cup 2025 prize pool of $70 million, distributed among 25 tournaments and club competitions. The prize pool for chess players amounted to $1.5 million.
This article was first published in Russian on 11 August 2025.