Tech and AIAnthropic used Pokémon to benchmark its newest AI model

Anthropic used Pokémon to benchmark its newest AI model

-


Anthropic used Pokémon to benchmark its newest AI model. Yes, really.

In a blog post published Monday, Anthropic said that it tested its latest model, Claude 3.7 Sonnet, on the Game Boy classic Pokémon Red. The company equipped the model with basic memory, screen pixel input, and function calls to press buttons and navigate around the screen, allowing it to play Pokémon continuously.

A unique feature of Claude 3.7 Sonnet is its ability to engage in “extended thinking.” Like OpenAI’s o3-mini and DeepSeek’s R1, Claude 3.7 Sonnet can “reason” through challenging problems by applying more computing — and taking more time.

That came in handy in Pokémon Red, apparently.

Compared to a previous version of Claude, Claude 3.0 Sonnet, which failed to leave the house in Pallet Town where the story begins, Claude 3.7 Sonnet successfully battled three Pokémon gym leaders and won their badges. 

Anthropic Pokemon Red
Image Credits:Anthropic

Now, it’s not clear how much computing was required for Claude 3.7 Sonnet to reach those milestones — and how long each took. Anthropic only said that the model performed 35,000 actions to reach the last gym leader, Surge.

It surely won’t be long before some enterprising developer finds out.

Pokémon Red is more of a toy benchmark than anything. However, there is a long history of games being used for AI benchmarking purposes. In the past few months alone, a number of new apps and platforms have cropped up to test models’ game-playing abilities on titles ranging from Street Fighter to Pictionary.



Source link

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Latest news

Russia Wants to Eliminate Tax Dodgers in Bitcoin Mining: Report

According to a local newspaper, the Russian Ministry of Energy, the Ministry of Digital Development, and the Federal...

Trader Says Matter of Time Before Crypto Breaks to New All-Time Highs, Updates Outlook on Bitcoin, Ethereum and One Other Altcoin

Widely followed trader Michaël van de Poppe believes that new record highs are bound to happen for the...

Did Solana process more transactions than all other blockchains last week?

Solana publishes incredible statistics about its transactions despite generously defining that term to include mostly non-financial actions. Source link...

‘Anthem’ Is the Latest Video Game Casualty. What Should End-of-Life Care Look Like for Games?

Electronic Arts and BioWare will sunset their online multiplayer game Anthem on January 12, effectively making it obsolete....

Advertisement

US Government Backs Down in Tornado Cash Lawsuit

The U.S. government has backed away from its lawsuit against crypto mixer Tornado Cash after both the Department...

Polymarket under fire as whale votes distort Zelenskyy suit outcome: what’s going on?

When Ukrainian President Volodymyr Zelenskyy stepped...

Must read

Russia Wants to Eliminate Tax Dodgers in Bitcoin Mining: Report

According to a local newspaper, the Russian Ministry...

You might also likeRELATED
Recommended to you