Search a title or topic

Over 20 million podcasts, powered by 

Player FM logo
Artwork

Content provided by LessWrong. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by LessWrong or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://podcastplayer.com/legal.
Player FM - Podcast App
Go offline with the Player FM app!

“Insights into Claude Opus 4.5 from Pokémon” by Julian Bradshaw

17:41
 
Share
 

Manage episode 524054794 series 3364758
Content provided by LessWrong. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by LessWrong or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://podcastplayer.com/legal.
Credit: Nano Banana, with some text provided. You may be surprised to learn that ClaudePlaysPokemon is still running today, and that Claude still hasn't beaten Pokémon Red, more than half a year after Google proudly announced that Gemini 2.5 Pro beat Pokémon Blue. Indeed, since then, Google and OpenAI models have gone on to beat the longer and more complex Pokémon Crystal, yet Claude has made no real progress on Red since Claude 3.7 Sonnet![1]
This is because ClaudePlaysPokemon is a purer test of LLM ability, thanks to its consistently simple agent harness and the relatively hands-off approach of its creator, David Hershey of Anthropic.[2] When Claudes repeatedly hit brick walls in the form of the Team Rocket Hideout and Erika's Gym for months on end, nothing substantial was done to give Claude a leg up.
But Claude Opus 4.5 has finally broken through those walls, in a way that perhaps validates the chatter that Opus 4.5 is a substantial advancement.
Though, hardly AGI-heralding, as will become clear. What follows are notes on how Claude has improved—or failed to improve—in Opus 4.5, written by a friend of mine who has watched quite a lot of ClaudePlaysPokemon over the past year.[3]
[...]
---
Outline:
(01:28) Improvements
(01:31) Much Better Vision, Somewhat Better Seeing
(03:05) Attention is All You Need
(04:29) The Object of His Desire
(06:05) A Note
(06:34) Mildly Better Spatial Awareness
(07:27) Better Use of Context Window and Note-keeping to Simulate Memory
(09:00) Self-Correction; Breaks Out of Loops Faster
(10:01) Not Improvements
(10:05) Claude would still never be mistaken for a Human playing the game
(12:19) Claude Still Gets Pretty Stuck
(13:51) Claude Really Needs His Notes
(14:37) Poor Long-term Planning
(16:17) Dont Forget
The original text contained 9 footnotes which were omitted from this narration.
---
First published:
December 9th, 2025
Source:
https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-into-claude-opus-4-5-from-pokemon
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Credit: Nano Banana, with some text provided.
Choosing your starter Pokémon in Professor Oak's lab.
  continue reading

706 episodes

Artwork
iconShare
 
Manage episode 524054794 series 3364758
Content provided by LessWrong. All podcast content including episodes, graphics, and podcast descriptions are uploaded and provided directly by LessWrong or their podcast platform partner. If you believe someone is using your copyrighted work without your permission, you can follow the process outlined here https://podcastplayer.com/legal.
Credit: Nano Banana, with some text provided. You may be surprised to learn that ClaudePlaysPokemon is still running today, and that Claude still hasn't beaten Pokémon Red, more than half a year after Google proudly announced that Gemini 2.5 Pro beat Pokémon Blue. Indeed, since then, Google and OpenAI models have gone on to beat the longer and more complex Pokémon Crystal, yet Claude has made no real progress on Red since Claude 3.7 Sonnet![1]
This is because ClaudePlaysPokemon is a purer test of LLM ability, thanks to its consistently simple agent harness and the relatively hands-off approach of its creator, David Hershey of Anthropic.[2] When Claudes repeatedly hit brick walls in the form of the Team Rocket Hideout and Erika's Gym for months on end, nothing substantial was done to give Claude a leg up.
But Claude Opus 4.5 has finally broken through those walls, in a way that perhaps validates the chatter that Opus 4.5 is a substantial advancement.
Though, hardly AGI-heralding, as will become clear. What follows are notes on how Claude has improved—or failed to improve—in Opus 4.5, written by a friend of mine who has watched quite a lot of ClaudePlaysPokemon over the past year.[3]
[...]
---
Outline:
(01:28) Improvements
(01:31) Much Better Vision, Somewhat Better Seeing
(03:05) Attention is All You Need
(04:29) The Object of His Desire
(06:05) A Note
(06:34) Mildly Better Spatial Awareness
(07:27) Better Use of Context Window and Note-keeping to Simulate Memory
(09:00) Self-Correction; Breaks Out of Loops Faster
(10:01) Not Improvements
(10:05) Claude would still never be mistaken for a Human playing the game
(12:19) Claude Still Gets Pretty Stuck
(13:51) Claude Really Needs His Notes
(14:37) Poor Long-term Planning
(16:17) Dont Forget
The original text contained 9 footnotes which were omitted from this narration.
---
First published:
December 9th, 2025
Source:
https://www.lesswrong.com/posts/u6Lacc7wx4yYkBQ3r/insights-into-claude-opus-4-5-from-pokemon
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Credit: Nano Banana, with some text provided.
Choosing your starter Pokémon in Professor Oak's lab.
  continue reading

706 episodes

All episodes

×
 
Loading …

Welcome to Player FM!

Player FM is scanning the web for high-quality podcasts for you to enjoy right now. It's the best podcast app and works on Android, iPhone, and the web. Signup to sync subscriptions across devices.

 

Copyright 2025 | Privacy Policy | Terms of Service | | Copyright
Listen to this show while you explore
Play