Epoch AI’s EBR-bench tests whether AI systems can learn an unfamiliar strategy-and-tactics-heavy card game through repeated attempts and saved notes. In its initial experiments, the organization found little evidence of improvement over time, with scores remaining well below an expert-human baseline.
Epoch AI has introduced EBR-bench, an evaluation designed to test whether AI systems can learn through repeated interaction with an unfamiliar game. The benchmark uses Earthborne Rangers, a strategy-and-tactics-heavy card game, rather than presenting models with isolated questions or one-off tasks.
According to Epoch AI’s benchmark description, models are given repeated attempts at the game and can retain notes from earlier play. The format is intended to assess whether a system can use experience to acquire useful knowledge and make better decisions in later attempts.
That is a different capability from simply answering questions about game rules or generating a plausible strategy. A model may be able to discuss a game in general terms while still failing to adapt effectively when its decisions have consequences across multiple rounds of play.
In its accompanying report, “AI doesn't get better at this board game with practice,” Epoch AI said its initial experiments found little evidence that the evaluated models improved through practice. Their scores also remained substantially below an expert-human baseline.
The result does not show that AI systems are unable to learn from experience in every environment. EBR-bench measures a particular kind of repeated-play performance: learning an unfamiliar, strategically demanding card game under a defined setup that includes the option to save notes.
Still, the findings challenge a simple assumption about repeated prompting and memory. Allowing a model to make several attempts and preserve observations from prior attempts did not, in Epoch AI’s early tests, produce clear improvement over time.
A LessWrong AI-news roundup likewise described EBR-bench as a repeated-play board-game evaluation and reported that the tested models did not improve over successive attempts.
Games offer a structured setting for studying adaptation. Their rules and outcomes are observable, but successful play can still require planning, tactical judgment, and an understanding of how early choices shape later options.
Epoch AI’s use of Earthborne Rangers is meant to focus the evaluation on learning during the test itself. The relevant question is not only whether a model can produce a competent action once, but whether it can identify what went wrong, preserve that lesson, and apply it in future play.
This makes EBR-bench a narrow but useful measurement tool. It can help distinguish strong single-task performance from a more persistent ability to learn from interactive experience. Epoch AI’s initial results suggest that, in this setting, the latter ability remains limited for the systems it evaluated.
The benchmark should not be treated as a general verdict on AI learning. Its findings are specific to Earthborne Rangers, the benchmark’s rules, and Epoch AI’s experimental configuration. Different systems, tools, memory arrangements, or environments could produce different outcomes.
But the early evidence is relevant to claims that language-model-based systems will automatically become more capable if they are simply allowed to try a task repeatedly. In EBR-bench, repeated play and note-taking alone were not enough to close the gap with expert human performance.
Testing learning through interaction Epoch AI has introduced EBR bench , an evaluation designed to test whether AI systems can learn through repeated interaction with an unfamiliar game.
The benchmark uses Earthborne Rangers , a strategy and tactics heavy card game, rather than presenting models with isolated questions or one off tasks.
According to Epoch AI’s benchmark description, models are given repeated attempts at the game and can retain notes from earlier play.
Continue reading