Teaching an AI to Play PAC-MAN 256

AI EXPERIMENT · REINFORCEMENT LEARNING · VERY QUESTIONABLE DECISIONS

Pixels In. Decisions Out.

I’m teaching an AI to play PAC-MAN 256 using nothing but the pixels on the screen—no access to the game’s internal state, no secret map data, and no little voice whispering where the ghosts are.

It sees what I see. Then it tries to survive. Sometimes it even looks like it knows what it’s doing.

A Little PAC-MAN History

Before PAC-MAN 256 became its own game, “256” was already a famous piece of arcade history.

The original PAC-MAN was released by Namco in 1980 and, for practical purposes, was designed to keep cycling through mazes as long as the player could survive. But after 255 completed boards, the game runs into an 8-bit programming limitation. When it tries to generate board 256, a counter overflows and part of the screen becomes filled with corrupted symbols and graphics. The result became known as the Level 256 glitch, split screen, or kill screen.

More than three decades later, Hipster Whale and Bandai Namco turned that famous bug into the central mechanic of a new game. PAC-MAN 256, released in 2015 for PAC-MAN’s 35th anniversary, reimagines the maze as an endless upward journey. Instead of encountering the glitch only after hundreds of boards, the corruption is always chasing Pac-Man from below, forcing the player to keep moving while collecting pellets, avoiding ghosts, and using power-ups.

So the colorful wall of corrupted characters in PAC-MAN 256 isn’t just a random hazard. It’s a direct reference to one of the most famous limitations, and accidental visual effects, in the original 1980 arcade game.

And for this project, that historical glitch creates one of the most interesting problems for the AI: it can’t simply survive. It has to understand that the world behind it is disappearing.

Why I’m Building It

Partly because reinforcement learning is fascinating. Partly because screen-only control forces the system to operate under the same uncertainty as a person. Mostly because this is exactly the kind of unreasonable engineering project I enjoy: one familiar game, a ridiculous amount of instrumentation, and an endless supply of surprising problems hiding inside it.

PAC-MAN 256 gameplay
Pac-Man runs through a neon-lit maze while a digital glitch effect cascades from the top of the screen.

Why PAC-MAN 256?

Because it is familiar, fast, visually readable—and secretly mean. The maze never ends. Ghosts have different personalities. Power pellets temporarily reverse the food chain. And the bottom of the screen is constantly being eaten by a pixelated glitch.

That makes it a great reinforcement-learning problem: the rules are easy to explain, but survival requires perception, memory, timing, planning, and a healthy fear of colorful things with eyes.

How the Trainer Works

The trainer captures the game window, converts the image into observations, and asks a PPO policy what to do next. The policy chooses a direction, the trainer presses the corresponding control, and the whole loop repeats—frame after frame, game after game.

PPO, or Proximal Policy Optimization, is the part that turns experience into a better policy. It compares what happened with what the agent expected, nudges useful actions upward, discourages catastrophic ones, and tries not to change everything so violently that yesterday’s good ideas disappear overnight.

  • Capture pixels from the live game
  • Build an observation the model can understand
  • Choose and execute an action
  • Measure the result and assign a reward
  • Update the policy, checkpoint it, and do it all again

Reward Shaping: How to Accidentally Teach Weirdness

An agent only knows what the reward function tells it. Give points for pellets, survival, progress, power pellets, and ghosts; subtract points for dying, stalling, or getting swallowed by the glitch. Sounds sensible. Then the agent finds the loopholes.

One version learned that moving was risky and hesitation was cheap. Another discovered locally profitable loops that looked busy without accomplishing much. A reward can produce exactly the behavior it asks for and still be completely wrong for the behavior I meant.

The hard part is not making the number go up. It is making “number goes up” mean “PAC-MAN is actually playing better.”

The Langoliers Are Coming

PAC-MAN 256 has a moving wall of corruption that climbs from the bottom of the maze and erases everything behind it. I think of it as the Langoliers: it is always coming, it does not negotiate, and any plan that ignores it eventually becomes a very short plan.

The trainer estimates where that boundary is from the screen image and turns proximity into urgency. Hanging around below the safe zone should feel increasingly expensive. Climbing buys time. Getting caught ends the episode.

IMAGE PLACEHOLDER — Debug overlay with the detected glitch boundary, safe zone, and agent position.

The Advisor System

The neural policy is the driver, but it is not the only voice in the car. An advisor layer looks for immediate hazards and opportunities the policy may be about to miss: a ghost in the chosen corridor, a nearby escape, a power pellet, or the glitch getting uncomfortably close.

The goal is not to hard-code the game. It is to give the learned policy a small amount of tactical common sense, then log when the advisor agrees, disagrees, or overrides. That makes failures much easier to explain than a single opaque “left” or “right.”

A Checkpoint, Not a Victory Lap

Current evaluation across 100 games:

84.95

Average score

45.5

Median score

514

Best score

42%

Scored 50+

23%

Scored 100+

13%

Scored 200+

Evaluation run, checkpoint comparison, and score distribution.

Newer Does Not Always Mean Better

Training can drift. A newer checkpoint may score worse, become timid, forget how to exploit power pellets, or trade consistent survival for one spectacular run surrounded by disasters.

That is why I evaluate saved checkpoints over many complete games instead of trusting the latest training graph. The best policy is the one that holds up when it is asked to play again and again—not simply the file with the newest timestamp.

How Did It Die?

A low score is only useful if I know what caused it. The evaluator classifies deaths into two broad buckets so I can tune the right part of the system.

Ghost death

The agent made a tactical mistake: entered a bad corridor, reacted too late, ignored an escape, or simply drove directly into danger with enormous confidence.

Glitch death

The strategy failed at a larger scale: it stayed too low, stalled, chose a dead end, or underestimated how quickly the bottom of the world was disappearing.

PAC-MAN PPO Command Center

It turns a collection of scripts into something closer to an instrument panel: start a run, see what the model sees, watch decisions happen, and investigate the difference between “the AI is learning” and “the AI got lucky once.”

The Command Center is the dashboard I built to run the experiment without living in a pile of terminal windows. It launches training, tracks checkpoints, runs evaluations, compares policies, surfaces reward components, and gives me the visual debug views I need when the agent invents a new way to fail

WOULD YOU LIKE TO PLAY A GAME?.

What’s Next

The next version moves beyond the current desktop setup toward Android-based commercial media-player hardware. The idea is a compact, repeatable box that can capture the game, run inference, send controls, and report results without needing my main computer attached to the experiment.

That means rethinking capture latency, input injection, model size, thermal limits, remote monitoring, and how much of the Command Center belongs on the device. In other words: the project is becoming hardware, which is usually where things get more interesting.

Current Status

The system can train from screen pixels, evaluate checkpoints over complete games, recognize different failure modes, and produce runs that are starting to look intentional. There is still plenty to improve—especially consistency, long-horizon escape planning, and knowing when a locally attractive move is about to become a terrible idea.

It still occasionally confidently drives straight into a ghost.