AI EXPERIMENT · REINFORCEMENT LEARNING · VERY QUESTIONABLE DECISIONS
Pixels In. Decisions Out.
I’m teaching an AI to play PAC-MAN 256 using nothing but the pixels on the screen—no access to the game’s internal state, no secret map data, and no little voice whispering where the ghosts are.
It sees what I see. Then it tries to survive. Sometimes it even looks like it knows what it’s doing.

A Little PAC-MAN History
Before PAC-MAN 256 became its own game, “256” was already a famous piece of arcade history.
The original PAC-MAN was released by Namco in 1980 and, for practical purposes, was designed to keep cycling through mazes as long as the player could survive. But after 255 completed boards, the game runs into an 8-bit programming limitation. When it tries to generate board 256, a counter overflows and part of the screen becomes filled with corrupted symbols and graphics. The result became known as the Level 256 glitch, split screen, or kill screen.
More than three decades later, Hipster Whale and Bandai Namco turned that famous bug into the central mechanic of a new game. PAC-MAN 256, released in 2015 for PAC-MAN’s 35th anniversary, reimagines the maze as an endless upward journey. Instead of encountering the glitch only after hundreds of boards, the corruption is always chasing Pac-Man from below, forcing the player to keep moving while collecting pellets, avoiding ghosts, and using power-ups.
So the colorful wall of corrupted characters in PAC-MAN 256 isn’t just a random hazard. It’s a direct reference to one of the most famous limitations, and accidental visual effects, in the original 1980 arcade game.
And for this project, that historical glitch creates one of the most interesting problems for the AI: it can’t simply survive. It has to understand that the world behind it is disappearing.
Why I’m Building It
Partly because reinforcement learning is fascinating. Partly because screen-only control forces the system to operate under the same uncertainty as a person. Mostly because this is exactly the kind of unreasonable engineering project I enjoy: one familiar game, a ridiculous amount of instrumentation, and an endless supply of surprising problems hiding inside it.


Why PAC-MAN 256?
Because it is familiar, fast, visually readable—and secretly mean. The maze never ends. Ghosts have different personalities. Power pellets temporarily reverse the food chain. And the bottom of the screen is constantly being eaten by a pixelated glitch.
That makes it a great reinforcement-learning problem: the rules are easy to explain, but survival requires perception, memory, timing, planning, and a healthy fear of colorful things with eyes.
How the Trainer Works
The trainer captures the game window, converts the image into observations, and asks a PPO policy what to do next. The policy chooses a direction, the trainer presses the corresponding control, and the whole loop repeats—frame after frame, game after game.
PPO, or Proximal Policy Optimization, is the part that turns experience into a better policy. It compares what happened with what the agent expected, nudges useful actions upward, discourages catastrophic ones, and tries not to change everything so violently that yesterday’s good ideas disappear overnight.
- Capture pixels from the live game
- Build an observation the model can understand
- Choose and execute an action
- Measure the result and assign a reward
- Update the policy, checkpoint it, and do it all again

Reward Shaping: How to Accidentally Teach Weirdness
An agent only knows what the reward function tells it. Give points for pellets, survival, progress, power pellets, and ghosts; subtract points for dying, stalling, or getting swallowed by the glitch. Sounds sensible. Then the agent finds the loopholes.
One version learned that moving was risky and hesitation was cheap. Another discovered locally profitable loops that looked busy without accomplishing much. A reward can produce exactly the behavior it asks for and still be completely wrong for the behavior I meant.
The hard part is not making the number go up. It is making “number goes up” mean “PAC-MAN is actually playing better.”

The Langoliers Are Coming
PAC-MAN 256 has a moving wall of corruption that climbs from the bottom of the maze and erases everything behind it. I think of it as the Langoliers: it is always coming, it does not negotiate, and any plan that ignores it eventually becomes a very short plan.
The trainer estimates where that boundary is from the screen image and turns proximity into urgency. Hanging around below the safe zone should feel increasingly expensive. Climbing buys time. Getting caught ends the episode.

The Advisor System
The neural policy is the driver, but it is not the only voice in the car. An advisor layer looks for immediate hazards and opportunities the policy may be about to miss: a ghost in the chosen corridor, a nearby escape, a power pellet, or the glitch getting uncomfortably close.
The goal is not to hard-code the game. It is to give the learned policy a small amount of tactical common sense, then log when the advisor agrees, disagrees, or overrides. That makes failures much easier to explain than a single opaque “left” or “right.”

A Checkpoint, Not a Victory Lap
Current evaluation across 100 games:
84.95
Average score
45.5
Median score
514
Best score
42%
Scored 50+
23%
Scored 100+
13%
Scored 200+


Newer Does Not Always Mean Better
Training can drift. A newer checkpoint may score worse, become timid, forget how to exploit power pellets, or trade consistent survival for one spectacular run surrounded by disasters.
That is why I evaluate saved checkpoints over many complete games instead of trusting the latest training graph. The best policy is the one that holds up when it is asked to play again and again—not simply the file with the newest timestamp.
How Did It Die?
A low score is only useful if I know what caused it. The evaluator classifies deaths into two broad buckets so I can tune the right part of the system.

Ghost death
The agent made a tactical mistake: entered a bad corridor, reacted too late, ignored an escape, or simply drove directly into danger with enormous confidence.

Glitch death
The strategy failed at a larger scale: it stayed too low, stalled, chose a dead end, or underestimated how quickly the bottom of the world was disappearing.
PAC-MAN PPO Command Center
It turns a collection of scripts into something closer to an instrument panel: start a run, see what the model sees, watch decisions happen, and investigate the difference between “the AI is learning” and “the AI got lucky once.”
The Command Center is the dashboard I built to run the experiment without living in a pile of terminal windows. It launches training, tracks checkpoints, runs evaluations, compares policies, surfaces reward components, and gives me the visual debug views I need when the agent invents a new way to fail



What’s Next
The next version moves beyond the current desktop setup toward Android-based commercial media-player hardware. The idea is a compact, repeatable box that can capture the game, run inference, send controls, and report results without needing my main computer attached to the experiment.
That means rethinking capture latency, input injection, model size, thermal limits, remote monitoring, and how much of the Command Center belongs on the device. In other words: the project is becoming hardware, which is usually where things get more interesting.
Current Status
The system can train from screen pixels, evaluate checkpoints over complete games, recognize different failure modes, and produce runs that are starting to look intentional. There is still plenty to improve—especially consistency, long-horizon escape planning, and knowing when a locally attractive move is about to become a terrible idea.
It still occasionally confidently drives straight into a ghost.
