← Blog

Out of memory, on purpose

5 September 2026

sitebehind-the-scenes

This morning someone hit the rugged box generator with a big divider layout and got a 502 back. I know because I went digging after a burst of them showed up in the logs, and the trail ended in the kernel log: Memory cgroup out of memory: Killed process (python), anon-rss: 806336kB. The box generator lives inside a container with an 800 MB memory limit. Someone asked it for a model that needed 806.

That is the third time a generator has died this way in production, and I want to explain why I keep letting it happen.

One machine, thirty-six tenants

Everything on this site runs on a single VPS with 7.6 GB of RAM. Each of the 36 generators sits in its own container and idles at about 87 MB, which together eats under half the machine. The rest is headroom for actual work, and the question is what happens when one request wants all of it.

Without limits, the answer is ugly. A single oversized lithophane could balloon one Python process until the kernel starts killing whatever it can find, and "whatever it can find" is rarely the process that caused the problem. So every container carries a hard cap, 500 MB by default, and when a request blows past it, that one generator dies, restarts in a few seconds, and the other 35 never notice. The person who sent the request gets a 502, which is bad. Everyone else gets nothing at all, which is the point.

Three kills, three numbers

The first was the lithophane generator. A visitor kept feeding it a 12 megapixel photo, and turning that into a thickness mesh peaks at about 1.09 GB. The cap was 500 MB. They got 502s for three days before I traced it, and the fix was not a guess: I ran their exact photo size locally, watched the peak, and set the new cap at 1400 MB.

The second and third came together, this morning. The box at 806 MB against 800, killed four times in a row by one persistent user with a many-compartment build. And the keychain generator at 475 MB against its 500 MB default, taken down by a single page load that fired 18 previews at once, which smells like somebody generating a whole batch of name tags. The kernel log timestamps match the 502 bursts to the second.

Box now gets 1200 MB and keychain gets 800. Both numbers sit in a comment in the config next to the measurement that produced them, because a limit you cannot explain is a limit you will cargo-cult forever.

Raising the limit is the lazy half of the fix

I want to be honest about what happened here: the users were not doing anything wrong. A big divider layout is a legitimate box. Eighteen keychains is a legitimate batch. When a legitimate request kills the process, the memory limit is not the bug, the memory appetite is.

The newest generator, the bolt generator, already does this properly. It has a hard vertex budget per part, 350,000, and when a giant M60 bolt would exceed it, the mesh gets coarser instead of bigger. The chord error at that scale is under half a millimeter, invisible in plastic. The box generator has no such budget yet, and that is the real item on my list. Raising a cap buys headroom. A budget makes the worst case a design decision instead of a kernel decision.

Until then, the caps stand guard. If you ever get a 502 from a huge model, wait five seconds and try again, the container is already back up. And if it keeps happening, the report button tells me exactly which build to go measure next.