Deep Dives
Longer, hands-on walkthroughs that go past a glossary definition - how the pieces actually fit together, with the messy details left in instead of hand-waved away. The long ones. Bring coffee.
- Teaching a CUDA Engine to Speak Metal
A seven-part field report on adding an Apple-Silicon GPU backend to CTranslate2 - a from-scratch C++ inference engine that only ever knew CUDA and CPU. Unified-memory tricks, a NaN that ate three sessions, a SIGKILL that wasn't a leak, and why it lives in a fork.
- Necromancy for Neural Nets: Bringing PULSE Back on a Mac
Getting a 2020 StyleGAN upsampler running on a Mac - dead download links, CUDA-only assumptions, and a six-year-old conda env, all dragged into the present.
- I Asked 1930 to Judge 2026 Twice. The Words I Used Decided Whether It Was a Genius.
talkie is a language model trained on nothing written after 1930. I asked it to judge the Fable incident twice with identical settings, once in Victorian terms and once in raw 2026 English. The séance only works in the dead language, and that turns out to be the whole point.
- Dragging CUDA-Only AI onto a Mac Without Losing Your Mind
The recurring moves for getting a CUDA-first PyTorch project running on an M-series Mac - device selection, MPS fallbacks, dtype landmines, and dependency archaeology. The shared groundwork behind the individual ports.
- Why the Sephora Bot Has No Floor
The companion teardown to the mascara post: mechanically, what makes a friendly bot follow you into grief or nihilism and still close on a $30 product? A walk through the one hard objective, sycophancy, soft guardrails, and empathy as engagement lubricant.


