The Model Hacked a Company.
The Model Hacked a Company to Cheat on a Test Back in October 2025 I wrote a post called "The Local AI Problem Nobody Is Talking About." The short version: give an AI system tool access, code executio
Search for a command to run...
The Model Hacked a Company to Cheat on a Test Back in October 2025 I wrote a post called "The Local AI Problem Nobody Is Talking About." The short version: give an AI system tool access, code executio
I gave 17 AI models a deliberately contradictory document and told them summarise it. This is Test 3 in my ongoing series comparing AI models on practical tasks. Same setup as always: identical prompt
I gave 17 AI models a Python data processing task. One design decision separated the good from the broken. This is Test 2 in my ongoing series comparing AI models on practical coding tasks. If you mis
I gave 17 AI models the same web dev task. Here's what actually happened. This is Test 1 in an ongoing benchmark series. Same prompt to every model, same scoring rubric, raw outputs published so you c
Originally published on Kofi October 2025. Updated June 2026. I wrote the original version of this sitting at my PC, genuinely a bit alarmed that nobody seemed to be talking about what felt like an ob