Opinion
We tested our self-hosted LLMs and accidentally built a game
It sounds like the start of a joke: a cybersecurity company sets out to test its AI models and ends up with a game. But that is genuinely what happened, and it raises a fair question I expect from partners within a week of this blog going live. Why are you building a game when your AI agent isn’t finished yet?
Let me answer it before anyone asks. We weren’t building a game at all. We were testing large language models, and a game is what came out.
That distinction matters more than it sounds. Most MSPs and IT leaders I speak to are working through the same questions we are: which large language models can we trust, should we host them ourselves, and what should our people be doing while AI coding agents do the typing? Our accidental game, the by-product of a simple one-shot prompting test, turned out to be a surprisingly useful lens on all three.
How did a model test turn into a game?
Our Lead SRE has a standard test for every new model. He gives it one fixed prompt: build a complete platform game in a single go. This is one-shot prompting; one instruction, no follow-up corrections, no hand-holding. How close a model gets tells you a great deal about its reasoning, the quality of its code and its ability to hold a large task together. We run that test on the self-hosted open-source LLMs we use, and because we replace those models regularly as better ones appear, a test that stays exactly the same is worth more than any leaderboard.
This time he went further. Rather than stopping at the one-shot result, he let the model draw on our public website and the information it exposes through the Model Context Protocol (MCP), and asked it to weave that into the game. In side sessions, while other AI coding sessions were running, the experiment grew into Lighthouse Run: six levels and 24 zones full of references to our sector, end bosses, a scan report at the end of every level and a semi-randomly generated Daily Scan level that changes every day. The entire game lives in a single self-contained HTML file, and the project around it includes tests so the AI can run and verify its own changes.
I have played a few levels. They are not easy, which I mean as a compliment. You can try it at https://guardian360.net/lighthouse-run/; turn the sound up, and it works on a phone or with a controller too.
Why test an LLM with a one-shot prompt?
Because our impressions of AI are unreliable. In a survey of 349 technical workers published in May 2026, METR found that participants self-reported a median 1.4 to 2 times change in the value of their work due to AI tools, and a median speed change of 3 times. METR itself warns against taking that at face value: its early 2025 study found that people overestimated AI’s effect on their time spent on tasks by 40 percentage points on average.
If experienced technical people misjudge the tools they use every day, a polished vendor demo is not a sound basis for choosing a model. A fixed, repeatable prompt is crude, but it is honest. Same task, every model, results side by side. A garage that test-drives every car on the same route learns more than one that relies on the brochure.
What do engineers do while AI agents do the typing?
Agent-driven work is no longer an experiment. JetBrains’ 2026 Developer Ecosystem Survey of more than 15,000 professional developers found that as of May to July 2026, 90% were using AI coding agents at work at least weekly, and 68% were using them daily.
That changes the shape of an engineer’s day. Less time is spent typing, more is spent instructing, reviewing and waiting. METR ran into this when it tried to repeat its productivity research. It saw a significant increase in developers choosing not to take part because they did not want to work without AI, and found that its time measurements were unreliable for developers using multiple AI agents concurrently.
That is exactly how our Lead SRE works. Several sessions run in parallel, and there are gaps while each one finishes. What happens in those gaps is a leadership question. You can pretend they do not exist, fill them with more of the same, or let people use them to test ideas. We have a culture of rapid prototyping: build small, test quickly, keep what works and throw away what does not. Those gaps are where hypotheses get tested cheaply.
Why is model choice a security decision?
Choosing an AI model is no longer just a question of capability; it is a security question. Veracode’s 2026 GenAI Code Security Report tested more than 100 models and found that they averaged a 56% security pass rate, little changed from 2025, even as syntax pass rates approached 100%. Coding-focused models were no safer than general-purpose ones. Veracode’s own conclusion is that model selection has quietly become a security decision.
For a company working in cybersecurity, that settles it. We cannot choose models on reputation or on benchmarks published by their makers. We need to see for ourselves how a model behaves on a substantial task, which is precisely what a one-shot test shows. Hosting models ourselves gives us control over where our data goes. And building tests into the project means the AI can check its own work, rather than us having to trust it. As Veracode puts it, AI-generated code should be treated like any unreviewed code.
What can MSPs and IT leaders take from this?
The first lesson is to test before you trust. Pick one substantial, repeatable task that reflects your own work and run every model you consider through it. It does not need to be a game; it needs to be the same every time.
The second is to treat AI output as unreviewed code. Build verification into the process, ideally so the AI can run tests against its own work, and keep a human accountable for what goes into production.
The third is to give experimentation boundaries rather than banning it. Small, self-contained, time-boxed and with a clear purpose. Within those boundaries, an unexpected result is a welcome bonus rather than a worrying distraction.
Good experiments deliver more than you were looking for
We were looking for a reliable answer on which models we can trust. We got that, and a game we are slightly too proud of. Give it a go; there are a few hidden extras for those who know where to look, and for those who know their internet memes.
And then tell me: how do you test the AI models your team relies on before you trust them?
Sources
- JetBrains Research, AI Coding Agents: Adoption Trends (Developer Ecosystem Survey 2026)
- METR, Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity
- METR, We are Changing our Developer Productivity Experiment Design
- Veracode, LLMs Are Getting Smarter, But Not Safer: 2026 GenAI Code Security Report