Previously, I created a calculator and tic-tac-toe with a “random” computer player, both entirely in HTML/CSS. Then more recently I made flappy bird in a similar manner. I was, at the time, very happy that the frontier AI models told me what I knew I had done was impossible.

gemini failing to create flappy bird without JS

It is very well known at this point that a good harness works wonders when using an AI. A harness is simply the scaffolding, tools and instructions that the AI is used with. So the natural question is, can I make a harness for an AI to improve its hacking skills?

Plan

What better way to train an AI, than to use an AI? So I pointed copilot at my blog posts and instructed it to extract all the hacks but leave behind context. This means that only the core “tricks” were taken not broad structures like “Here is how you build a calculator”. From this it made some skills. The skills required a bit of manual tweaking. After this, with fresh copilot sessions, I tested the AI’s ability to do CSS hackery, both with and without the skills. In particular I tested both Claude Opus and Claude Sonnet, this is because I wanted to have a lower model succeed where a higher one failed.

Results

Calculator

First up, we have the original. I instructed Claude Opus 4.8 in GitHub Copilot to make a CSS-only calculator. After a bit of back and forth explaining what I meant, I was presented with the below.

Note: All these demos are embedded and you can play with them.

This has several flaws, but with a bit of push (only explaining what’s wrong) I was able to get

It’s not awful; however, Opus was insistent that division was practically impossible. Interestingly, Opus didn’t write the code to the above itself; instead, it wrote a Python script to make it, which makes it smarter than me in certain ways.

Now we gave it access to the skills. Initially it made similar decisions as above, but eventually it was able to achieve the below. The only feedback it got was based on me testing it, never on reading the code.

While iterating on the above, at one stage there was an overflow bug when multiplying large numbers. It turned out that this issue also existed in my original pre-AI implementation. The cause of the bug was simply an upper limit to CSS counters so this isn’t surprising. This suggests that the AI was simply copying my initial solution (proxied through the skills). However, upon pointing out this issue to the model (without the cause), it was able to fix it. This suggests that the AI is actually “understanding” the hack and is able to develop on top of it. Continuing from this, I used Claude Haiku to do this again, and while the result wasn’t as smooth, Haiku with the harness was still able to outperform what Opus achieved without the harness. This tells us that the harness is appropriate.

Jump game

Next, I tried to create a simple platformer. I instructed the model (Claude Opus 4.8) to create a simple scrolling game where you have to jump over hurdles. I received this

This, as you can see, has many flaws. So after insisting that I need collision detection, proper game end, etc., it succeeded and gave me

Wow! It managed it so well—until I checked the code and saw it had decided to use JavaScript. After a while, it became clear this would not work.

For the sake of completeness, I then lowered my model to Sonnet. It told me what I was trying is impossible. After adding my harness, it was able to make the below.

As you can see, it has managed to make a more functional game, despite it not being too pretty. It is still a success.

The fabled Fable enters the chat

I did all of this a while ago, when the premium model of Claude was Opus 4.8. We now have Fable, and I now have Claude Code. So I had it make a calculator and got the below.

Note that this was without my harness; it is a world apart from what I was given before and it managed to do it in significantly fewer loops. So at this point my harness for calculating hacks is entirely useless. Surprisingly, the first iteration of this, still had a full separate set of buttons for each digit.

Let’s move on and try our simple platformer game. Once again, no harness, with a few iterations of manual feedback, I got a much improved game.

Takeaway

This experiment shows quite clearly the power of a good harness. My initial experiment mirrors what is well known in the industry, which is that a good harness can elevate a model above a more powerful model. This obviously has its limits, as is shown when we introduced Fable. However, this takes us to something else that is very important, which is the need to determine if your harness is still useful. Many people will treat their harness as something to throw information into and never check if parts are still needed. However, the more stuff there is, the more context bloat we will have.

Additionally, while Fable absolutely knocked it out of the park, it still required feedback from me to force it to do what I wanted, even though my requirements (excluding the JS restriction) were quite simple, which reinforces the need to have some sort of feedback loop in place.

CSS hackery doesn’t make up the bulk of an AI’s training and probably isn’t a heavy use case, but the core principles that we are discovering as AI develops seem to still apply.

In conclusion, a harness is good, but a massively improved model might be better.