Notes

A deterministic cover generator

ai, tools
Six off-white vertical bands of varying width on a sand field.

I wanted every note here to have a cover, so I tried an image model first. It gave two different notes the same laptop-and-mug setup, then added text I never asked for.

That wasn’t going to work, so I wrote a spec for a generator instead and had Claude build it. The first thing it does is hash the filename, which means turning it into a number. Every letter has a number behind it, and the code combines them in a fixed way, so cover-art-generator always comes out as 283832237. That number picks the colour and the layout.

The filename is the only input. Nothing random, nothing that depends on when I run it. That’s what deterministic means: same input, same output, every time. I can rebuild the whole site and not one cover changes.

Then I put it on a page, which meant deciding what other people get to change. As a script it needed no controls, since I was the only one running it. The obvious answer is two colour pickers, one for the background and one for the marks. I didn’t do that. Every background already comes with the mark colour I picked to sit on it, and two pickers break the pair.

So you never type a colour. You pick one of four palettes, the filename picks a swatch inside it, and one slider shifts the hue of the whole palette. Saturation and lightness stay fixed.

It’s at /lab/cover-art. Type in the name of something you’ve written and see what it gives you.

More notes

A model sits at the centre of a loop. On each turn it runs a web search, files what it found in memory, and waits at a checkpoint for a person to press Continue. The loop sits inside a boundary of limits with a step meter.

The AI setup that decides what your designs look like

The harness around an AI tool is usually treated as engineering work. I built one to find out what a designer could bring to it. I first heard the word "harness" at an adventure camp at school, when the instructors explained how one would keep us safe on the ropes. I heard the word again this year on an AI project at work, and it stuck with me. The screens looked like ours On that project, we used AI at almost every stage. Design had stayed in Figma for months, until the design team got to try prototyping in Claude Code, which at the time ran in the terminal. Some designers on the team found that restrictive. I'd worked as a UX engineer before, so it felt like coming back to something familiar. We each built a module and raised pull requests. What surprised me was that the screens looked like the ones we'd designed in Figma. The reason was the harness our engineering team had set up. Our design system's components and a base template were already wired in, so the model built with our parts instead of inventing its own. That was the one piece of the harness designers already owned. Harnesses were new to me, and I wanted to understand how one worked. I wanted to find out if a designer could build one too. What a harness is An AI model on its own answers one request and stops. A harness is the scaffolding around it that keeps it working through a task in steps. After each step, it decides what to run next based on what it has already found. OpenAI calls it harness engineering; Anthropic writes about harness design. Both mean the scaffolding that turns a model into something that can finish a job. It's close to what the camp instructors meant. A harness sets how far you can go and keeps you safe while you do. The harnesses I've read about share five parts: a loop that decides what runs next tools the model can use, like search memory that carries findings from one step to the next checkpoints where a person steps in limits the code enforces, whatever the model asks for Anatomy of a harness, drawn as a climbing one. Illustration generated with Claude Picking a problem I could judge The next question was what to build. It had to help with my own work while leaving the important calls to me. I also needed to be able to tell whether the output was any good. Competitive research fit both. Most of my design work starts there, it takes real time, and I know what a useful analysis looks like. It should map out who each competitor serves and where there's room to differentiate. Some of that work can't be handed to a tool. I still want to go through each product myself, signing up or browsing it on Mobbin, to see how it works and feels. What a harness can take on is the desk research: gathering what's known about each competitor and comparing them. So that's what I built, and I called it Compset. I built it in Claude chat, one piece at a time, with Claude writing all the code. My job was deciding what the tool should do and how it should look. When it got one of those wrong, I pushed back. How it works Each research session, which I call a run, has four stages and touches all five parts of a harness. You start by writing a brief describing what you want to research. The model finds competitors, you review the list, and a research loop runs until it produces a report. If the brief names a product to compare against, that product becomes the baseline. In the examples here, that's Figma. The loop follows a plan the model writes at the start, researching each competitor in more depth with every round. After each step, the model suggests what to do next. The tool runs as an artifact inside Claude chat, and an artifact can't browse the web on its own. So I connected Parallel Search as a search tool inside Claude. The model writes its own queries, and whatever it finds is saved with a link back to the source. Those findings go into memory, which keeps the model from going in circles. For each competitor, memory holds a short summary, key facts with sources, and an open question. The model reads this before every step, so each step starts from what it has already found. The final report is written only from what's in memory. Left to itself, the model would keep researching. My first version had no step limit, and I spent ten minutes staring at the screen waiting for it to stop before I stopped it myself. Limits make sure it stops. Before a run, you choose how many competitors to study and how deep to go, and that decides how many steps the run gets. Once those steps are used up, the run ends, even if the model wants to keep going. The harness also caps the model at two web searches per research pass. The part that enforces these limits is called the Guardrail, and it shows up in the log whenever it steps in. Where I stay in the loop The Guardrail handles limits, but some decisions should be made by a person. The first comes right after the model finds competitors, when every run pauses automatically. You see the list with a short reason for each, and you can remove any, add your own, or stop. Research only starts when you continue. The run pauses after finding competitors. I removed Adobe XD, which Adobe no longer actively develops, and added Framer, which the model missed. During research, the run only stops when it matters. Moving on to the next competitor is expected, so the model just does it. But sometimes the model wants to change the plan, whether that means researching a competitor out of order, returning to one it has already covered, or writing the report early. Then it asks first, with Allow and Don't allow buttons in the log. If you say no, it sticks to the plan. Back in 1999, Eric Horvitz argued that a system should weigh the cost of interrupting you against the risk of acting on its own. That gave me a rule: routine moves happen silently, and anything that changes the plan gets a button. The model asks to research Sketch ahead of plan. Doing it now changes the order, not the total number of steps. If you'd rather not be asked, a "Run unattended" switch lets the model make those calls itself. It's off by default, and the limits and competitor review still apply either way. I've kept it off, because I want to see those decisions. Designing the interface Everything I've described shaped how the tool runs. The next set of decisions shaped what a person sees. Colour The first interface Claude produced had a purple button. Of course it did. Purple gradients over white cards have become a telltale sign of AI-generated design. Claude was writing the code, but I wanted to make the visual choices myself. I switched to a neutral palette with black buttons, which left colour free to mean something. Coral shows when the Guardrail overrules the model, and teal shows when the run is waiting for you. The first version's purple button, and the neutral palette that replaced it. The setup screen The setup screen started as a form whose labels described the code. One field, "Deep-dive depth," had the hint "Iterations per competitor in the deep-dive loop." Nobody wants to think that hard before starting a research session. The only thing someone actually needs to write is what they want to research, so I started from that. The settings took a few tries to simplify into two choices. You pick how many competitors and how deep, and each depth describes what the report would cover. Before, a form whose labels described the code. After, a brief and two choices: how many competitors and how deep. Progress Once a run was going, the screen only said "Researching...", so I couldn't tell how far along it was or what it was doing. To fix that, I made the screen show what the run is doing, where it is in the loop, how many steps are left, and how far along each competitor is. Mid-run: what's happening now, the loop with the web search tool in use, steps used against the limit, and each competitor's progress. The log For a research tool, I needed to be able to check how it reached its conclusions, so I added a log. Each line says who made the decision, whether that's the model, the Guardrail, or you. Routine steps are marked Harness, and every web search is listed with the queries the model wrote. If someone questions the report, I want them to be able to trace any finding back to the search that produced it. The part I most wanted to get right was when the Guardrail overrules the model. In the log, the model's request is crossed out. On the loop diagram, the Guardrail step turns coral and shakes, and the path changes to show what the run does instead. The Guardrail overrules the model's request for another pass on Sketch. Shown in preview mode with example data. The report The finished report, with Figma as the baseline column. The report ends with a record of what you changed, how often the Guardrail stepped in, and how many searches ran. The record at the end of every report. To see how this experiment fared, I checked the pricing row against each company's pricing page. One price was wrong, and the report also included a competitor I hadn't approved. During the deep dive on Canva, the model added Affinity, which Canva owns. So for now, the report is a good first draft of the desk research, and I would check it before relying on it. What I'm taking back to work Building this helped me understand the harness I'd been working inside. It decides more than which components the model builds with. It decides what the model can reach, when it stops, and when it asks a person. Every limit and checkpoint in my tool is a guess about where the model needs me, and the Affinity slip showed that some of those guesses are wrong already. They'll keep changing as models improve. On our project, those calls were made by the engineers who set up the harness. After building one myself, I think deciding where a person steps in is design work too. References OpenAI, Harness engineering: leveraging Codex in an agent-first world (2026). https://openai.com/index/harness-engineering/ Anthropic, Harness design for long-running application development (2026). https://www.anthropic.com/engineering/harness-design-long-running-apps Horvitz, Principles of mixed-initiative user interfaces (CHI 1999). https://dl.acm.org/doi/10.1145/302979.303030 Gibbons, The Four Design Jobs AI Created (Nielsen Norman Group, 2026). https://www.nngroup.com/articles/design-jobs-ai-created/

September 29, 2026
Four table screens rendered four different ways without a shared convention file, and the same four rendered identically with a committed CLAUDE.md.

Claude Code, and where team conventions actually live

I decided a page needed tabs. The design file didn't have them yet, so when I prompted for them the tool drew the pattern from scratch, and it drew it wrong. That took a minute to fix. The interesting part came later. Several of us were generating pages against the same app. Each of us kept hitting gaps like that one, places where the design file had no answer and the tool had to invent something. Every session invented differently. Nobody was working carelessly. We were each getting a reasonable guess from a tool with no way of knowing what the other three had already decided. So I started correcting mine. Follow the pattern the rest of the app uses, and remember it for the next instance. That worked, and my later pages held the conventions without me asking again. It just never reached anyone else. Claude Code carries knowledge between sessions two ways, and the difference between them turned out to be the whole story. Auto memory is what I'd been using. You say remember this, and the tool writes itself a note, scoped to the repository but living outside it. Which means it's yours. Your corrections, your sessions, nobody else's. A CLAUDE.md is the other option: a markdown file you write and commit, read by every session that opens that repo, including sessions belonging to other people. Same instruction, much larger blast radius. I'd hit the second use case with the first mechanism, and put a team convention somewhere only I could see. So I did a reconciliation pass. I went back through pages other designers had generated and fixed the same three things by hand. Numbers right aligned in tables. Timestamps as a date and then a time. Page names matching their label in the navigation menu. Nobody finds these decisions interesting. They're the kind of thing a team agrees once and then stops thinking about, and we were paying for them one page at a time. Here's roughly what I'd commit now: ``markdown UI conventions Navigation Underlined tabs are reserved for page-level navigation. Page names must match their label in the navigation menu. Tables Numeric columns are right aligned. Text columns are left aligned. Timestamps render as date, then time. Before generating a new screen Check an existing page for the pattern before inventing one. If no existing page has it, ask rather than choosing. `` I'd copy two things from that shape before I copied any of the rules. It says what's reserved rather than what's preferred, because a preference invites a judgment call and every judgment call is another place two sessions can diverge. And the last block is the one I'd want in there on day one. Most convention documents describe the cases you thought of. The drift came from the cases nobody had thought of yet, where the tool filled a gap because filling gaps is what it does. Telling it to ask when the pattern is missing turns a silent guess into a question. None of this is an argument against generating UI this way. The pages were good and they arrived fast. So before I correct the tool again, I ask where the correction is going to land.

August 28, 2026

Designing AI-first product interfaces

The first wave of AI products treated the model as a feature. A chat box bolted onto an existing flow. A magic button next to a regular one. Users did their normal work, and off to the side, AI was there if they wanted to summon it. Most of those never made it past the demo. The products that stuck had a different shape. The model wasn't a feature you could add or remove. It was baked into how the work got done. The flow assumed AI was in the loop. The interface let people correct or override the model, not just trigger it. On one of the manufacturing projects I worked on, the question we kept circling wasn't "where do we put the AI?" It was: if an operator has to trust this output enough to sign off on it, what does that signoff actually look like? Once we framed it that way, a lot of decisions got easier. We made the model's output visible. We let operators edit it inline. We tied every action to a clear name and timestamp. None of it was glamorous, but it was the thing that made the product real. A lot of AI design still works the old way. The product keeps doing what it always did, and the AI gets added on top. That rarely holds up. When you design with AI from the start, the questions change, and the interface gets simpler, because the model is doing the right work in the right place instead of sitting in a sidebar waiting to be useful.

June 18, 2026