Working AI-first reminds me of the early days of the web.
Nobody had a playbook. You just learned by building.
The mobile shift in the early 2010s felt different. When we built Foundation to help teams design responsively, we mostly took what we knew and figured out what really needed to fit on a smaller screen. As Luke Wroblewski reminded us, mobile-first was about the decisions that went into a smaller device, beyond the screens themselves. You had to be brutal about what you removed.
AI asks something different. The work moves fast. Staying critical of it, and having conviction about where you're taking it, is the real problem.
Creating things got cheap
A year ago, ten versions of a homepage took a week. Now it takes an hour. Pages, tests, emails, and decks can all get created at once.

Speed can expose gaps in your decision making. You move fast and you don't always know what's happening. AI gets you to 80% fast. The last 20% is where the harder work starts. That stretch is where you work out what a piece is for, who it serves, and what it should help someone decide.
You'll recognize this if…
We've been working this way for about a year now, across our own work and our clients'. The same things keep showing up:
- Your team makes ten versions of a screen and the conversation turns into opinion.
- Someone senior says the page isn't ready and can't say why, partly because they weren't in the work.
- A regeneration quietly puts back a section you removed on purpose two weeks ago.
- The reasons behind a change sit in individual chat threads nobody can find.
- An executive asks how you arrived at these ideas, and the answer is in a thread with the data attached to it.
- You report visitors because visitors are easy to count.
That last one matters a lot now, as most teams still don't have a full stack of measures that ties back to business goals. They reach for what's easy to grab and get lost in the product aspect of all of this. We built Glare to use with Helio to help sort through much of this to create more transparency.
None of these are AI problems by nature. We've always had them when it comes to tacking and data. Especially if you don’t own the data. What's different is that you can get somewhere much faster now and have no idea why you're there.
What we learned on a live rebuild
We spent August rebuilding the marketing site for an AI technology company. Small team, fast movers, real traffic, real revenue on the line. Their engineer ships regularly and their marketing lead runs several products at once. Everyone is generating.
Early on we tested the homepage with about 100 people from their audience, even benchmarked it across 5 competitors. Roughly 70% knew what the site was for, and specifically that it was built for developers.
Then a full page regeneration was created.

It looked fine, and on the surface it could pass. We tested it and audience clarity dropped almost 30 points. People could no longer say who the product served.
Nothing was really technically broken, but the generation had put back a section we removed on purpose, removed important copy, and softened the language that made the audience obvious. Their engineer said it plainly: "I didn't notice that we removed it." The reasons behind the previous version weren't written down anywhere for the next generation to use.
The loop we’ve created from the experience.

An idea goes into prompts, skills and agents. Concepts come out of that, which is the part we're generating. Real people react to those concepts in Helio. Their behavior and feedback shape how we evolve the idea.
Three habits keep that loop from spinning out.
1. Know what you need to decide before you prompt
You don't need all the answers up front. You need the right question.
On that project, the hero changed between two versions. When we asked why, the answer was honest… "it was just a choice that I choose, but doesn't mean that it was a process behind it." A week later, the UX metrics were telling us something different, and the instinct was to roll back to the old version.
Rolling back throws away the thinking along with the change. And indecision kills design impact. So we asked something else. What were you trying to do? Once that was on the table, we could test it.
The same gap also showed up in measurement. We were moving people into the site through sign-up, and the A/B test wasn't set up right. The same people were seeing both variants, which tells us nothing. We fixed it and rebuilt it into two separate groups so the comparison held, and then we fixed the sign-up itself. This happens when you move quickly. AI sets up the files, someone wires the test by hand, and the two don't always line up.
What we're building into the way a team works is a shared understanding of intent. Everyone should know what a generation is supposed to produce and what rules it's working from, before it runs.
Once you start reworking an idea, the change itself matters less than how it lands across the rest of the site- that part needs to get written down. Explore wide with the AI tools, then narrow.
We’ve found when you mix inventing and refining in the same step, you get outputs that aren't quite right. We've watched that happen plenty of times.
2. Keep your reasons in one place
AI spreads your thinking across chats, tools and files. Whatever isn't written down gets overwritten.

After comprehension and clarity dropped in testing, we stopped generating pages and built a baseline instead. A style guide and a component library, living in the same repo the site builds from. About 30 components, an hour or two each in Figma, with nothing visible to show for a few days.

That work exists because Claude puts a lot of extra stuff into a screen. Eyebrow labels, extra buttons, more copy than anyone will read. You have to work meticulously toward what you actually want.
Their marketing lead named it better than we did, "A second brain of design."
Once the components existed, we redid the core pages against them. The library gives us a place to generate from instead of starting over each time. We still work in Figma when that's the better tool. What changed is that the generation now has a through line to production.
Your record needs four things.
- what you chose
- why
- evidence behind it
- when you decided
Then fold it back into the tools, so the next round can reuse it. Otherwise it stays stuck in a prompt.
Memory matters more than ever, because it has to be shared across a team. That's how you stop going backwards with every generation.
3. Find the human element
This one overlaps with the first two. When you know what you're deciding, and you keep the reasons, what's left is the part AI can't give you.

AI will mimic. It gets close to what you're going for in look and sound. People still notice when the feeling is missing. That's why you test with a lot of actual people.

We ran a round on their enterprise page with about 100 people in the target audience. The page matched expectations for 68% of them, down from 81%. Plenty of them called it confusing, and a lot of that traces back to the extra material a generation adds. More words, more boxes, less clarity.
The testimonials failed too. When content gets generated and nobody is sure where it came from, readers pick up on it. One person put it simply: "I cannot know if it's genuine."
We saw the same thing in the navigation. We showed people two versions. Two capabilities the team was proud of got dropped, because people didn't reach for them. The real-time work moved up, because people said it mattered more. On pricing, about one in five objected to the page for a single reason. They couldn't tell what it would cost them.
You don't get any of this from a model. You don't get it from a stakeholder review either. People rarely volunteer what's missing until you put something in front of them and watch.
A good test lets you keep moving fast. What it adds is knowing what you learned and how to apply it next.
Where this stands
The marketing pages are built on the style guide and the component library now. All core pages generate consistently, and one came out fully generated with no hand work in it at all.
Most of the work so far went into the process, not the pixels. Getting the style guide into the generations, and putting enough structure around it that the output holds up.
What skipping this costs
You can generate a set of pages in an hour. Working out what actually matters in them takes ten. Skipping the second part doesn't save you the time. It moves it. You pay it back in rework, in launches that slip, and in decisions that come down to whoever sounds most sure.
Conviction, and the halfway problem
Working this way takes commitment and discipline. That was always true. The workflow is what changed.
If you use AI halfway, waiting for it to prove itself, you get outputs that don't work and a path that stays blocked. You get the speed and none of the compounding, because you start at the front door every time.
Commit and learn by doing. Build the baseline. Write down the reasons. Put the work in front of real people every week. The loop gets faster each time around.
Start with one page
Pick the page that matters most to your business and let us audit it.

You'll get three things back:
- What the page asks a visitor to decide, and where it loses them.
- One test with your audience, so the read comes from people instead of opinions.
- The gaps we see between what you're generating and what you're deciding.
It takes a week. It works best when you have traffic to learn from and someone who can act on what we find.
So, where's your last 20% going?