When I rebuilt my website with AI agents, they wrote almost all the code.
The more interesting question came afterwards.
If implementation is becoming cheap, where should the engineering effort go?
For me, the answer has increasingly been: into the environment the agents work within.
Agents need constraints, not micromanagement
When I’m working with coding agents, I deliberately give them quite a lot of freedom.
If I want something changed, I generally don’t tell the agent exactly how to implement it. I want it to work that out. If I’m specifying every technical detail myself, I’ve simply found a slower way of writing the code.
But implementation freedom isn’t the same as architectural freedom.
I don’t want a coding agent deciding afresh whether accessibility matters on this project, or choosing an entirely different CSS architecture because this particular model happens to favour one. And I don’t want the quality of a build depending on whichever agent or model happened to work on it that day.
I don’t want the coding agent making architectural decisions I’ve already made.
That isn’t about limiting AI. It’s about engineering the environment in which it works.
The delivery system sits above the agent
That thinking led to the Astro Site Playbook.
It provides a public, MIT-licensed Astro foundation with the architectural principles, design-system approach, content modelling and quality checks I want available from the beginning of a project. But the repository itself is only part of the delivery system.
Alongside it, I now have an Astro Build Playbook: a companion skill in my wider Agency Skills setup that guides agents through the actual delivery process, from design ingestion and content modelling through implementation, CMS choices, hosting and verification.
The distinction matters. The Site Playbook contains things that should be generally true of an Astro build. The build skill contains knowledge about how the agent should conduct the work. The individual project contains what is specific to that client or site.
That separation gives the agent room to solve the implementation without mixing project-specific decisions into the reusable system.
Useful knowledge shouldn’t belong to one agent
I don’t always use the same coding agent. Different agents can work on different projects, but they access the same shared skills. That means useful knowledge isn’t trapped inside one conversation or tied to one particular model.
If an agent encounters something during a build that it thinks should change the way we work, I’ll usually ask it to suggest the change. At that point, one of three things normally happens:
- it belongs in the public Astro Site Playbook and needs to be expressed generically;
- it belongs in the Astro Build Playbook because it’s about how the delivery process should work;
- or it’s specific to that project and shouldn’t be generalised at all.
The agent is usually quite good at working out the distinction. I review the proposed change and decide whether to approve it.
Once approved, the updated skill is immediately available to the other agents using the same system. So the useful knowledge sits above the individual agent rather than being locked inside one session or model.
Each build should improve the next one
That creates a simple feedback loop:
build → learn → generalise → improve the delivery system → next build
The Playbook isn’t a static rulebook. Each build can reveal something that should be clarified, simplified or handled better next time, and those lessons can then improve the environment used for the next project.
The code from one project may never be reused. The lesson from it can be.
Expertise becomes reusable
My background gives me an advantage when working this way. I can inspect generated code, understand the architectural choices being made and recognise something I don’t like even if the site appears to work perfectly well.
But once some of that judgement is encoded into a shared delivery system, it no longer has to be explained from scratch each time. A future agent can benefit from decisions made on earlier builds, and a different agent doesn’t need to rediscover what the previous one learned.
That doesn’t eliminate the value of expertise.
It makes some of that expertise reusable.
More instructions aren’t necessarily better
There is an obvious danger with this approach. Every time an agent encounters something new, you could add another instruction. Before long, you’ve created a huge body of context that every agent has to work through before it can do anything useful.
I don’t want that either.
From time to time I use an agent to review the skills portfolio itself, looking for duplication, unnecessary complexity and opportunities to consolidate. A recent review reduced the Etch material from five skills to two, and across the wider Agency Skills portfolio the number of Markdown files fell from 140 to 103.
That’s not just tidying up. Context has a cost, and I don’t want an agent burning a load of tokens reading instructions before it can start doing the work.
The goal isn’t to capture the maximum amount of knowledge.
It’s to provide enough context for the agent to make good decisions consistently, without making the system heavier than it needs to be.
Sometimes the right improvement is adding something. Sometimes it’s deleting thirty-seven files.
Repeatable doesn’t mean repetitive
One risk with any opinionated delivery system is that every website starts looking the same. That’s not what I want.
The project brief should determine the design direction. Typography, spacing, colour, scale and the other choices that give a site its character can change completely from one project to the next.
The Playbook’s job isn’t to dictate those choices. Its job is to keep the build within the engineering standards I’ve already established.
The brief determines what the site looks like. The Playbook determines the standards it has to meet.
That gives the agent freedom where freedom is useful, without asking it to reinvent decisions that don’t need to be made again.
Standards are more useful when they can be tested
A written set of principles is useful. An executable one is better.
If I say accessibility matters, some aspects of accessibility should be checked automatically. If the delivery system expects particular conventions, some of those can be tested too.
Not everything can be automated, of course. A passing test doesn’t prove that a site is good. Accessibility can’t be reduced to a perfect automated score, and neither can design quality or usability.
Human judgement remains important, but routine checks should happen consistently without somebody having to remember them.
Automate what can be automated.
Pay attention to what can’t.
This isn’t really about Astro
The Astro Site Playbook is where this approach started, but the underlying idea is broader.
The same thinking is already influencing the way I structure and simplify the skills I use for Etch and WordPress work, even though those aren’t packaged as a separate playbook in the same way.
That’s important because I don’t see Astro and WordPress as opposing choices. Some projects are better suited to a static Astro build. Others genuinely need the editing experience, plugin ecosystem or application behaviour that WordPress provides.
The technology should fit the problem. The delivery system should help make whichever route is appropriate repeatable.
Frameworks will change. Browser capabilities will change. AI agents will change. The software engineering principles behind the way I work are considerably older than any of them.
The engineering principles are what make the delivery repeatable.
What does software engineering look like when AI writes the code?
I don’t think we have the complete answer yet, but my own answer is becoming clearer.
I’m spending less time telling an agent exactly how to implement something, and more time deciding what environment it should work within. That means thinking about which decisions have already been made, where the agent should have freedom, what knowledge belongs in the reusable system, what should remain specific to the project, and what can be verified automatically.
Those feel like software engineering questions to me.
The codebase still matters, but increasingly I’m also engineering the delivery capability that produces the codebase.
AI agents are becoming very good at writing code.
Our job isn’t simply to give them more code to write.
It’s to build the conditions in which different agents can produce good software repeatedly, and improve the system as they go.