How we use AI to build software: our AI-assisted lifecycle
What does an AI assisted development lifecycle look like inside an engineering company that uses coding agents every day?
AI drafts specs, scaffolds, writes tests and boilerplate; engineers own architecture, security and the decisions that matter. We use coding agents at every stage, but every line they produce passes the same review, evaluation and security gates as human-written code, and the engineer who merges it is accountable.
An AI assisted development lifecycle is how we build everything, including the AI systems we deliver. Coding agents draft specifications from interview notes, scaffold services, write the first pass of tests, generate boilerplate, and propose fixes from failing evaluations. Engineers decide the architecture, own security and data handling, review every change and make the calls that a model cannot be held accountable for. This article walks through the lifecycle stage by stage and is candid about where the tools help a great deal, where they help a little, and where we keep them out.
The principle: assistance is not delegation
The productivity gain from coding agents is real, and it is easy to squander. A team that accepts generated code without reading it ships faster for a month and then spends a quarter finding out what it shipped. Our rule is that a model may produce any artefact, but a named engineer merges it and is accountable for it, exactly as if they had typed it. That accountability is what keeps the gains from turning into debt. It is also why our fixed-price programs can be short without being reckless: the speed comes from the tools, the safety comes from the gates.
Where AI works in our lifecycle
| Stage | What the AI does | What the engineer does |
|---|---|---|
| Discovery | Transcribes and summarises interviews, drafts the scope document and risk register from notes | Runs the interviews, challenges the draft, decides what is out of scope |
| Architecture | Proposes options with trade-offs, drafts diagrams and interface definitions | Chooses, and documents why; owns data flows and trust boundaries |
| Scaffolding | Generates project structure, CI configuration, infrastructure definitions, API stubs | Reviews against our templates, sets secrets and permissions by hand |
| Implementation | Writes first-pass code for well-specified units, refactors, explains unfamiliar code | Specifies precisely, reviews line by line, handles anything touching auth, payments or PII |
| Testing | Drafts unit tests, generates edge cases, proposes golden-set cases for evals | Decides what must be tested, checks tests assert the right thing, owns the evaluation threshold |
| Review | Summarises diffs, flags likely defects and style violations before human review | Performs the review; the AI summary is an input, never a substitute |
| Security | Runs dependency and static checks, explains findings | Threat-models, decides severity, fixes and verifies |
| Documentation | Drafts runbooks, API docs and handover notes from the code and commit history | Edits for truth; runs the runbook to prove it works |
| Operations | Clusters incidents and logs, drafts post-incident notes | Diagnoses, decides, and changes the system |
Coding agents in practice: what actually happens
Specification before generation
The quality of generated code tracks the quality of the specification, so the engineer's first job has shifted from typing to specifying. A unit of work is described with its inputs, outputs, error cases and the tests it must pass before an agent is asked to write it. When the specification is vague, we do not ask the agent to guess; we go back to the client or the design. The most useful side effect is that specifications are now written down, which they often were not.
Tests and evals are the contract
An agent that writes both the code and the tests will happily write tests that pass. So the engineer writes or reviews the tests first, and for AI features the evaluation suite exists before the prompt. The agent then iterates against those tests, and the engineer reviews the result. This is the same evals-over-demos stance we hold clients to, applied to ourselves; our prompt versioning and evaluation article describes the mechanics for prompts specifically.
Where we keep the agent out
Authentication and authorisation logic, payment flows, anything that handles personal or regulated data, cryptography, and infrastructure permissions are written and reviewed by people, with the agent limited to explaining and checking. Not because models cannot write that code, but because the cost of a subtle error there is out of proportion to the minutes saved. We also do not let agents act autonomously on production systems; they propose, a person applies. The same boundary applies to database migrations and anything that deletes data: the agent may draft the script and the rollback, but a person runs it, on a rehearsal environment first.
What it changes for the client
Three things. Programs are shorter, which is why a six-week MVP is a realistic fixed scope rather than a marketing claim; our note on AI-accelerated development is honest about what does and does not speed up. Documentation is better, because drafting it costs little and the engineer's time goes on making it true. And the client's own team inherits a working setup: the agent configuration, the review rules and the evaluation harness are part of the handover, so the productivity does not leave with us. What does not change is who is responsible. The engineer who merged the code is the person you can ask about it.
Security and confidentiality of client code
Client code and data are used with model providers under terms that exclude training on the inputs, through enterprise or API agreements rather than consumer tools, and agent access is scoped to the repository and environment in question. Where a client's policy prohibits sending code to any external provider, we use self-hosted open-weight models for assistance, at some cost to speed, and say so in the plan. Secrets are never in the context an agent sees. Our security page lists the controls, and the same private deployment patterns we sell in private agentic AI are the ones we use internally.
The agent writes the draft. The engineer signs it. Nothing ships on a draft.
A worked example
A university needed its ERP modernised without a rewrite: a large, old codebase with sparse tests and no one left who remembered why parts of it existed. Agents were used to read and explain modules, to draft characterisation tests that captured current behaviour, and to propose the API layer that new services would use. Engineers decided which modules to wrap and which to leave, reviewed every test to confirm it asserted real behaviour rather than restating the code, and wrote the data-migration steps by hand. The agent-drafted documentation became the first accurate map of the system the university had owned in years, after the engineers had corrected it. The approach is outlined in the legacy ERP modernisation case study and is the standard shape of a ReCore program.
Team and timeline
Our teams are small because the tooling lets them be: a Launch 6 MVP is typically two engineers and a product lead for six weeks at $26,500 to $45,500, and the price on the pricing page reflects that. The client's engineers are welcome in the repository from day one and receive the agent configuration, review rules and CI gates as part of the handover pack. For clients who want to adopt the lifecycle in their own teams, a short enablement engagement sits inside our product engineering practice and runs alongside a build rather than as a training course.
Before you start: a checklist
- Ask any vendor which stages they use AI in and which they deliberately do not
- Ask who is accountable for a merged change, and whether the review is human
- Confirm your code will not be used to train provider models, and under which agreement
- Decide whether any of your code must never leave your infrastructure
- Require tests and evaluations to exist before generated code is accepted
- Ask for the agent configuration and review rules in the handover
- Check that documentation has been executed, not just generated
Glossary
- Coding agent: a model-driven tool that reads a repository, proposes changes and runs tests, working from an engineer's instruction
- Scaffolding: the generated skeleton of a project: structure, build configuration, CI, infrastructure definitions and stubs
- Characterisation test: a test that records what existing code currently does, written before changing it
- Review gate: the rule that a named engineer reviews and merges every change, generated or not
- Evaluation harness: the runner, graders and golden set that score an AI feature on every change
- Self-hosted model: an open-weight model run on infrastructure the client controls, used when code may not leave it
Questions clients ask
- Are we paying for AI-generated code at human rates? You are paying for a fixed outcome at a fixed price. The tooling is why the price is what it is; the review is why the outcome holds.
- Is generated code lower quality? Unreviewed, often. Reviewed against tests written by a person, no; and it is usually better documented.
- Which models do you use for coding? Several, chosen per task and switched as they improve. The same model-agnostic stance applies to our tools as to yours.
- Can our team keep using the setup after handover? Yes; it is in the repository and the runbook, and we walk your engineers through it in the final week.
Related reading
Read how the stance fits the rest of our principles on the about page, and the companion piece on evals over demos. For the provider-side terms on data use in API products, Anthropic's documentation and OpenAI's platform documentation are the primary sources to read before agreeing a policy.
AI makes our engineers faster; it does not make them less responsible, and the difference is the whole method.
Frequently asked questions
How much faster is AI assisted development?
▾
It depends on the stage: scaffolding, tests, boilerplate and documentation speed up a great deal; architecture, security and integration with messy existing systems speed up far less. We do not quote a percentage because it varies by project.
Do coding agents write your production code?
▾
They draft much of it. A named engineer specifies the work, reviews the draft against tests, and merges it. Code touching authentication, payments, regulated data or infrastructure permissions is written by people.
Is our source code sent to AI providers?
▾
Only under API or enterprise terms that exclude training on it, scoped to the repository in question, and never with secrets. If your policy prohibits it, we use self-hosted models and plan for the slower pace.