Regenerating documentation for undocumented systems with AI
How does AI code documentation work for systems nobody documented?
AI-assisted documentation reads the code and produces reviewable docs for modules nobody documented, with your team validating. The model traces entry points, data flows and business rules, drafts module summaries and diagrams, flags what it cannot infer, and the people who run the system correct and approve each page.
AI code documentation is the practice of pointing a language model at a codebase and having it produce a first draft of what each module does, how data moves through it, and which business rules are buried in the conditionals, then having people who know the system correct it. For an old codebase whose authors have left, that first draft is the difference between a modernisation that starts in week one and one that spends a quarter in archaeology. The model does not replace the reviewer; it removes the blank page. This article explains how the process runs, what it gets right and wrong, how it feeds a modernisation, and what it costs.
Why documenting old codebases matters before anything else
Every modernisation decision depends on knowing what the system does: which modules to replace first, where the business rules live, which jobs run at night, what the data means. Without documentation, that knowledge is reconstructed by reading code, interviewing whoever is left, and guessing. The guesses become bugs at month end. Legacy code documentation produced early, even imperfectly, shortens every later step: characterisation tests know which inputs to sample, the API façade knows which capabilities to expose, and the team knows which parts of the system nobody should touch during a fee window.
What AI generated docs can and cannot produce
| Documentation type | AI draft quality | Human role | Notes |
|---|---|---|---|
| Module summaries | Good for cohesive modules; weaker for tangled ones | Correct scope and naming | The fastest win |
| Entry points and call graphs | Reliable when derived from static analysis | Confirm which are still used | Combine with request logs |
| Data flow and schema descriptions | Good for column meaning inferred from code; weak for meaning known only to users | Supply business meaning | Finance and operations must review |
| Business rules | Finds conditionals; cannot say why they exist | Explain intent, flag obsolete rules | The highest-value review |
| Sequence and architecture diagrams | Good as text-based diagrams | Check against reality | Regenerate as code changes |
| Runbooks and operational notes | Poor; depends on knowledge outside the code | Write from experience | AI can structure, not source |
| API reference | Excellent when generated from a contract | Approve descriptions | Keep in the pipeline |
How the process runs
Inventory and static analysis first
The model works best when it is given structure, not a folder. Static analysis produces the list of files, functions, classes, database tables, cron entries and external calls, plus a dependency graph. Request logs, where they exist, show which entry points are live. That inventory is fed to the model alongside the code, so its drafts describe the system as it is used rather than as it was once written.
Draft bottom-up, then summarise
Documentation is drafted from the leaves upward: functions, then modules, then subsystems, then the system overview. Each level uses the approved level below it as context, which keeps the model grounded in what the code does rather than in what similar systems usually do. Long-context models make this practical for modules of tens of thousands of lines, but the leaf-first order is still what keeps summaries accurate.
Mark confidence and gaps
Every generated page carries two things: a confidence note per section, and an explicit list of what the model could not infer from the code. Column meanings that only a user would know, the reason a rule exists, whether a dead-looking job is actually dead. That gap list becomes the interview agenda for the remaining experts, which is a far better use of their hour than asking them to describe the whole system from memory.
Review, correct, approve
Pages are reviewed by an engineer for technical accuracy and by a business owner for meaning. Corrections are made in the document, and the corrected page is fed back as context for the next drafts, so the model's later output inherits the reviewers' knowledge. Nothing is published as documentation until a named person has approved it. AI generated docs without that step are plausible fiction, and plausible fiction is more dangerous than no documentation at all.
Keeping the documentation alive
Documentation that is generated once and left rots at the same rate as the hand-written kind. The fix is to keep it in the repository beside the code and regenerate the affected pages in the pipeline when a module changes, with the diff sent to a reviewer. API reference is generated from the contract on every build. Architecture diagrams are kept as text so they can be diffed and regenerated. The approved corrections live in the repository too, so a regeneration never loses what the reviewers taught the model. Anthropic's documentation on long-context prompting is a useful primary reference for feeding large modules to a model without losing accuracy.
What this feeds in a modernisation
- The characterisation test plan: which entry points and inputs to sample, from the call graph and the live-endpoint list
- The API façade design: which capabilities exist and which are still used
- The slice sequence for a strangler migration, ordered by dependency and value
- The runtime upgrade inventory: removed functions, extensions and libraries per module
- The risk register: modules with business rules nobody can explain get a parallel run, not a quick swap
- Onboarding for the engineers who will do the work, cutting weeks from the first phase
Each of those is a first step in our legacy-to-AI modernization programme, and the documentation pass is usually the first deliverable of a Sprint Zero when a client is deciding whether to modernise at all. The strangler pattern and characterisation tests both depend on it.
Limits and honest caveats
The model can only document what is in the code and what reviewers tell it. It will confidently describe a function that is never called, unless the live-endpoint list says otherwise. It will name a column by its code usage and miss that finance treats it differently. It cannot write a runbook for an outage nobody has recorded. And on very tangled modules it will produce summaries that are true at the sentence level and misleading at the paragraph level. None of these is a reason not to use it; each is a reason the review step is not optional and the confidence notes are part of the output.
A worked example
A university's administrative system had been built over fifteen years by successive in-house developers, the last of whom had retired. Nobody could say which of its several hundred scripts still ran. Static analysis and a term of request logs produced the inventory; the model drafted module summaries, a call graph and a first pass at the data dictionary, marking every column whose meaning it could not infer. The registrar's office and the finance office spent two afternoons on the gap list, which resolved most of it and revealed one nightly job the results process silently depended on. The approved documentation set the sequence for the modernisation that followed, described in the university ERP case study, and it now regenerates in the pipeline as modules are replaced.
Team and timeline
A documentation pass for a mid-sized system is an engineer who sets up the analysis and generation pipeline, a technical writer or senior engineer who runs the review, and two to four hours a week from each remaining expert on your side over three to four weeks. It fits inside a Sprint Zero at $3,250 / ₹2,00,000, credited against the build that follows, or as the opening weeks of a ReCore programme; the pricing page has the ranges. The documentation, the pipeline and the review history are yours, and keeping them regenerating is part of a Care Plan. For systems that also need an API layer, the API development and integrations team picks up where the documentation leaves off.
Before you start: a checklist
- Read access to every repository, including the scripts outside the main application
- Static analysis output: files, functions, tables, cron entries, external calls and a dependency graph
- Request or job logs covering at least one business cycle, to separate live from dead code
- A list of the people who still know parts of the system, with a few hours each booked
- A decision on where documentation lives and how it is regenerated on change
- A review rule: no page is published without a named approver
- A confidence and gap convention that every generated page follows
- Data handling agreed before any code or data leaves your environment, with a self-hosted model where required
Questions clients ask
- Will our source code be sent to a model provider? Only with your agreement; for regulated clients we run an open-weight model inside your environment.
- Can it document a database with no comments? It infers meaning from how the code uses each column and marks the rest for your finance and operations teams.
- Is the output good enough to hand to auditors? After review and approval, yes; before review, no.
- Does it work for COBOL, classic ASP or old PHP? Yes; the process is language-agnostic, though static analysis tooling varies.
- How do we stop it going stale? Regenerate affected pages in the pipeline and send reviewers the diff.
Related reading
See application modernization vs rewrite for the decision the documentation informs, adding an API layer to a legacy monolith for the next step, and self-hosted LLMs for keeping code inside your walls.
Let the model draft, let your people correct, keep it regenerating, and the undocumented system becomes one you can safely change.
Frequently asked questions
Can AI document a legacy codebase accurately?
▾
It produces an accurate first draft of what the code does, especially module summaries, call graphs and entry points. It cannot know why rules exist or what columns mean to the business, so reviewers correct and approve every page before it counts as documentation.
How long does AI-assisted documentation take for an old system?
▾
Three to four weeks for a mid-sized system, including analysis, generation, review and approval. It fits inside a ten-day Sprint Zero for scoping, or the opening weeks of a modernisation programme.
Is generated documentation safe for regulated organisations?
▾
Yes, when the model runs inside your environment and every page is approved by a named person. Our legacy-to-AI modernization service uses self-hosted open-weight models where code cannot leave your infrastructure.