Onboarding documentation stays accurate when its facts (commands, variables, structure, services) are read from the codebase instead of copied by hand, and only judgment (why, in what order, whom to ask) stays written by people. An onboarding guide does not go wrong through neglect. It describes the code as it was the day it was written, and the code kept moving.
This article follows one onboarding page at two dates, splits it into three kinds of content, and says which kind the code can carry.
Why the arrival guide is the first document to go stale
Onboarding documentation: the pages a person joining an engineering team follows to get a working environment and understand how the product is organized.
The reader of an onboarding guide is the worst possible judge of its accuracy: they are new to the subject. When a command fails, they cannot tell whether the fault is theirs or the page's. They ask a colleague, the colleague fixes it out loud, and the page stays wrong for the next person.
Two things make it worse. The page is read rarely (a few times a year in a small team), so nobody rereads it between hires. And it was written once, by someone who already knew everything, at a moment when everything felt obvious. The result is a document the team does not notice aging until the next arrival.
It is one instance of documentation drift, with a costlier consequence than usual: the error lands on the person whose ramp-up time was the very thing the page was meant to shorten.
The same page, at two dates
Here is a "Run the project locally" section as it might exist. Names, commands and dates are illustrative, not taken from any customer.
| The day it was written (March) | Six months later (September) | |
|---|---|---|
| Start command | make dev | make dev (unchanged) |
| Database | "Start Postgres with docker compose up db" | The service is now called postgres; db no longer exists |
| Environment variables | "Copy .env.example" | Three variables added to the file; the page mentions none of them |
| Sample data | "Load seed.sql" | The script was replaced by a make seed command |
| Runtime version | "Node 18" | Production runs Node 22 |
| Whom to ask | "See Claire for access" | Claire moved to another team |
Five of six rows went wrong without any decision about them. They are side effects of changes made elsewhere, for good reasons. The last two rows stand apart: one is a fact you can verify in the code, the other is information about people, which the code does not contain.
Three kinds of content, three treatments
An onboarding guide mixes three kinds of statement. Separating them is half the job.
Facts verifiable in the code. The command that starts the project, the variables read at startup, the services declared in configuration, the runtime version, the folder layout. Each has a single source that changes when it changes. Copying them by hand creates a second copy that is bound to diverge.
Facts the code does not state, but history keeps. Why this service was split off, since when this module has been deprecated, which change introduced that variable. Git history (commits and diffs) carries them, often without anyone having put them into words. This is the territory of architecture decisions recovered from git.
Judgment. In what order to read the code, what can be ignored in the first month, whom to reach for what, what annoys people in code review. No tool derives it, and it is the part worth writing with care, because it is the only part a person does better than a reading of the code.
A practical rule follows: write by hand only the third kind. The other two are read, not written.
What "generated from the code" means, and what it does not
It does not mean a document produced once by a language model, reviewed, then left in place. That would be the same static page with a different author, and it would go stale at the same rate.
It means a page whose facts are read from the code every time it is asked for. DeployIt works this way: it reads the real code (resynced on every push) and the git history (commits and diffs), then checks the code before asserting anything. A new hire's question ("how do I run the project?") gets an answer built on the repository's current state, not on a page frozen in March.
Two concrete consequences. When docker compose up db becomes docker compose up postgres, the answer changes at the next push, without anyone remembering to open a page. And when the code cannot answer (the information simply is not there), you are entitled to expect the tool to say so instead of filling the gap. A read-only tool describes, it does not invent.
What is left to write by hand
Four pages are worth keeping in human hands, and they are short.
- The first-week path: what to do in which order, and what is expected of a person by the end of each day.
- The responsibility map: who decides what. It goes stale too, but at the speed of teams, not of code.
- The team's traps: the three mistakes everyone makes once.
- The glossary: the house words a newcomer cannot guess.
Each fits on half a page. That is the opposite of a forty-page wiki, and it is deliberate: the more derivable facts a document carries, the more lines can become wrong without anyone knowing.
A test for the page you already have
Take your current onboarding guide and ask of each line: if the code changed tomorrow, would this line know?
- If not, and the line is a fact (a command, a name, a version), it is a line to derive.
- If not, and the line is advice, it is where it belongs.
- If you cannot tell, you have no source for it. Note it down: it is the next one to go stale.
Count the lines of each kind. The ratio of copied facts to judgment tells you how much of your guide depends on the team's memory.
"Our README already does this"
Keep it. A well-kept README is an excellent source for the first kind, and nothing here asks you to drop it. Its limit is that of any hand-written file: it is accurate as long as whoever changes the code remembers to edit it in the same pull request. A review rule that requires it helps, a CI check that catches the omission helps more, and neither covers the facts the README never mentions. The README stays one input among others.
What DeployIt needs to see, and what it exposes
A read-only GitHub connection, one server per project, revocable at any time. Answers cover what the code establishes (a command, a variable's name, a service). Secrets do not belong in a well-kept repository, and an onboarding guide has no need of them. The code itself stays in the repository.
Three frequent questions
Should we delete the onboarding wiki? Not at once. Find the pages that carry only derivable facts and retire those first; keep the ones that carry judgment.
What about a team of three? The problem is quieter there, not smaller: everyone assumes "everybody knows". The first hire demonstrates otherwise, within days.
Can a new hire query the tool directly? Yes, that is the intended use: the question goes to the product's expert instead of to a page to be found. An AI agent can query it as well, through the MCP server.
Where to start
Pick the most-visited section of your guide, probably "run the project locally". Separate its facts from its advice. Delete the facts, keep the advice, and let the question "how do I run this?" go to the code. Do the next page the following month.
DeployIt's MCP server is in early access / design partner program: if your new hires ask the same questions at every arrival, the program is open. DeployIt is free during early access; paid plans arrive in January 2027. For the same mechanism on the customer side, see how a help center stays in sync with releases, how to detect documentation drift and the DeployIt help center.
Continue reading
What should an API changelog for integrators contain?
An integrator reads an API changelog to decide one thing: do I have to touch my code? Six fields answer it, and most changelogs carry two.
Can you generate Notion documentation from your code without overwriting what your team writes?
Generating Notion documentation from code is the easy part; making it coexist with what your team writes by hand is the real problem. The four writing rules that decide whether the page survives.
How do you build a "What's new" page generated from git?
A "What's new" page is not a changelog with better styling: its window belongs to the reader, not to the release. What you have to derive from diffs to make it work, and the triage that remains.