Because the agent reads a description of the API, and the API is whatever the code did at the last push. After a release the two diverge, and an agent grounded only on documentation has no way to notice: it answers with the contract as last written, in the same confident tone as before.
The fix is not a better model. It is a different source: the current state of the code, the changes it received since the documentation was last edited, and a way for the agent to say "I cannot tell" when neither is available.
What "wrong about the API" actually means
API drift: the gap between the contract as documented (reference pages, OpenAPI spec, guides) and the contract as shipped (what the code accepts, returns and enforces today).
Drift is ordinary. Nobody neglects the documentation on purpose; a release simply changes things faster than someone rewrites a page. A field renamed during a cleanup, a default changed in a fix, a validation tightened in a controller: the code carries each one on the day it merges, the page carries it when someone gets to it.
A person reading a stale page often senses something is off: the example looks old, the page is dated. An agent senses nothing. It receives the passage, treats it as true, and phrases the answer without a trace of doubt. If you have read why agents invent product features, this is the same failure from the other side: here the agent invents nothing. It quotes a source that stopped being right.
Three layers, three ages
When an agent answers a question about your API, the text comes from one of three places, and they do not age at the same speed.
- The reference and the spec. Precise, structured, and as old as their last edit. If the spec is generated from the code, it follows the shape of the contract closely. It still says nothing about behavior.
- The guides, the changelog, the forum answers. Written by people at a given date, about a given version. They age at the pace of the team's attention.
- The model's memory. Frozen at a training date, and mixed with every other API that looks like yours.
Illustrative example. A team renames a query parameter and tightens one validation in the same release. The spec is regenerated, so the parameter name is right. The "getting started" guide still shows the old name, a forum answer from last year recommends it, and the validation change appears nowhere in writing. Ask the agent three questions about the endpoint and you get three answers from three different layers, none of them flagged as older than the others.
The same question, asked twice
The exchange below is a schematic illustration, not a real session. An integrator's agent is asked whether it can still filter invoices by the old parameter name.
Agent grounded on documentation
Question: can I still call
GET /invoices?customer_id=…? Yes. Thecustomer_idparameter filters invoices by customer, as described in the "Invoices" guide.
Agent with access to the code and its recent changes
Same question. The guide still shows
customer_id, but the filter was renamed toclient_idin a change merged six days ago (commit cited). The code no longer reads the old name. The guide has not been updated since.
The first answer is not careless. It is faithful to its source, which is the problem. The second differs on three points, and only three: it cites the change with its date, it says where the documentation disagrees with the code, and it does not pretend the disagreement is resolved. That is the whole difference between a source that was current and a source that is current.
What to give the agent instead
Three things, and the order matters.
The current state of the code. Not a copy indexed last month, but something that reflects the last push. An answer about GET /invoices should come from the handler that serves it today.
The window of change. Drift is a question of time: "what moved since this page was last edited?" An agent that can see the diffs of the last days can tell whether the passage it holds is still safe. This is the same mechanism described in release awareness for AI agents, applied to an API instead of a product.
An explicit way to be silent. When the code does not settle a question, the right result is "not established", delivered as a result and not as an omission. An agent that must always produce an answer will produce the most plausible one, and plausible is exactly what a stale page is.
DeployIt covers these three from one place. It reads the code (resynced on every push) and the git history, checks the code before asserting anything, and exposes the result as an MCP server that an integrator's agent can call. The question integrators ask every week and the question their agents ask at 3 a.m. go to the same source.
"Our spec is generated from the code, so it cannot drift"
A generated spec is a real improvement, and you should keep it. It removes the most common drift: a field that exists in one place and not the other.
What it cannot carry is behavior. A default that changed, a stricter validation, a sort order that moved: the shape of the contract is identical, and the integration breaks anyway. The spec also does not explain itself. It says what the contract is now, not what it was last week, so an agent holding it cannot tell whether its earlier answer is still good.
Use the spec for what it does well, and add the code and its history for the rest.
"Does the agent then see our source code?"
A fair question when the repository is private. The connection to GitHub is read-only, a project gets its own server, and access is revocable at any time. What comes back are answers, sourced from commits and diffs: the caller learns that a parameter was renamed, not how your handler is written. The agent never holds repository credentials, and neither does the integrator behind it. For the model of access in detail, see giving AI agents access to your code safely.
Three frequent questions
Does this replace the documentation? No. Documentation remains what humans read first and what explains intent. The agent's answers come from the code; when the two disagree, that is a signal about the documentation, and it is worth fixing there.
What if the question is about a version that is not the latest? The history keeps earlier states, so "as of last month" is a legitimate question. What counts is that the answer names the version it speaks for.
Can I check what the agent saw? That is the purpose of sourced answers: each claim points to a commit or a file. A claim without a source should be treated as a claim you have not verified. The mechanics of such calls are laid out in asking your codebase questions via an API.
A source that moves with the code
An agent that answers about your API is only as current as what it reads. Documentation will lag behind releases for as long as people write it by hand, as every knowledge base does. The agent does not need the lag to disappear. It needs to see it.
DeployIt's MCP server is in early access / design partner program: if your integrators' agents already ask about your API, the program is open. DeployIt is free during early access; paid plans arrive in January 2027.
Continue reading
Where does the data go when an AI agent asks questions about your product?
A question to a product-knowledge server travels through four places: the Git provider, the vendor's infrastructure, a language-model provider, and the platform that asked. Self-hosting your portal settles the first two at most. Here is what DeployIt's public DPA says about each, and what it leaves open.
How does an AI agent authenticate to an MCP server?
An agent authenticates to a remote MCP server with an OAuth access token, obtained through a flow the server starts by answering 401. What that token can do, and what revoking it actually cuts, depends on which of three credentials you mean.
How do you ask your codebase questions via an API?
An MCP server in front of your codebase turns a plain-language question into a sourced answer. The real shift: the caller is no longer a person, and nobody proof-reads before the answer gets used.