Code Production Has Changed. How Should Software Teams Collaborate?
Acti on decision boundaries, documentation, and AI development

Software teams have long used automation for builds, tests, and deployments. Writing code, though, has remained largely a craft. Engineers work through requirements line by line, making decisions and checking with colleagues along the way. AI is beginning to industrialize code writing itself. That leaves teams with a question: how do we handle the judgment and coordination that used to happen while someone was writing the code?
Suppose we ask two AI agents to build a photo app, one handling the mobile client and the other the desktop client. When a user deletes a photo on their phone, should it disappear from their computer too?
Both agents can implement deletion. But before they begin, several things need to be decided: does deletion affect only the current device, or the entire account? When do offline devices synchronize? Can an accidental deletion be undone?
If each agent adopts its own reasonable interpretation, the mobile client may consider the photo deleted while the desktop client still keeps a copy. At the next sync, the photo might even reappear on the phone.
The two agents can write their code independently, but the team has to agree on what deletion means for the user. As code gets written faster, deciding who needs to settle these questions, and when, becomes more pressing.
AI writes the code, but collaboration still revolves around people
When the Acti team first used Cursor, we used it mainly for autocomplete. We chose where to make a change and what it should do; AI filled in the next few lines. As models improved, we gave them larger tasks, eventually handing over entire modules. We reached a point where every code change was written by AI, whether it added a feature or changed existing behavior.
For a long time, our way of working barely changed. We broke down requirements, explained the background, checked the result, and gave the next instruction. Someone still had to coordinate changes across modules. A new session meant explaining earlier decisions again. AI wrote the code; people kept the work connected, one handoff at a time.
Cursor describes a similar shift in The third era of AI software development: from autocomplete, through agents guided at each step, to agents that can work independently for longer stretches. As more work happens between our instructions, we have to choose where to intervene.
We call this "human-assisted AI development." We set the goals and take responsibility for the important tradeoffs, while AI carries the work forward. Our documentation needs to tell an agent which choices it can make on its own and which would change something others have agreed to rely on.
A sound choice does not settle who gets to decide
Decision-making is at the heart of engineering. A feature spanning a few hundred lines of code may contain dozens of choices: whether to retry a failure or return an error, preserve old data or migrate it, favor responsiveness or data freshness, or let the client or the server determine a particular state.
An engineer used to work through these choices while writing the code: spot a problem, check with a colleague, then carry on. Some decisions made it into design documents. Others survived only in the code, with the reasoning left in someone's memory or a chat that would be hard to find later.
An agent can write a lot of code in one task, making the choices we left unspecified. An unresolved question may become an implementation before anyone else notices it. Interfaces and responsibilities have always needed discussion. Those discussions do not happen automatically when AI writes the code, so we have to make them an explicit part of the workflow.
A choice can be reasonable without the agent being allowed to make it. Reorganizing internal functions may be up to the agent. Changing what deletion means affects the other client and the user, even if the proposed design is sound.
A better model may suggest a better design. It still needs to know whose agreement is required to change an existing commitment. The team remains responsible for the consequences.
We try to give AI autonomy according to what its decisions affect. It can work independently within agreed constraints. To change those constraints, it needs to involve the people responsible for them and those who depend on them.
Making collaboration explicit through documentation
Existing code alone cannot fully express these boundaries.
An established codebase contains compromises and unfinished migrations alongside the current design. An agent may understand what a piece of code does without knowing why it is still there. Is that behavior a promise to another client, or a workaround nobody has removed yet? Reading more code does not always answer that question.
For now, we prefer to write down important agreements, give the agent the relevant documents, and then have it inspect the code it needs to change. A client needs to know what a field means and how caching and errors work. It usually does not need to understand the whole backend. Nor should the backend have to keep reading client code to discover what behavior those clients rely on.
An index points the agent to the relevant sections, including their sources and the versions they apply to. Anthropic makes the case for selecting and loading context as needed in its article on context engineering. Even with a larger context window, we still have to decide which information applies.

We ran into this while reviewing a request to manage demo content. Our operations team wanted to edit a demo or stop making it available. The existing agreement allowed clients to keep using a cached copy indefinitely. Adding edit and disable controls to the admin interface would leave those clients free to keep playing the old version.
An admin action could succeed while the user kept seeing the old content. The backend could decide how to store a record of that action. But iOS, Android, and the backend had to agree on how clients would detect a change and whether a disabled demo could still play from cache. We made this distinction explicit in the document: internal backend design did not need client sign-off; delivery and caching changes did.
The revised design called for a new cache key to identify content changes. It also required clients to stop playing cached copies when the server reported the content unavailable. Each side needed to agree to and implement those rules. Its agent could still choose how to structure the code. The behavior that crosses the boundary is what needs a shared decision.
An Architecture Decision Record (ADR) records the choice and why we made it. Contracts describe the behavior each side must follow. Role and module documents set out responsibilities, and acceptance criteria say what the result should be. A new agent session can use these records to work out what is settled, what is still open, and whom to involve in a change.
AI can draft the ADR and investigate what a change would affect. People review the tradeoffs and take responsibility. A local fix should stay lightweight; changing shared behavior or an important constraint needs discussion.
Clear boundaries bring new questions
Then the documents began to accumulate. Which decision should an agent follow if a later ADR changed part of an earlier one, but a module document still referred to the original? Loading every document would leave the agent to resolve that conflict for itself.
We settled on one authoritative source for shared documentation, with branch manifests recording which designs applied where. A design could be approved and implemented on a feature branch while released clients still followed the old agreement. Approval and applicability had to be recorded separately.
AI can produce plans and checklists faster than we can read them. The discussion of spec-driven development in practice describes familiar problems: too much process for small changes, specifications that are hard to maintain, and agents that miss instructions anyway. A document needs to earn the time it takes to review. Does it clear up an ambiguity or preserve a useful reason for a decision?
Even with clear instructions, an agent will eventually encounter something the documents do not cover. Asking a person about every detail holds up the work. Letting every agent quietly fill in the gaps can leave different parts of the system working from conflicting assumptions.
We want the agent to return with enough detail for someone to make a decision: what conflicts, what would change, and what options the evidence supports. We are still working out how to catch these moments without interrupting every task.
Decision boundaries need maintenance too. A division of responsibility that works today may need to change as the product and the models develop.
Agreements need evidence that they are in effect
An accepted ADR does not mean the feature is running. The code may be merged but still waiting on a dependency or a deployment. And a deployed feature may still need to be checked with users.
Harness engineering for coding agent users distinguishes guidance before an action from feedback afterward. Our documents tell an agent what it should do. Reviews and checks of the running system help us see whether it did.
Documentation expresses intent; runtime behavior provides evidence. When they disagree, we have to investigate. A document cannot prove that a feature works, and an old implementation cannot prove that its behavior should be preserved. "Unknown" and "not yet verified" need to be acceptable answers.
We want to follow a decision through to the released software: which clients use it, where it is deployed, and what has been verified. An agent could then find the right constraints before starting, check what a proposed change would affect, and say what remains unverified when it finishes.
Some constraints are clear and stable enough to check automatically. Product tradeoffs and user experience still need judgment informed by the running software. Passing tests does not rule out a misunderstanding: the implementation and the tests may share the same mistaken assumption.
We will judge this work by whether it makes collaboration easier. Are there fewer misunderstandings between client and server work? Less rework and repeated explanation? Once we make a shared decision, can each side get on with its part?
Organizing for a different way of producing code
In the photo app, both agents can write a delete function. The team still has to settle what deletion means across devices and check that users get that behavior. As AI takes on more implementation, that is where more of our attention needs to go.
At Acti, handing 100% of code writing to AI means changing how we work together. We use documents to show agents where they can proceed on their own and where they need agreement. We then check that the software does what was agreed.
Software's industrialization will be judged by what teams can reliably deliver. As AI takes over writing code, we need clear responsibilities and standards, with ways to check the result. Faster code generation only helps if teams can agree on what to build—and verify that they built it.