Module 3: SDD at Org Scale
Learn to decide before the code exists: when to spec, what to enforce in code, what counts as done, and who owns it.
- 3.1When to use SDD, and when to skip it
- 3.2What improves your teams' AGENTS.md files
- 3.3How to enforce engineering rules for agents
- 3.4How to define done when agents ship more code than your reviewers can read
- 3.5How to split spec ownership between PMs, architects, engineers, and leadership
- 3.6How Microsoft adopted SDD across teams
When to use SDD, and when to skip it
Engineering Leadership newsletter author Gregor Ojstersek says SDD is overhead for a small change, like a refactor of one file or a small adjustment to a feature. Prompt the agent and review the code instead.
On a bigger project, especially one of uncertain feasibility, the spec saves a lot of time on rewrites and reviews of bad AI-generated code. Decisions on the fly are usually worse than decisions made beforehand.
Microsoft reached the same conclusion across its teams. Not every change needs the full lifecycle, so adoption should be right-sized.
If the work fits one session, skip the spec
Matt Pocock, who runs AI Hero, says the whole trigger for a spec is a build too big for one agent session, one that has to survive a split across several.
The spec exists because context windows end; it is what survives when you clear the conversation. On a single-session change it buys you nothing and adds a synthesis step where the model can drift.
The bug fix that became 16 acceptance criteria
Böckeler asked a spec tool to fix a small bug, and the workflow turned it into 4 user stories with 16 acceptance criteria. Her verdict was a sledgehammer to crack a nut.
One generated user story read, “As a developer, I want the transformation function to handle edge cases gracefully, so that the system remains robust when new category formats are introduced.”
Her second test was a 3 to 5 point feature that built on a lot of existing code. The steps and the markdown files felt like overkill, and she never finished.
In the same time, she believes, plain AI-assisted coding would ship the feature and leave her more in control.
Three places where SDD pays off
GitHub names three scenarios where the approach works especially well:
- Greenfield, zero to one. A small amount of upfront work on a spec and a plan makes the agent build what you intend. Without it you get a generic solution from common patterns.
- Feature work in existing systems, N to N+1. This is where SDD is most powerful. The spec forces clarity on how the feature interacts with the existing system. The plan encodes the architectural constraints so the new code feels native.
- Legacy modernization. The original intent is often lost to time. The spec captures the essential business logic, the plan designs a fresh architecture, and the agent rebuilds without inherited technical debt.
The core benefit, in GitHub’s words, is to separate the stable “what” from the flexible “how.”
Four rules for when to use SDD, side by side
| Source | Use SDD when | Skip it when |
|---|---|---|
| Ojstersek, Larridin | A bigger project, especially one of uncertain feasibility | A small refactor or a small feature adjustment: prompt, then review |
| Pocock, AI Hero | The work spans several agent sessions | The work is decided and fits one context window |
| Böckeler, Thoughtworks | Many situations; she writes some form of spec first herself | A small bug fix or a 3 to 5 point feature turns into a pile of markdown |
| GitHub | Greenfield, a feature in an existing system, or a legacy rebuild | Quick prototypes where vibe coding is enough |
What improves your teams’ AGENTS.md files
An AGENTS.md is the first thing your agents read on every task, and every line in it costs tokens on every run. A line earns its place only if it saves the agent something. That can be a fact it would find expensively, an intent it cannot see in the code, or a trap it falls into.
This lesson covers the three kinds of line that belong in the file and the four that do not. It also covers how to turn the file into a map to your docs instead of an encyclopedia.
PostHog’s rule to writing AGENTS.md
PostHog says your AGENTS.md is like a backpack you carry everywhere. Its goal is to save the agent tokens when it reads the file.
Three things belong in it:
- Discovery shortcuts are facts an agent will find eventually, at a high token cost. PostHog’s example is its own unguessable CLI for tests, lints, and builds.
- Undiscoverable intent is direction of travel, taste, and policy. An example is which of two storage systems to prefer when the team plans to remove the other.
- Landmines are things agents get confidently wrong and only learn from CI or people. An example is a workflow that breaks CI on backward compatibility.
Four things stay out:
- Lint-enforced content, such as naming case or spelling.
- Obvious model defaults, such as to follow existing patterns and write tests.
- Anything one ls or grep answers, such as directory trees, package lists, or 200-line reference components.
- Docs meant for humans, such as mission statements or contributing etiquette.
How to point agents at your docs instead of pasting them in
A short file that points to deeper documents beats a long file that tries to hold every rule. The agent starts from a small entry point and reads further only where the task needs it.
An OpenAI member of technical staff described what a large agent-written codebase does with its context file. The rule: give Codex a map, not a 1,000-page instruction manual. The AGENTS.md is roughly 100 lines.
The team tried one big file first, and it failed in predictable ways. It crowded out the task. It marked everything as important, so nothing was. It rotted into a graveyard of stale rules, and no tool could check it mechanically.
Treat AGENTS.md as the table of contents, not the encyclopedia. A structured docs directory is the system of record. Agents start from a small, stable entry point and learn where to look next.
The team governs the file like code and enforces it mechanically. Dedicated linters and CI jobs validate that the knowledge base is current, cross-linked, and structured correctly. A recurring doc-gardening agent opens fix-up pull requests for documentation that no longer matches the code.
From the agent’s point of view, anything it cannot access in context while it runs does not exist. Knowledge in Google Docs, chat threads, or people’s heads is not accessible to the system.
Goal: Cut your AGENTS.md to what earns its place, and measure whether the cut changed anything, in one week.
- Paste your current AGENTS.md or CLAUDE.md into your coding agent with this prompt:
Prompt
Classify every line or bullet of this context file into exactly one of: - KEEP / discovery shortcut: a fact the agent would find eventually, expensively - KEEP / undiscoverable intent: direction, taste, or policy not visible in the code - KEEP / landmine: something agents get confidently wrong and only learn from CI or people - CUT / lint-enforced: already enforced by a linter, formatter, or type checker - CUT / model default: behavior a current model does anyway - CUT / one grep away: directory trees, package lists, reference code - CUT / for humans: mission, etiquette, onboarding prose Return the file with each line tagged, then the count of lines in each bucket and the share of the file marked CUT. Do not rewrite anything. - Write the KEEP lines into a new file by hand. Do not let the agent draft it.
- Pick twenty recent, closed tasks. Run ten with the old file and ten with the new one. Record pass rate and token cost per task.
- Keep whichever file passed more. If they tie, keep the shorter one.
Expected result: a context file with only human-written shortcuts, intent, and landmines, and a before-and-after pass rate and cost that justify it.
How to enforce engineering rules for agents
At agent volume, your engineering rules hold only where a machine enforces them. Documents get skimmed. Linters, type checks, and structural tests run on every change, and they can tell the agent how to correct itself.
This lesson covers what that looks like in a codebase with zero hand-written code, and how to sort your controls into guides and sensors. It also covers where automated loops make code worse instead of better.
1. How OpenAI ran a million-line codebase on enforced rules
An OpenAI member of technical staff built a million-line product with zero hand-written code on that difference. In five months from August 2025, a team that grew from three engineers to seven produced about a million lines. The code spans application logic, infrastructure, tooling, and documentation.
They merged roughly 1,500 pull requests, an average of 3.5 per engineer per day. That took about one tenth of the time hand-written code takes. Hundreds of people use the product internally.
Documentation alone does not keep a fully agent-generated codebase coherent. Enforce invariants instead of micromanaging implementations, and agents ship fast without undermining the foundation.
The architecture is rigid on purpose. Agents work best in environments with strict boundaries and predictable structure, so the team built the application around a rigid architectural model.
Within each business domain, code can only depend forward through a fixed set of layers: Types → Config → Repo → Service → Runtime → UI. Cross-cutting concerns such as auth, telemetry, and feature flags enter through a single explicit interface, Providers. Custom linters and structural tests disallow anything else.
This is the kind of architecture you usually postpone until you have hundreds of engineers. With coding agents it is an early prerequisite. The constraints are what allow speed without decay or architectural drift.
Lints enforce taste too. Custom lints check structured logging, naming conventions, file size limits, and reliability requirements. Their error messages inject remediation instructions into agent context.
In a human-first workflow, these rules might feel pedantic. With agents they become multipliers: once encoded, they apply everywhere at once.
The team promotes rules into code. Human taste feeds back into the system continuously. Review comments, refactoring pull requests, and user-facing bugs become documentation updates or go directly into tooling. When documentation falls short, the rule moves into code.
Background agents on a cadence scan for deviations, update quality grades, and open targeted refactoring pull requests.
The team states the limit itself. The autonomy depends heavily on the structure and tooling of this one repository. Do not assume it generalizes without similar investment.
2. How to sort every agent control as guide or sensor, computational or inferential
Thoughtworks distinguished engineer Birgitta Böckeler built the mental model: “Agent = Model + Harness.” The harness raises the odds that the agent gets it right the first time. It also provides a feedback loop that self-corrects as many issues as possible before they reach human eyes.
Two axes sort every control:
- Guides are feedforward. They anticipate the agent’s behavior and steer it before it acts. Sensors are feedback. They observe after the agent acts and help it self-correct. They work best when the team writes their output for the model. Linter messages that carry self-correction instructions are a positive kind of prompt injection.
- Computational controls are deterministic and fast, run by the CPU: tests, linters, type checkers, structural analysis. They run in milliseconds to seconds, and their results are reliable. Inferential controls are semantic analysis, AI code review, and LLM as judge. They are slower, costlier, and non-deterministic.
Without both directions, you get one of two failures. A feedback-only agent keeps repeating the same mistakes. A feedforward-only agent encodes rules but never finds out whether they worked.
The human’s job is the steering loop. Whenever an issue happens more than once, improve the feedforward and feedback controls so it becomes less likely, or impossible.
Sensors have limits. Computational sensors catch the structural problems reliably: duplicate code, cyclomatic complexity, missing test coverage, architectural drift, style violations. Inferential ones partially catch semantic duplication and over-engineering, but expensively and probabilistically.
Neither catches the higher-impact problems reliably: misdiagnosed issues, unnecessary features, misunderstood instructions. Correctness sits outside any sensor’s remit when the human did not specify what they wanted in the first place.
3. Why agent loops amplify defensive, complex code, and where loops still work
A loop repeats whatever the agent does, so it repeats the agent’s habits along with its output. Sensors catch structure. They do not catch a design that is too defensive or too complex, and each pass adds more of it.
Armin Ronacher, creator of Flask and of the Pi coding-agent harness, reports little success with loops on code he deeply cares about. That turns out to be quite a lot of code. Part of that is taste and part is control.
The failure he sees is one no lint catches. Present-day models tend to produce code that is too defensive, too complex, and too local in its reasoning. They avoid strong invariants.
Then the harness makes it worse. Put that behavior behind loops and you amplify it.
If each iteration adds another small defense, the system slowly becomes less understandable while it appears more robust. The more hands-off you are, the more that happens.
He is specific about where loops do work: code ports between languages, performance experiments, security scans, research. These tasks share one property. They either transform code that already exists, or they produce code with an intentionally short shelf life.
Goal: Turn the three rules your reviewers repeat most into enforcement, in one working session.
- Pull the last month of review comments from one repository. Ask your coding agent to cluster them and return the three most repeated rules, with example comments.
- For each rule, run this prompt:
Prompt
Here is a code review rule our team keeps repeating, with example comments. 1. Say whether it can be enforced computationally (linter, type check, structural test) or only inferentially (AI review). Name the tool. 2. If computational, draft the lint or structural test, and write the error message so it tells an agent exactly how to fix the violation. 3. If inferential only, write the review instruction as a short skill. 4. Classify the result as guide or sensor. Do not add rules we did not give you. - Merge the computational ones into CI. Add the inferential ones to the review agent’s instructions.
- Delete the corresponding lines from AGENTS.md. They now live in code.
Expected result: three rules that reviewers no longer have to type. Each sits in one of Böckeler’s four quadrants, with an error message an agent can act on.
How to define done when agents ship more code than your reviewers can read
Review cannot scale with agent output, so done can no longer mean a human read every line. It has to mean something a machine can check and an engineer can answer for.
This lesson covers who owns what when agents write the code, how high the quality bar sits for code that reaches production, and the twelve engineering practices that agents reward.
1. How to apply Anthropic engineer’s outer-loop principles to agent work
Anthropic engineer Addy Osmani says engineers need to own the outer loop, which acts as the accountability for these systems.
Agents have leverage, and leverage creates obligations. Someone must be able to explain exactly what changed, why it was safe, and what happens if they are wrong.
- Quality is every check you install before you let the system loose. Those checks produce evidence.
- Verdict is the production decision: ship, block, redirect, narrow the response, add a guardrail, or reject outright. “The model may write the line, but the Verdict is mine.”
- Answerability is the guarantee that if someone asks, you can explain why.
Agents now run the inner execution loop. Engineers own the outer loop.
The speed of generation moved faster than the speed of control.
Quality as back pressure. Do not grant agents as much autonomy as they can exercise. Grant just enough that you keep the back pressure to stop them, regulate them, and check their work. The signals: type checks, tests, hooks, sandbox limits, audit logs, monitors.
The four loops a human stays in:
- The constraints loop decides what inputs, architectures, instructions, or invariants to set.
- The sampling loop decides how much output to sample and review.
- The audit loop decides what evidence to keep and how to make sure the audit log is effective.
- The ownership loop decides what part of the production boundary to own.
The human does not need to be in the inner loop.
2. Why Claude Code lead Boris Cherny sets a higher bar for agent-written production code
The quality bar depends on what the code is for. Throwaway code with a low blast radius needs no bar. Production code needs a bar higher than a human’s, because volume hides the mess until it is hard to maintain.
Boris Cherny, who leads Claude Code at Anthropic, puts it this way:
“I think there is room for both.
1. Prototypes and other throw-away code can be treated as totally black box. If you’re going to throw it away anyway, and if the blast radius of it breaking is low, it doesn’t need to be perfect.
2. Production code written by Claude should have a higher bar than if it was written by a human. At Anthropic, we have many guardrails in place to make sure this is happening: lots of lint rules, lots of tests, Claude-driven end to end tests, Claude-powered fuzzers running daily, automated code reviews and security reviews, automated code refactoring, and so on. Without these, you can end up with a mess that is hard to maintain down the line.
Your job is to hold the bar on code quality.”
When code misses the bar, his remedies run from a stronger model and higher effort settings, through the context file and skills, to more steering. The last resort is to have the agent pay down the accumulated debt.
3. Why agents reward Simon Willison’s twelve practices that senior engineers already have
The teams that get the most from coding agents already had the discipline. Simon Willison, creator of Datasette, lists twelve practices that agents reward, and notes that “LLMs actively reward existing top tier software engineering practices.”
- Automate your tests. Without them, an agent can claim that something works without a single test run.
- Plan in advance. Iterate on the plan first, then hand it to the agent to write the code.
- Write comprehensive documentation first. The model can build the matching implementation from that input alone.
- Keep good version control habits. LLMs are fiercely competent at Git and better than most developers at git bisect.
- Put effective automation in place. That means CI, formatting, linting, and continuous deployment to a preview environment.
- Build a culture of code review. A team that reviews code fast and well gets much more out of LLMs.
- Learn a very weird form of management. Good results from a coding agent come the same way they come from a human collaborator.
- Do really good manual QA. Predict the edge cases and dig into them.
- Build strong research skills. Someone still has to find the best options and prove an approach before an agent starts.
- Make sure you can ship to a preview environment.
- Develop an instinct for what to hand to AI and what to do yourself.
- Update your sense of estimation. Work that used to take a long time is much faster, but estimates now depend on new factors that nobody fully understands yet.
Almost all of these are characteristics of senior software engineers already. AI tools amplify existing expertise.
Goal: Write a definition of done for one repository that an agent can satisfy and a human can answer for, in one session.
- Add this block to the repository’s PR template:
Template
## Accountability contract - Checks understood: which of the repo's gates (types, tests, lint, structural, AI review) ran on this change - Evidence: link or paste the outputs that back the merge decision - Verdict owner: @name of the human who decides ship / block / narrow - Tier: throwaway (black box, low blast radius) or production (higher bar than human-written) - Rollback: the command or the flag that reverts this, and who runs it - Paste Willison’s twelve practices into your coding agent with this prompt:
Prompt
For this repository, mark each of the twelve practices below as PRESENT, PARTIAL, or ABSENT, with one line of evidence from the repo (a file, a CI job, a doc). Then list the ABSENT ones in the order an agent would most benefit from them. Do not add practices of your own. - Take the first ABSENT practice and open a ticket for it, owned by a named engineer.
- Add the four outer loops to the team’s operating notes. Name who sets constraints, who samples and how much, what evidence the team keeps, and who owns which part of the production boundary.
Expected result: a PR template that requires evidence and an owner, a scored checklist of the twelve practices, and one ticket for the biggest gap.
The Code: Your daily unfair advantage in software engineering.
Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.
How to split spec ownership between PMs, architects, engineers, and leadership
When the spec drives development instead of supporting it, one question decides whether it works. Who owns the spec, the constraints, the intent, and the mandate?
This lesson answers it with the four roles Microsoft’s Inside Track team set across its own teams. It covers what each role now does that it did not do before.
- Product managers own the spec. Everything starts with the spec, so they become the starting point for the entire system. If the specs are right, everything falls into place.
They define success criteria, resolve ambiguities, document business requirements, and keep priorities visible through implementation. - Architects own the constitution. They establish the guardrails that govern the development process.
Because those decisions come first, teams find issues before implementation begins, not during code review or testing. - Developers validate intent. Their first step is a complete understanding of the problem, including a definition of success, before any code exists.
Developers get excited and want to jump straight into coding. That habit was the hardest to change. - Leadership makes it standard. SDD delivered the greatest value once it became a standard way of working, not an optional process that individuals can bypass.
The adoption has to be deliberate.
Two more things did not work at Microsoft: vibe coding, because “speed without direction is just expensive chaos,” and static planning documents, which go out of date once development starts.
Goal: Draft a one-page constitution for one team and name its three owners, in one meeting.
- Book ninety minutes with the team’s PM, its architect or most senior engineer, and its manager. Microsoft’s rule is that this is a team exercise, not an individual one.
- Fill four headings, one paragraph each, using the Inside Track definition: architectural principles, governance requirements, security standards, development limitations.
- Paste the draft into your coding agent with this prompt:
Prompt
Here is a draft constitution for one engineering team. For each statement: 1. Mark it ENFORCEABLE if a linter, type check, structural test, or CI job could check it, and name the check. 2. Mark it GUIDANCE if only a human or an AI reviewer could judge it. 3. Flag any statement that is already true by default in our stack and adds nothing. Return the constitution with tags and the count in each bucket. Do not add principles we did not write. - Commit the file to the team’s main repository beside the code. Record three names in it: who owns the spec, who owns the constitution, who makes the process standard.
Expected result: a version-controlled constitution with every statement tagged as enforceable or guidance, and three named owners.
The lesson on how Microsoft adopted SDD across teams shows the roles above at work in one organization.
How Microsoft adopted SDD across teams
Microsoft Digital, the company’s IT organization, finds that SDD delivered the greatest value once it became a standard way of working, not an optional process that individuals could bypass. The adoption has to be deliberate. It starts with the hard lesson, where ad-hoc adoption made individuals faster and teams no faster.
“It’s not something where you decide, today I woke up and I’ll start using SDD. It doesn’t work like that. It’s a mindset shift and a learning curve.”
Microsoft serves as its own Customer Zero, so the process below is the one it offers customers as a model.
The six-stage workflow, run on Spec Kit
Microsoft Digital’s developers and PMs begin with specifications that capture business goals, user requirements, edge cases, and acceptance tests. The specs are version-controlled in the same repository as the code and updated as the project evolves.
Before any spec, the team agrees on a constitution. Then the teams run six stages on Spec Kit.
Each stage builds on the spec, so teams trace implementation decisions back to business requirements. Agents enhance the spec and generate code, tests, and documentation from it; engineers validate intent, review outputs, and refine requirements.
“Effective SDD adoption relies on well-scoped specifications.”
The five pillars behind the value
Microsoft names five pillars that drive the value of SDD, from its own experience:
- Spec as the source of truth. The spec is the single authoritative reference for all stakeholders.
- Living, executable artifacts. Specs evolve with the project, stay synchronized with the code and tests, and are co-authored by the stakeholders.
- AI-assisted automation. Agents generate, test, and validate code from the spec.
- Human validation and collaboration. Human expertise reviews specs, refines requirements, and validates outcomes.
- Predictability and measurability. Teams track progress, measure outcomes, and improve predictability.
How roles changed: leaders, developers, PMs, architects
The biggest change at Microsoft was behavioral. Specifications stopped being documents that supported development and became the artifact that drove it. Roles coupled to the old lifecycle had to change.
Leadership. Leaders reward clarity, alignment, and collaboration before implementation begins, and they drive adoption, since SDD only paid off as a standard.
Developers. The first step is now a complete understanding of the problem, with success defined, before any code. “As developers, we get excited and want to jump directly into coding,” says Mridul Verma, Senior Software Engineer at Microsoft Digital. “Changing that habit was really tricky.”
The developer must make sure the agent fully understands the task before the coding starts. “SDD makes you clarify and refine your thoughts first for agents, not for yourself,” says Sudhakar Sadasivuni, Principal Group Engineering Manager at Microsoft Digital. “This is agent-first solutioning, not human-first solutioning.”
Engineers now spend more time on requirements, edge cases, plan evaluation, and validation of AI output, and less on every line of code.
Program and product managers. PMs keep stakeholder management, prioritization, and roadmap planning. What changed is ownership.
“In spec-driven development, PMs become the owners of the spec. Everything starts with the spec, so they become the starting point for the entire system. If the specs are right, then everything falls into place.”
PMs define success criteria, resolve ambiguities, document business requirements, and keep priorities visible through implementation. Gopal Panigrahy, Principal Product Manager at Microsoft Digital, describes the payoff. You move from an idea to a prototype, show it to customers and leaders, and get feedback much faster.
Architects. Architects own the constitution, so architectural decisions come early and teams find issues before implementation rather than in review.
Outcomes from three projects
Gupta reports three projects in his Microsoft for Developers post, and sums them up as spec quality equals output quality.
- Repeated onboarding turned into a pattern. In a brownfield project, each new asset type needed the same UI, API, and test changes. The team captured the pattern in parameterized specs and documented only the deviations per asset. Onboarding fell from 2 to 3 weeks to a few days.
- A multi-service platform aligned before the build. A large greenfield project covered attendees, facilities, security, vendors, logistics, and compliance. The constitution, specs, and plans served as the source of truth. The team improved cross-service consistency, made constraints explicit, and reduced churn as implementation scaled.
- A prototype became a product faster. In another brownfield project, the team moved a React and TypeScript prototype to a working product. The product had multiple agents, health monitoring, and admin dashboards. Custom prompts and quality-gate scripts made the process repeatable across contributors.
The adoption playbook and Microsoft’s five takeaways
Goal: Start SDD in your organization with one pilot, and check the pilot against the guidelines Microsoft learned.
Gupta’s four-step playbook, for teams that do not need to adopt the full lifecycle at once:
1. **Pilot.** Start with one feature or workflow where alignment problems are visible.
2. **Formalize.** Write a lightweight spec with scenarios, constraints, and acceptance criteria.
3. **Iterate.** Use AI to generate implementation artifacts from that shared context.
4. **Refine and scale.** Review the output against the spec and refine the workflow as you learn.
Keep the process lightweight at first. Treat specs as living artifacts.
Avoid over-specifying too early. Expand the workflow only where it adds clear value.
Microsoft Digital’s five key takeaways for an organization that considers SDD:
1. Make the specification the source of truth. A living artifact, not static documentation.
2. Resolve ambiguity before implementation begins. Requirements, constraints, edge cases, success criteria.
3. Assign clear ownership. PMs, architects, developers, and designers all maintain the spec.
4. Establish architectural guardrails early. Governance, security, and design principles before code.
5. Use AI to accelerate execution, not replace judgment. Humans validate outputs and alignment with intent.
Expected result: one pilot feature with a lightweight spec, a review of its output against that spec, and a decision on where to expand next.
By this point you should have:
- Your last twenty tickets sorted into where a spec would have paid and where it would have been overhead.
- A context file cut to human-written shortcuts, intent, and landmines, with a before-and-after pass rate.
- Your three most repeated review rules turned into lints or review skills, placed in Böckeler’s quadrants.
- A PR template that requires evidence and a named owner, and a scored checklist of Willison’s twelve practices.
- A one-page constitution for one team with three named owners.
Module 4: Guardrails for Coding Agents
The guardrails themselves. It covers what a coding agent must never do and how the incidents of 2026 happened. It then covers how to set permissions, sandboxes, and review gates so that the constraints loop is real.