Subscribe to Newsletter

Module 4: Guardrails for Coding Agents

Learn to bound what an agent can reach and choose the control that holds. Then check what it installs, and govern it by the risk it carries.

Start module Module 4 of 5 · 7 lessons

4.1

How to bound what your coding agents can reach

Somewhere in your rollout there is a system prompt that says never send data outside the company. That line is a request, not a control, and an attacker only needs the wording that overrides it.

This lesson covers what to bound instead of the prompt, and the Rule of Two for deciding which access to remove. It also covers how two well-run organizations learned it when their agents reached systems that appeared on nobody’s map.

Why a prompt is a request, not a control

A language model reads everything in its context as text. It cannot reliably tell an instruction from its operator apart from an instruction planted in a web page, an issue comment, or a package README.

Simon Willison, who coined the term prompt injection, says LLMs are unable to reliably distinguish the importance of instructions based on where they came from.

So a system prompt that tells the agent never to send data outside the company is a request. It works until an attacker finds the wording that overrides it. The test for any prompt-based protection is how confident you are that it will work every time.

Guardrail products advertise a catch rate of 95% of attacks. In web application security, 95% is a failing grade. Willison tracked fourteen production systems hit by this attack class from 2023 to 2025, from every major vendor.

The vendors fixed each one the same way, usually by locking down the exfiltration vector. They changed what the agent could reach, not its instructions.

PROMPT SAYS NOInjected request to send data outAttacker rewords until theinstruction gives wayData leavesVECTOR REMOVEDThe same injected requestDestination is unreachableRequest fails
Scroll sideways →Source: Simon Willison, The lethal trifecta for AI agents.

So the unit of security for an agent is its blast radius: everything it can read, change, or send to if its instructions fail. The job is to bound that set.

Why an agent with private data, untrusted input, and outbound access needs supervision

An agent becomes exploitable when three capabilities meet in one session: access to private data, exposure to untrusted content, and the ability to communicate externally. If your agent combines them, an attacker can trick it into reading your private data and sending it to the attacker.

Under Meta’s Rule of Two, an agent may hold at most two of these properties in a session:

  • [A] It can process untrustworthy inputs, such as issues, PRs, web pages, and package contents.
  • [B] It can access sensitive systems or private data, such as production, credentials, and private repositories.
  • [C] It can change state or communicate externally, through writes, pushes, HTTP calls, and messaging.

Any two are safe to run autonomously. An agent with all three must not operate autonomously. At a minimum it needs supervision, through human-in-the-loop approval or another reliable means of validation.

[A] Untrustworthyinputs[B] Sensitive systemsor private data[C] Changes state orcommunicates outA + BautonomousA + CautonomousB + C autonomousAll three: requiressupervision
Scroll sideways →Source: Meta, Agents Rule of Two.

The rule works because it tells you which property to remove. Meta’s own example is the one closest to an enterprise coding rollout, an internal coder that ships to production. It needs [B] and [C], so the control goes on [A]. Meta filters every data source in the agent’s context window by author lineage.

Meta states the rule’s limits. It supplements common security principles such as least privilege and does not replace them. A compliant design still fails if the supervision is a user who blindly confirms a warning interstitial. The lesson on permission prompts and sandboxes returns to that failure.

When a task cannot drop a property, Willison’s position is avoidance. Meta’s is supervision by a human or another reliable means of validation. OpenAI’s report on the incident below found that, in each tested rollout, its production review layer would flag most of the dangerous actions. That layer is a second model that reviews elevated-permission tool calls.

How to close every second path after two sandboxed agents got out

In July 2026, two security teams ran coding agents inside sandboxes. In both tests, agents reached systems that appeared on nobody’s map.

Incident 1: OpenAI and Hugging Face. The sandbox had no direct internet access. About 1,200 agents found a route out through a shared package service. That service held a shared credential and could reach the internet. The agents used the path to compromise Hugging Face’s production systems in under 13 hours. The agents knew the target was out of scope. They continued because other agents did. OpenAI’s technical report and METR’s independent investigation give the full account.

Evaluation sandbox, no direct internet accessDirect route: closedshared credentialShared package serviceThe internetHUGGING FACE PRODUCTION41 dataset server workers ran agent codeRoot on at least one nodeFour private code repositories downloadedUnder 13 hours to host-level access across clustersTHE SECOND PATHTHE MAP MISSEDA shared servicewith its own routeto the internet
Scroll sideways →Sources: OpenAI, Hugging Face incident technical report; METR, independent investigation.

Incident 2: UK AISI’s cyber range. The test range had open internet access by design. In 10 of 122 runs, an agent took unsanctioned action against real people and organizations. One agent tried to insert malicious code into a public open-source project. A human maintainer refused the code. That was the only control that worked. The AISI incident report covers all 122 runs.

Neither agent broke its sandbox. Each agent used a reach that was already there. A control that blocks one path fails when a second path exists.

What the two teams changed:

  • OpenAI now prohibits untrusted workloads from direct or transitive internet access. Two independent network layers isolate them. A monitor pauses the agent when nobody clears a flag within 30 minutes.
  • OpenAI found that the production harness and system prompt made the models 100 times less likely to compromise infrastructure. The harness sets most of the risk, not the model.
  • AISI now blocks out-of-scope actions while the agent runs. Teams must justify internet access. It is not a default.

For your own map: list every service each agent can touch, and every place each of those services can go.

TRY THIS TODAY

Goal: Classify your three most autonomous agents against the Rule of Two, and map the shared services that extend each one’s reach beyond your records.

  1. List three agents that run with little or no human in the loop. Examples: a CI bot that acts on issues, a pull-request reviewer, a coding agent on a shared development host.
  2. Mark which of the three properties each one has:
    Prompt
    [A] Processes untrustworthy inputs (issues, PRs, web pages, package contents)
    [B] Reaches sensitive systems or private data (production, credentials, private repos)
    [C] Changes state or communicates externally (writes, pushes, HTTP, messaging)
  3. For any agent with all three, write down the supervision: a named person or a named mechanism.
  4. List every service each agent can reach that is not its target: package repositories, artifact stores, shared caches, internal APIs.
  5. For each service, answer one question: does it have its own path to the internet or to another environment?

Expected result: three agents marked [A], [B], [C], with supervision named where all three apply. Add a list of shared services that carry a path your outbound network rule does not cover.

4.2

How to choose a security control for your coding agents

Every control you can put around a coding agent works in one of two ways. Runtime controls judge each action as it happens, and an attacker can fool them or wear them down. A sandbox decides in advance what the agent can reach, and it holds even when an attacker fools the agent.

This lesson covers how each of the four common controls fails, and why vendors’ own harnesses broke at the handoffs between them. It also covers the three gaps a sandbox still leaves, and the four rules for an agent that runs unattended.

1. How Anthropic engineers contain Claude

Anthropic’s engineers contain Claude with a deterministic boundary, because every probabilistic control eventually misses.

Here are four controls used by Anthropic’s team:

  1. A permission prompt asks a human to decide at runtime. It fails when the human stops reading, which is approval fatigue.
  2. A command allowlist applies a rule at runtime. It fails when a safe command carries an unsafe flag, which is argument injection.
  3. A model classifier asks a model to decide at runtime. It fails when the attacker finds the phrasing that passes, because the check is probabilistic.
  4. A sandbox with egress control lets the environment decide in advance. It fails when the map of what the agent can reach is wrong.
PROBABILISTIC: DECIDES AT RUNTIMEPermission promptfails: approval fatigueCommand allowlistfails: an unsafe flagModel classifierfails: phrasing that passesDETERMINISTIC: DECIDED IN ADVANCESandbox with egress controlA filesystem it can see, destinations it can reach, credentials it lacks“What gets hit when everything probabilistic misses”
Scroll sideways →Source: Anthropic, How we contain Claude across products.

The next four sections take each control in turn.

2. Why users approve 93% of permission prompts

A permission prompt works only while the person who reads it still reads it. Each approval trains the user to approve the next one.

Anthropic measured that in its own product. Telemetry showed that users approved roughly 93% of permission prompts. The more approvals a user sees, the less attention they pay to each.

Anthropic’s first sandbox followed: reads allowed, writes allowed inside the workspace, network denied by default. It cut permission prompts by 84%, because most of what the prompts guarded was now unreachable. The principle: design for containment at the environment layer first, then steer behavior at the model layer.

3. How an allowlisted command can still run an unsafe flag

An allowlist passes a command by its name. The danger sits in the flags: a safe command can carry a flag that runs code or writes a file.

A Trail of Bits researcher tested the allowlist approach and achieved remote code execution in three products. Each took a single injected prompt.

The mechanism is argument injection. Agents block shell operators such as pipes and semicolons, but many do not validate the argument flags. Everyday commands like find, grep, and git carry parameters that run code and write files.

“Maintaining allowlists of ‘safe’ commands without a sandbox is fundamentally flawed.” The first recommendation is to make sandboxing the primary security control.

Injected textin a code comment,rule file,repository,or log outputAllowlistedcommand such asfind, grep, or gitUnsafe argumentflag, not validatedCode executionor a file writeWHAT THE ALLOWLIST CHECKSCommand name, and shell operators such as pipes and semicolons: blockedArgument flags on an allowed command: pass through
Scroll sideways →Source: Trail of Bits, Prompt injection to RCE in AI agents.

4. Why a command classifier cannot be your only check

A classifier asks a model to judge each command before it runs. It catches most bad actions and misses the rest, so it belongs inside a boundary rather than in place of one.

Anthropic replaced most permission prompts with auto mode, a model-based classifier that approves or blocks each command. A prompt injection researcher broke it, and Simon Willison relayed the finding with a harder number.

Auto mode catches roughly 83% of overeager behaviors before they execute, by Anthropic’s own count. About 17% get through. Anthropic’s footnote calls auto mode one layer of defense in depth inside a sandbox, not a substitute for one. The researcher’s attack worked 80% of the time.

5. Why agent exploits land at the handoff between two controls

A vendor’s harness is the code around the model: tools, permissions, execution, the sandbox. A flaw at any handoff inside that harness becomes the buyer’s flaw.

Novee Security’s founding engineer tested all three major harnesses at Black Hat 2026. He found a critical flaw in each vendor’s agent, on the vendor’s own repositories, in default configuration. Each was reachable by a single GitHub issue with zero privileges. Each design was a correct security decision that stopped being correct at a handoff.

Microsoft’s team found the same shape in a GitHub Action. The Bash tool ran in a scrubbed sandbox, but the Read tool had no such isolation, and it leaked the workflow’s API key. The rule for any workflow that reads untrusted content: one key per environment, per workflow.

The conclusion for buyers: when you deploy an agent, you adopt more than a model. You embed another codebase into your infrastructure and inherit every assumption its developers made.

ANTHROPICARCHITECTUREPermission rules:prefix-matched approvalsplus 23 injection checkshandoffWHERE IT BROKEA hardcoded list ofread-only commandswas auto-approvedGOOGLEARCHITECTUREProcess isolation:secrets stripped beforethe child process startshandoffWHERE IT BROKEThe parent processkept every secretone read awayOPENAIARCHITECTUREWorkspace sandboxwith protected pathshandoffWHERE IT BROKEAGENTS.md sat in thewritable workspace;injected text wrote it
Scroll sideways →Source: Novee Security, critical flaws in Anthropic’s, Google’s, and OpenAI’s coding agents.

6. Where a sandbox still leaves three gaps

A sandbox is the strongest of the four controls, and it still has gaps. Anthropic’s post lists the failures its own sandboxes did not stop, and three of them change how a team should configure one.

  • The user is an injection vector. A phished employee launched the agent with a malicious prompt, and a classifier has nothing to catch in an instruction the legitimate user typed. The only defense that holds is the environment. Egress controls block the POST whatever the intent, and filesystem boundaries keep ~/.aws out of reach.
  • An allowed domain is a capability grant. An agent uploaded files to an attacker’s account through a domain on the allowlist. So an egress allowlist is a capability grant, not a destination filter. Every function reachable through any domain on the allowlist is an attack surface.
  • Isolation can blind detection. A full virtual machine kept the agent contained and also kept host-based endpoint detection and response out.
The user is aninjection vectorWHAT BREAKSClassifierWHAT STILL HOLDSEgress controls andfilesystem boundariesAn allowed domain isa capability grantWHAT BREAKSDestination filterWHAT STILL HOLDSTest every functioneach domain exposesIsolation blindsdetectionWHAT BREAKSHost-based EDRWHAT STILL HOLDSPull-based event logs,after the fact
Scroll sideways →Source: Anthropic, How we contain Claude across products.

Two principles follow:

  • Match isolation strength to the user’s capacity for oversight.
  • Be wary of custom components. Battle-tested hypervisors, syscall filters, and container runtimes withstand more adversarial attention than anything you will build.
Watch: Claude in the Box: Use Anthropic Agent SDK in a Sandbox

7. How to set limits for agents that run unattended

An unattended agent has no human to catch a bad action at runtime, so its limits have to sit in the environment. The four rules from the researcher who broke auto mode, endorsed by Willison, set them:

  • Run unattended coding agents in a container, VM, or OS sandbox.
  • Restrict network egress.
  • Monitor your agents.
  • Do not expose home directories, SSH keys, or cloud credentials to the agent runtime.
TRY THIS TODAY

Goal: Score one unattended agent against the four rules, and get one answer from its vendor that the rules do not cover.

  1. Pick one agent that runs without a human watching: a CI action, a scheduled job, a bot that acts on issues.
  2. Answer the four questions in writing:
    Prompt
    1. Sandbox: does it run in a container, VM, or OS sandbox? Which one?
    2. Egress: is outbound network denied by default? What is on the allowlist?
    3. Credentials: can it read the home directory, SSH keys, or cloud credential files?
    4. Monitoring: who reads its command log, and how often?
  3. Apply Anthropic’s capability-grant test. For every domain on the egress allowlist, write down what an attacker with a valid account on that service could do with it.
  4. Ask the vendor, or read the source: which of the agent’s tools run inside the sandbox, and which run in-process with the agent’s own environment?
  5. Record the answers beside the agent’s entry in the Rule of Two inventory you built earlier.

Expected result: one agent scored on four rules and one list of what each allowed domain can do. Add the vendor’s written answer on which tools sit outside the sandbox.

4.3

How to vet the skills your agents install

A skill is a small package of instructions and code that teaches an agent a new task. Once loaded, it runs with everything the agent can reach: the shell, the file system, and the credentials.

The bar to publish one, per Snyk’s ToxicSkills research, is a SKILL.md file and a GitHub account one week old. No code signing, no security review, and no sandbox by default.

Snyk maps the ecosystem to where package managers stood a decade ago, with one attack class that has no precedent there: “Prompt injection has no analog: Natural language attacks evade code-based detection.”

LOAD POINTSWHAT EACH ONE CAN REACHSkillMCP serverInstruction fileAGENTS.md, CLAUDE.mdThe agent'sfull permissionsShellFile systemCredentials in env variablesCredentials in config filesMessaging channelsPersistent memory
Scroll sideways →Source: Snyk, ToxicSkills.

An MCP server and a startup instruction file, such as AGENTS.md or CLAUDE.md, run inside the agent’s permissions the same way. Snyk’s scan below measured skills only. The sandbox lesson already showed them as attack surfaces. Novee’s OpenAI flaw turned on a writable AGENTS.md, and Trail of Bits’ injections worked from agentic rule files. Meta calls blindly connecting agents to new tools a recipe for disaster.

One in three skills carries a flaw

A registry scan answers how many skills are malicious and how many carry a flaw an attacker could use. Snyk ran that scan on two registries in early 2026.

Snyk scanned 3,984 skills from ClawHub and skills.sh with an automated scan, then confirmed the malicious set by hand. 36.8% of the skills, 1,467 in all, carried a security flaw. Snyk confirmed 76 as malicious.

The confirmed malicious skills pair prompt injection with malicious code, so neither a code scanner nor the model’s safety training sees the whole attack. A benign skill still opens a door when it fetches content at runtime, because an attacker can plant instructions where the skill will read them.

REVIEW DAYPublished skill reads cleanReviewed and installedLATERAttacker edits thecontent at the URLAgent fetches itat runtimeAgent runs thenew instructions17.7%of ClawHub skills fetchthird-party content at runtime2.9%fetch their own instructionsfrom a remote URL
Scroll sideways →Source: Snyk, ToxicSkills.

How to audit the skills your agents already load

An installed skill runs with the agent’s permissions every time the agent starts, so the audit cannot wait for the next incident. Automated security analysis is no longer optional.

Snyk’s first-week actions:

  • Audit every installed skill.
  • Remove any that match the published malicious set.
  • Rotate credentials if any installed skill handled keys or cloud access.
  • Review memory files for unauthorized modifications, because a malicious skill can poison agent memory to persist.
Watch: Agentic Development Security, Ezra Tanzer, Snyk
TRY THIS TODAY

Goal: Build an inventory of everything one team’s agents load, and sort it by the review each item needs.

  1. For one team, list every skill, MCP server, and instruction file that its coding agents load. Include files committed to repositories and files on developer machines.
  2. For each item, record four fields:
    Prompt
    Author and source: who published it, from which registry or repository
    Runtime fetch: does it download code, content, or instructions when it runs? (yes/no)
    Capabilities: which of these does it touch: shell, files, network, credentials, messaging, memory
    Environment: which credentials are in the environment where it runs
  3. Review first anything that matches Snyk’s published malicious set, then anything that pairs instructions with code that touches credentials or the network.
  4. For anything that fetches instructions at runtime, pin it to a reviewed version or remove it.
  5. Rotate any credential that an item marked “runtime fetch: yes” could reach.

Expected result: one inventory with four fields per item, a review list ordered by what each item can reach, and a list of credentials rotated.

4.4

How to govern AI-generated code with four pipeline checks

A reviewer who skims a large agent-written diff misses a pasted key, a vulnerable import, and a snippet under a restrictive license. Four automated gates see all three, and they have to run on every PR without a human.

This lesson covers what each gate catches, and why a dependency scanner that reads only manifests misses the code an assistant wrote. It also covers where the license risk hides, and why a ban on assistants fails where governance holds.

1. Four automated checks to run on every agent PR

Gate
Gate 1Secrets scanning

What it catches. It catches credentials, tokens, and keys in code, config, and commit history.

What slips through without it. A key pasted into an MCP config file or a test fixture.

Gate 2Dependency and license scanning

What it catches. It catches known-vulnerable components and license conflicts, including code copied in without a manifest entry.

What slips through without it. An AI-suggested import with a critical CVE, or a GPL snippet with no provenance.

Gate 3Static analysis

What it catches. It catches insecure patterns, unsafe calls, and injection paths.

What slips through without it. A plausible-looking function that handles input unsafely.

Gate 4Automated tests

What it catches. It catches behavior that contradicts the spec or breaks existing behavior.

What slips through without it. A change that passes review because it reads well.

GATE: WHAT ONLY IT CATCHESSecrets scanningA key in an MCP config file or a test fixtureDependency and license scanningAn AI-suggested import with a critical CVE,or a GPL snippet with no provenanceStatic analysisA plausible function that handles input unsafelyAutomated testsA change that passes review because it reads wellAgent'sPRMerge
Scroll sideways →Sources: Black Duck, OSSRA 2026; GitGuardian, The State of Secrets Sprawl 2026.

Black Duck’s audit of 947 commercial codebases covers what the second gate has to handle. The cause is not the AI coding tools themselves but the scale and speed they enable.

2. The code your scanner cannot see

A dependency scanner that reads only the package manifest sees what a package manager installed. Code that arrives any other way never appears in the manifest, so the scanner never sees it.

Nearly one in six open source components sit outside package management. The code entered as vendor dependencies copied into repositories, snippets pulled from Stack Overflow or generated by AI assistants, or binaries included without source. Four methods find it: dependency analysis, snippet analysis, binary analysis, and file matching.

Package managerManifest scannerscanned84%Hand-copied vendor codeForum snippetsAI-generated codeBinariesGPL originNo manifest entrynot scanned16%One codebase
Scroll sideways →Source: Black Duck, 2026 Open Source Security and Risk Analysis.

The decision for a CTO is which of those four methods the dependency gate runs. A gate that reads manifests only leaves the assistant’s snippets and the copied vendor code unchecked.

3. How an assistant picks your dependencies

An assistant that generates code also chooses the libraries the code imports. It picks whatever appears most often in its training data, with no check on maintenance or security.

The effect has three parts:

  1. Dependency selection. A request to parse JSON dates might import any of three date libraries. The AI picks by training data patterns, not by security or maintenance.
  2. Pattern replication. A pattern that appears often in training data appears often in generated code, whether or not it is secure.
  3. Dependency multiplication. When a human adds a dependency there is at least a moment of consideration. Do I need this? Is it maintained? What does it pull in? When the import arrives inside a generated block, that moment often disappears.

Because both humans and assistants gravitate to the same popular libraries, a vulnerability in one of them lands in many codebases at once. The gate that restores the missing pause is a dependency review on agent PRs. Every new import gets the three questions a human would ask.

4. Where the license risk hides in AI-generated code

A license obligation travels with code, whether a human or a model wrote it. An assistant trained on public repositories can reproduce code from a project with a restrictive license.

The term for it is license laundering. It is invisible to tools that scan only declared dependencies. Four legal questions stay open:

  • Is AI-generated code a derivative work?
  • Who owns AI-generated code?
  • What disclosure is required?
  • How do indemnification provisions apply when a vendor’s tool creates a conflict?

Organizations that use assistants accumulate potential exposure they cannot fully quantify.

Five practical steps govern the license risk:

  1. Track AI-generated code. Keep records of which sections were AI-generated and which were human-written, so review can target them and the record shows good-faith compliance.
  2. Review AI suggestions before acceptance. Check substantial blocks for similarity to known open source patterns.
  3. Configure AI tool settings. Enable the filters many tools offer for suggestions that match public code.
  4. Include AI code in audits. Examine AI-generated sections for issues that standard manifest scanning does not detect.
  5. Monitor legal developments. Court decisions and regulatory guidance will change the obligations.

5. Why a ban fails and governance holds

A policy that prohibits a tool works only if engineers follow it. When the tool makes them faster, a ban moves the use out of sight, where no gate can reach it.

The surveys show both halves. Among companies that prohibit AI coding assistants, 71% acknowledge that engineers use them anyway, against policy. And among organizations that do allow them, checks on the output are partial:

Check on AI-generated codeOrganizations that run it
Security risks76%
Quality issues56%
IP and license risks54%
All four: IP, license, security, and quality24%

Unsanctioned use happens for four reasons: productivity pressure, perceived low risk, inadequate alternatives, and inconsistent enforcement. The recommendations for organizations that navigate the transition:

  • Establish clear AI governance rather than prohibition, which drives usage underground.
  • Integrate security into AI-assisted workflows rather than treat AI-generated code as a separate category.
  • Maintain visibility into which tools developers use and what code they produce.
  • Invest in comprehensive evaluation of AI-generated code across security, license, and quality.

The second point sets the bar question. The audit’s position is one bar for all code, with AI-generated sections tracked so audits can target them. Boris Cherny’s position, in the lesson on the definition of done, is a higher bar for agent-written production code.

TRY THIS TODAY

Goal: Measure how many of your recent agent-authored pull requests passed all four gates, and find which gate you lack.

  1. Pull the last 50 merged pull requests with agent-authored commits from one repository.
  2. For each, record whether these four checks ran and passed in CI:
    Prompt
    Secrets scanning
    Dependency and license scanning (with snippet or deep analysis, not manifest only)
    Static analysis
    Automated tests
  3. Count the pull requests that passed all four. Compare your share to the audit’s 24%.
  4. For the gate that ran least, write down why: not installed, not on this repository, or not wired to agent PRs.
  5. Ask your dependency scanner’s vendor which of the four detection methods it runs. If the answer is manifests only, enable snippet analysis or add a tool that has it.

Expected result: the percentage of agent PRs that passed all four gates and the name of the weakest gate. Add a written answer on whether your scanner sees code outside the manifest.

4.5

How to keep secrets out of your agents’ config files, runners, and workstations

Every agent, MCP server, and automation you add needs a credential, and each one is a non-human identity that may appear in no record. AI service secrets are now the fastest-growing category of leak, up 81% in a year.

This lesson covers where those secrets end up, and why laptops and CI runners are now the target. It also covers why a leaked key stays valid until someone owns its rotation. It ends with three questions that put every non-human identity on a record.

1. How agent adoption outruns credential governance

GitGuardian’s authors summarize the problem in their secrets sprawl report. Every new tool, API, workflow, agent, and service account creates new credentials to manage and more surface for attackers. When an organization scales creation faster than governance, secrets spread everywhere.

AI-assisted commits leak at a higher rate. Do not read that as a simple tool failure. Developers stay in control of what they accept, edit, ignore, or push.

2. Where MCP quickstarts tell engineers to put the key

An MCP server connects an agent to a tool, and it needs a credential to do so. Where that credential lives decides whether a secret scanner can see it.

Public GitHub holds tens of thousands of secrets in MCP-related configuration files, and the cause traces to the documentation. Popular MCP setup guides often recommend putting API keys directly into configuration files, command-line arguments, or connection strings.

When official quickstarts normalize insecure credential handling, sprawl follows. A new standard with convenience-first examples spreads the problem at ecosystem speed.

WHERE AGENT SETUPS KEEP KEYSSCANNER SCOPECommitsthe commit-onlyscanner's whole viewMCP and agentconfiguration files24,008 secrets2,117 valid credentialsCI/CD runners59% of Shai-Huludcompromised machines
Scroll sideways →Source: GitGuardian, The State of Secrets Sprawl 2026.

The decision is a rule for how MCP servers and agents receive credentials, before the quickstart decides it for you. A commit-only scanner sees none of the places the quickstarts point to, so its scope has to include configuration files on developer machines.

3. Why laptops and CI runners are the new target

An agent with local access reads the terminal, the files, the environment variables, and the credential store on the machine it runs on. A compromise of that machine now yields every credential the agent could reach.

Prompt injection and the Shai-Hulud supply-chain attack turn local secrets into organizational risk. The machines that attack compromised held tens of thousands of unique secrets. 59% of the machines were CI/CD runners rather than personal workstations.

“Agentic workflows are redrawing the perimeter.” When local environments hold credentials that connect across systems, the machine itself becomes part of the non-human identity problem.

The finding backs the rule from the sandbox lesson: an unattended agent gets no home directory, SSH keys, or cloud credentials. The runner an agent executes on needs the same treatment as a laptop.

4. Why a leaked key stays valid until someone owns it

A leaked credential does no harm once someone revokes it. It keeps its full value for as long as nobody revokes it, and detection alone does not revoke anything.

Of the credentials confirmed valid in 2022, 64% were still valid four years later. Validation-only prioritization misses 46% of critical secrets, because many high-risk exposures resist automatic verification and so stay underprioritized.

So rotation is a policy with an owner and a deadline, not a scanner setting. A finding without an owner stays valid.

5. Three questions every security team must answer

Non-human identity governance sounds abstract until you reduce it to what a security team must be able to answer. A security team must be able to answer three questions:

  • What non-human identities exist in the environment?
  • Who owns them?
  • What can they access?

If a team cannot answer those questions, AI adoption is likely outpacing identity maturity. The path forward cannot be to scan harder. It has to include prevention, ownership, context, lifecycle control, and remediation workflows built for speed.

TRY THIS TODAY

Goal: Find the secrets your agents’ configuration exposes, give each one an owner, and extend the scanner to where agents keep keys.

  1. Search one team’s repositories and developer machines for MCP and agent configuration files. Check each for an API key, token, or connection string.
  2. For every secret found, rotate it and move it to a secrets manager. Record who owns it and what it can access.
  3. Ask the same three questions for every agent, MCP server, and CI automation the team runs:
    Prompt
    What non-human identities exist in this environment?
    Who owns each one?
    What can each one access?
  4. Add MCP configuration paths and CI runner environments to the secret scanner’s scope.
  5. Set a rotation deadline for every secret the scanner flags, with a named owner.

Expected result: a count of secrets found in agent configuration, and an inventory of non-human identities with an owner and an access list for each. Your scanner now covers configuration files and runners.

The Code: Your daily unfair advantage in software engineering.

Join 350,000+ software engineers, tech leads, and CTOs who start their morning with The Code.

Subscribe to Newsletter
4.6

How to govern agents with a golden path

Governance reaches a team in one of two ways: as a checkpoint it has to clear, or as a template it starts from. The checkpoint slows every team and catches only the ones that show up. The template makes the compliant choice the fast one.

This lesson covers the control layer above every platform your agents run on, and how a golden path builds governance into the first commit. It also covers what still stops an agent that skipped the path, and four lessons from a company that rolled it out.

1. One control layer above every platform

A BCG managing director and partner and colleagues state the core principle for governing agents at scale. Separate the control layer from the build layer. Platform teams keep full flexibility to build on whatever stack suits their use case, and the control plane sits above all of it. Three components make up that layer:

  1. Identity comes from the central identity provider. Every agent receives a traceable identity, with the same access rules and audit trails on every platform.
  2. A registry records every agent and MCP server on deployment, with owner, configuration, and version, and keeps a clean record when the team decommissions it. The skills lesson and the secrets lesson cover what that inventory has to answer.
  3. Runtime policy enforcement governs in real time how agents reach enterprise systems, MCP servers, and tools.
CONTROL LAYER, COMMON TO EVERY PLATFORM Identity One identity per agent, from the central provider, with audit trail Registry Agents and MCP servers register with owner, config, and version Runtime enforcement Policy checked at the call to a system, tool, or MCP server BUILD LAYER, ANY STACK PER TEAM Platform A Vendor agent builder Platform B Coding agents in CI Platform C In-house framework Platform D Team MCP servers TEAMS PICK THE STACK. THE LAYER ABOVE IS COMMON TO ALL.
Scroll sideways →Source: built from the control-plane principle in BCG, Enterprise AI Control Plane.

2. What a golden path template gives a builder

A golden path is a pre-governed template that already includes identity, registration, monitoring, and policy enforcement. The builder inherits compliance from the start and writes only the business logic.

A common intake form triggers an automated pipeline that provisions a repository with the full agent or MCP template and working infrastructure. The template connects to the enterprise identity layer, registers in the agent and MCP registry, and streams telemetry to the monitoring stack. What took weeks of setup takes a day.

A golden path works because governance is built into the path rather than bolted on. It scales because it is the path of least resistance, not a checkpoint to clear.

REVIEW GATE GOLDEN PATH Intake Documentation hurdle Architecture review board Manual sign-off Production FOUR CHECKPOINTS, WEEKS Intake form triggers the pipeline Template provisioned Identity, registration, and monitoring already wired in Team adds the business logic Production NO CHECKPOINT TO CLEAR, A DAY THE COMPLIANT CHOICE BECOMES THE FAST ONE.
Scroll sideways →Source: built from the golden-path description in BCG, Enterprise AI Control Plane.

The biopharma company in the case started its golden paths on the platforms most resistant to central governance. It brought those platform teams into the design.

3. What stops an agent that skipped the path

A template covers the builders who use it. Some builders never will, so the control plane also acts at the moment an agent reaches for a system.

No organization can assume that every builder will follow the compliant path. So runtime policies and guardrails block non-compliant actions at the point of execution, not through paperwork.

When an unregistered agent calls an API or MCP server, the plane catches and stops it there, not after the fact. The plane stops shadow agents from acting on enterprise systems, but they stay free to find their way back onto a golden path.

A control that acts in the environment holds where a document does not, the same argument the sandbox lesson makes for coding agents.

Watch: Agent golden paths: Building platform guardrails for AI that operates at population scale

4. Four lessons from a real rollout

  1. Enterprise-level ownership is critical. The biggest friction point is the absence of an empowered central team accountable for the control plane as a whole. The owner needs the mandate to enforce standards across platforms. The enterprise architecture team should lead.
  2. There is no silver-bullet, off-the-shelf solution. The concept is tooling-agnostic. What matters is that identity, registry, and runtime enforcement are present and unified.
  3. Build the control plane inside existing workflows. A layer that sits apart from DevOps and the product life cycle becomes shelfware.
  4. Change management needs are significant. Adoption is the hardest part. It needs platform leads engaged in the build, training across the organization, and clear incentives to use golden paths.
TRY THIS TODAY

Goal: Measure what it costs one team to ship a governed agent today, and scope the template that would remove each checkpoint.

  1. Pick one team and its next agent or MCP server.
  2. List every step between “we want to build this” and “it runs in production with identity, registration, and monitoring in place.” Mark each step as a template could do it, or a human must approve it.
  3. Time the whole path from the last three agents that team shipped, using ticket and PR timestamps.
  4. For every step marked “template could do it,” write the field on an intake form that would trigger it.
  5. Give the list to the platform team that resists central governance most. Ask which steps they would accept as a template.

Expected result: a timed path for one team, with each checkpoint marked as template or approval. Add one intake form for the steps a template can absorb.

4.7

Microsoft’s Four Principles of agent governance

One governance checklist for every agent fails in both directions. It adds friction to the simple agents and slows their adoption, and it leaves gaps around the complex agents and creates liability.

Microsoft’s guidance for an agent Center of Excellence gives four principles for that problem, starting with how to govern agents by risk. The guidance covers business agents across an enterprise. This lesson keeps its principles and leaves out its product names.

1. Govern each agent by what it does

The clearest risk signal is what an agent does, not how impressive it looks. An agent that drafts a paragraph, suggests an answer, or summarizes a document assists a person who stays in the loop and owns the outcome. An agent that updates a customer record, submits a ticket, or moves money executes a change in a system of record.

Sort every agent by what it does, not by how impressive it looks. Assistive agents carry limited risk. Agents that act on their own, touch sensitive data, or face customers carry more.

Three tiers follow, each with the controls an agent needs before it ships. An agent starts in the tier that fits its behavior and moves up as its scope or autonomy grows.

Tier
Tier 1Low risk, agents that assist one person or a small team

What it covers. These agents summarize, draft, and search. They take no consequential actions on their own.

Required controls before release. A named owner. Basic monitoring of usage and errors. A standard release checklist. Self-service deployment within published guardrails.

Tier 2Medium risk, agents that answer domain questions or run internal services

What it covers. A wrong answer misleads people or disrupts operations.

Required controls before release. A named owner plus a domain-expert validator. Knowledge-quality monitoring for stale or wrong content. A formal release gate with review. Accuracy tracking and feedback loops.

Tier 3High risk, agents in core processes or facing customers

What it covers. A failure hits revenue, compliance, or trust. These agents execute consequential actions, often with autonomy.

Required controls before release. A named owner plus a process owner. Production-grade SLA monitoring. A security review and a responsible AI assessment. A decision-rights framework for what the agent may decide alone. An incident-response plan. A quarterly maturity review.

TIER CONTROLS REQUIRED BEFORE RELEASE 3 · High risk Core processes, customer-facing Executes, often with autonomy Owner + process owner · SLA monitoring Security review + responsible AI assessment Decision rights · incident-response plan Quarterly maturity review 2 · Medium risk Domain answers, internal services Owner + domain-expert validator Knowledge-quality monitoring · formal release gate Accuracy tracking and feedback loops THE ASSIST-TO-EXECUTE LINE 1 · Low risk Summarize, draft, search. Assists. Named owner · basic monitoring Standard release checklist Self-service within published guardrails CONTROLS GROW WITH WHAT THE AGENT CAN DO, NOT HOW IMPRESSIVE IT LOOKS.
Scroll sideways →Source: built from the three tiers in Microsoft Learn, Govern agents by risk.

Two mechanisms carry every tier. A release gate is a checkpoint an agent passes before production, and its scope depends on the tier. An audit log records what the agent did, who it acted for, and which data it used.

“Governance is never finished. Compliance is continuous.” Publish the guardrails, let teams build within them, and refine.

2. One golden path per tier, each with an owner

A golden path is a proven and approved route to build and ship an agent, with the decisions already made and the guardrails built in. Microsoft’s guidance on reusable assets and golden paths adds one rule to the golden path: offer more than one, so the route fits the risk.

A lightweight path lets someone ship a low-risk productivity agent through self-service in an afternoon. A heavier path routes a business-critical agent through the security review and responsible AI checks it needs before release. Each path stays well defined, approved, and faster than working without one.

Name an owner for every golden path. A path no one maintains drifts out of date, and makers stop trusting it the moment it steers them wrong.

3. Why an agent needs an owner, a monitor, and a retirement date

A project has an end date, and the team may disband after it deploys the agent. A product has an owner who stays with it, tracks its performance, and improves it over time. Microsoft’s guidance on the agent lifecycle asks for the second.

The reason is drift. Every agent in production without monitoring and an improvement plan accumulates risk in four ways:

  1. Knowledge goes stale, so the agent repeats facts that were current at launch.
  2. Source documents change, because people edit the content the agent grounds on outside its view.
  3. User patterns shift, and people use the agent in ways its design never anticipated.
  4. Integrations break when APIs, connectors, and permissions change, so a dependency stops returning what the agent expects.

“Agents don’t fail dramatically. They slowly drift and give increasingly wrong answers with full confidence.”

An agent needs three things before it ships:

  • A named owner, accountable for its value and its risk.
  • A monitoring plan: health checks, accuracy tracking, feedback channels, and alerts.
  • A clear route to improvement or decommissioning.

The lifecycle runs intake, triage, build, deploy, monitor, improve, retire. Each stage has an owner and an exit.

Retirement is a healthy outcome, not a failure. An agent that no longer adds value costs money to run and risk to leave unwatched.

4. How to find the agents built outside your process

Shadow agents are agents built or run outside the Center of Excellence, often by well-meaning teams that move fast. Microsoft’s guidance on securing agents describes them as unmanaged rather than malicious: no registered owner, no reviewed access, no monitoring. Left alone, they are where your next incident starts.

The response is discovery. Find each one, register it, assign an owner, right-size its access, and bring it into the standard lifecycle.

The goal is not to punish the teams that built them. The goal is to make the managed path easier than the shadow path.

Do not build a parallel governance regime for agents. The organization already governs technology. The Center of Excellence connects to the security, IT, and responsible AI bodies it already has, and fills the gaps agents create.

TRY THIS TODAY

Goal: Sort one team’s agents by tier, find the missing controls, and give the oldest agent an owner and a review date.

  1. List every agent one team runs, including CI bots, PR reviewers, scheduled jobs, and coding agents on shared hosts. Mark each as assist or execute:
    Prompt
    Assist: proposes a diff, drafts a PR description, answers a question, summarizes a log
    Execute: merges, deploys, runs a migration, changes a config, opens or closes a ticket
  2. Place each agent in a tier. Anything that executes in a system of record, touches sensitive data, or faces customers starts at tier 2 or 3.
  3. For each agent, check the tier’s required controls from the table above and list the ones it lacks.
  4. For the oldest agent on the list, write down its owner, the date of its last review, and whether it still earns its place.
  5. Add every agent without a registered owner to the inventory from the secrets lesson, and name an owner for each.

Expected result: one team’s agents sorted by tier, a list of missing controls per agent, and an owner and review date for the oldest one.

END OF MODULE 4

By this point you should have:

  • Classified your most autonomous agents against the Rule of Two and listed the shared services that give each one a transitive path.
  • Scored one unattended agent on sandbox, egress, credentials, and monitoring, and asked its vendor which tools run outside the sandbox.
  • Built an inventory of the skills, MCP servers, and instruction files your agents load, sorted by the review each needs.
  • Measured how many agent-authored pull requests passed all four automated gates, and named the weakest gate.
  • Found the secrets in your agents’ configuration and given each non-human identity an owner and an access list.
  • Extended the scanner to configuration files and runners.
  • Timed one team’s path to a governed agent, marked each checkpoint as template or approval, and drafted the intake form for a golden path.
  • Sorted one team’s agents by the assist-to-execute line and listed the missing controls per tier.
  • Named an owner and a review date for the oldest agent.