Engineering the Loop: How Enterprise Teams Make Agents Actually Work

Engineering the Loop: How Enterprise Teams Make Agents Actually Work
•
10 min read
•

I spent a session recently on loop engineering, and it reframed something I had been circling for months. The shift happening in software delivery is not that AI helps you write code faster. It is that the shape of the work is changing, and the way most teams operate was designed for a world where a human typed every line and reviewed every change. That world is ending.

Most teams still run the loop by hand. Task, prompt, review, repeat. The problem was never the human in the loop. It is the human as the loop. Move the person up to the goals and the guardrails, let the cycle run in between, and you have a different kind of system. Leave them cranking every prompt and you have autocomplete with a longer leash.

The teams pulling ahead are building something different. This post is about what that actually looks like, the patterns that work, and why most teams stall before they get there.


The thing most teams miss

Deploying a coding agent is not the hard part. Getting one to produce output that experienced engineers trust and act on is the hard part.

The gap almost always comes down to context. An agent that gets a vague task and a codebase it has never seen produces output that looks reasonable and falls apart on review. An agent that knows the existing contracts, understands the decisions that shaped the codebase, and has clear guidance on what good looks like produces output engineers can work with.

The teams that get this right invest in the context layer before the agent layer:

  • Keep API specs current so the agent works against real contracts instead of guessing

  • Write architecture decisions down in the repo so the reasoning behind the code travels with it

  • Add AGENTS.md and CLAUDE.md files that tell agents which conventions matter and what to avoid

The agent reads all of it before touching a file. Give it something worth reading.

The teams that skip this end up in a pattern everyone recognizes. Agents produce plausible output, engineers spend time correcting it, confidence erodes, usage drops, the project gets quietly shelved. The failure was not the agent. It was the assumption that agents could work around the same documentation debt that makes human onboarding miserable.

What the loop actually is

Stop handing out instructions and start handing over an outcome. You write down the goal and what "done" means, and a system runs the cycle to get there. It picks up the next piece of work, does it, checks its own result, records what happened, and decides what comes next. Then it goes around again, and keeps going until the goal is met. Nobody is sitting there typing "now do this."

Two dials decide what any given loop looks like. One is what starts each turn: a failing CI run, a fresh dependency alert, a scheduled sweep, an issue landing in the queue. The work sets itself off instead of waiting for someone to ask. The other is how far the loop can go on its own before a human steps in. That second dial is the whole game, and it is why the patterns below read as a climb instead of a list.

The machinery underneath is nothing new:

  • Worktrees keep several agents working at once without treading on each other

  • Skills package a recurring task, like cutting a release or running a migration check, so the agent runs it the same way every time instead of improvising

  • Connectors let it act on real systems instead of only describing them

  • Sub-agents mean the one writing the code is not the only one checking it

  • Memory carries context across sessions, so each turn builds on the last instead of starting cold

None of this is exotic. Most of it is already in your stack. What is usually missing is the wiring that turns a pile of one-off scripts into something that runs on its own.

Put together it looks like this: triggers feed a loop on the platform runtime, the loop pulls from a context layer and acts through connectors, and a tiered autonomy gate decides what reaches production.

Loop architecture

Everything below is this one loop pointed at a different job. Spec-to-PR, PR review, incident response, dependency fixes, config drift: same cycle every time. What changes is the trigger that starts it and how far it gets to go before a human says yes.

Five patterns that show up in mature teams

Pattern 1: Spec-to-PR

A developer writes a clear GitHub issue with acceptance criteria and assigns it to the Copilot coding agent, or to Claude Code when the task needs deeper reasoning. The agent spins up an ephemeral GitHub Actions environment, reads the repo, builds an implementation plan in the draft PR, writes the code, and runs tests iteratively until they pass.

Worktrees mean several issues can be in flight at once without conflict. The agent retrieves past patterns and ADR decisions from the repo on every run, not just when a human thinks to mention them.

The developer reviews decisions, not scaffolding. Did the agent reuse the right pattern or invent something? Does the error handling match the service? Does coverage actually exercise the acceptance criteria? That is a different job than writing the implementation, and it is the higher-value half of it.

Spec-to-PR flow

Pattern 2: Autonomous PR review

Every pull request triggers a review sub-agent that runs independently of whatever wrote the code. It pulls the diff, the linked issue, the CI results, the CodeQL output, the service spec, and the ADR history. Then it checks the things that matter:

  • Does the implementation match the contract?

  • Security gaps: missing auth, exposed secrets, injection vectors

  • Performance: N+1 queries, unbounded list operations

  • Test quality: are the failure modes actually covered?

  • Observability: enough instrumentation to debug this in production?

Findings post inline, tagged by severity. Copilot's agentic code review does this natively and gathers full project context before it forms a judgment. When it finds something, it can hand a suggested fix to the coding agent, which opens a fix PR. That closed loop runs without anyone asking.

The naive version of this dies fast. Pass just the diff to a model and senior engineers stop reading the comments after two PRs, because the model does not know what the service is supposed to do. The version that sticks gives the agent the history.

Autonomous PR review flow

Pattern 3: Incident response

An alert fires at 2 a.m. Before the on-call engineer has found their laptop, the agent has pulled correlated logs, mapped the blast radius against the service dependency graph, lined the degradation up against recent deployments, and drafted both the on-call page and the status update.

Autonomy here is tiered, and the tiering is non-negotiable:

Tier How it acts Examples
Tier 1 · auto No confirmation Log retrieval, blast radius map, draft status page, page on-call pre-filled
Tier 2 · confirm Human approves in under a minute Pod restart, feature flag toggle, traffic re-route
Tier 3 · human only Agent prepares, engineer executes DB rollback, secrets rotation, cross-region failover

The agent does not get to bypass this. It is enforced in the infrastructure, not the prompt. Instructions in a prompt are a request. What tool calls are available to the agent is a control.

After resolution, a postmortem sub-agent drafts the timeline and action items and opens a PR against the runbooks directory. Engineers review and merge, and that merged postmortem becomes part of what the incident agent reads next time. The loop learns.

Pattern 4: Dependency remediation

GitHub Dependabot handles detection. The loop handles the reasoning. The agent triages each finding: is the vulnerable path actually called, what is the CVSS score against the org's risk policy, what does the patch break.

  • Patch-level, no breaking changes, no auth or crypto libraries → open PR, run tests, auto-merge if green

  • Major bump or failing tests → open PR for human review with the triage reasoning attached

  • Auth or crypto library, any version → always human, no exceptions

Every finding lands in the compliance log: CVE, exposure window, action taken, timestamp. Automating this does not make the risk zero. It replaces the usual alternative, which is findings sitting untouched in a review queue for six weeks. That queue is the real situation at most organizations.\

Pattern 5: Config drift remediation

Someone made a change in the AWS console during an incident six months ago. It was supposed to be temporary. It is still there. Multiply that across every engineer and every emergency over three years and the gap between what your Terraform says and what is actually running becomes a real problem.

A scheduled loop compares declared state in the repo against live cloud state and classifies what it finds:

  • Expected drift, like autoscaling adjustments → log and skip

  • Suspicious drift, an unrecognized manual change → correlate CloudTrail to find who and when, open a revert PR

  • Policy violation, a security group open to the world → page security on-call immediately, before the investigation finishes

The repo stays the source of truth. The loop enforces it continuously instead of waiting for the next audit.

The GitHub platform is where agents operate

GitHub is the platform. The agents run on it. Most teams have that order backwards: they pick an agent first, then build the workflow around it.

The platform provides the surface every agent plugs into: the trigger, the ephemeral Actions environment, the worktree that isolates changes, the PR, the branch protections and required reviews, the audit log. That contract is the same no matter which agent shows up. GitHub Copilot runs on it natively, the path of least resistance for well-scoped tasks. Claude runs on the same surface through Agentic Workflows, clearing the same triggers, permissions, and review gates. The agent that writes the code does not even have to be the one that reviews it.

So the decision is per task, not per team. Route well-scoped work to whichever agent is fastest and cheapest, ambiguous work to whichever reasons best. Because the guardrails live in the platform and not the agent, you run Copilot and Claude side by side without multiplying your security surface. Bet on the platform, stay flexible on the agents: when the next better one ships, you adopt it instead of re-architecting around it.

Where to start

Not all at once, and not with the workflow that demos best. Climb one rung of autonomy at a time. Do not climb until the rung below is boring.

  • Let it read. The context layer. AGENTS.md, current specs, decisions written down. Everything above reads from it.

  • Let it comment. PR review, read-only. It speaks. It changes nothing. This is the safe place to fail.

  • Let it propose. Spec-to-PR and patch-level dependency bumps. It writes the change. A human merges it.

  • Let it act. Auto-merge the narrow cases. Tier 1 incident comms. Now it changes things on its own.

Conclusion

The agents were never the hard part. The models are already good enough, and they keep improving on their own. What you control is the foundation underneath them: the context you write down, the verification you trust, and the limits on what runs unattended. That is ordinary engineering work, and it stays valuable as the agents get stronger. Get it right and every new model that ships makes your system better without you re-architecting a thing.