Dev Loop: A Sandboxed Multi-Agent Delivery Platform — AI System Brief | Maneesh Maddala Skip to main content

Dev Loop: A Sandboxed Multi-Agent Delivery Platform

From a Jira comment to an isolated agent that opens a pull request, with governed model access and agent-to-agent interop.

A2A · LiteLLM · Okta · OpenCode · AWS · ECS Fargate · Python

Across our company, teams were building their own agents, each on the stack that suited them. Every one ran alone, with no shared way to execute safely, reach a model, or hand work to another team’s agent. We watched team after team re-solve the same four problems from scratch: isolation, credentials, model access, and interop.

Our team built Dev Loop as one shared path from a Jira comment to a reviewed pull request. The comment starts an isolated agent, the agent works the ticket in its own sandbox, and every model call goes through one governed route.

A ticket enters through a WAF that only trusts Jira, then passes API Gateway and the dev-loop-api Lambda. The Lambda takes a per-ticket lock in DynamoDB so the same ticket never runs twice, then puts an event on SQS. Anything that fails to start lands in a dead-letter queue we can inspect.

A dispatcher Lambda reads the queue and launches one ephemeral ECS Fargate task per ticket. Each task runs an OpenCode agent in its own sandbox, pulls its image from ECR, and reads short-lived Jira, Git, and LiteLLM credentials from Secrets Manager. An EventBridge reaper shuts down any task that runs past its budget.

From inside the sandbox the agent calls MCP tool servers for its tools and context, commits to a branch, and opens a pull request. It speaks A2A when it needs an agent from another team. Every model call goes through LiteLLM, which authorizes it against an Okta-provisioned identity, routes to OpenRouter or AWS Bedrock, and falls back to the default provider when one is down. Persisting run memory in DynamoDB is the next piece.

Architecture diagram
Dev Loop platform architecture A Jira ticket comment passes a WAF Jira allowlist and API Gateway to the dev-loop-api Lambda, which takes a per-ticket lock in DynamoDB and puts an event on an SQS queue. A dispatcher Lambda launches one ephemeral ECS Fargate sandbox per ticket, running an OpenCode agent that pulls its image from ECR and short-lived credentials from Secrets Manager, while an EventBridge reaper enforces a time limit. From the sandbox the agent calls MCP tool servers, opens a pull request on the Git host, makes model calls through LiteLLM which authorizes each call against an Okta-provisioned identity before routing to OpenRouter or AWS Bedrock with a default fallback, and delegates to other teams' agents over A2A. The agent posts status and the pull request link back to the Jira ticket. AWS · DEV LOOP VPC · PRIVATE SUBNETS PER-TICKET EPHEMERAL · ECS FARGATE request acquire lock enqueue consume run task pull image short-lived creds time limit tools open PR model call delegate posts status + PR link back to the ticket Jira ticket agent comment WAF Jira allowlist API Gateway HTTPS endpoint DynamoDB per-ticket lock + state planned: run memory dev-loop-api Lambda SQS + DLQ event queue Dispatcher Lambda ECR agent image Secrets Manager Jira · Git · LiteLLM EventBridge reaper Sandbox OpenCode agent one Fargate task per ticket MCP tool servers tools + context Git host branch · commit · PR Okta authorizes each call LiteLLM routing + fallback OpenRouter · AWS Bedrock A2A agents other teams · any stack request / data flow result posted back to Jira Okta authorization ephemeral · one per ticket

A stuck agent could hold a Fargate task open with nothing to show for it. Each run is an ephemeral task with its own lifecycle, and an EventBridge reaper shuts down anything that passes its time budget.

Two events for one ticket could start the work twice. The dev-loop-api Lambda takes a per-ticket lock in DynamoDB before it enqueues anything, so it drops any second event for a ticket already in flight.

One model provider having an outage could stall every run. Every call goes through LiteLLM rather than a provider SDK, so a request can route to OpenRouter or AWS Bedrock and fall back to the default provider when the first choice is down.

Provider keys spread across repositories would be hard to govern or revoke. LiteLLM authorizes every call against an Okta-provisioned identity, so access is one policy the platform team manages, and each sandbox only ever reads short-lived credentials from Secrets Manager.

Agents from different teams now take part in the same workflow, and no team had to move to a shared stack to join.

Model access is one policy the platform team owns: LiteLLM checks every call against an Okta identity, so there are no provider keys scattered across repositories.

Every run is isolated, scoped to the credentials it needs, and ends at a pull request a person reviews before merge.