Cloud Agent Cost

Cost per Merged PR Across Cloud Agent Platforms

Different harnesses and models drive wildly different costs to ship the same code.

Senior Analyst · · 11 min read
Cover illustration for “Cost per Merged PR Across Cloud Agent Platforms”
Cloud Agent Pricing · October 7, 2026 · 11 min read · 2,499 words

Most engineering teams budget for AI coding tools using two numbers: per-seat license cost and per-token API cost. Neither one tells a VP of Engineering whether any code actually shipped. The invoice arrives, leadership asks what the return has been on a few thousand dollars a month in subscriptions, and the honest answer most teams can give is a shrug dressed up in adoption metrics.

A flat monthly seat fee is fixed overhead. It costs the same whether an agent merges a dozen pull requests that month or zero. Seat price measures nothing about whether the tool worked. Token price has the opposite problem: it measures the attempt, not the result. A cheap model that burns five review cycles before it produces mergeable code can end up costing more in total than an expensive model that gets the job right on the first pass. Cost per token and cost per outcome are two different numbers, and teams that track only the first one are flying blind on the second.

The distortion gets worse as agentic workflows scale, because agents consume tokens at a different rate than inline code completion, and billing models are shifting under teams' feet while they try to track it. GitHub Copilot moved to token-based AI Credits billing on June 1, 2026, and promotional credits covered the gap through the summer, masking real baseline costs for Business and Enterprise plans until that window closed on September 1, 2026. A team watching its Copilot bill in July was looking at a number that had little to do with what it would pay in October.

The number that actually connects spend to outcome is cost per merged PR: the total spend attributable to a single pull request that passed review and entered the codebase. Computing it honestly means tracing every dollar back to the PR it produced, and that attribution is harder to build than it sounds.

What cost per merged PR actually measures and what goes into calculating it

Cost per merged PR is the sum of every attributable spend, token consumption, seat allocation, environment compute, and review cycles, that can be traced back to a PR a human actually accepted and merged.

Token cost covers input tokens (the repository map, file reads, the system prompt, accumulated context) and output tokens (reasoning, generated code, commit messages), plus the overhead of retries. In a real agentic session, the code itself is often a small fraction of total tokens consumed; the bulk goes to loading context and iterating toward a working answer. Harness cost is the subscription or seat fee prorated to the session that used it. Environment cost is the compute time for the VM or sandbox that ran the agent, which matters on platforms that bill by the runtime block. Review cycles count how many times the agent's output got sent back for revision before it was mergeable, and each cycle re-sends the full context to the model, which multiplies the token bill every time it happens.

The word "merged" in this metric is doing real work. It excludes PRs that were opened and then rejected, abandoned mid-review, or technically merged but only after so much human rewriting that the agent's actual contribution was nominal. That exclusion ties the metric to a verified engineering outcome.

Most teams miss a cost sitting right next to this one: rework. Unmeasured rework eats into a meaningful share of the time savings teams report to leadership, and it inflates the true cost per merged PR well past what the billing dashboard shows. A PR that "merged" after an engineer spent two hours rewriting half the diff did not cost what the token meter says it cost.

None of this is computable without attribution. For every agent run, a team needs to know which harness ran it, which model it called, which environment it executed in, how long it ran, and which PR (if any) came out the other end. Skip any link in that chain and cost per merged PR turns from a measurement into a guess with a decimal point.

Diagram: What Actually Drives Cost Per Merged PR. Visualizes: Show the five components that sum to cost per merged PR, as defined in the article: token cost (input tokens: repo map, file reads, system prompt, context; output tokens: reasoning…

Why the harness around the model moves cost per PR more than the model itself

Harnesses differ in how much context they load before writing a single line of code, how they handle a failed attempt, how many review cycles they typically need to land a mergeable PR, and how well they use caching to avoid re-paying for the same context twice. Every one of those differences multiplies or compresses the total token bill independent of what the model's per-token price actually is.

That's where the "cheaper model" fallacy breaks down. A model with a lower sticker price per token that needs several extra review cycles to converge can end up with a higher cost per merged PR than a pricier model that resolves the task in one pass. Speed and cost aren't the same axis either: the fastest harness to produce a merge isn't always the cheapest one running, and the cheapest per-token rate doesn't always belong to the fastest path to a merge. A team optimizing for one number alone is very likely optimizing for the wrong thing.

Benchmarks obscure this because they're built to. A benchmark runs a clean prompt against a clean file and reports a cost figure that looks stable and comparable across models. A real PR arrives with a repository map, a pile of file reads, a system prompt, and whatever iteration history has built up from prior attempts. Real-world costs run meaningfully higher than the headline benchmark figures suggest, and the gap is the harness, not the model card.

The same model, routed through different harnesses, can produce very different costs to reach the same merged PR, a finding that should change how teams shop for AI coding tools. Shopping by model name and per-token rate answers a question that doesn't determine the bill. Shopping by harness behavior, how it loads context, how it recovers from a bad first attempt, how many cycles it needs, answers the one that does.

Comparing Replicas, Claude Code, Codex, and OpenCode on Cost Per Merged PR

Replicas takes a different position in this comparison because it is harness-agnostic, letting a team route each task to whichever agent produces the best cost per merged PR for that particular kind of work, rather than committing every task to one harness's cost profile and hoping it fits. That structural flexibility answers the harness spread described above directly: instead of guessing which harness will converge faster on a given repository, a team can run the comparison itself.

Every agent run on Replicas executes inside its own sandboxed VM, pre-loaded with the team's dependencies and tooling, so the agent can install packages, start services, drive a browser, and check its own work before handing back a result. That environment is functionally equivalent to a developer's own machine, and it removes an entire category of failure that comes from agents sharing or fighting over a misconfigured environment. Full analytics attach to every run: the source, the harness, the model, and the credential that paid for it, tracked down to the minute. The attribution chain, harness, model, environment, run time, PR, is recorded by the platform rather than reconstructed after the fact, which is what turns cost per merged PR from something a team estimates into something it can actually compute. Replicas also plugs into Slack, Linear, GitHub, and GitLab, so engineers delegate tasks where their work already happens and get back reviewable PRs, with the human-in-the-loop review step intact that ties the spend to a verified outcome.

Claude Code offers the deepest harness on the market today. It supports more than 31 hook lifecycle events, subagents, agent teams, Dynamic Workflows for parallel orchestration, and a headless CI mode. Doctolib used that headless mode to automatically open PRs for routine maintenance work. It also used Claude Code to replace its entire legacy visual regression testing infrastructure in hours. Claude Code became the most-used AI tool on Doctolib's engineering team, and engineers there shipped features meaningfully faster as a result. The cost driver to watch is that Claude Code reads more files and plans more thoroughly before it writes anything, so token consumption per session runs high, and running several agents in parallel multiplies quota consumption in proportion to how many are running at once. A team used to single-session use can watch daily spend climb faster than expected the moment it starts running several agents side by side.

OpenAI Codex comes included with Free, Go, Plus, and Pro subscriptions, with Pro running at $100 a month and higher tiers that expand usage limits and add open-source, system-level sandboxing. Codex switched to token-based pricing in April 2026, and cost per merged PR now varies with task complexity and context size in ways flat-seat pricing used to hide from view.

OpenCode takes the opposite approach from a managed harness. It's free, open-source, supports more than 75 model providers including local models, and can run headless as a self-hosted server. Because OpenCode carries no bundled model billing, cost per merged PR on it is set almost entirely by whichever model and provider a team configures underneath it. That makes it the most transparent option available for a team that wants to see the raw model cost with nothing else layered on top, but it also puts the full weight of cost management on that team, since the harness itself does nothing to optimize convergence speed or cut down review cycles the way a purpose-built harness does.

GitHub Copilot is market context even though it isn't built around agentic PR workflows in the same way. Pro runs $10 a month, with higher per-user tiers for Business and Enterprise when bundled with GitHub Enterprise Cloud. Copilot moved to token-based AI Credits billing on June 1, 2026: code completions stayed free, but agent mode, chat, Copilot CLI, and code review now draw from a shared credit pool. Promotional credits covered Business and Enterprise usage through August 2026, so teams whose usage patterns hold steady saw their real cost for the first time only in September. An August 2026 Slack integration brought Copilot's agentic features into Slack in public preview, letting teams delegate coding tasks straight from a Slack conversation, and repository administrators can require an extra approval step for any PR attributed to the Copilot app identity before it merges.

What workflow integration and sandbox configuration do to the cost equation

Two teams running the identical harness and the identical model can still land on very different costs per merged PR, because the configuration around the model, not the model, is often what determines the outcome. Where an agent runs and where in the workflow it gets triggered are both cost decisions, whether a team treats them that way or not.

An isolated sandbox is a security property, but it's also a cost and reliability property. An agent running in a properly configured VM, with the team's actual dependencies already installed, is less likely to fail partway through a task, need a retry, or hand back a PR that fails CI the moment it's opened, and each of those failure modes adds directly to spend. Codex's own execution model illustrates the pattern: it provisions or resumes a cached, isolated container, checks out the repository's default branch, and runs its agent loop with no internet access by default, routing any permitted traffic through a proxy. That network restriction closes off a class of prompt injection risk from external content, which in turn cuts down on the mid-task failures that would otherwise mean a wasted run and a second attempt.

Where an agent gets triggered in the workflow matters just as much. An agent kicked off from a well-scoped issue in Linear or GitHub arrives with structured context already in hand: a clear task description, a linked repository, known acceptance criteria. An agent triggered from a loose message in Slack has to do more work just to figure out what's being asked, raising planning token spend and the odds that the first attempt misses the mark. Linear's own agent platform, which reached general availability in June 2026, ties a coding session directly to a specific issue inside a managed sandbox, which is the kind of structured handoff that keeps planning costs down before the agent writes a line of code.

Memory files embedded in the repository, Claude Code's CLAUDE.md, the equivalent AGENTS.md pattern used elsewhere, work as standing instructions that shape how an agent reads the codebase. Teams that keep these files current cut down the context the agent has to reconstruct from scratch on every run, and that compression in context cost applies across every future session against that repository, not just the one in front of it.

Long-running sessions carry their own cost risk through context degradation. A session that starts small can accumulate tens of thousands of tokens across tool outputs, retrieved files, and the back-and-forth of iteration, and once that pile gets large enough, output quality can slip in ways that are hard to trace back to a cause. That slippage appears downstream as an extra review cycle or two, and each one adds straight to the cost of the PR that eventually merges.

Measuring Cost Per Merged PR Honestly

Diagram: Three Prerequisites Before Cost Per Merged PR Means Anything. Visualizes: Visualize the three sequential prerequisites the article names for computing cost per merged PR honestly: (1) Per-run attribution — record harness, model…

Most engineering organizations can't compute cost per merged PR today because they lack the attribution chain the arithmetic depends on: agent runs aren't linked to the specific PRs they produced, harness and model aren't recorded per run, and environment compute sits lumped into general cloud spend rather than broken out by task.

Fixing that takes three things in place before the metric means anything. Per-run attribution comes first: every agent run needs its harness, model, environment (the specific VM or sandbox instance), wall-clock runtime, and resulting PR recorded, including the runs that produced no merged PR. Review cycle tracking comes second, logging how many rounds of revision a given PR went through before it was accepted, since this is the variable most often missing entirely and the one most likely to throw the whole number off. Rework accounting comes third: any PR that required substantial human rewriting after the agent submitted it needs to be flagged as such, because the agent's share of that outcome was partial, and the human time that went into finishing it is a real cost that a token meter will never show.

Platforms built with this attribution as a first-class feature, Replicas among them, record harness, model, environment, and run time for every execution, which turns cost per merged PR into a number a team can pull from actual data rather than reconstruct after the fact from scattered invoices and guesswork. Teams that build this instrumentation now will be the ones able to answer, with a real number, whether their investment in AI coding tools is actually producing shipped work. Teams that keep budgeting by seat price and token price will keep asking the question and keep getting an answer that isn't one.

More in Cloud Agent Pricing