TL;DR

Copilot is fastest on routine, in-file work; Cursor and Claude Code handle cross-file changes better but cost more review time. Standardise the repository first, gate every AI-touched pull request with review, strict static analysis, tests on changed lines and a security pass, keep AI out of auth, payments and crypto, and judge tools on merge rate over twelve weeks.

What does production-grade mean for an AI coding assistant?

An AI coding assistant is production-grade when it passes three measurable checks on your codebase:

  • AI-authored pull requests merge at the team’s overall merge rate or better.
  • Defects escape from AI-touched code no more often than from human-written code, within statistical noise.
  • The time senior engineers spend cleaning up AI suggestions stays smaller than the productivity gain the tool was bought for.

Vendor benchmarks, demo videos and excitement in the team chat say little about any of the three. After two years of rollouts, the debate has moved from whether developers should use assistants to whether the code that ships is any good. Our answer for 2026, from our own production codebases, is: sometimes, with discipline.

Published trials show the possible upside. In Accenture’s randomised controlled trial of Copilot, reported by GitHub, developers opened 8.69% more pull requests each, the merge rate rose 11% and successful builds rose 84%. GitHub Research puts the acceptance rate of Copilot’s inline suggestions in enterprise deployments at about 30%. In our experience such gains survive only where review and testing hold up, so we build the gate before buying seats.

Where do AI coding assistants save time?

AI coding assistants save the most time on work with repeated structure and little product ambiguity. Web development has plenty of it, and the product decisions stay with the developers. The wins we see most often:

  • Boilerplate, component scaffolding and new endpoints in well-known frameworks.
  • Tests generated from existing function signatures or drafted around patterns the repository already uses.
  • Converting one component style to another and refactoring repetitive logic.
  • Accessibility checks, content transformation and first-draft migration notes.
  • Pull request summaries, explanations of legacy files and internal guides that would otherwise be delayed or skipped.
  • Design-to-code handoff, when the tool works from reusable UI patterns rather than raw screenshots.

On contained tasks the published speed-up is large. In GitHub’s controlled trial, developers using Copilot finished a set coding task 55% faster, in 1 hour 11 minutes against 2 hours 41 minutes (95% confidence interval 21-89%), and completed it more often, 78% against 70%, according to GitHub Research. A production backlog is messier than one controlled task, so we treat that figure as a ceiling.

Vadim Leviev, who founded Levievs, uses a rule of thumb: assistants multiply output on routine work, change little on hard design decisions and turn dangerous on security-critical code. The reasoning behind putting routine work first is in our piece on how modern product teams ship.

GitHub Copilot in 2026: strong on routine code, weak across files

GitHub Copilot is our default for inline tab completion on routine work, and we turn it off for new-pattern code. After two years of model upgrades its strengths are unglamorous: boilerplate, tests from existing function signatures and scaffolding in well-known frameworks. Test scaffolding goes much faster when a similar test file is open in the same buffer.

Copilot stalls wherever the context it needs lives outside the current file, as in cross-file refactors or new patterns in an existing codebase. The model fills the gap with plausible code that compiles and passes basic tests while being wrong. In our reviews, routine Copilot suggestions were merged unchanged far more often than suggestions on cross-file refactors, most of which were heavily edited or rejected.

On security, the NYU study “Asleep at the Keyboard” (IEEE S&P 2022) found vulnerabilities in about 40% of Copilot completions in security-sensitive scenarios, and follow-up work in 2024 from Veracode and academic groups put the share closer to 22% for general code. Our own review findings sit between the two.

Copilot is also where most teams begin: in the JetBrains developer survey of January 2026, the split was 29% for Copilot and 18% each for Cursor and Claude Code (Opsera’s analysis of the JetBrains data).

Cursor and Claude Code on real pull requests

Cursor and Claude Code read more of the repository before they generate, so on cross-file work their suggestions fit the codebase far more often than Copilot’s. The price is a slower first response and a heavier review, because their diffs touch more files and the reviewer has to validate a larger surface.

We give them different jobs. Cursor is our main tool for cross-file refactors and migrations, the bulk pattern translation that comes with framework upgrades and platform consolidation. Claude Code takes architecture-level questions and large diffs; it has the heaviest review cost of the three and the highest share of suggestions merged unchanged. None of them replaces a human reviewer.

Assistant Where we use it Where it struggles Review cost
GitHub Copilot Inline completion on routine tests, scaffolds and boilerplate Cross-file refactors, new patterns in an existing codebase Lowest: small, in-file suggestions
Cursor Cross-file refactors and migrations Slower first response Higher: diffs span more files
Claude Code Architecture questions and large diffs Slower first response, large surface to check Highest of the three

Pick by workflow. Start with the assistant that runs inside the IDE your engineers already use, because in 2026 the friction of switching environments outweighs the capability gap between the top three. Add Cursor or Claude Code once migrations, large refactors or architecture exploration become regular work.

What should a team standardise before buying AI seats?

Standardise the repository and the merge rules before buying seats. An assistant learns from the conventions it can see, and a codebase where every module follows its own gives it no stable signal. Five things are worth fixing first:

  1. One linter and one formatter, enforced in CI instead of agreed at stand-up.
  2. A naming convention for files, components and functions, written down in the repository.
  3. The same folder structure in every module. Controller, service and model is fine; mixing patterns causes the trouble.
  4. A README for each significant module that explains in three sentences what it does.
  5. An architecture overview short enough to read in five minutes.

Then set up branch rulesets, a PR template and a CODEOWNERS file. Teams that buy seats without changing those get speed in the wrong direction and review debt that can take a quarter to clear. Write the acceptable-use policy at the same time: when AI-generated code is appropriate, which areas need extra review, and which secrets and private data must never go into a prompt.

After one client standardised module structure and added a short architecture overview, their new engineers reached a first merged pull request noticeably sooner. The co-pilot finally had a coherent codebase to learn from, and the gain came from that cleanup.

What is the four-step review gate for AI-generated code?

Every AI-authored or AI-touched pull request passes the same four steps before merge. GitHub branch rulesets and CODEOWNERS enforce them, so the gate does not depend on chat agreements or goodwill.

  1. Human review with a specific question. In place of “does this look good?”, the reviewer is asked “what does this code assume about the repository that the model could not have seen?” That framing changes what gets caught.
  2. Static analysis on strict settings: TypeScript strict, PHPStan level 8, mypy strict, with every type explicit. AI suggestions that pass loose linting often fail under strict types, and that is where most silent bugs live.
  3. Test coverage on the lines the AI changed. File-level coverage can look healthy while a clever new branch has no test at all.
  4. Security review for sensitive areas. Anything touching authentication, payments, user data or cryptography goes to a separate security reviewer, whoever wrote it, with CODEOWNERS doing the routing.

The PR template adds a nine-point AI code review checklist: input validation, ownership-based authorisation, parameterised queries, file path resolution, secret scanning, dependency vetting, error responses, rate limiting and audit logging. It is the smallest discipline we have seen catch the usual failure modes reliably. AI-generated markup also tends to forget keyboard focus, ARIA labels and colour contrast, so every pull request gets an automated axe or Lighthouse check in CI and UI changes get a manual accessibility review.

Without the gate, an assistant pushes the cost of bugs from writing to review and on to production. In the DORA 2024 report, 39% of respondents said they had little or no trust in AI-generated code, even though more than 80% said AI raised their productivity. A 2025 large-scale arXiv analysis of public GitHub repositories found CWE-mapped security issues in about 12% of AI-generated code, more than in human-written code. Our AI-Powered Web Development Tools work always ships this gate together with the assistant.

Where should AI coding assistants stay switched off?

Keep AI coding assistants out of code where a plausible but wrong answer can end the business. Our standing no-AI zones are:

  • Authentication, authorisation and session handling, because models have absorbed every tutorial pattern that ships exploitable.
  • Cryptographic primitives.
  • Payment-handling code paths, where the surface is small and the consequences are large.
  • Flows that carry personal or health data.
  • Smart-contract logic and custom rate limiters.
  • Performance-critical hot paths, since an assistant optimises for code that looks right and ignores the runtime profile.
  • Novel problems. The model averages across its training data, so a problem far from the average gets an answer that does not fit.

CODEOWNERS enforces the zones, and the PR template repeats them so the reviewer is reminded the moment a sensitive file opens. Outside the zones, AI assistance is normal, and every pull request still passes the nine-point checklist.

Total AI bans push usage underground, where there is zero review.

Vadim Leviev, founder of Levievs

That is why the list stays short instead of turning into a ban. Developers are wary of complex work themselves: in the Stack Overflow Developer Survey 2024, 43% said they trust the accuracy of AI tool output and 45% said AI handles complex tasks poorly. The human-judgement guardrails here are the same shape as those in our note on human-centred AI moderation at scale.

How do AI coding assistants shift work into code review?

AI coding assistants move labour from the IDE into the review queue. The least comfortable result of our own rollout: code-writing productivity went up, review time went up by roughly the same amount, and the net for senior engineers was close to flat. Junior engineers saved real time and their work arrived better shaped, so we kept the tools and started treating review bandwidth as a first-class metric.

Public data points the same way. The DORA Accelerate State of DevOps 2024 report associated high AI adoption with a 7.2% drop in delivery stability and a 1.5% drop in throughput, and attributed both to the larger batch sizes AI encourages. GitClear’s 2025 AI code quality report, built on 211 million lines of code, found copy-pasted lines rising from 8.3% to 12.3% between 2021 and 2024, churn rising from 5.5% to 7.9% and refactored lines falling from 25% to under 10%. More duplication and less refactoring both land on reviewers.

Measure the review-bandwidth tax as the median time to approve AI-touched pull requests minus the median for human-only pull requests. Track it per reviewer: a gap that keeps widening for one person is a burnout signal, and a team average hides it.

The cheapest fix we know: one of our teams narrowed its gap noticeably by adding a one-line “what does this assume?” comment to the PR template, with no new tooling.

The first returns come from internal work

The highest return from AI co-pilots usually comes from internal productivity before any customer-facing feature: faster onboarding, quicker bug diagnosis, easier test writing and better documentation. Co-pilots are strongest on context-rich work that needs broad file awareness and carries little product ambiguity, such as explaining how a legacy module fits a larger workflow or drafting tests around established patterns. The assistant cuts search time and writes the first draft of a change that the engineer still owns.

Context decides the size of the gain. An assistant that cannot see your conventions mostly saves keystrokes, while one that can read standardised modules, READMEs and an architecture overview saves engineering time. In our engagements, time-to-first-PR for new engineers fell sharply in repositories with standardised structure and clear documentation, and bug localisation on legacy modules sped up when the assistant could read the architecture overview. Most of the measurable gains we have recorded came from process discipline more than from the model.

We usually build these rules into a wider Custom Web & Software Development engagement, so the discipline ships with the code.

Which metrics show whether AI coding tools pay off?

Track what happens to AI-touched code and to the people who review it. Seat count measures licences, and hours saved make a story that rarely survives a budget conversation. These are the numbers we use:

Metric How it is measured Healthy signal
Merge rate AI-authored PRs against the team’s overall merge rate At or above the team rate
Defect-escape rate AI-touched code against human-written code No measurable difference
Review-bandwidth tax Median time to approve AI-touched PRs minus human-only PRs, per reviewer Flat or narrowing
Edit rate How much model output survives into the final commit Rising slowly
Time-to-first-PR Time until a new engineer’s first merged pull request Falling
Review-cycle length Average PR review cycle Falling or flat, never growing

Edit rate needs careful reading. Stuck near zero, it means the team rejects everything, which is a process problem. Stuck near 100%, it means nobody reads the output, which is a review problem. Changing the tool fixes neither. Watch the PR rejection rate alongside edit rate.

Give any tool twelve weeks before judging it: four weeks of adoption, four of noise, then four of the first usable data. Decisions made earlier almost always reverse. Repositories with reasonable structure show gains in about a quarter; legacy codebases with mixed conventions take two quarters or more, because their early gains come from standardisation. If neither time-to-first-PR nor review-cycle length has moved after a quarter, the tool is theatre, and one that fails the merge-rate test should go within that quarter.

What survived six months of production use?

After six months on our own codebases, three assistants kept a defined role and two were removed. We describe the results in words because the dashboards behind them are internal.

  • Copilot stayed, as inline tab completion only, and is disabled for new-pattern work.
  • Cursor stayed as the primary tool for cross-file refactors and migration work.
  • Claude Code stayed for architecture-level questions and large diffs, with the heaviest review cost and the highest share of suggestions merged unchanged.
  • Two other assistants went. Their merge rates were low and their defect escapes too high to justify the seat cost.
  • Every pull request now carries an “AI-touched code” label so the gate and the dashboard can find it. No other change we made had a bigger effect.

Across the six months, the assistants behaved as productivity tools in codebases with strict typing, real test coverage and CODEOWNERS files, and produced noise in codebases without them.

When should you write code without an AI assistant?

Write code by hand when the project has strict security requirements, a zero-dependency constraint, functionality that maps to no common pattern, or a need to audit every line, as in regulated industries or code another team will maintain. Lightweight apps and performance-sensitive work often go better without an assistant too.

Assistants sometimes suggest packages that do not exist, and attackers register those hallucinated names with malicious code. Most AI coding tools send code to external servers, so keys, credentials, proprietary algorithms and regulated data leave your environment. A confident-looking implementation also discourages review where it matters most, in authentication and encryption.

One client came to us after a security audit flagged an authentication bypass in their app. The code had been accepted from an AI suggestion six months earlier and had passed every test since, because the tests were AI-generated too and missed the edge case. Manual review of that section would have caught it. Another client needed a web app with zero third-party dependencies, a hard requirement from their security team, so vanilla code was the only real option, and their team could audit every line we delivered.

Our longer piece on when manual coding beats AI assistance covers the security cases in detail.

Frequently asked questions

Is AI-generated code safe to merge into production?

Only after it clears the checks any code should clear, plus a security review in sensitive areas. The NYU “Asleep at the Keyboard” study found vulnerabilities in about 40% of Copilot completions in security-sensitive scenarios. We merge AI-touched code after human review, strict static analysis and tests on the changed lines, and anything touching authentication, payments, user data or cryptography also goes to a separate security reviewer.

Does GitHub Copilot make developers faster?

On contained tasks, clearly. In GitHub’s controlled trial, developers finished a coding task 55% faster with Copilot, and 88% of the more than 2,000 developers in a GitHub survey said they felt more productive. On our own codebases, senior review time grew by about as much as writing time fell, while junior engineers kept a real saving.

Can we use AI coding assistants on proprietary or regulated code?

Yes, with limits set before the first prompt. Most AI coding tools send code to external servers, so keys, credentials, proprietary algorithms and health or financial records should never reach one. Some organisations run self-hosted models or air-gapped environments instead. Ask each vendor how long it retains code and whether it trains on it, the same questions behind the security layers enterprise buyers check.

Do AI coding assistants suggest packages that do not exist?

Yes. A 2024 arXiv analysis found Copilot suggesting non-existent npm packages about 15% of the time. Attackers register such hallucinated names with malicious code, a form of typosquatting. Never run an npm install or pip install command from an assistant without checking that the package exists and who maintains it. Dependency vetting is one of the nine points on our review checklist.

Can AI coding assistants write our tests?

They draft tests well from existing function signatures, but someone still has to check what the tests cover. We have seen an authentication bypass reach production because both the code and its tests came from an assistant, and the tests skipped the edge case. Require coverage on the exact lines the AI changed and review test cases like code.

Do AI coding assistants help with debugging?

For some bugs. They help an engineer find their way around a legacy module and generate test suites around a fault, and bug localisation gets faster when the assistant can read an architecture overview. Race conditions and performance bugs are a different matter. Those still need engineers with profilers and patience, so we keep that work with people.

Should every engineer on the team get an AI coding seat?

Yes, once the no-AI-zones policy and the review checklist exist, and not before. Licences bought ahead of the policy are how teams build up review debt that can take a quarter to clear. Start with the assistant that runs inside your team’s existing IDE, then add Cursor or Claude Code for engineers who regularly run migrations or large refactors.