AI agents enhance productivity in writing production software for major companies
What's this about?
People disagree about whether AI tools help big firms make software much faster.
The key question is whether they help with the whole job, not just small code tasks.
What supporters say
- AI can help real work teams make better code, feel more pleased, or work better.
- Tests show that AI can help coders finish set tasks much faster.
What critics say
- AI may make code fast, but shift more work to checks, fixes, and care later.
- One real-world test found that skilled coders became 19% slower with AI tools.
- Some tool-maker tests used small, neat tasks, so their gains may not fit all big firms.
How to read this
The number of points on each side does not show who is right; look at the strength of the proof.
The bottom line
AI can speed up some coding tasks, and it may help some real work teams.
But the proof does not show that AI agents make the full software job much better at big firms, so the claim is too broad.
AI agents can make software development faster, but the evidence does not show that they reliably improve the full production process at major companies. Results vary sharply depending on the task, the developers, the tools and the amount of checking and maintenance required.
The case for
Controlled experiments provide strong evidence that AI can speed up some coding work. A randomized Copilot study found that developers completed a standardized programming task substantially faster. Three field experiments also reported productivity gains in some settings and higher task completion rates. Together, these studies show more than anecdotal enthusiasm: AI assistance can improve performance when tasks are bounded or otherwise favorable. 1
There is also strong evidence that AI can improve results inside real organizations. Studies involving professional developers, including work at TiMi and an evaluation at ANZ Bank, found effects on outcomes such as code quality, satisfaction or workplace performance. DORA research has linked AI adoption with software-delivery outcomes. But that research is based largely on associations, so it cannot prove that AI caused the improvements. 2
The most commonly cited vendor result found developers completing a task 55.8% faster and reporting greater satisfaction (see Figure 1). That finding is consistent with the wider experimental evidence, but the task was tightly controlled and the study was connected to the product’s maker. It therefore shows that the tool can be useful under particular conditions, not that major companies will see the same gains across the software lifecycle.
The case against
Some evidence from a realistic trial points in the opposite direction. In a randomized study of experienced developers working in their own repositories, early-2025 AI tools made participants 19% slower on average. The work involved the kinds of complications common in production software, including navigating repositories, handling dependencies, testing, debugging and integrating changes. The study’s small sample and the age of the tools limit how widely the result can be applied, but it shows that real-world complexity can erase or reverse apparent speed gains. 3
There is strong evidence that faster code generation can move work elsewhere. AI-written code still needs review, testing and security checks, and it may create technical debt or additional maintenance. Large-scale security research has found weaknesses in AI-generated code, although the risks vary by prompt, model, programming language and verification process. Short-term measures of output can therefore miss the cost of fixing, securing and maintaining that code. 4
The strongest positive studies also have limits. Vendor-linked research often uses constrained tasks and measures little about maintainability, security or long-term lifecycle costs. Benchmarks such as SWE-bench show that agents can solve some real software issues, but they do not capture architecture decisions, review burdens, deployment risks, incident response or years of maintenance. Weak evidence from vendor studies is therefore not enough to establish a general effect across major companies. 5
Surveys showing that workers use or plan to use AI demonstrate interest and perceived usefulness, but they also record substantial concerns about accuracy and trust. The central problem is not a lack of research; it is a lack of long-term, broadly representative studies measuring net productivity from design through maintenance.
The bottom line
The evidence supports a conditional version of the claim, but not the broad claim as stated. AI agents can produce substantial productivity gains for routine, well-defined tasks when developers and organizations can effectively verify the results. Controlled studies and some workplace experiments provide strong support for that conclusion.
However, the evidence that AI agents generally and durably improve production-software productivity across major companies is not yet established. The realistic slowdown among experienced developers, along with review, security and maintenance costs, provides a serious counterweight. The overall assessment is balanced with high confidence, while confidence in any single average effect for major companies is lower. Task mix, developer experience, tool maturity, repository familiarity and organizational safeguards remain the main unresolved factors.
Pros — Supporting Arguments
Figures & data
All contributions are reviewed for clarity, balance, and evidence. The strongest insights are elevated into the argument graph — with credit to you.
Help improve this analysis →