← All blog posts

Is AI increases developer’s productivity? Maybe not!

July 12, 2025 · AI

Generated by Gemini

Today, with my morning coffee I got stumbled upon this paper, titled “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”.

Since we are doing some trials in my tech organization, by using code assistants and several IDEs, like Windsurf and Cursor, and so far we have good results … not exciting but good. So this makes the above paper even more interesting, because it says that access to early-2025 frontier AI tools increased the time it took for experienced open-source developers to complete software development tasks by 19%! This surprising result contradicts the developers’ own forecasts and estimates, as well as predictions from machine learning and economics experts.

The Study

The study was a randomized controlled trial (RCT) involving 16 experienced developers working on mature open-source projects they were highly familiar with. The developers had, on average, five years of experience with their respective repositories. They completed 246 real-world tasks, such as bug fixes and feature requests, which were randomly assigned to either allow or disallow the use of AI tools. The primary AI tools used were Cursor Pro, a popular AI-native code editor, and Claude 3.5/3.7 Sonnet.

The key findings were:

Unexpected Slowdown: Instead of speeding up development, the use of AI tools led to a 19% increase in task completion time. This slowdown was observed across a wide range of tasks and developers.

Misjudged Productivity: Before the study, developers predicted that AI would reduce their completion time by 24%. After the study, they still estimated a 20% reduction in time, demonstrating a significant disconnect between perceived and actual productivity.

Inaccurate Expert Forecasts: Experts in economics and machine learning also failed to predict the slowdown, forecasting time reductions of 39% and 38%, respectively.

Time Allocation: When using AI, developers spent less time on active coding and reading/searching for information. This time was reallocated to prompting AI systems, waiting for generations, and reviewing AI-generated outputs.

Contributing Factors

he analysis found evidence suggesting five factors likely contributed to the slowdown:

Over-optimism about AI usefulness: Developers’ belief in AI’s benefits may have led to overuse.

High developer familiarity with repositories: The developers’ deep existing knowledge of their codebases may have limited the relative benefit of AI assistance.

Large and complex repositories: The AI tools struggled with the complexity and scale of the mature projects, which averaged over a million lines of code.

Low AI reliability: Developers accepted less than 44% of AI generations and spent significant time (9% of their effort) reviewing and cleaning up AI-generated code.

Implicit repository context: AI tools lacked the tacit knowledge and understanding of undocumented conventions that experienced developers possess.

Epilogue

The authors stress that these findings are specific to the context of experienced developers working on large, familiar codebases. The results do not imply that AI tools are not beneficial in other settings, such as for less experienced developers or on smaller, “greenfield” projects. The paper concludes by highlighting the importance of conducting field experiments with robust outcome measures to accurately gauge the real-world impact of AI, rather than relying on benchmarks or subjective reports.

I have a gut feeling that the tooling we are building for the AI-world is not in the right direction. Instead of use AI as tool to automate more stuff and the English language as the uber-DSL (Domain-specific Language), we are trying to make it replace us. But it shows that it is not the time yet.

Of course, there is hope. People that are going to the right direction.