MEASURING THE IMPACT OF GENERATIVE AI CODING ASSISTANTS ON SOFTWARE QUALITY, DEVELOPER PRODUCTIVITY, AND TECHNICAL DEBT

Authors

  • Wafa Hassan Author
  • Anees Qumar Abbasi Author
  • Nazia Azim Author
  • Asiya Ali Khan Author
  • Hameed Hussain Author

Keywords:

generative AI; coding assistants; GitHub Copilot; developer productivity; software quality; technical debt; maintainability; software security; empirical software engineering

Abstract

Generative artificial intelligence (AI) coding assistants are now embedded in professional software-development workflows, yet evidence about their engineering value remains fragmented across productivity experiments, code-quality trials, security studies, and repository mining. This article integrates recent empirical evidence to evaluate three connected outcomes: developer productivity, software quality, and technical debt. The evidence base includes controlled and field experiments involving professional developers, code-quality studies using automated tests and blinded review, security analyses of AI-generated code, and a 2026 large-scale study of approximately 302,600 verified AI-authored commits from 6,299 GitHub repositories. Productivity effects are heterogeneous: a controlled experiment reported 55.8% faster task completion with GitHub Copilot [1], three field experiments covering 4,867 developers reported a 26.08% increase in completed tasks [2], and an enterprise randomized trial estimated approximately 21% less time on task [3], whereas a randomized study of experienced open-source developers found that allowing frontier AI tools increased completion time by 19% [4]. Quality evidence is similarly multidimensional. A 2024 randomized study of 202 valid submissions found higher functional success and small but statistically significant improvements in readability, reliability, maintainability, and conciseness with Copilot [5], while security studies continue to identify important weaknesses in AI-generated code [6–8]. At repository scale, AI-authored commits can introduce persistent maintenance liabilities: 89.3% of identified AI-introduced issues were code smells and 22.7% of tracked issues remained at the latest revision [9]. The synthesis shows that generative AI does not produce a uniform productivity-quality trade-off. Its effects depend on task context, developer expertise, evaluation horizon, and verification practices. Lifecycle-adjusted productivity provides a broader organizational measure by incorporating review, rework, security remediation, and technical-debt costs.

 

Downloads

Published

2026-09-30