Evaluating The Reliability and Robustness of Ai-Augmented Code Generation Tools in Large-Scale Software Projects

Authors

  • Jasvant Mandloi Government Polytechnic Daman UT of DNH & DD Author
  • Rakesh Bhujade Government Polytechnic Daman UT of DNH & DD Author

Keywords:

AI-augmented code generation, software engineering, reliability, robustness, functional correctness, code security, maintainability, evaluation framework, large-scale software, prompt engineering.

Abstract

The rise of AI-powered code generation tools, such as GitHub Copilot, ChatGPT, and Amazon CodeWhisperer, has revolutionised software development by accelerating the process and enabling developers to think more creatively. These tools exhibit significant productivity improvements and perform well on benchmarks such as HumanEval and MBPP; however, their reliance on statistical prediction makes them highly risky for use in large-scale, real-world projects. Current evaluation methods focus on functional correctness in isolation. Still, they don't take into account the problems that arise with large codebases, architectural limits, domain-specific rules, and non-functional requirements such as security, maintainability, and performance. This paper presents a comprehensive evaluation framework designed to assess the reliability and robustness of AI-generated code within realistic project contexts. The methodology integrates quantitative metrics, such as Compilation Success Rate, Functional Correctness Rate, and Prompt Efficiency, with qualitative expert evaluations that emphasize security vulnerabilities, maintainability, and robustness across diverse prompt specificity and context availability. An ecological validity benchmark of 120 tasks drawn from large open-source projects, security datasets, and bug-fix corpora ensures the validity of the results. Real-world tests show that tools can be up to 78% accurate when they have sufficient project context. Still, they remain highly sensitive to unclear prompts and limited contextual information, which can result in code that is not secure or difficult to maintain. CodeWhisperer is more effective than other tools at identifying injection flaws, but no tool is perfect for production-grade systems. The research finds that AI code assistants are not yet capable of working independently. Instead, they need to become context-aware, quality-driven partners. Future research should focus on standardized evaluation frameworks, security-integrated generation, and longitudinal studies concerning technical debt

Downloads

Published

10.09.2026

How to Cite

Evaluating The Reliability and Robustness of Ai-Augmented Code Generation Tools in Large-Scale Software Projects. (2026). BRICS Journal of Applied Mathematics and Engineering Sciences, 1(1), 21-35. https://bricsjournals.com/index.php/bjames/article/view/8