An Exploratory Study on Code Smells Detection and Refactoring using LLMs

Authors

  • Giorgia Paisi University of Milano-Bicocca
  • Francesca Arcelli Fontana University of Milano-Bicocca
  • Bartosz Walter Poznań University of Technology

DOI:

https://doi.org/10.64552/wipiec.v12i2.145

Keywords:

code smells refactoring, code smells detection, LLM

Abstract

With the observed progress in machine learning (ML), and particularly the introduction of Large Language Models (LLMs), several activities related to code maintenance could be automated. That includes not only detection and evaluation of design flaws, but also code transformation and refactoring. However, the general-purpose LLMs, while being commonly used and popular, have not been specifically trained for code analysis, and may not be suitable for conducting software maintenance tasks due to biases, and inherent shortcomings of the models. In this paper, we explore if the widely available LLMs could aid the detection and the refactoring of code smells. We focus on four common smells (God Class, Long Method, Feature Envy, and Refused Bequest) and consider five prompts of diverse complexity, asking the model for detecting and removing the identified code smells. Results suggest that general-purpose LLMs cannot be reliably used for that. They can effectively detect or remove code smells only in simple cases, and frequently produce invalid code. However, their performance depends on various factors, e.g., the model, the specific code smell or the prompt objective and composition.

Author Biographies

Giorgia Paisi, University of Milano-Bicocca

University of Milano-Bicocca, Milano, Italy.

Francesca Arcelli Fontana, University of Milano-Bicocca

University of Milano-Bicocca, Milano, Italy.

Bartosz Walter, Poznań University of Technology

Poznań University of Technology, Poznań, Poland.

References

M. Fowler, Refactoring: Improving the Design of Existing Code. Addison-Wesley, 1999.

M. Lanza and R. Marinescu, Object-Oriented Metrics in Practice: Using Software Metrics to Characterize, Evaluate, and Improve the Design of Object-Oriented Systems. Springer Science & Business Media, 2006.

M. M¨antyl¨a, J. Vanhanen, and C. Lassenius, “A taxonomy and an initial empirical study of bad smells in code,” in Proceedings of the 19th International Conference on Software Maintenance (ICSM). IEEE, 2003, pp. 381–384.

R. Marinescu, “Detection strategies: Metrics-based rules for detecting design flaws,” in Proceedings of the 20th IEEE International Conference on Software Maintenance (ICSM). IEEE, 2004, pp. 350–359.

S. Vidal, H. Vazquez, A. Diaz-Pace, C. Marcos, A. Garcia, and W. Oizumi, “Jspirit: a flexible tool for the analysis of code smells,” 2015.

N. Moha, Y.-G. Gu´eh´eneuc, L. Duchien, and A.-F. Le Meur, “DECOR: A method for the specification and detection of code and design smells,” IEEE Transactions on Software Engineering, vol. 36, no. 1, pp. 20–36, 2009.

M. Fokaefs, N. Tsantalis, and A. Chatzigeorgiou, “Jdeodorant: Identification and removal of feature envy bad smells,” in 2007 IEEE International Conference on Software Maintenance, 2007.

F. Palomba, G. Bavota, M. Di Penta, R. Oliveto, A. De Lucia, and D. Poshyvanyk, “Detecting bad smells in source code using change history information,” in Proceedings of the 28th IEEE/ACM International Conference on Automated Software Engineering (ASE), 2013, pp. 268–278.

F. A. Fontana, M. V. M¨antyl¨a, M. Zanoni, and A. Marino, “Comparing and experimenting machine learning techniques for code smell detection,” Empir. Softw. Eng., vol. 21, no. 3, pp. 1143–1191, 2016. [Online]. Available: https://doi.org/10.1007/s10664-015-9378-4

F. Palomba, R. Oliveto, and A. De Lucia, “Investigating code smell cooccurrences using association rule learning: A replicated study,” in 2017 IEEE Workshop on Machine Learning Techniques for Software Quality Evaluation (MaLTeSQuE), 2017, pp. 8–13.

B. Walter, F. A. Fontana, and V. Ferme, “Code smells and their collocations: A large-scale experiment on open-source systems,” J. Syst. Softw., vol. 144, pp. 1–21, 2018. [Online]. Available: https://doi.org/10.1016/j.jss.2018.05.057

L. L. Silva, J. R. d. Silva, J. E. Montandon, M. Andrade, and M. T. Valente, “Detecting code smells using chatgpt: Initial insights,” in Proceedings of the 18th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, 2024.

J. Cordeiro, S. Noei, and Y. Zou, “Llm-driven code refactoring: Opportunities and limitations,” 2025 IEEE/ACM Second IDE Workshop (IDE), pp. 32–36, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID:277677973

——, “An empirical study on the code refactoring capability of large language models,” 2024. [Online]. Available: https://arxiv.org/bs/2411.02320

D. Wu, F. Mu, L. Shi, Z. Guo, K. Liu, W. Zhuang, Y. Zhong, and L. Zhang, “ismell: Assembling llms with expert toolsets for code smell detection and refactoring,” in Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, ser. ASE ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 1345–1357. [Online]. Available: https://doi.org/10.1145/3691620.3695508

G. Shetty and T. Sharma, “Mapping code smells and refactorings accurately: Insights from an empirical study,” Preprint, 2025. [Online]. Available: https://tusharma.in/preprints/ESEM2025 Smells Refactoring Mapping.pdf

A. R. Sadik and S. Govind, “Benchmarking llm for code smells detection: Openai gpt-4.0 vs deepseek-v3,” 2025. [Online]. Available: https://arxiv.org/abs/2504.16027

A. Midolo and E. Tramontana, “Refactoring loops in the era of llms: A comprehensive study,” Future Internet, vol. 17, no. 9, 2025.

B. Liu, Y. Jiang, Y. Zhang, N. Niu, G. Li, and H. Liu, “An empirical study on the potential of llms in automated software refactoring,” 2024. [Online]. Available: https://arxiv.org/abs/2411.04444

M. A. Karabiyik, “Refactorgpt: a ChatGPT-based multi-agent framework for automated code refactoring,” PeerJ Computer Science, vol. 11, 2025. [Online]. Available: https://doi.org/10.7717/peerj-cs.3257

A. Alazba, H. Aljamaan, and M. Alshayeb, “Smellybot: An ai-powered software bot for code smell detection,” Software: Practice and Experience, 2025.

N. Alomari, A. Alazba, H. Aljamaan, and M. Alshayeb, “Smellycode++: Multi-label dataset for code smell detection,” Scientific Data, 2025.

N. Anquetil, A. Etien, G. Andreo, and S. Ducasse, “Decomposing god classes at siemens,” in 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME), 2019.

A. Kovacevic, J. Slivka, D. Vidakovi´c, K.-G. Gruji´c, N. Luburi´c, S. Prokic, and G. Sladic, 2022.

M. Shahidi, M. Ashtiani, and M. Zakeri, “An automated extract method refactoring approach to correct the long method code smell,” Journal of Systems and Software, 2022.

D. Yu, Y. Xu, L. Weng, J. Chen, X. Chen, and Q. Yang, “Efficient feature envy detection and refactoring based on graph neural network,” Automated Software Engineering, vol. 32, 12 2024.

E. Ligu, A. Chatzigeorgiou, T. Chaikalis, and N. Ygeionomakis, “Identification of refused bequest code smells,” 09 2013.

M. Zakeri-Nasrabadi, S. Parsa, E. Esmaili, and F. Palomba, “A systematic literature review on the code smells datasets and validation mechanisms,” ACM Computing Surveys, 2023.

D. Cruz, A. Santana, and E. Figueiredo, “Detecting bad smells with machine learning algorithms: an empirical study,” 2020.

K. Prete, N. Rachatasumrit, and M. Kim, “Catalogue of template refactoring rules,” Univ. of Texas at Austin, Tech. Rep. UTAUSTINECETR-

, 2010.

C. Tessa, M. Bochicchio, and F. Arcelli Fontana, Exploring Architectural Smells Detection Through LLMs, ser. Lecture Notes in Computer Science. Springer, 2025, vol. 15929, pp. 90–98. [Online]. Available: https://doi.org/10.1007/978-3-032-02138-0_6

G. Pandini, A. Martini, A. N. Videsjorden, and F. A. Fontana, “An exploratory study on architectural smell refactoring using large languages models,” in 2025 IEEE 22nd International Conference on Software Architecture Companion (ICSA-C), 2025.

B. Rozi`ere, Gehring, Gloeckle, Sootla, Gat, Tan, Adi, Liu, Remez, J. Rapin, Kozhevnikov, I. Evtimov, Bitton, Bhatt, Ferrer, A. Grattafiori, Xiong, D’efossez, Copet, Azhar, Touvron, Martin, N. Usunier, Scialom, and Synnaeve, “Code llama: Open foundation models for code,” ArXiv, vol. abs/2308.12950, 2023. [Online]. Available: https://api.semanticscholar.org/CorpusID:261100919

Downloads

Published

2026-08-25

How to Cite

Paisi, G., Arcelli Fontana, F., & Walter, B. (2026). An Exploratory Study on Code Smells Detection and Refactoring using LLMs. WiPiEC Journal - Works in Progress in Embedded Computing Journal, 12(2), 7. https://doi.org/10.64552/wipiec.v12i2.145