How Well Do Large Language Models Detect Bugs in Code Changes?
| dc.contributor.author | Yakubu, Ayinde | |
| dc.date.accessioned | 2026-09-29T19:07:19Z | |
| dc.date.issued | 2026-09-29 | |
| dc.date.submitted | 2026-09-25 | |
| dc.description.abstract | This thesis evaluates how well general-purpose open-weight large language models detect bugs in software code changes. The evaluation uses historical development data from the Apache Kafka project obtained through the ApacheJIT dataset. From approximately 12,000 commit records, the dataset was filtered to obtain 524 one-to-one bug-inducing commit (BIC) and bug-fixing commit (BFC) relationships and 530 non-bug-inducing commits. An automated framework was developed to retrieve commit patches, submit code changes for LLM-based review, and record predictions and review comments. Three open-weight LLMs—gpt-oss-120b, gemma-4-31B-it, and Qwen3.6-35B-A3B—were evaluated under a common zero-shot prompting strategy across three repeated experimental runs. Performance was measured using precision, accuracy, recall, F1-score, balanced accuracy, Matthews correlation coefficient, and processing coverage. In addition, an LLM-as-a-Judge procedure assessed whether generated defect reports were semantically consistent with evidence from corresponding bug-fixing commits and Apache Kafka JIRA issue records. The results show that the evaluated LLMs have limited reliability as autonomous defect detectors. Although the models identified subsets of historically labelled bug-inducing changes, substantial numbers of false positives and false negatives were observed. The first gpt-oss-120b run achieved the highest reported recall of 0.5163, while the highest individual-run accuracy was 0.4872. However, comparison with a trivial always-NOBUG baseline showed that model accuracies did not exceed the corresponding baseline accuracies on successfully processed records. Across the reported runs, balanced accuracy remained below 0.5 and Matthews correlation coefficient (MCC) remained negative, indicating weak overall discrimination between BIC and non-BIC benchmark examples. The results further demonstrate that conventional classification metrics alone provide an incomplete characterisation of LLM bug-detection reliability, because useful review requires semantic correctness, actionable explanations, and sufficient project context. Classification metrics alone do not establish explanation quality. The assigned judges rated 16–23% of first-run true-positive explanations as matching the historically documented defect. The findings suggest that assistant-style use is a more appropriate direction for further evaluation than autonomous defect detection. The thesis contributes a real-world evaluation framework, a comparative empirical assessment of three open-weight LLMs, and an evidence-based methodology for assessing generated explanations against historical defect evidence. The results also highlight repository context, semantic grounding, and hallucination reduction as important directions for improving future LLM-based bug detection systems. | |
| dc.identifier.uri | https://hdl.handle.net/10012/24442 | |
| dc.language.iso | en | |
| dc.pending | false | |
| dc.publisher | University of Waterloo | en |
| dc.relation.uri | https://git.uwaterloo.ca/mmath-thesis-2025-ayinde/onboarding-chatbot-platform | |
| dc.title | How Well Do Large Language Models Detect Bugs in Code Changes? | |
| dc.type | Master Thesis | |
| uws-etd.degree | Master of Mathematics | |
| uws-etd.degree.department | David R. Cheriton School of Computer Science | |
| uws-etd.degree.discipline | Computer Science | |
| uws-etd.degree.grantor | University of Waterloo | en |
| uws-etd.embargo.terms | 0 | |
| uws.contributor.advisor | Nagappan, Mei | |
| uws.contributor.affiliation1 | Faculty of Mathematics | |
| uws.peerReviewStatus | Unreviewed | en |
| uws.published.city | Waterloo | en |
| uws.published.country | Canada | en |
| uws.published.province | Ontario | en |
| uws.scholarLevel | Graduate | en |
| uws.typeOfResource | Text | en |