{"id":2017,"date":"2025-02-13T14:21:00","date_gmt":"2025-02-13T05:21:00","guid":{"rendered":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/?p=2017"},"modified":"2026-07-17T13:12:35","modified_gmt":"2026-07-17T04:12:35","slug":"improving-the-accuracy-of-source-code-generated-by-large-language-models-through-prompt-engineering-fy-2024-graduation-research","status":"publish","type":"post","link":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/en\/archives\/2017","title":{"rendered":"Improving the Accuracy of Source Code Generated by Large Language Models through Prompt Engineering (FY 2024 Graduation Research)"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">This study focused on source code generation using Large Language Models (LLMs) and investigated how different prompting methods affect the accuracy of generated source code.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Recently, the rapid development of LLMs, including GPT-4 developed by OpenAI, has made it possible to automatically generate source code, improving the efficiency of software development. However, the quality of the generated code largely depends on the prompts given to the model, and it is not always possible to generate correct source code. Therefore, it is important to identify prompting methods that can improve the accuracy of code generation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In this study, GPT-4 developed by OpenAI was used to solve 37 programming problems from AtCoder(<a href=\"https:\/\/atcoder.jp\/\">https:\/\/atcoder.jp\/<\/a>), a competitive programming platform. Two prompting methods, Zero-Shot Prompting and Zero-Shot Chain of Thought (Zero-Shot CoT), were compared. Zero-Shot Prompting generates source code directly from the problem statement, whereas Zero-Shot CoT encourages the model to perform reasoning before generating the code. The generated code was evaluated using AtCoder&#8217;s online judge system, and its accuracy was analyzed based on the number of accepted solutions and various types of errors.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The results showed that Zero-Shot Prompting correctly solved 9 problems, while Zero-Shot CoT correctly solved 8 problems. Although Zero-Shot CoT reduced the number of compilation errors and time limit exceeded errors, it produced more wrong answers than Zero-Shot Prompting. In addition, both methods had difficulty generating correct source code for relatively difficult programming problems.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In the future, we aim to analyze the causes of incorrect code generation and investigate other prompting methods, such as Few-Shot Prompting, to further improve the accuracy of source code generation by LLMs.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><img loading=\"lazy\" decoding=\"async\" width=\"551\" height=\"265\" src=\"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wordpress\/wp-content\/uploads\/2026\/07\/20260710_\u30b3\u30f3\u30c6\u30f3\u30c4B.png\" alt=\"\" class=\"wp-image-2015\" srcset=\"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wordpress\/wp-content\/uploads\/2026\/07\/20260710_\u30b3\u30f3\u30c6\u30f3\u30c4B.png 551w, https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wordpress\/wp-content\/uploads\/2026\/07\/20260710_\u30b3\u30f3\u30c6\u30f3\u30c4B-300x144.png 300w\" sizes=\"auto, (max-width: 551px) 100vw, 551px\" \/><\/figure>\n<\/div>\n\n\n<p class=\"has-text-align-center wp-block-paragraph\">Figure1. Results from the online judging system<\/p>\n","protected":false},"excerpt":{"rendered":"<p>This study focused on source code generation using Large Language Models (LLMs) and investigated how&hellip;<\/p>\n","protected":false},"author":19,"featured_media":2015,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_lmt_disableupdate":"","_lmt_disable":"","_locale":"en_US","_original_post":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/?p=2014","footnotes":""},"categories":[24],"tags":[],"class_list":["post-2017","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-24","en-US"],"modified_by":"\u5e0c\u76f4\u846d\u539f","_links":{"self":[{"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/posts\/2017","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/users\/19"}],"replies":[{"embeddable":true,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/comments?post=2017"}],"version-history":[{"count":1,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/posts\/2017\/revisions"}],"predecessor-version":[{"id":2019,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/posts\/2017\/revisions\/2019"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/media\/2015"}],"wp:attachment":[{"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/media?parent=2017"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/categories?post=2017"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.comm.tcu.ac.jp\/masuda-lab\/wp-json\/wp\/v2\/tags?post=2017"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}