

FOLLOWUS
School of Cyberspace Security, Zhongyuan University of Technology, Zhengzhou 450007, China
School of Computer and Information Security, Guilin University of Electronic Science and Technology, Guilin 541004, China
Henan Key Laboratory of Computational Intelligence and Intelligent System, Zhengzhou 450007, China
✉ Bing XIA, xiabing@zut.edu.cn
Received:19 January 2026,
Revised:2026-06-30,
Published:01 August 2026
Scan QR Code
Yunxiang GE, Bing XIA, Yong DING, et al. Benchmarking large language models for binary function name prediction[J]. ENGINEERING Information Technology & Electronic Engineering, 2026, 27(8): 1-11.
Yunxiang GE, Bing XIA, Yong DING, et al. Benchmarking large language models for binary function name prediction[J]. ENGINEERING Information Technology & Electronic Engineering, 2026, 27(8): 1-11. DOI: 10.1631/ENG.ITEE.2026.0017.
In practical reverse engineering
stripped and optimized binaries lack high-level semantics
hindering automated function understanding. This paper adopts binary function name prediction as a benchmark to evaluate open-source large language models (LLMs) for function-level semantic inference under realistic conditions. We systematically examine key factors affecting performance
including the pretraining domain
model scale
architecture and optimization settings
prompting
and contextual information. Experiments on a large-scale multi-project binary dataset reveal clear limitations in recovering function semantics from stripped binaries. Code-oriented LLMs consistently outperform general-purpose models
while increasing the model size alone does not yield monotonic gains. We further show that realistically recoverable contextual signals
especially structurally recoverable cross-function contexts
substantially mitigate semantic sparsity and improve prediction quality. These results delineate current capability boundaries of LLM-based binary semantic understanding and suggest future directions in contextual modeling and domain-adaptive techniques for practical reverse engineering.
Ahmad W , Chakraborty S , Ray B , et al. , 2021 . Unified pre-training for program understanding and generation . Proc Conf North American Chapter of the Association for Computational Linguistics: Human Language Technologies , p. 2655 - 2668 . https://doi.org/10.18653/v1/2021.naacl-main.211 https://doi.org/10.18653/v1/2021.naacl-main.211
Allamanis M , Barr ET , Bird C , et al. , 2015 . Suggesting accurate method and class names . Proc 10 th Joint Meeting on Foundations of Software Engineering , p. 38 - 49 . https://doi.org/10.1145/2786805.2786849 https://doi.org/10.1145/2786805.2786849
Alon U , Brody S , Levy O , et al. , 2019a . code2seq: generating sequences from structured representations of code . 7 th Int Conf on Learning Representations .
Alon U , Zilberstein M , Levy O , et al. , 2019b . code2vec: learning distributed representations of code . Proc ACM Programm Lang , 3 : 40 . https://doi.org/10.1145/3290353 https://doi.org/10.1145/3290353
Bhattacharya P , Chakraborty M , Palepu KNSN , et al. , 2023 . Exploring large language models for code explanation . https://doi.org/10.48550/arXiv.2310.16673 https://doi.org/10.48550/arXiv.2310.16673
Boronat A , Mustafa J , 2025 . MDRE-LLM: a tool for analyzing and applying LLMs in software reverse engineering . IEEE Int Conf on Software Analysis, Evolution and Reengineering , p. 850 - 854 . https://doi.org/10.1109/SANER64311.2025.00090 https://doi.org/10.1109/SANER64311.2025.00090
Chen M , Tworek J , Jun H , et al. , 2021 . Evaluating large language models trained on code . https://doi.org/10.48550/arXiv.2107.03374 https://doi.org/10.48550/arXiv.2107.03374
David Y , Alon U , Yahav E , 2020 . Neural reverse engineering of stripped binaries using augmented control flow graphs . Proc ACM Programm Lang , 4 ( OOPSLA ): 225 . https://doi.org/10.1145/3428293 https://doi.org/10.1145/3428293
Gao H , Cheng SY , Xue YX , et al. , 2021 . A lightweight framework for function name reassignment based on large-scale stripped binaries . Proc 30 th ACM SIGSOFT Int Symp on Software Testing and Analysis , p. 607 - 619 . https://doi.org/10.1145/3460319.3464804 https://doi.org/10.1145/3460319.3464804
GitHub , 2008 . GitHub: a Platform for Hosting and Collaborating on Software Development . https://github.com/ https://github.com/ [Accessed on Jan. 11, 2026 ] .
Guo DY , Ren S , Lu S , et al. , 2020 . GraphCodeBERT: pre-training code representations with data flow . https://doi.org/10.48550/arXiv.2009.08366 https://doi.org/10.48550/arXiv.2009.08366
Han K , Lim JH , Im EG , 2013 . Malware analysis method using visualization of binary files . Proc Research in Adaptive and Convergent Systems , p. 317 - 321 . https://doi.org/10.1145/2513228.2513294 https://doi.org/10.1145/2513228.2513294
Harzevili NS , Belle AB , Wang JJ , et al. , 2023 . A survey on automated software vulnerability detection using machine learning and deep learning . https://arxiv.org/abs/2306.11673 https://arxiv.org/abs/2306.11673
Høst EW , Østvold BM , 2009 . Debugging method names . 23 rd European Conf on Object-Oriented Programming , p. 294 - 317 . https://doi.org/10.1007/978-3-642-03013-0_14 https://doi.org/10.1007/978-3-642-03013-0_14
Hu XY , Fu ZW , Xie SC , et al. , 2025 . SoK: potentials and challenges of large language models for reverse engineering . https://doi.org/10.48550/arXiv.2509.21821 https://doi.org/10.48550/arXiv.2509.21821
Hugging Face , 2016 . Hugging Face . https://huggingface.co/ https://huggingface.co/ [Accessed on Jan. 15, 2026 ] .
Jaffal NO , Alkhanafseh M , Mohaisen D , 2025 . Large language models in cybersecurity: a survey of applications, vulnerabilities, and defense techniques . AI , 6 ( 9 ): 216 . https://doi.org/10.3390/ai6090216 https://doi.org/10.3390/ai6090216
Jiang LX , Jin X , Lin ZQ , 2025 . Beyond classification: inferring function names in stripped binaries via domain adapted LLMs . Proc ACM SIGSAC Conf on Computer and Communications Security .
Jin X , Pei KX , Won JY , et al. , 2022 . SymLM: predicting function names in stripped binaries via context-sensitive execution-aware code embeddings . Proc ACM SIGSAC Conf on Computer and Communications Security , p. 1631 - 1645 . https://doi.org/10.1145/3548606.3560612 https://doi.org/10.1145/3548606.3560612
Kaplan J , McCandlish S , Henighan T , et al. , 2020 . Scaling laws for neural language models . https://arxiv.org/abs/2001.08361 https://arxiv.org/abs/2001.08361
Kim H , Bak J , Cho K , et al. , 2023 . A Transformer-based function symbol name inference model from an assembly language for binary reversing . Proc ACM Asia Conf on Computer and Communications Security , p. 951 - 965 . https://doi.org/10.1145/3579856.3582823 https://doi.org/10.1145/3579856.3582823
Lawrie D , Morrell C , Feild H , et al. , 2006 . What’s in a name? A study of identifiers . 14 th IEEE Int Conf on Program Comprehension , p. 3 - 12 . https://doi.org/10.1109/ICPC.2006.51 https://doi.org/10.1109/ICPC.2006.51
Meng XZ , Miller BP , 2016 . Binary code is not easy . Proc 25 th Int Symp on Softwar e Testing and Analysis , p. 24 - 35 . https://doi.org/10.1145/2931037.2931047 https://doi.org/10.1145/2931037.2931047
OpenAI , 2023 . OpenAI API Documentation . https://developer.neureality.ai/docs/user-guide/openai-api-doc.html https://developer.neureality.ai/docs/user-guide/openai-api-doc.html [Accessed on Jan. 11, 2026 ] .
Patrick-Evans J , Cavallaro L , Kinder J , 2020 . Probabilistic naming of functions in stripped binaries . Proc 36 th Annual Computer Security Applications Conf , p. 373 - 385 . https://doi.org/10.1145/3427228.3427265 https://doi.org/10.1145/3427228.3427265
Sha ZH , Wang H , Gao ZY , et al. , 2025 . llasm: naming functions in binaries by fusing encoder-only and decoder-only LLMs . ACM Trans Softw Eng Methodol , 34 ( 4 ): 93 . https://doi.org/10.1145/3702988 https://doi.org/10.1145/3702988
Shang XW , Cheng SY , Chen GQ , et al. , 2024 . How far have we gone in binary code understanding using large language models . IEEE Int Conf on Software Maintenance and Evolution , p. 1 - 12 . https://doi.org/10.1109/ICSME58944.2024.00012 https://doi.org/10.1109/ICSME58944.2024.00012
Shin ECR , Song D , Moazzezi R , 2015 . Recognizing functions in binaries with neural networks . 24 th USENIX Conf on Security Symp , p. 611 - 626 .
Song YF , Zhang DD , Wang J , et al. , 2025 . Application of deep learning in malware detection: a review . J Big Data , 12 ( 1 ): 99 . https://doi.org/10.1186/s40537-025-01157-y https://doi.org/10.1186/s40537-025-01157-y
Tan HZ , Luo Q , Li J , et al. , 2024 . LLM4Decompile: decompiling binary code with large language models . Proc Conf on Empirical Methods in Natural Language Processing , p. 3473 - 3487 . https://doi.org/10.18653/v1/2024.emnlp-main.203 https://doi.org/10.18653/v1/2024.emnlp-main.203
Tian JF , Xing WJ , Li Z , 2020 . BVDetector: a program slice-based binary code vulnerability intelligent detection system . Inform Softw Technol , 123 : 106289 . https://doi.org/10.1016/j.infsof.2020.106289 https://doi.org/10.1016/j.infsof.2020.106289
vLLM Project , 2023 . vLLM: a high-throughput and memory-efficient inference and serving engine for LLMs . https://github.com/vllm-project/vllm https://github.com/vllm-project/vllm [Accessed on Jan. 13, 2026 ] .
Wang Y , Wang WS , Joty S , et al. , 2021 . CodeT5: identifier-aware unified pre-trained encoder-decoder models for code understanding and generation . Proc Conf on Empirical Methods in Natural Language Processing , p. 8696 - 8708 . https://doi.org/10.18653/v1/2021.emnlp-main.685 https://doi.org/10.18653/v1/2021.emnlp-main.685
Wong WK , Wu D , Wang H , et al. , 2025 . DecLLM: LLM-augmented recompilable decompilation for enabling programmatic use of decompiled code . Proc ACM Softw Eng , 2 : ISSTA081 . https://doi.org/10.1145/3728958 https://doi.org/10.1145/3728958
Xia B , Pang JM , Wang J , et al. , 2021 . Study on binary code evolution with concrete semantic analysis . 7 th Int Conf of Pioneering Computer Scientists, Engineers and Educators , p. 30 - 43 . https://doi.org/10.1007/978-981-16-5943-0_3 https://doi.org/10.1007/978-981-16-5943-0_3
Xia B , Ge YX , Yang RN , et al. , 2023 . BContext2Name: naming functions in stripped binaries with multi-label learning and neural networks . IEEE 10 th Int Conf on Cyber Security and Cloud Computing and IEEE 9 th Int Conf on Edge Computing and Scalable Cloud , p. 167 - 172 . https://doi.org/10.1109/CSCloud-EdgeCom58631.2023.00037 https://doi.org/10.1109/CSCloud-EdgeCom58631.2023.00037
Xu HX , Wang SN , Li NK , et al. , 2025 . Large language models for cyber security: a systematic literature review . ACM Trans Softw Eng Methodol , in press . https://doi.org/10.1145/3769676 https://doi.org/10.1145/3769676
Yao YF , Duan JH , Xu KD , et al. , 2024 . A survey on large language model (LLM) security and privacy: the good, the bad, and the ugly . High-Confid Comput , 4 ( 2 ): 100211 . https://doi.org/10.1016/j.hcc.2024.100211 https://doi.org/10.1016/j.hcc.2024.100211
Zou MQ , Cai HY , Wu HW , et al. , 2025 . D-LiFT: improving LLM-based decompiler backend via code quality-driven fine-tuning . https://doi.org/10.48550/arXiv.2506.10125 https://doi.org/10.48550/arXiv.2506.10125
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024621