Publication: Code Detection for Hardware Acceleration Using Large Language Models
Authors
Martínez Sánchez, Pablo Antonio ; Bernabé García, Gregorio ; García Carrasco, José Manuel
item.page.secondaryauthor
item.page.director
Publisher
publication.page.editor
publication.page.department
DOI
https://doi.org/10.1109/ACCESS.2024.3372853
item.page.type
info:eu-repo/semantics/article
Description
©2024. This manuscript version is made available under the CC-BY-NC-ND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/
This document is the published, version of a Published Work that appeared in final form in IEEE Access. To access the final edited and published work see https://doi.org/10.1109/ACCESS.2024.3372853
Abstract
Large language models (LLMs) have been massively applied to many tasks, often surpassing state-of-the-art approaches. While their effectiveness in code generation has been extensively studied (e.g., AlphaCode), their potential for code detection remains unexplored. This work presents the first analysis of code detection using LLMs. Our study examines essential kernels, including matrix multiplication, convolution, fast-fourier transform and LU factorization, implemented in C/C++. We propose both a preliminary, naive prompt and a novel prompting strategy for code detection. Results reveal that conventional prompting achieves great precision but poor accuracy (67.5%, 22.5%, 79.5% and 64% for GEMM, convolution, FFT and LU factorization, respectively) due to a high number of false positives. Our novel prompting strategy substantially reduces false positives, resulting in excellent overall accuracy (91.2%, 98%, 99.7% and 99.7%, respectively). These results pose a considerable challenge to existing state-of-the-art code detection methods.
publication.page.subject
Citation
IEEE Access. Volumen 12, 2024
item.page.embargo
Collections
Ir a Estadísticas
Este ítem está sujeto a una licencia Creative Commons. http://creativecommons.org/licenses/by-nc-nd/4.0/