SecureLite: Lightweight Vulnerability Detection on CVEFixes with TF‑IDF, Gradient Boosting, and a Tiny Transformer plus LLM‑Assisted Triage
DOI:
https://doi.org/10.32734/jocai.v10i2.24898Keywords:
vulnerability detection, Python security, TF‑IDF, LightGBM, Transformer encoderAbstract
Static application security testing remains difficult to deploy at scale in Python repositories because practical tools must be accurate, fast, and interpretable under extreme class imbalance. This paper introduces SecureLite, a lightweight pipeline that combines (i) a compact vulnerability detector trained on statement‑annotated data and (ii) an LLM‑assisted triage stage that converts detector evidence into actionable fix guidance. We conduct experiments on the DetectVul/CVEFixes dataset, using its provided train/test splits (4,584/1,146 Python functions). Each function is represented as a sequence of statements, statement types, and per‑statement vulnerability labels. We convert these labels into a function‑level target and compare four efficient detectors: token TF‑IDF + logistic regression (SGD), type‑aware token TF‑IDF, character TF‑IDF, and a LightGBM model over 15 static features. We additionally train a tiny Transformer encoder (TinyVulFormer‑XS) to test whether a minimal self‑attention model can compete with linear baselines under small‑data constraints. On the test set, the best lightweight models (character TF‑IDF and type‑aware token TF‑IDF) achieve AUROC 0.925 and AUPRC 0.381 with an F1 score of 0.421 at a 0.5 decision threshold, outperforming both static‑feature boosting and the tiny Transformer. We further analyze threshold sensitivity, error modes, and how LLM triage can reduce analyst time by proposing fixes and unit tests for high‑risk predictions. The resulting system offers a reproducible, CPU‑friendly baseline for Python vulnerability screening and a practical blueprint for integrating lightweight detection with LLM‑guided remediation.
Downloads
References
[1] H.-C. Tran, A.-D. Tran, and K.-H. Le, “DetectVul: A statement-level code vulnerability detection approach for Python,” Future Generation Computer Systems, vol. 163, p. 107504, 2025, doi: 10.1016/j.future.2024.107504.
[2] Hugging Face, “DetectVul/CVEFixes dataset card,” accessed 2026-03-05.
[3] G. P. Bhandari, A. Naseer, and L. Moonen, “CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software,” arXiv:2107.08760, 2021.
[4] G. P. Bhandari, A. Naseer, and L. Moonen, “CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software,” in Proc. 17th Int. Conf. Predictive Models and Data Analytics in Software Engineering (PROMISE), ACM, 2021.
[5] G. Ke et al., “LightGBM: A Highly Efficient Gradient Boosting Decision Tree,” in Advances in Neural Information Processing Systems (NeurIPS), 2017.
[6] G. Salton and C. Buckley, “Term-weighting approaches in automatic text retrieval,” Information Processing & Management, vol. 24, no. 5, pp. 513–523, 1988.
[7] L. Bottou, “Stochastic Gradient Descent Tricks,” in Neural Networks: Tricks of the Trade, 2nd ed. Springer, 2012.
[8] A. Vaswani et al., “Attention Is All You Need,” in Proc. Advances in Neural Information Processing Systems (NeurIPS), 2017.
[9] Z. Feng et al., “CodeBERT: A Pre-Trained Model for Programming and Natural Languages,” in Findings of EMNLP, 2020.
[10] D. Guo et al., “GraphCodeBERT: Pre-training Code Representations with Data Flow,” arXiv:2009.08366, 2020.
[11] Y. Zhou et al., “Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks,” arXiv:1909.03496, 2019.
[12] M. Chen et al., “Evaluating Large Language Models Trained on Code,” arXiv:2107.03374, 2021.
[13] FIRST.org, “Common Vulnerability Scoring System v3.1: Specification Document,” 2019.
[14] NIST, “National Vulnerability Database (NVD),” accessed 2026-03-05.
[15] OpenStack Security Project, “Bandit: A security linter for Python,” documentation, accessed 2026-03-05.
[16] SonarSource, “SonarQube Security Rules documentation,” accessed 2026-03-05.
[17] T. Saito and M. Rehmsmeier, “The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets,” PLOS ONE, 2015.
[18] F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.
[19] A. Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library,” in Proc. NeurIPS, 2019.
[20] W. McKinney, “Data Structures for Statistical Computing in Python,” in Proc. SciPy, 2010.
[21] C. R. Harris et al., “Array programming with NumPy,” Nature, 2020.
[22] M. Lhoest et al., “Datasets: A Community Library for Natural Language Processing,” in Proc. EMNLP, 2021.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Qi Xin

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.










