# SoftwareVulnerabilityDetectionPaper **Repository Path**: lapulatos/software-vulnerability-detection-paper ## Basic Information - **Project Name**: SoftwareVulnerabilityDetectionPaper - **Description**: 收集整理软件缺陷检测相关论文 - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-06-03 - **Last Updated**: 2026-06-03 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # Software Vulnerability Detection — Paper Collection Curated literature anchored on two recent reasoning-based vulnerability detection / verification systems and the broader landscape that motivates them. **Anchor papers:** - **SAVD** (SANER 2026) — *Synergizing LLM-Driven Semantic Reasoning with Assertion-Guided Analysis for Enhanced Vulnerability Detection.* Wang, Su, Wen, et al. (Xidian / GPNU). Verifier-in-the-loop, LLM-assisted weakest-precondition reasoning over assertion-guided slices. Home: **D.2** below. - **FM-Agent** (arXiv 2604.11556, SJTU IPADS) — *Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning.* Caller-driven natural-language Hoare triples scaled to 277k LoC, 522 previously undiscovered bugs found in 2 days. Home: **D.3** below. Scope of the collection: - Classical symbolic execution, static analysis, slicing, formal verification — the methodological roots. - Pre-LLM deep-learning vulnerability detectors (Transformers, GNNs, hybrids). - General-purpose LLM reasoning (CoT, ToT, search, self-critique, o1-style RL reasoners). - LLMs applied to vulnerability detection / program verification (prompting, static/symbolic hybrids, formal-methods hybrids, agents, fine-tuning, RAG). - Adjacent techniques: fuzzing, automated repair, code LLMs. - Evaluation: benchmarks, datasets, empirical critiques. Each entry links to its open-access copy (arXiv / DOI / publisher). PDFs are mirrored locally under `paper/{TopCat}/{SubCat}/` but are git-ignored — only this README is tracked. --- ## Contents - **A. Surveys & Foundations** - [A.1 Surveys — LLM × SE / Vulnerability Detection](#a1-surveys-llm-se-vulnerability-detection) — 11 papers - [A.2 Surveys — Symbolic Execution & Verification](#a2-surveys-symbolic-execution-verification) — 3 papers - [A.3 Surveys — LLM Reasoning](#a3-surveys-llm-reasoning) — 5 papers - [A.4 Classical Symbolic & Concolic Execution](#a4-classical-symbolic-concolic-execution) — 8 papers - [A.5 Classical Static Analysis & Program Slicing](#a5-classical-static-analysis-program-slicing) — 7 papers - [A.6 Classical Formal Verification & Model Checking](#a6-classical-formal-verification-model-checking) — 8 papers - [A.7 Code Property Graphs & Code Representations](#a7-code-property-graphs-code-representations) — 5 papers - **B. Pre-LLM Learning-Based Vulnerability Detection** - [B.1 Token / Transformer Models for Code](#b1-token-transformer-models-for-code) — 8 papers - [B.2 GNN-Based Vulnerability Detection](#b2-gnn-based-vulnerability-detection) — 13 papers - [B.3 Hybrid Multi-View / Causal Detection](#b3-hybrid-multi-view-causal-detection) — 5 papers - **C. LLM Reasoning (General)** - [C.1 Chain-of-Thought & In-Context Reasoning](#c1-chain-of-thought-in-context-reasoning) — 10 papers - [C.2 Tree / Graph / Forest of Thoughts & Search](#c2-tree-graph-forest-of-thoughts-search) — 7 papers - [C.3 Self-Consistency, Self-Refine, Self-Critique](#c3-self-consistency-self-refine-self-critique) — 5 papers - [C.4 Process Supervision & RL-for-Reasoning (o1-style)](#c4-process-supervision-rl-for-reasoning-o1-style) — 10 papers - [C.5 Tool / Code-Aided / ReAct Reasoning](#c5-tool-code-aided-react-reasoning) — 5 papers - [C.6 LLM-as-Verifier, Critic & Debate](#c6-llm-as-verifier-critic-debate) — 3 papers - **D. LLMs for Vulnerability Detection & Verification** - [D.1 LLM Prompting & Zero/Few-shot Detection](#d1-llm-prompting-zerofew-shot-detection) — 12 papers - [D.2 LLM + Static / Symbolic Hybrids — **SAVD home**](#d2-llm-static-symbolic-hybrids-savd-home) — 6 papers - [D.3 LLM + Formal Methods / Hoare-style — **FM-Agent home**](#d3-llm-formal-methods-hoare-style-fm-agent-home) — 13 papers - [D.4 LLM Agents & Multi-Agent Workflows for Security](#d4-llm-agents-multi-agent-workflows-for-security) — 16 papers - [D.5 Fine-Tuning / Instruction Tuning of Code LLMs](#d5-fine-tuning-instruction-tuning-of-code-llms) — 15 papers - [D.6 Retrieval-Augmented / Knowledge-Augmented Detection](#d6-retrieval-augmented-knowledge-augmented-detection) — 2 papers - [D.7 CoT / Reasoning-Chain Variants for Vulnerability](#d7-cot-reasoning-chain-variants-for-vulnerability) — 1 papers - **E. Adjacent Techniques** - [E.1 Fuzzing — Classical & LLM-Augmented](#e1-fuzzing-classical-llm-augmented) — 13 papers - [E.2 Automated Program Repair](#e2-automated-program-repair) — 11 papers - [E.3 Code Generation / Code LLMs](#e3-code-generation-code-llms) — 4 papers - **F. Evaluation** - [F.1 Benchmarks & Datasets](#f1-benchmarks-datasets) — 10 papers - [F.2 Empirical Studies & Critiques](#f2-empirical-studies-critiques) — 9 papers --- ## A. Surveys & Foundations _Surveys, tutorials, and the classical foundations of vulnerability detection: symbolic execution, static analysis, formal verification, and code representations._ ### A.1 Surveys — LLM × SE / Vulnerability Detection _Systematic reviews of LLM-assisted software engineering, vulnerability detection, and code analysis._ - [A Survey on Large Language Model (LLM) Security and Privacy: The Good, The Bad, and The Ugly (High-Confidence Computing 2024)](https://arxiv.org/abs/2312.02003) - [Large Language Models for Software Vulnerability Detection: A Guide for Researchers on Models, Methods, Techniques, Datasets, and Metrics (Int. J. Inf. Sec. 2025)](https://doi.org/10.1007/s10207-025-00992-7) - [Large Language Models for Software Engineering: A Systematic Literature Review (TOSEM 2024)](https://arxiv.org/abs/2308.10620) - [Deep Learning Based Vulnerability Detection: Are We There Yet? (TSE 2022)](https://arxiv.org/abs/2009.07235) - [Beyond Self-Talk: A Communication-Centric Survey of LLM-Based Multi-Agent Systems (arXiv 2025)](https://arxiv.org/abs/2502.14321) - [A Survey on LLM-Based Code Generation for Low-Resource and Domain-Specific Programming Languages (arXiv 2024)](https://arxiv.org/abs/2410.03981) - [Code Vulnerability Detection: A Comparative Analysis of Emerging Large Language Models (arXiv 2024)](https://arxiv.org/abs/2409.10490) - [Large Language Models for Cyber Security: A Systematic Literature Review (arXiv 2024)](https://arxiv.org/abs/2405.04760) - [A Survey on Automated Software Vulnerability Detection Using Machine Learning and Deep Learning (arXiv 2023)](https://arxiv.org/abs/2306.11673) - [Pre-trained Models for Source Code: A Survey (arXiv 2022)](https://arxiv.org/abs/2205.11739) - [Software Vulnerability Detection via Deep Learning over Disaggregated Code Graph Representation (arXiv 2021)](https://arxiv.org/abs/2109.03341) ### A.2 Surveys — Symbolic Execution & Verification _Classical surveys of symbolic / concolic execution, verification, and model checking._ - [A Survey of Symbolic Execution Techniques (ACM Comput. Surv. 2018)](https://arxiv.org/abs/1610.00502) - [A Survey of New Trends in Symbolic Execution for Software Testing and Analysis (STTT 2009)](https://doi.org/10.1007/s10009-009-0118-1) - [The Art, Science, and Engineering of Fuzzing: A Survey (TSE 2021)](https://arxiv.org/abs/1812.00140) ### A.3 Surveys — LLM Reasoning _Surveys of reasoning paradigms, foundation-model reasoning, chain-of-X._ - [A Survey on Large Language Models with some insights on their capabilities and limitations (arXiv 2025)](https://arxiv.org/abs/2501.04040) - [From System 1 to System 2: A Survey of Reasoning Large Language Models (arXiv 2025)](https://arxiv.org/abs/2502.17419) - [Reasoning Beyond Limits: Advances and Open Problems for LLMs (arXiv 2025)](https://arxiv.org/abs/2503.22732) - [Beyond Chain-of-Thought: A Survey of Chain-of-X Paradigms for LLMs (arXiv 2024)](https://arxiv.org/abs/2404.15676) - [A Survey of Reasoning with Foundation Models (arXiv 2023)](https://arxiv.org/abs/2312.11562) ### A.4 Classical Symbolic & Concolic Execution _KLEE, DART, S2E, Angr, SymCC and other symbolic-execution engines that SAVD's verifier layer draws from._ - [S2E: A Platform for In-Vivo Multi-Path Analysis of Software Systems (ASPLOS 2011)](https://doi.org/10.1145/1950365.1950396) - [Symbolic Execution for Software Testing: Three Decades Later (CACM 2013)](https://doi.org/10.1145/2408776.2408795) - [EXE: Automatically Generating Inputs of Death (CCS 2006)](https://doi.org/10.1145/1180405.1180445) - [Enhancing Symbolic Execution with Veritesting (ICSE 2014)](https://doi.org/10.1145/2568225.2568293) - [KLEE: Unassisted and Automatic Generation of High-Coverage Tests for Complex Systems Programs (OSDI 2008)](https://llvm.org/pubs/2008-12-OSDI-KLEE.pdf) - [DART: Directed Automated Random Testing (PLDI 2005)](https://doi.org/10.1145/1065010.1065036) - [(State of) The Art of War: Offensive Techniques in Binary Analysis (Angr) (S&P 2016)](https://doi.org/10.1109/SP.2016.17) - [Symbolic Execution with SymCC: Don't Interpret, Compile! (USENIX Security 2020)](https://arxiv.org/abs/1907.01054) ### A.5 Classical Static Analysis & Program Slicing _Dataflow analysis, value-flow, taint, points-to, and program slicing — the assertion-guided slicing lineage._ - [SVF: Interprocedural Static Value-Flow Analysis in LLVM (CC 2016)](https://doi.org/10.1145/2892208.2892235) - [Lifestate: Event-Driven Protocols and Callback Control Flow (ECOOP 2019)](https://doi.org/10.4230/LIPIcs.ECOOP.2018.13) - [Practical Static Analysis of JavaScript Applications in the Presence of Frameworks and Libraries (FSE 2013)](https://doi.org/10.1145/2491411.2491417) - [Pinpoint: Fast and Precise Sparse Value Flow Analysis for Million LoC C Programs (PLDI 2018)](https://doi.org/10.1145/3192366.3192418) - [Precise Interprocedural Dataflow Analysis via Graph Reachability (POPL 1995)](https://doi.org/10.1145/199448.199462) - [Program Slicing (TSE 1984)](https://doi.org/10.1109/TSE.1984.5010248) - [WALA: T.J. Watson Libraries for Analysis (Tool) (Tool 2006)](https://github.com/wala/WALA) ### A.6 Classical Formal Verification & Model Checking _Bounded model checking, CEGAR, lazy abstraction, software verifier competitions. The verifier-in-the-loop lineage._ - [Bounded Model Checking (Advances in Computers 2003)](https://doi.org/10.1016/S0065-2458(03)58003-2) - [CPAchecker: A Tool for Configurable Software Verification (CAV 2011)](https://arxiv.org/abs/1301.6837) - [Counterexample-Guided Abstraction Refinement (CAV 2000)](https://doi.org/10.1007/10722167_15) - [Model Checking and the State Explosion Problem (LASER 2011)](https://doi.org/10.1007/978-3-642-35746-6_1) - [Lazy Abstraction (POPL 2002)](https://doi.org/10.1145/503272.503279) - [ESBMC v7.4: Harnessing the Power of Intervals (TACAS 2024)](https://arxiv.org/abs/2312.14746) - [Symbolic Model Checking without BDDs (TACAS 1999)](https://doi.org/10.1007/3-540-49059-0_14) - [Checking Safety Properties Using Compositional Reachability Analysis (TOSEM 1999)](https://doi.org/10.1145/295558.295570) ### A.7 Code Property Graphs & Code Representations _Joern CPG, ASTNN, code2vec / code2seq, GNN code embeddings._ - [code2seq: Generating Sequences from Structured Representations of Code (ICLR 2019)](https://arxiv.org/abs/1808.01400) - [Learning to Represent Programs with Graphs (ICLR 2018)](https://arxiv.org/abs/1711.00740) - [A Novel Neural Source Code Representation Based on Abstract Syntax Tree (ICSE 2019)](https://doi.org/10.1109/ICSE.2019.00086) - [code2vec: Learning Distributed Representations of Code (POPL 2019)](https://arxiv.org/abs/1803.09473) - [Modeling and Discovering Vulnerabilities with Code Property Graphs (Joern) (S&P 2014)](https://doi.org/10.1109/SP.2014.44) ## B. Pre-LLM Learning-Based Vulnerability Detection _Deep-learning detectors trained on code corpora before the LLM era: token / Transformer encoders, GNNs, and multi-view hybrids._ ### B.1 Token / Transformer Models for Code _CodeBERT, GraphCodeBERT, CodeT5, VulBERTa, LineVul — token / sequence encoders for vulnerability tasks._ - [Transformer-Based Language Models for Software Vulnerability Detection (ACSAC 2022)](https://arxiv.org/abs/2204.03214) - [CodeBERT: A Pre-Trained Model for Programming and Natural Languages (EMNLP Findings 2020)](https://arxiv.org/abs/2002.08155) - [Automated Vulnerability Detection in Source Code Using Deep Representation Learning (ICMLA 2018)](https://arxiv.org/abs/1807.04320) - [Large Language Model for Vulnerability Detection: Emerging Results and Future Directions (ICSE NIER 2024)](https://arxiv.org/abs/2401.15468) - [VulBERTa: Simplified Source Code Pre-Training for Vulnerability Detection (IJCNN 2022)](https://arxiv.org/abs/2205.12424) - [LineVul: A Transformer-based Line-Level Vulnerability Prediction (MSR 2022)](https://doi.org/10.1145/3524842.3528452) - [Software Vulnerability Prediction Using Text Analysis Techniques (MetriSec 2012)](https://doi.org/10.1145/2372225.2372226) - [TreeBERT: A Tree-based Pre-trained Model for Programming Language (UAI 2021)](https://arxiv.org/abs/2105.12485) ### B.2 GNN-Based Vulnerability Detection _Graph neural networks over CFG / DFG / CPG / AST: Devign, VulDeePecker, SySeVR, ReVeal, IVDetect, AMPLE, EPVD, PrimeVul._ - [When Less is Enough: Positive and Unlabeled Learning Model for Vulnerability Detection (ASE 2023)](https://arxiv.org/abs/2308.10523) - [IVDetect: Vulnerability Detection with Fine-Grained Interpretations (FSE 2021)](https://arxiv.org/abs/2106.10478) - [Vulnerability Detection with Code Language Models: How Far Are We? (PrimeVul) (ICSE 2025)](https://arxiv.org/abs/2403.18624) - [AMPLE: Enhancing Vulnerability Detection via Simplified Code Representation (ICSE 2023)](https://arxiv.org/abs/2302.04675) - [MVD: Memory-Related Vulnerability Detection Using Flow-Sensitive Graph Neural Networks (ICSE 2022)](https://arxiv.org/abs/2203.02660) - [VulCNN: An Image-Inspired Scalable Vulnerability Detection System (ICSE 2022)](https://arxiv.org/abs/2204.09124) - [ReGVD: Revisiting Graph Neural Networks for Vulnerability Detection (ICSE Companion 2022)](https://arxiv.org/abs/2110.07317) - [LineVD: Statement-level Vulnerability Detection using Graph Neural Networks (MSR 2022)](https://arxiv.org/abs/2203.05181) - [VulDeePecker: A Deep Learning-Based System for Vulnerability Detection (NDSS 2018)](https://arxiv.org/abs/1801.01681) - [Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks (NeurIPS 2019)](https://arxiv.org/abs/1909.03496) - [SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities (TDSC 2022)](https://arxiv.org/abs/1807.06756) - [DeepWukong: Statically Detecting Software Vulnerabilities Using Deep Graph Neural Network (TOSEM 2021)](https://doi.org/10.1145/3436877) - [Vulnerability Detection via Multiple-Graph-Based Code Representation (TSE 2024)](https://doi.org/10.1109/TSE.2024.3402543) ### B.3 Hybrid Multi-View / Causal Detection _Multi-granularity, program-metric, causal-DL, dependence-aware hybrids._ - [Pre-Training by Predicting Program Dependencies for Vulnerability Analysis Tasks (ICSE 2024)](https://arxiv.org/abs/2402.00657) - [TRACED: Execution-Aware Pre-Training for Source Code (ICSE 2024)](https://arxiv.org/abs/2306.07487) - [Towards Causal Deep Learning for Vulnerability Detection (ICSE 2024)](https://arxiv.org/abs/2310.07958) - [LEOPARD: Identifying Vulnerable Code for Vulnerability Assessment Through Program Metrics (ICSE 2019)](https://arxiv.org/abs/1901.11479) - [mVulPreter: A Multi-Granularity Vulnerability Detection System with Interpretations (TDSC 2022)](https://doi.org/10.1109/TDSC.2022.3225525) ## C. LLM Reasoning (General) _General-purpose LLM reasoning capabilities — chain-of-thought, search-based reasoning, self-critique, RL-based reasoners, code-aided reasoning, LLM-as-verifier. These are the engine behind SAVD and FM-Agent._ ### C.1 Chain-of-Thought & In-Context Reasoning _The CoT lineage: CoT, zero-shot reasoners, self-consistency, complexity-based prompting, STaR, Quiet-STaR._ - [Towards Reasoning in Large Language Models: A Survey (ACL Findings 2023)](https://arxiv.org/abs/2212.10403) - [Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking (COLM 2024)](https://arxiv.org/abs/2403.09629) - [Automatic Chain of Thought Prompting in Large Language Models (ICLR 2023)](https://arxiv.org/abs/2210.03493) - [Complexity-Based Prompting for Multi-Step Reasoning (ICLR 2023)](https://arxiv.org/abs/2210.00720) - [Self-Consistency Improves Chain of Thought Reasoning in Language Models (ICLR 2023)](https://arxiv.org/abs/2203.11171) - [Chain-of-Thought Prompting of Large Language Models for Discovering and Fixing Software Vulnerabilities (ICSE 2025)](https://arxiv.org/abs/2402.17230) - [Faithful Chain-of-Thought Reasoning (IJCNLP-AACL 2023)](https://arxiv.org/abs/2301.13379) - [Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (NeurIPS 2022)](https://arxiv.org/abs/2201.11903) - [STaR: Bootstrapping Reasoning With Reasoning (NeurIPS 2022)](https://arxiv.org/abs/2203.14465) - [Chain-of-Thought in Neural Code Generation: From and For Lightweight Language Models (TSE 2024)](https://arxiv.org/abs/2312.05562) ### C.2 Tree / Graph / Forest of Thoughts & Search _Search-augmented reasoning: ToT, GoT, RAP, Buffer-of-Thoughts, Forest-of-Thought, MindStar, RAR._ - [Graph of Thoughts: Solving Elaborate Problems with Large Language Models (AAAI 2024)](https://arxiv.org/abs/2308.09687) - [Reasoning with Language Model is Planning with World Model (RAP) (EMNLP 2023)](https://arxiv.org/abs/2305.14992) - [Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models (NeurIPS 2024)](https://arxiv.org/abs/2406.04271) - [Tree of Thoughts: Deliberate Problem Solving with Large Language Models (NeurIPS 2023)](https://arxiv.org/abs/2305.10601) - [Forest-of-Thought: Scaling Test-Time Compute for Enhancing LLM Reasoning (arXiv 2024)](https://arxiv.org/abs/2412.09078) - [Mindstar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time (arXiv 2024)](https://arxiv.org/abs/2405.16265) - [RAR: Retrieval-Augmented Reasoning with Trustworthy Process Rewarding (arXiv 2024)](https://arxiv.org/abs/2412.13746) ### C.3 Self-Consistency, Self-Refine, Self-Critique _Iterative self-correction: Self-Refine, Reflexion, CRITIC, Self-Verification, Self-RAG, self-correction limits._ - [CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing (ICLR 2024)](https://arxiv.org/abs/2305.11738) - [Large Language Models Cannot Self-Correct Reasoning Yet (ICLR 2024)](https://arxiv.org/abs/2310.01798) - [Reflexion: Language Agents with Verbal Reinforcement Learning (NeurIPS 2023)](https://arxiv.org/abs/2303.11366) - [Self-Refine: Iterative Refinement with Self-Feedback (NeurIPS 2023)](https://arxiv.org/abs/2303.17651) - [Self-Correction Bench: Uncovering and Addressing the Self-Correction Blind Spot in Large Language Models (arXiv 2025)](https://arxiv.org/abs/2507.02778) ### C.4 Process Supervision & RL-for-Reasoning (o1-style) _Process reward models and reasoning-RL: Let's Verify Step by Step, Math-Shepherd, o1 system card, DeepSeek-R1, Marco-o1, Skywork-OR1, Kimi K1.5, rStar-Math, LLaMA-Berry._ - [Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations (ACL 2024)](https://arxiv.org/abs/2312.08935) - [Let's Verify Step by Step (ICLR 2024)](https://arxiv.org/abs/2305.20050) - [DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv 2025)](https://arxiv.org/abs/2501.12948) - [Kimi K1.5: Scaling Reinforcement Learning with LLMs (arXiv 2025)](https://arxiv.org/abs/2501.12599) - [Reasoning with Reinforced Functional Token Tuning (arXiv 2025)](https://arxiv.org/abs/2502.13389) - [Skywork-OR1: New Frontier in Open Reasoning Models (arXiv 2025)](https://arxiv.org/abs/2505.22312) - [rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking (arXiv 2025)](https://arxiv.org/abs/2501.04519) - [LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning (arXiv 2024)](https://arxiv.org/abs/2410.02884) - [Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions (arXiv 2024)](https://arxiv.org/abs/2411.14405) - [OpenAI o1 System Card (arXiv 2024)](https://arxiv.org/abs/2412.16720) ### C.5 Tool / Code-Aided / ReAct Reasoning _Program-aided LMs (PAL, PoT), Toolformer, ReAct, code-prompting for conditional reasoning._ - [Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs (EMNLP 2024)](https://arxiv.org/abs/2401.10065) - [ReAct: Synergizing Reasoning and Acting in Language Models (ICLR 2023)](https://arxiv.org/abs/2210.03629) - [PAL: Program-aided Language Models (ICML 2023)](https://arxiv.org/abs/2211.10435) - [ToolFormer: Language Models Can Teach Themselves to Use Tools (NeurIPS 2023)](https://arxiv.org/abs/2302.04761) - [PoT: Program-of-Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks (TMLR 2023)](https://arxiv.org/abs/2211.12588) ### C.6 LLM-as-Verifier, Critic & Debate _Verifier-style reasoning supervision: Generative Verifiers, LLM critics, debate, adversarial critic._ - [Debate Helps Supervise Unreliable Experts (arXiv 2024)](https://arxiv.org/abs/2311.08702) - [Generative Verifiers: Reward Modeling as Next-Token Prediction (arXiv 2024)](https://arxiv.org/abs/2408.15240) - [LLM Critics Help Catch LLM Bugs (arXiv 2024)](https://arxiv.org/abs/2407.00215) ## D. LLMs for Vulnerability Detection & Verification _LLMs applied directly to vulnerability detection and program verification — prompting, static/symbolic hybrids (SAVD home), formal/Hoare-style reasoning (FM-Agent home), agents, fine-tuning, RAG, and CoT variants._ ### D.1 LLM Prompting & Zero/Few-shot Detection _Direct prompting of GPT / Claude / DeepSeek / Code LLMs for vulnerability identification and triage._ - [A Comprehensive Study of the Capabilities of Large Language Models for Vulnerability Detection (FSE 2024)](https://arxiv.org/abs/2403.17218v1) - [Static Application Security Testing (SAST) Tools for Smart Contracts: How Far Are We? (FSE 2024)](https://arxiv.org/abs/2404.18186) - [PROZE: Generating Parameterized Unit Tests Informed by Runtime Data (ICSE 2025)](https://arxiv.org/abs/2407.00768) - [LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning (ISSTA 2025)](https://arxiv.org/abs/2401.16185) - [How Secure is Code Generated by ChatGPT? (SMC 2023)](https://arxiv.org/abs/2304.09655) - [When ChatGPT Meets Smart Contract Vulnerability Detection: How Far Are We? (TOSEM 2024)](https://arxiv.org/abs/2309.05520) - [Exploring ChatGPT's Capabilities on Vulnerability Management (USENIX Security 2024)](https://arxiv.org/abs/2311.06530) - [An Empirical Study of Automated Vulnerability Localization with Large Language Models (arXiv 2024)](https://arxiv.org/abs/2404.00287v1) - [LLM4Fuzz: Guided Fuzzing of Smart Contracts with Large Language Models (arXiv 2024)](https://arxiv.org/abs/2401.11108) - [SeCodePLT: A Unified Platform for Evaluating the Security of Code GenAI (arXiv 2024)](https://arxiv.org/abs/2410.11096) - [Understanding the Effectiveness of Large Language Models in Detecting Security Vulnerabilities (arXiv 2024)](https://arxiv.org/abs/2311.16169) - [How Far Have We Gone in Vulnerability Detection Using Large Language Models (VulBench) (arXiv 2023)](https://arxiv.org/abs/2311.12420) ### D.2 LLM + Static / Symbolic Hybrids — **SAVD home** _Tight integration of LLM semantic reasoning with static analysis, code property graphs, or symbolic execution. The direct neighborhood of SAVD._ - [GRACE: Empowering LLM-Based Software Vulnerability Detection with Graph Structure and In-Context Learning (JSS 2024)](https://doi.org/10.1016/j.jss.2024.112061) - [Enhancing Static Analysis for Practical Bug Detection: An LLM-Integrated Approach (LLift) (OOPSLA 2024)](https://arxiv.org/abs/2404.17886) - [LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graphs and Large Language Models (USENIX Security 2025)](https://arxiv.org/abs/2507.16585) - [IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities (arXiv 2024)](https://arxiv.org/abs/2405.17238) - [Vul-LMGNN: Combining Language Models and Graph Neural Networks for Vulnerability Detection (arXiv 2024)](https://arxiv.org/abs/2404.14719) - [DefectHunter: A Novel LLM-Driven Boosted-Conformer-Based Code Vulnerability Detection Mechanism (arXiv 2023)](https://arxiv.org/abs/2309.15324) ### D.3 LLM + Formal Methods / Hoare-style — **FM-Agent home** _Hoare-triple reasoning, weakest precondition, theorem proving, Dafny / Coq / Lean / Verus + LLM hybrids. The direct neighborhood of FM-Agent._ - [Selene: Pioneering Automated Proof in Software Verification with LLMs (ACL 2024)](https://arxiv.org/abs/2401.07663) - [Finding Inductive Loop Invariants using Large Language Models (ACM SIGAda Ada Letters 2024)](https://arxiv.org/abs/2311.07948) - [Lemur: Integrating Large Language Models in Automated Program Verification (ICLR 2024)](https://arxiv.org/abs/2310.04870) - [Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming (ICSE 2025)](https://arxiv.org/abs/2405.01787) - [LeanDojo: Theorem Proving with Retrieval-Augmented Language Models (NeurIPS 2023)](https://arxiv.org/abs/2306.15626) - [Lean Copilot: Large Language Models as Copilots for Theorem Proving in Lean (NeurIPS Workshop 2024)](https://arxiv.org/abs/2404.12534) - [FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning (arXiv 2026)](https://arxiv.org/abs/2604.11556) - [HoarePrompt: Structural Reasoning About Program Correctness in Natural Language (arXiv 2025)](https://arxiv.org/abs/2503.19599) - [Automated Proof Generation for Rust Code via Self-Evolution (arXiv 2024)](https://arxiv.org/abs/2410.15756) - [Cobblestone: A Divide-and-Conquer Approach for Automating Formal Verification (arXiv 2024)](https://arxiv.org/abs/2410.19940) - [Laurel: Unblocking Automated Verification with Large Language Models (arXiv 2024)](https://arxiv.org/abs/2405.16792) - [Towards AI-Assisted Synthesis of Verified Dafny Methods (arXiv 2024)](https://arxiv.org/abs/2402.00247) - [Clover: Closed-Loop Verifiable Code Generation (arXiv 2023)](https://arxiv.org/abs/2310.17807) ### D.4 LLM Agents & Multi-Agent Workflows for Security _SWE-agent, AutoCodeRover, RepairAgent, MASAI, AutoSafeCoder, VulnBot, autonomous hacking / penetration agents._ - [RepairAgent: An Autonomous LLM-Based Agent for Program Repair (ICSE 2025)](https://arxiv.org/abs/2403.17134) - [AutoCodeRover: Autonomous Program Improvement (ISSTA 2024)](https://arxiv.org/abs/2404.05427) - [MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution (NeurIPS 2024)](https://arxiv.org/abs/2403.17927) - [SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering (NeurIPS 2024)](https://arxiv.org/abs/2405.15793) - [VulnBot: Autonomous Penetration Testing for A Multi-Agent Collaborative Framework (arXiv 2025)](https://arxiv.org/abs/2501.13411) - [Agentless: Demystifying LLM-Based Software Engineering Agents (arXiv 2024)](https://arxiv.org/abs/2407.01489) - [AutoSafeCoder: A Multi-Agent Framework for Securing LLM Code Generation (arXiv 2024)](https://arxiv.org/abs/2409.10737) - [CodeAgent: Autonomous Communicative Agents for Code Review (arXiv 2024)](https://arxiv.org/abs/2402.02172) - [Hyperagent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale (arXiv 2024)](https://arxiv.org/abs/2409.16299) - [Iterative Experience Refinement of Software-Developing Agents (arXiv 2024)](https://arxiv.org/abs/2405.04219) - [LLM Agents Can Autonomously Exploit One-Day Vulnerabilities (arXiv 2024)](https://arxiv.org/abs/2404.08144) - [LLM Agents Can Autonomously Hack Websites (arXiv 2024)](https://arxiv.org/abs/2402.06664) - [MASAI: Modular Architecture for Software-engineering AI Agents (arXiv 2024)](https://arxiv.org/abs/2406.11638) - [MarsCode Agent: AI-native Automated Bug Fixing (arXiv 2024)](https://arxiv.org/abs/2409.00899) - [Multi-role Consensus through LLMs Discussions for Vulnerability Detection (arXiv 2024)](https://arxiv.org/abs/2403.14274) - [Teams of LLM Agents Can Exploit Zero-Day Vulnerabilities (arXiv 2024)](https://arxiv.org/abs/2406.01637) ### D.5 Fine-Tuning / Instruction Tuning of Code LLMs _Supervised / instruction / contrastive fine-tuning of Code Llama, StarCoder, DeepSeek-Coder, CodeT5+ for security._ - [Generalization-Enhanced Code Vulnerability Detection via Multi-Task Instruction Fine-Tuning (ACL Findings 2024)](https://arxiv.org/abs/2406.03718) - [An Empirical Study on Fine-Tuning Large Language Models of Code for Automated Program Repair (ASE 2023)](https://arxiv.org/abs/2308.11518) - [CodeT5+: Open Code Large Language Models for Code Understanding and Generation (EMNLP 2023)](https://arxiv.org/abs/2305.07922) - [CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation (EMNLP 2021)](https://arxiv.org/abs/2109.00859) - [Are Decoder-Only Large Language Models the Silver Bullet for Code Search? (FSE 2024)](https://arxiv.org/abs/2410.22240) - [ContraBERT: Enhancing Code Pre-trained Models via Contrastive Learning (ICSE 2023)](https://arxiv.org/abs/2301.09072) - [MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning (KDD 2024)](https://arxiv.org/abs/2311.02303) - [PLBART: Unified Pre-Training for Program Understanding and Generation (NAACL 2021)](https://arxiv.org/abs/2103.06333) - [On the Reliability and Explainability of Language Models for Program Generation (TOSEM 2024)](https://arxiv.org/abs/2302.09587) - [Vulnerability Detection by Learning from Syntax-Based Execution Paths of Code (EPVD) (TSE 2023)](https://doi.org/10.1109/TSE.2023.3286586) - [DeepSeek-Coder: When the Large Language Model Meets Programming (arXiv 2024)](https://arxiv.org/abs/2401.14196) - [Finetuning Large Language Models for Vulnerability Detection (arXiv 2024)](https://arxiv.org/abs/2401.17010) - [RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair (arXiv 2024)](https://arxiv.org/abs/2312.15698) - [StarCoder 2 and The Stack v2: The Next Generation (arXiv 2024)](https://arxiv.org/abs/2402.19173) - [Code Llama: Open Foundation Models for Code (arXiv 2023)](https://arxiv.org/abs/2308.12950) ### D.6 Retrieval-Augmented / Knowledge-Augmented Detection _RAG-style approaches and retrieval-based prompt selection for vulnerability tasks._ - [Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (ICLR 2024)](https://arxiv.org/abs/2310.11511) - [Retrieval-Based Prompt Selection for Code-Related Few-Shot Learning (ICSE 2023)](https://arxiv.org/abs/2208.11626) ### D.7 CoT / Reasoning-Chain Variants for Vulnerability _Chain-of-thought, step-by-step, reasoning-augmented detectors tailored to vulnerability tasks._ - [Vul-RAG: Enhancing LLM-Based Vulnerability Detection via Knowledge-Level RAG (arXiv 2024)](https://arxiv.org/abs/2406.11147) ## E. Adjacent Techniques _Techniques that share infrastructure with vulnerability detection: fuzzing, automated program repair, and general code-LLM modeling._ ### E.1 Fuzzing — Classical & LLM-Augmented _Greybox / hybrid fuzzing and LLM-generated seeds, drivers, mutations: AFL, AFLFast, Driller, TitanFuzz, Fuzz4All, ChatAFL, WhiteFox, FuzzCoder._ - [PromptFuzz: Harnessing the Power of LLMs to Generate Effective and Robust Fuzz Drivers (CCS 2024)](https://arxiv.org/abs/2312.17677) - [Coverage-Based Greybox Fuzzing as Markov Chain (AFLFast) (CCS 2016)](https://doi.org/10.1145/2976749.2978428) - [When Fuzzing Meets LLMs: Challenges and Opportunities (FSE Companion 2024)](https://arxiv.org/abs/2404.16297) - [Fuzz4All: Universal Fuzzing with Large Language Models (ICSE 2024)](https://arxiv.org/abs/2308.04748) - [Fuzzing Symbolic Expressions (ICSE 2021)](https://arxiv.org/abs/2102.06580) - [Large Language Models are Zero-Shot Fuzzers: Fuzzing Deep-Learning Libraries via Large Language Models (TitanFuzz) (ISSTA 2023)](https://arxiv.org/abs/2212.14834) - [Large Language Model Guided Protocol Fuzzing (ChatAFL) (NDSS 2024)](https://doi.org/10.14722/ndss.2024.24556) - [Driller: Augmenting Fuzzing Through Selective Symbolic Execution (NDSS 2016)](https://www.ndss-symposium.org/wp-content/uploads/2017/09/driller-augmenting-fuzzing-through-selective-symbolic-execution.pdf) - [Large Language Models are Zero-Shot Reasoners (NeurIPS 2022)](https://arxiv.org/abs/2205.11916) - [WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language Models (OOPSLA 2024)](https://arxiv.org/abs/2310.15991) - [Magma: A Ground-Truth Fuzzing Benchmark (SIGMETRICS 2020)](https://arxiv.org/abs/2009.01120) - [QSYM: A Practical Concolic Execution Engine Tailored for Hybrid Fuzzing (USENIX Security 2018)](https://www.usenix.org/system/files/conference/usenixsecurity18/sec18-yun.pdf) - [FuzzCoder: Byte-Level Fuzzing Test via Large Language Model (arXiv 2024)](https://arxiv.org/abs/2409.01944) ### E.2 Automated Program Repair _LLM-based and search-based repair: GenProg, CURE, InferFix, ChatRepair, RepairLLaMA, PyTy, RepairAgent._ - [InferFix: End-to-End Program Repair with LLMs (FSE 2023)](https://arxiv.org/abs/2303.07263) - [Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-Shot Learning (FSE 2022)](https://arxiv.org/abs/2207.08281) - [ITER: Iterative Neural Repair for Multi-Location Patches (ICSE 2024)](https://arxiv.org/abs/2304.12015) - [Out of Sight, Out of Mind: Better Automatic Vulnerability Repair by Broadening Input Ranges and Sources (ICSE 2024)](https://arxiv.org/abs/2401.15459) - [PyTy: Repairing Static Type Errors in Python (ICSE 2024)](https://arxiv.org/abs/2401.06619) - [CURE: Code-Aware Neural Machine Translation for Automatic Program Repair (ICSE 2021)](https://arxiv.org/abs/2103.00073) - [Keep the Conversation Going: Fixing 162 out of 337 Bugs for $0.42 Each Using ChatGPT (ISSTA 2024)](https://arxiv.org/abs/2304.00385) - [How Effective Are Neural Networks for Fixing Security Vulnerabilities (ISSTA 2023)](https://arxiv.org/abs/2305.18607) - [LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?) (S&P 2024)](https://arxiv.org/abs/2312.12575) - [Examining Zero-Shot Vulnerability Repair with Large Language Models (S&P 2023)](https://arxiv.org/abs/2112.02125) - [GenProg: A Generic Method for Automatic Software Repair (TSE 2012)](https://doi.org/10.1109/TSE.2011.104) ### E.3 Code Generation / Code LLMs _Code-LLM modeling, code search, code summarization — the engineering layer under D.5._ - [Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation (AAAI 2024)](https://arxiv.org/abs/2308.10335) - [GraphCodeBERT: Pre-training Code Representations with Data Flow (ICLR 2021)](https://arxiv.org/abs/2009.08366) - [Code Search Is All You Need? Improving Code Suggestions with Code Search (ICSE 2024)](https://arxiv.org/abs/2403.10059) - [Large Language Models for Code Analysis: Do LLMs Really Do Their Job? (USENIX Security 2024)](https://arxiv.org/abs/2310.12357) ## F. Evaluation _Benchmarks, datasets, and empirical critiques — how the field measures progress and where the gaps are._ ### F.1 Benchmarks & Datasets _Vulnerability and security benchmarks: SV-COMP, FormAI, BigVul, D2A, DiverseVul, CrossVul, CVEfixes, PrimeVul, Magma, Cybench, SARD/Juliet._ - [CrossVul: A Cross-Language Vulnerability Dataset with Commit Data (FSE 2021)](https://doi.org/10.1145/3468264.3473122) - [Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models (ICLR 2025)](https://arxiv.org/abs/2408.08926) - [D2A: A Dataset Built for AI-Based Vulnerability Detection Methods Using Differential Analysis (ICSE SEIP 2021)](https://arxiv.org/abs/2102.07995) - [A C/C++ Code Vulnerability Dataset with Code Changes and CVE Summaries (BigVul) (MSR 2020)](https://doi.org/10.1145/3379597.3387501) - [The FormAI Dataset: Generative AI in Software Security Through the Lens of Formal Verification (PROMISE 2023)](https://arxiv.org/abs/2307.02192) - [CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software (PROMISE 2021)](https://arxiv.org/abs/2107.08760) - [DiverseVul: A New Vulnerable Source Code Dataset for Deep Learning Based Vulnerability Detection (RAID 2023)](https://arxiv.org/abs/2304.00409) - [VulnBench / SARD Catalog (NIST Software Assurance Reference Dataset) (Tool 2017)](https://nvlpubs.nist.gov/nistpubs/ir/2025/NIST.IR.8561.pdf) - [CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge (arXiv 2024)](https://arxiv.org/abs/2402.07688) - [A Ground-Truth Dataset of Real Security Patches (arXiv 2021)](https://arxiv.org/abs/2110.09635) ### F.2 Empirical Studies & Critiques _Reality checks on DL / LLM detectors, data quality, severity inconsistency, hallucination, false-positive analyses._ - [Shallow or Deep? An Empirical Study on Detecting Vulnerabilities Using Deep Learning (ICPC 2021)](https://arxiv.org/abs/2103.11933) - [An Empirical Study of Deep Learning Models for Vulnerability Detection (ICSE 2023)](https://arxiv.org/abs/2212.08109) - [Data Quality for Software Vulnerability Datasets (ICSE 2023)](https://arxiv.org/abs/2301.05456) - [Machine Learning for Source Code Vulnerability Detection: What Works and What Isn't There Yet (IEEE Security & Privacy 2022)](https://doi.org/10.1109/MSEC.2022.3176058) - [An Investigation into Inconsistency of Software Vulnerability Severity Across Data Sources (SANER 2022)](https://arxiv.org/abs/2112.10356) - [Limits of Machine Learning for Automatic Vulnerability Detection (USENIX Security 2024)](https://arxiv.org/abs/2306.17193) - [Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview (arXiv 2025)](https://arxiv.org/abs/2503.10784) - [How Far Have We Gone in Stripped Binary Code Understanding Using Large Language Models (arXiv 2024)](https://arxiv.org/abs/2404.09836) - [VulnLLMEval: A Framework for Evaluating Large Language Models in Software Vulnerability Detection and Patching (arXiv 2024)](https://arxiv.org/abs/2409.10756) --- ## Candidate Research Directions Seeded by SAVD & FM-Agent Hooks for follow-up work. Each direction maps to gaps surfaced by the papers above. 1. **Verifier-feedback as fine-tuning signal.** SAVD's verifier-in-the-loop produces counterexamples and constraint repairs. Treat these as a ConstraintRL / SFT corpus so the model internalizes reachability discipline (link: §C.4 process supervision, §D.5). 2. **Hoare-style natural-language reasoning meets assertion-guided slicing.** FM-Agent shows natural-language Hoare scales to 277k LoC but is unsound; SAVD's WP-based verifier-in-the-loop adds soundness signals. Combining the two — NL Hoare on top, WP + SMT verifier underneath — directly attacks both papers' admitted weaknesses (§D.2 + §D.3). 3. **Multi-agent CEGIS for assertion triggerability.** Recast triggerability as a prover-vs-counterexample agent game (§C.6 debate, §D.4 agents). 4. **Inter-procedural slicing with LLM-induced function summaries.** SAVD is intra-function; FM-Agent uses caller-driven top-down specs. Hybridize the two for whole-program slicing (§A.5 + §D.3). 5. **o1-style process supervision for path-constraint generation.** Use a per-step process reward model on Reasoning-Chain Rethinking (§C.4 + §D.7). 6. **Triggerability benchmarks beyond classification.** Most §F.1 benchmarks are vulnerable/not; build benchmarks with ground-truth exploit inputs (extend FormAI, PrimeVul, Magma). 7. **Real-CVE replay datasets** — actual exploit, observed assertion, observed patch — to expose the gap between synthetic safety properties and live exploit chains. 8. **Cross-LLM ensembling for path-constraint generation.** GPT-4o / Qwen-2.5 / DeepSeek-V3 disagree in informative ways (SAVD Table II); a quorum-checked WP-generation step could trade modest cost for monotone precision (§C.6 verifier ensembles). 9. **Bug-validator-as-reward.** FM-Agent uses differential testing against a reference implementation. Reuse that bug-validator as a reward model for fine-tuning the reasoning LLM. 10. **Cost-accuracy frontier study.** Sweep (LLM, slicer aggressiveness, verifier budget) across SAVD and FM-Agent benchmarks. Both papers explicitly defer this to future work. --- ## Category × Venue Matrix Counts of papers per (sub-category, venue-tier). Tiers grouped roughly by CCF SE/security catalog. Preprints = arXiv; *Other* includes workshops, journals, and tools without a clean CCF tier mapping. ### Tier summary | Sub-category | CCF-A | CCF-B | Preprint | Other | Total | |---|---:|---:|---:|---:|---:| | A.1 Surveys — LLM × SE / Vulnerability Detection | 2 | 2 | 7 | | **11** | | A.2 Surveys — Symbolic Execution & Verification | 1 | 2 | | | **3** | | A.3 Surveys — LLM Reasoning | | | 5 | | **5** | | A.4 Classical Symbolic & Concolic Execution | 5 | 2 | | 1 | **8** | | A.5 Classical Static Analysis & Program Slicing | 4 | 1 | | 2 | **7** | | A.6 Classical Formal Verification & Model Checking | 2 | 4 | | 2 | **8** | | A.7 Code Property Graphs & Code Representations | 5 | | | | **5** | | B.1 Token / Transformer Models for Code | 1 | 2 | | 5 | **8** | | B.2 GNN-Based Vulnerability Detection | 11 | 1 | | 1 | **13** | | B.3 Hybrid Multi-View / Causal Detection | 5 | | | | **5** | | C.1 Chain-of-Thought & In-Context Reasoning | 8 | 1 | | 1 | **10** | | C.2 Tree / Graph / Forest of Thoughts & Search | 4 | | 3 | | **7** | | C.3 Self-Consistency, Self-Refine, Self-Critique | 4 | | 1 | | **5** | | C.4 Process Supervision & RL-for-Reasoning (o1-style) | 2 | | 8 | | **10** | | C.5 Tool / Code-Aided / ReAct Reasoning | 4 | 1 | | | **5** | | C.6 LLM-as-Verifier, Critic & Debate | | | 3 | | **3** | | D.1 LLM Prompting & Zero/Few-shot Detection | 6 | | 5 | 1 | **12** | | D.2 LLM + Static / Symbolic Hybrids — **SAVD home** | 2 | 1 | 3 | | **6** | | D.3 LLM + Formal Methods / Hoare-style — **FM-Agent home** | 4 | | 7 | 2 | **13** | | D.4 LLM Agents & Multi-Agent Workflows for Security | 4 | | 12 | | **16** | | D.5 Fine-Tuning / Instruction Tuning of Code LLMs | 10 | | 5 | | **15** | | D.6 Retrieval-Augmented / Knowledge-Augmented Detection | 2 | | | | **2** | | D.7 CoT / Reasoning-Chain Variants for Vulnerability | | | 1 | | **1** | | E.1 Fuzzing — Classical & LLM-Augmented | 10 | | 1 | 2 | **13** | | E.2 Automated Program Repair | 11 | | | | **11** | | E.3 Code Generation / Code LLMs | 4 | | | | **4** | | F.1 Benchmarks & Datasets | 2 | 2 | 2 | 4 | **10** | | F.2 Empirical Studies & Critiques | 3 | 1 | 3 | 2 | **9** | | **Total** | **116** | **20** | **66** | **23** | **225** | ### Year distribution (per top-level cat) | Top-level | ≤2017 | 2018–2019 | 2020 | 2021 | 2022 | 2023 | 2024 | 2025 | 2026 | Total | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | A. Surveys & Foundations | 21 | 7 | 1 | 2 | 2 | 2 | 7 | 5 | | **47** | | B. Pre-LLM Learning-Based Vulnerability Detection | 1 | 4 | 1 | 3 | 9 | 2 | 5 | 1 | | **26** | | C. LLM Reasoning (General) | | | | | 2 | 13 | 18 | 7 | | **40** | | D. LLMs for Vulnerability Detection & Verification | | | | 2 | | 11 | 44 | 7 | 1 | **65** | | E. Adjacent Techniques | 3 | 1 | 1 | 3 | 2 | 4 | 14 | | | **28** | | F. Evaluation | 1 | | 1 | 5 | 2 | 4 | 4 | 2 | | **19** | | **Total** | **26** | **12** | **4** | **15** | **17** | **36** | **92** | **22** | **1** | **225** | ### Full venue breakdown Venues with 3+ papers shown as columns. Rare venues are aggregated into *Other*. | Sub-category | arXiv | ICSE | ICLR | NeurIPS | FSE | TSE | USENIX Security | ISSTA | TOSEM | EMNLP | S&P | CCS | MSR | NDSS | POPL | Other | Total | |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:| | A.1 Surveys — LLM × SE / Vulnerability Detection | 7 | | | | | 1 | | | 1 | | | | | | | 2 | **11** | | A.2 Surveys — Symbolic Execution & Verification | | | | | | 1 | | | | | | | | | | 2 | **3** | | A.3 Surveys — LLM Reasoning | 5 | | | | | | | | | | | | | | | | **5** | | A.4 Classical Symbolic & Concolic Execution | | 1 | | | | | 1 | | | | 1 | 1 | | | | 4 | **8** | | A.5 Classical Static Analysis & Program Slicing | | | | | 1 | 1 | | | | | | | | | 1 | 4 | **7** | | A.6 Classical Formal Verification & Model Checking | | | | | | | | | 1 | | | | | | 1 | 6 | **8** | | A.7 Code Property Graphs & Code Representations | | 1 | 2 | | | | | | | | 1 | | | | 1 | | **5** | | B.1 Token / Transformer Models for Code | | | | | | | | | | | | | 1 | | | 7 | **8** | | B.2 GNN-Based Vulnerability Detection | | 4 | | 1 | 1 | 1 | | | 1 | | | | 1 | 1 | | 3 | **13** | | B.3 Hybrid Multi-View / Causal Detection | | 4 | | | | | | | | | | | | | | 1 | **5** | | C.1 Chain-of-Thought & In-Context Reasoning | | 1 | 3 | 2 | | 1 | | | | | | | | | | 3 | **10** | | C.2 Tree / Graph / Forest of Thoughts & Search | 3 | | | 2 | | | | | | 1 | | | | | | 1 | **7** | | C.3 Self-Consistency, Self-Refine, Self-Critique | 1 | | 2 | 2 | | | | | | | | | | | | | **5** | | C.4 Process Supervision & RL-for-Reasoning (o1-style) | 8 | | 1 | | | | | | | | | | | | | 1 | **10** | | C.5 Tool / Code-Aided / ReAct Reasoning | | | 1 | 1 | | | | | | 1 | | | | | | 2 | **5** | | C.6 LLM-as-Verifier, Critic & Debate | 3 | | | | | | | | | | | | | | | | **3** | | D.1 LLM Prompting & Zero/Few-shot Detection | 5 | 1 | | | 2 | | 1 | 1 | 1 | | | | | | | 1 | **12** | | D.2 LLM + Static / Symbolic Hybrids — **SAVD home** | 3 | | | | | | 1 | | | | | | | | | 2 | **6** | | D.3 LLM + Formal Methods / Hoare-style — **FM-Agent home** | 7 | 1 | 1 | 1 | | | | | | | | | | | | 3 | **13** | | D.4 LLM Agents & Multi-Agent Workflows for Security | 12 | 1 | | 2 | | | | 1 | | | | | | | | | **16** | | D.5 Fine-Tuning / Instruction Tuning of Code LLMs | 5 | 1 | | | 1 | 1 | | | 1 | 2 | | | | | | 4 | **15** | | D.6 Retrieval-Augmented / Knowledge-Augmented Detection | | 1 | 1 | | | | | | | | | | | | | | **2** | | D.7 CoT / Reasoning-Chain Variants for Vulnerability | 1 | | | | | | | | | | | | | | | | **1** | | E.1 Fuzzing — Classical & LLM-Augmented | 1 | 2 | | 1 | | | 1 | 1 | | | | 2 | | 2 | | 3 | **13** | | E.2 Automated Program Repair | | 4 | | | 2 | 1 | | 2 | | | 2 | | | | | | **11** | | E.3 Code Generation / Code LLMs | | 1 | 1 | | | | 1 | | | | | | | | | 1 | **4** | | F.1 Benchmarks & Datasets | 2 | | 1 | | 1 | | | | | | | | 1 | | | 5 | **10** | | F.2 Empirical Studies & Critiques | 3 | 2 | | | | | 1 | | | | | | | | | 3 | **9** | | **Total** | **66** | **25** | **13** | **12** | **8** | **7** | **6** | **5** | **5** | **4** | **4** | **3** | **3** | **3** | **3** | **58** | **225** | _*Other* venues (≤2 papers each): AAAI, ACL, ACL Findings, ASE, CAV, OOPSLA, PLDI, PROMISE, TACAS, TDSC, Tool, ACM Comput. Surv., ACM SIGAda Ada Letters, ACSAC, ASPLOS, Advances in Computers, CACM, CC, COLM, ECOOP, EMNLP Findings, FSE Companion, High-Confidence Computing, ICML, ICMLA, ICPC, ICSE Companion, ICSE NIER, ICSE SEIP, IEEE Security & Privacy, IJCNLP-AACL, IJCNN, Int. J. Inf. Sec., JSS, KDD, LASER, MetriSec, NAACL, NeurIPS Workshop, OSDI, RAID, SANER, SIGMETRICS, SMC, STTT, TMLR, UAI_ --- ## Source - Git mirror: - Zotero collection: *SoftwareVulnerabilityDetection-2026-06* (key XMKNMNCX, 28 sub-collections) - Only `README.md` is tracked in git. PDFs, plan scripts, and notes live locally and are regenerable from `plan/candidates_v2.json` + `plan/categorized.json`.