2025 to 2026 · MSc thesis
Automated Patch Generation for Software Vulnerabilities Using Large Language Models
My MSc dissertation (19/20), part of the AI-SSD project at CISUC: a five-phase methodology for generating and validating security patches with LLMs, implemented as an open-source, configuration-driven pipeline for C/C++ systems software and evaluated on glibc, OpenSSL and tcpdump across eight models.

- CVEs aggregated across glibc, OpenSSL and tcpdump
- 361CVEs aggregated across glibc, OpenSSL and tcpdump
- reproduced in era-matched containers
- 178reproduced in era-matched containers
- validated patches
- 924validated patches
- models compared, hosted and self-hosted
- 8models compared, hosted and self-hosted
The problem
LLMs readily produce code that reads like a fix, but syntactic correctness is not semantic correctness, and models struggle to repair security flaws without introducing new defects, especially in memory-unsafe C and C++. A repair has to be shown to close the vulnerability by execution, not by inspection.
The approach
No repair is attempted for a defect that has not first been demonstrated, and no patch counts until it defeats the artefact that demonstrated it. Phase 0 curates a dataset in which every vulnerability carries a runnable reproducer; Phase 1 replays it in an era-matched container and records a deterministic baseline; Phase 2 generates candidate patches from the vulnerable code and its context; Phase 3 accepts a candidate only if it compiles, silences the reproducer against that baseline and adds no new high or critical static-analysis finding; Phase 4 reports the evidence behind every verdict. Failed candidates go back to the model with the validator's diagnostics. Nothing is specific to a project or a vulnerability: a new target is described in YAML, not programmed.
Results
- 01Of 361 CVEs aggregated, 246 reached reproduction and 178 were confirmed; an independent differential oracle scored those verdicts at a lenient precision of 0.961 and a recall of 0.980.
- 02924 validated patches across eight models, six of them in three independent repetitions: 48.5% passed on first validation, 30.5% came from best-of-four resampling and 21.0% from validator-driven retries.
- 03Success was defect-dependent far more than model-dependent: 60.8% of tcpdump's validated CVEs, 37.6% of glibc's and 6.7% of OpenSSL's.
- 04The strongest hosted model, gpt-5-mini, repaired at about 3 cents per vulnerability; self-hosted 32B models reached just under half that yield at no marginal cost.
- 05No positive evidence of training-data memorisation: CVE-2026-6368, published in August 2026, was repaired in two of three repetitions, more than two years after the model's training cut-off.
One vulnerability, five phases
Phase 0
Dataset
The NVD record for CVE-2023-0286, its fixing commit, the vulnerable function and the test the fix ships.
Phase 1
Reproduction
An era-matched image in which that test fails on the vulnerable build; the failure is recorded as the baseline.
Phase 2
Generation
A prompt carrying the description, the weakness class and the verbatim function; the model returns one edit block.
Phase 3
Validation
The patched build compiles, the same test now passes and no new high-severity finding appears.
Phase 4
Reporting
The verdict and the evidence behind it, archived with the run that produced them.
Had validation failed for a recoverable reason, the diagnostics would have gone back to Phase 2.
CVE-2023-0286, a type-confusion defect in OpenSSL, followed through the pipeline. After Thesis, Figure 4.2.


