Skip to content
Tiago Almeida
All work

2025 to 2026 · MSc thesis

Automated Patch Generation for Software Vulnerabilities Using Large Language Models

My MSc dissertation (19/20), part of the AI-SSD project at CISUC: a five-phase methodology for generating and validating security patches with LLMs, implemented as an open-source, configuration-driven pipeline for C/C++ systems software and evaluated on glibc, OpenSSL and tcpdump across eight models.

Methodology diagram: Phase 0 CVE aggregation, Phase 1 reproduction, Phase 2 patch generation, Phase 3 validation with feedback loops, Phase 4 reporting
The methodology, phases 0 to 4. A patch reaches reporting only when the exploit is blocked and static analysis finds nothing new; otherwise the result goes back to the model. Thesis, Figure 4.1.
CVEs aggregated across glibc, OpenSSL and tcpdump
361CVEs aggregated across glibc, OpenSSL and tcpdump
reproduced in era-matched containers
178reproduced in era-matched containers
validated patches
924validated patches
models compared, hosted and self-hosted
8models compared, hosted and self-hosted

The problem

LLMs readily produce code that reads like a fix, but syntactic correctness is not semantic correctness, and models struggle to repair security flaws without introducing new defects, especially in memory-unsafe C and C++. A repair has to be shown to close the vulnerability by execution, not by inspection.

The approach

No repair is attempted for a defect that has not first been demonstrated, and no patch counts until it defeats the artefact that demonstrated it. Phase 0 curates a dataset in which every vulnerability carries a runnable reproducer; Phase 1 replays it in an era-matched container and records a deterministic baseline; Phase 2 generates candidate patches from the vulnerable code and its context; Phase 3 accepts a candidate only if it compiles, silences the reproducer against that baseline and adds no new high or critical static-analysis finding; Phase 4 reports the evidence behind every verdict. Failed candidates go back to the model with the validator's diagnostics. Nothing is specific to a project or a vulnerability: a new target is described in YAML, not programmed.

Results

  1. 01Of 361 CVEs aggregated, 246 reached reproduction and 178 were confirmed; an independent differential oracle scored those verdicts at a lenient precision of 0.961 and a recall of 0.980.
  2. 02924 validated patches across eight models, six of them in three independent repetitions: 48.5% passed on first validation, 30.5% came from best-of-four resampling and 21.0% from validator-driven retries.
  3. 03Success was defect-dependent far more than model-dependent: 60.8% of tcpdump's validated CVEs, 37.6% of glibc's and 6.7% of OpenSSL's.
  4. 04The strongest hosted model, gpt-5-mini, repaired at about 3 cents per vulnerability; self-hosted 32B models reached just under half that yield at no marginal cost.
  5. 05No positive evidence of training-data memorisation: CVE-2026-6368, published in August 2026, was repaired in two of three repetitions, more than two years after the model's training cut-off.

One vulnerability, five phases

  1. Phase 0

    Dataset

    The NVD record for CVE-2023-0286, its fixing commit, the vulnerable function and the test the fix ships.

  2. Phase 1

    Reproduction

    An era-matched image in which that test fails on the vulnerable build; the failure is recorded as the baseline.

  3. Phase 2

    Generation

    A prompt carrying the description, the weakness class and the verbatim function; the model returns one edit block.

  4. Phase 3

    Validation

    The patched build compiles, the same test now passes and no new high-severity finding appears.

  5. Phase 4

    Reporting

    The verdict and the evidence behind it, archived with the run that produced them.

Had validation failed for a recoverable reason, the diagnostics would have gone back to Phase 2.

CVE-2023-0286, a type-confusion defect in OpenSSL, followed through the pipeline. After Thesis, Figure 4.2.

Bar chart of the corpus at each stage, split by glibc, OpenSSL and tcpdump: 361 aggregated, 333 with a reproducer, 246 submitted, 178 reproduced, 171 to generation, 124 repaired
What became of the corpus at each stage, split by project. Thesis, Figure 5.10(a).
Grouped bar chart of the share of validated CVEs repaired per model and project, highest for gpt-5-mini on tcpdump at about 61%
End-to-end success rate per model and project, mean of three repetitions with the standard deviation as whiskers. Models marked * ran once. Thesis, Figure 5.5.
Scatter plot of successful patches per repetition against cost on a log scale, with self-hosted models in a shaded band at zero cost
Patch quality against cost, one point per model. The cost axis is logarithmic and cannot show zero, so it is broken at the dashed line: the shaded band to its left holds the self-hosted models, which have no per-token charge. Thesis, Figure 5.7.