Skip to content
Tiago Almeida
All work

2026 · Research tooling

EXPLOITFORGE

An open, configuration-driven methodology and pipeline that turns published vulnerabilities into executable instances: it links NVD entries to their fix commits, vulnerable and patched code, and candidate triggers from ExploitDB and the projects' own test suites, rebuilds the vulnerable software in era-matched containers, and runs a selected trigger against it.

Overview diagram: dataset construction modules M1 to M7, reproduction modules M8 to M12, and differential validation feeding automated repair
Overview of EXPLOITFORGE: dataset construction, reproduction, differential validation and downstream use for automated repair. SANER 2027 submission, Figure 1.
CVEs collected
1,144CVEs collected
public exploit artefacts
240public exploit artefacts
reproduced through developer tests
177reproduced through developer tests
confirmed by differential validation
148confirmed by differential validation

The problem

Automated vulnerability repair needs executable evidence that a vulnerability is present and that a candidate patch suppresses it, such as a public exploit or a vulnerability-specific developer test. Building that evidence at scale is hard: candidate artefacts have to be found, linked to the right vulnerable and fixed revisions, and shown to execute reliably.

The approach

Twelve modules in two stages. Dataset construction (modules 1 to 7) queries the NVD, recovers fix commits and changed code, maps candidate PoCs from ExploitDB and developer tests from the fixing commits, and curates the public exploits through syntax validation and LLM-assisted repair under fidelity guards. Reproduction (modules 8 to 12) ranks the triggers, builds one base image per era and a per-CVE image, runs the trigger, and passes the result through correctness gates: it must exercise the project build, behave sanely and agree with itself in two of three runs. The resulting reproduction claims are then checked by separately implemented differential validation, anchored on the fix and, for public exploits, on the exploit. Projects are described in YAML with language adapters; there is no per-CVE configuration.

Results

  1. 01Instantiated for glibc, OpenSSL and tcpdump: 1,144 CVEs collected, 532 with a recoverable fix commit, and only 150 with at least one public PoC match, giving 240 exploit artefacts.
  2. 02Just 41.7% of those exploits parse exactly as retrieved.
  3. 03246 CVEs pass the gates for execution and 177 are reproduced, all of them through developer tests.
  4. 04Differential validation confirms 148 reproductions and contradicts 6; the exploit-anchored differential confirms the public PoCs of 4 of 15 evaluable CVEs.
  5. 05Used as repair tasks, 124 of 170 reproduced instances get at least one validated repair across six LLM configurations.
Detailed diagram of the seven dataset-construction modules, from NVD fetching to output generation, with the LLM-assisted PoC repair loop
Dataset construction in detail: the seven aggregation modules, the LLM-assisted PoC repair loop with its fidelity guards, and the in-tree test path that skips repair. Thesis, Figure 4.3.