Building a 300k-record vulnerability database that stays current.
A vulnerability database is only useful if it is fresh, accurate, and matches the way real code is shipped. Here is how we approached ingestion, deduplication, and version matching at the scale of public threat feeds, and what we learned about keeping false positives below 4%.
Most vulnerability tools share a common failure mode. They produce a long list of findings, most of which do not actually apply to your code, and a handful of which are critical and buried in the noise. The team gives up on the list within a week, and a year later something on it gets exploited.
The root cause is almost always the same. The vulnerability database is either stale, or matched too loosely, or matched too narrowly. Each of those produces a different failure mode, and getting all three right at the same time is what makes a usable scanner. This is a write-up of the principles we used to get there.
What "vulnerability database" actually means
A vulnerability database is not one feed. It is the merge of several, each with different coverage, different latency, and different accuracy. The version that ships in CyberDebunk pulls from a mix of public sources, internal curation, and customer-reported false positives. The numbers below describe the public-source side.
The merged record set is what gets to 300k. Roughly two-thirds of those records are duplicates across sources, with the same advisory described slightly differently. Deduplication is a project in itself.
The freshness problem
For a vulnerability database, "fresh" has a very specific meaning. A new advisory is published somewhere, and the question is how long before a scan against your code reflects it. Anything over a few hours is unacceptable in modern dependency security, because that is the window in which proof-of-concept exploit code typically appears.
The simplest answer is to poll. Every source has a feed of recently-changed entries. We pull those on a tight cadence, deduplicate against what we already have, and index the changes. The total roundtrip from a feed update to a record being matchable against customer code is usually a small number of minutes.
Eighty-eight hours.
That was the median time between public disclosure of CVE-2024-23334 (aiohttp path traversal) and the first internet-wide exploit attempts we observed. A daily scan would have missed three full cycles. A continuous one caught it in the first.
The matching problem
Knowing about a vulnerability is not the same as knowing it applies to your code. A vulnerability says something like:
"aiohttp versions>=3.9.0, <3.9.5are vulnerable to path traversal whenfollow_symlinks=true."
Matching that against your code means answering three questions correctly:
- Do you use this package at all? Easy. Read the manifest.
- What version do you actually have installed? Harder. Manifests express ranges. Lockfiles express resolutions. Transitive dependencies do not appear in either.
- Does the version range in the advisory overlap with the version range in your code? Hardest. Version range syntax varies by ecosystem.
^1.2.3means very different things in npm and Composer.
Why most scanners get this wrong
Two failure modes dominate.
The first is over-matching. Treating any presence of the package name as a match, regardless of version. This produces a long list of "you might be affected", most of which are false. The user learns to ignore the list.
The second is under-matching. Requiring exact version equality, missing anything where the range is expressed loosely. This produces a clean list, but it has gaps. The user has no way to know about the gaps until something gets exploited.
Getting the middle ground right requires ecosystem-aware version range parsing. Semver for npm. PEP 440 for PyPI. Maven's range syntax. Cargo's range syntax. We treat each ecosystem's version model as a first-class citizen, parse the range from the advisory in that ecosystem's grammar, and compute true range overlap. Not "string contains", not "starts with".
The accuracy budget
Industry-wide, false-positive rates for dependency scanners hover between 15% and 40%, depending on who is measuring and which corpus they use. The number we hold ourselves to is below 4%. That target is what drives most of the design decisions above.
Concretely, we measure false positives by running the scanner against a curated corpus of repositories where the right answer is known. Every release shifts the corpus, and the number gets re-measured. Two specific tactics matter:
- Resolved-version matching, not declared-version matching. A manifest that says
"react": "^18.0.0"tells you very little. The lockfile that resolves to18.3.1tells you exactly what is installed. We match on the resolved version where one exists. - Reachability-aware filtering. A vulnerability in a transitive dependency that is only used by a code path you do not invoke is technically present but not exploitable. We do not call it a false positive (it is real), but we surface it differently from the criticals on your hot path.
Deduplication
NVD might describe a vulnerability under one CVE ID. GHSA might describe the same thing under a different identifier. OSV might describe it under both. If we naively summed feeds, our 300k count would balloon to 800k, and customers would see the same finding three times.
The merge key is the underlying package coordinates plus the affected version range, not the identifier. Two records that describe the same package and the same range get folded into one. The identifiers stay attached as cross-references so an auditor can trace it.
What we are not solving
Public vulnerability data has known limitations, and we do not pretend to fix all of them.
- Unreported vulnerabilities. By definition, we cannot match against what nobody has disclosed. The database is bounded by what is public.
- Misattributed packages. When an advisory misnames the affected package, or names a parent package instead of the actual vulnerable child, our match goes wrong. We catch these via curated overrides.
- Configuration-dependent vulnerabilities. Some advisories only apply when a specific option is set. We surface the configuration requirement, but we do not currently inspect your runtime to confirm it.
These are open problems. They are open everywhere.
The takeaway for someone evaluating a scanner
If you are evaluating a vulnerability tool, the three questions worth asking are not the ones on the marketing page. They are:
- How fast does a newly disclosed advisory reach your dashboard? Measured in hours, not days.
- What is the false-positive rate against a corpus you control? Anything above 10% will be ignored by your team within a month.
- How are version ranges parsed? If the answer is anything other than "per ecosystem", over-match and under-match are baked in.
The mechanics are not glamorous. But they are what separates a tool teams actually use from a tool that quietly gets muted.
Connect a repository and see the matching in action.
The first scan usually completes within a few minutes, and the findings come with the reasoning attached. You can see exactly which advisory matched, why, and where the version range overlap is.