Blog / Engineering
Engineering

Building a 300k-record vulnerability database that stays current.

A vulnerability database is only useful if it is fresh, accurate, and matches the way real code is shipped. Here is how we approached ingestion, deduplication, and version matching at the scale of public threat feeds, and what we learned about keeping false positives below 4%.

300k+
Vulnerability records indexed
Across public threat feeds
<4%
False-positive rate on matches
Measured against curated test corpus
~minutes
From public disclosure to dashboard
Median across last 90 days

Most vulnerability tools share a common failure mode. They produce a long list of findings, most of which do not actually apply to your code, and a handful of which are critical and buried in the noise. The team gives up on the list within a week, and a year later something on it gets exploited.

The root cause is almost always the same. The vulnerability database is either stale, or matched too loosely, or matched too narrowly. Each of those produces a different failure mode, and getting all three right at the same time is what makes a usable scanner. This is a write-up of the principles we used to get there.

What "vulnerability database" actually means

A vulnerability database is not one feed. It is the merge of several, each with different coverage, different latency, and different accuracy. The version that ships in CyberDebunk pulls from a mix of public sources, internal curation, and customer-reported false positives. The numbers below describe the public-source side.

NVD
National Vulnerability Database. The canonical source, but lags real disclosure by hours to days.
~28k entries per year
GitHub Advisory Database
Ecosystem-aware, includes packages NVD has not picked up. Good for npm and PyPI.
~22k entries per year
OSV.dev
Open-source schema with consistent version-range encoding. Easiest to deduplicate against.
~35k entries per year
CISA KEV
Known Exploited Vulnerabilities. Small list, very high signal. Prioritisation source, not bulk.
~1k entries total
Vendor advisories
Direct from Red Hat, Microsoft, Apple, Cisco. Earliest signal, least structured.
Tens of thousands of bulletins
Curated overrides
Internal corrections for ecosystem-specific edge cases, false positives, and renamed packages.
Growing weekly

The merged record set is what gets to 300k. Roughly two-thirds of those records are duplicates across sources, with the same advisory described slightly differently. Deduplication is a project in itself.

The freshness problem

For a vulnerability database, "fresh" has a very specific meaning. A new advisory is published somewhere, and the question is how long before a scan against your code reflects it. Anything over a few hours is unacceptable in modern dependency security, because that is the window in which proof-of-concept exploit code typically appears.

The simplest answer is to poll. Every source has a feed of recently-changed entries. We pull those on a tight cadence, deduplicate against what we already have, and index the changes. The total roundtrip from a feed update to a record being matchable against customer code is usually a small number of minutes.

What it costs to be wrong about freshness

Eighty-eight hours.

That was the median time between public disclosure of CVE-2024-23334 (aiohttp path traversal) and the first internet-wide exploit attempts we observed. A daily scan would have missed three full cycles. A continuous one caught it in the first.

The matching problem

Knowing about a vulnerability is not the same as knowing it applies to your code. A vulnerability says something like:

"aiohttp versions >=3.9.0, <3.9.5 are vulnerable to path traversal when follow_symlinks=true."

Matching that against your code means answering three questions correctly:

Why most scanners get this wrong

Two failure modes dominate.

The first is over-matching. Treating any presence of the package name as a match, regardless of version. This produces a long list of "you might be affected", most of which are false. The user learns to ignore the list.

The second is under-matching. Requiring exact version equality, missing anything where the range is expressed loosely. This produces a clean list, but it has gaps. The user has no way to know about the gaps until something gets exploited.

Getting the middle ground right requires ecosystem-aware version range parsing. Semver for npm. PEP 440 for PyPI. Maven's range syntax. Cargo's range syntax. We treat each ecosystem's version model as a first-class citizen, parse the range from the advisory in that ecosystem's grammar, and compute true range overlap. Not "string contains", not "starts with".

The accuracy budget

Industry-wide, false-positive rates for dependency scanners hover between 15% and 40%, depending on who is measuring and which corpus they use. The number we hold ourselves to is below 4%. That target is what drives most of the design decisions above.

Concretely, we measure false positives by running the scanner against a curated corpus of repositories where the right answer is known. Every release shifts the corpus, and the number gets re-measured. Two specific tactics matter:

Deduplication

NVD might describe a vulnerability under one CVE ID. GHSA might describe the same thing under a different identifier. OSV might describe it under both. If we naively summed feeds, our 300k count would balloon to 800k, and customers would see the same finding three times.

The merge key is the underlying package coordinates plus the affected version range, not the identifier. Two records that describe the same package and the same range get folded into one. The identifiers stay attached as cross-references so an auditor can trace it.

What we are not solving

Public vulnerability data has known limitations, and we do not pretend to fix all of them.

These are open problems. They are open everywhere.

The takeaway for someone evaluating a scanner

If you are evaluating a vulnerability tool, the three questions worth asking are not the ones on the marketing page. They are:

The mechanics are not glamorous. But they are what separates a tool teams actually use from a tool that quietly gets muted.

Connect a repository and see the matching in action.

The first scan usually completes within a few minutes, and the findings come with the reasoning attached. You can see exactly which advisory matched, why, and where the version range overlap is.