Vexta's cross-scan finding deduplication answers whether a bug found today is the same one a previous scan already reported, and it answers that with a stable fingerprint rather than a rough string match. Vexta is VITI Security's agentless, AI-augmented vulnerability scanner and pentest platform, and every finding gets a FingerprintHash, a SHA-256 hash of the target, vulnerability type, URL, parameter, and payload signature, with random tokens normalized out first.
How does Vexta fingerprint a finding?
The hash is built from five inputs: the target, the vulnerability type, the URL, the parameter, and a payload signature. Before hashing, random tokens, the kind of thing a session ID or a cache-buster query string throws into a URL, get normalized out. Without that step, the same bug would hash differently every time just because the URL happened to carry a different random value, which would defeat the point of fingerprinting.
The choice of inputs matters. Target and URL narrow the fingerprint down to a specific spot in the application. Vulnerability type and parameter narrow it down to a specific bug at that spot, so a SQL injection and a reflected XSS on the same parameter don't collide into one fingerprint. The payload signature adds one more layer on top of that, tying the fingerprint to the specific way the bug was triggered rather than just its location.
Why does this matter across repeat scans?
Run the same target twice, a week apart, and most of what comes back the second time is the same bugs still sitting there, plus maybe a handful of genuinely new ones. Without a stable identity for each finding, a dashboard either shows every result from every scan as new (drowning the real new findings in noise) or asks a human to eyeball two finding lists and figure out which entries match. Vexta's fingerprint hash lets the dashboard do that matching automatically and mark a repeat as a re-discovery instead of reporting it a second time.
This shows up most on any target that gets scanned on a recurring basis: a retainer client, an internal application under continuous assessment, or a bug bounty program a hunter keeps coming back to. The value of a repeat scan is in what changed, not in re-confirming everything that didn't. Deduplication is what makes that comparison practical instead of a manual side-by-side between two exported finding lists.
What keeps the fingerprint stable over time?
The fingerprinting algorithm is pinned by a golden-corpus regression test, meaning there is a fixed set of known findings whose hashes are checked on every change to the codebase. That matters because stored history is only useful if it stays comparable. A quiet refactor to the hashing logic that changed every hash overnight would silently reset every deduplication baseline and make months of scan history stop lining up. Pinning the algorithm against a regression test is what keeps that from happening between releases.
This is a detail most scanning tools never have to explain, because most never keep a hashing scheme stable long enough for it to matter. A hunter who has watched a tool's own de-duplication quietly break after an update, suddenly reporting a year's worth of known findings as brand new, understands why a golden-corpus test on the hashing algorithm specifically is worth calling out rather than assuming.
Deduplication runs on every scan regardless of plan tier in the public feature list; check the pricing page if you want the full breakdown of what's included.
What this looks like in practice
On a re-scan, a finding that fingerprints identically to one already in the dashboard shows up flagged as a re-discovery rather than as a brand-new entry demanding fresh triage. That keeps a hunter's attention on what actually changed since the last pass: genuinely new findings, plus whatever fix verification (reverify) has confirmed as fixed or still vulnerable, rather than re-reading a list that is mostly the same bugs from last time.
For a hunter reporting to a client or a program on a recurring cadence, this also changes what the report itself looks like. A monthly summary can say plainly how many findings are new this period and how many are prior findings still open, because the underlying data already distinguishes the two. Without stable fingerprints, that distinction would have to be reconstructed by hand from two separate exports every time a report goes out.
Over a series of scans, this is what makes trend data honest in the first place. Metrics like defect density and recurrence rate depend on knowing which findings are the same bug reappearing and which are genuinely new, and that distinction only holds up if the fingerprinting underneath it is stable. A hunter reviewing a month of recurring scans against one target is really reviewing a timeline of the same fingerprinted bugs moving through open, fixed, and in some cases still-vulnerable states, not a pile of unrelated finding lists stitched together by hand.
Frequently asked questions
What is cross-scan finding deduplication in Vexta?
How does Vexta build a finding's fingerprint?
Why normalize random tokens before hashing?
Can the fingerprint hash change between Vexta releases?
What happens when a finding is a re-discovery?
Why does the fingerprint include a payload signature and not just the URL?
Run repeat scans without duplicate noise
See how Vexta's fingerprinting keeps re-discoveries separate from genuinely new findings on your next authorized scan.

