Skip to content
VITI Security

Reliability Badges and False-Positive Tracking in Vexta

by VITI Security TeamSep 26, 2026

Vexta labels every finding with a reliability badge and tracks each detector's false-positive rate over time, so a hunter knows how much to trust a result before spending time on it.

Reliability Badges and False-Positive Tracking in Vexta - VITI Security

Vexta shows how reliable a finding is by attaching a visual badge to every row in the findings list and by tracking, per detector, how often that detector's findings turn out to be real. Vexta is VITI Security's agentless, AI-augmented vulnerability scanner and pentest platform, and this pairing of per-finding badges with a per-detector track record is how it answers the question every hunter asks before writing a report: is this one worth my time?

What do the reliability badges on a finding mean?

Every finding row in Vexta can carry up to six small visual flags, and none of them require expanding the row to read. PROVEN, shown in green, means the proof-of-exploitation engine ran a safe, non-destructive check and confirmed the bug behaves the way the finding claims. UNVERIFIED, amber with a double border, means the verifier could not reach a decision either way; before that finding can go out the door, Vexta gates the "Submit to bounty platform" button behind an explicit "I have manually reproduced this" override, so an unverified result cannot slip into a submission by accident. REFUTED, shown grey and dimmed, means the verifier actively ruled the finding out, and the submit button refuses outright in that case.

The remaining three badges add context rather than a verdict. EXPERIMENTAL marks a finding whose detector class sits on a curated known-flaky list, currently things like cache deception, git-history exposure, and SAML XSW checks, as a reminder to verify even when that same finding also shows PROVEN. FROM AI flags a finding whose discovery path involved the AI triage loop rather than a deterministic detector, since AI-assisted discoveries are treated as non-reproducible until checked. A small version tag such as v1.1-fp-fix identifies the detector revision that produced the finding, so if a detector gets tuned in a later release, you can tell at a glance whether a given finding came from the old version or the fixed one.

How does Vexta measure a detector's false-positive rate?

Behind the badges, Vexta keeps a running outcome counter for every detector: how many of its findings ended up proven, how many refuted, how many stayed unproven, and how many hit a transient error. Those counters persist across daemon restarts, so the picture doesn't reset every time the scanner restarts. The Settings menu's PoE Metrics dashboard (PoE is proof-of-exploitation, the non-destructive verifier that tries to safely confirm each finding) turns that history into a sortable table, one row per verifier, with the raw counts and a computed false-positive rate: refuted findings divided by proven plus refuted findings.

Each detector in that table also carries a color-coded health badge. Healthy means a false-positive rate under 5%. Noisy covers the 5% to 50% range. DEGRADED applies once a detector's false-positive rate reaches 50% or higher, and only once it has accumulated at least 20 verified findings, so a detector doesn't get flagged degraded off a handful of unlucky results. The table isn't empty on a fresh install either; it ships pre-populated with baseline figures from VITI Security's own internal validation runs, and real findings from your own scans layer on top of that baseline as they come in.

What happens when a detector goes DEGRADED?

Once a detector trips the DEGRADED threshold, Vexta's orchestrator quietly stops running that detector class on subsequent scans. It stays skipped until enough newly proven evidence accumulates to flip its rate back down, at which point it resumes running automatically. Nothing here needs a hunter to notice the dashboard and manually turn a check off; the auto-downrank happens on its own, and its only visible effect is that a chronically wrong detector stops contributing noise to your next scan's report.

That matters because a false-positive-heavy detector is worse than a slow one. A slow detector costs scan time. A noisy detector costs your credibility with a triager the next time you submit, and it costs you the hours spent chasing down a bug that was never there. Auto-downrank is Vexta correcting for that on its own, using the same evidence trail (proven, refuted, unproven counts per detector) that's sitting in the PoE Metrics table for you to check.

Reading the badges during triage

In practice, the badges change how you work through a findings list. A row that shows PROVEN and nothing else is close to submission-ready, pending your own read of the evidence. A row that shows PROVEN plus EXPERIMENTAL is worth a manual look anyway, because the detector class behind it is known to occasionally get this wrong even on a technically confirmed result. A row that shows UNVERIFIED tells you the verifier tried and couldn't decide, not that the bug is fake; that distinction matters because Vexta's proof-of-exploitation engine is built so network errors never produce a false REFUTED, which keeps unverified findings from being quietly buried alongside the ones it actively ruled out.

The version tag has its own quiet use during a long-running engagement. If you scanned a target three weeks ago and are rescanning it now, a changed detector revision tag on a repeat finding tells you the detector logic itself moved, not just the target, which is a useful thing to know before you assume nothing changed.

Why this pairing matters more than either piece alone

A badge on a single finding tells you about that one result. A false-positive rate on a detector tells you about every result that detector has ever produced. Put together, they answer two different questions a hunter actually has at different points in a workflow. Mid-scan, looking at one row, the question is should I spend time on this right now, and the badge answers it directly. Before committing to a detector class at all, especially one flagged EXPERIMENTAL, the question is how often has this class of check been right historically, and that's what the PoE Metrics table is for.

The auto-downrank behavior ties both together over time. A detector doesn't earn a DEGRADED label from one bad finding; it takes at least 20 verified findings crossing a 50% refuted rate. That threshold is high enough that a detector isn't punished for a small unlucky streak, but low enough that a genuinely broken check gets pulled from future scans before it burns through many more hours of a hunter's attention across many more targets.

Key takeaways

  • Every finding can carry up to six badges: PROVEN, UNVERIFIED, REFUTED, EXPERIMENTAL, FROM AI, and a detector version tag.
  • UNVERIFIED findings are gated behind a manual reproduction override before submission; REFUTED findings block submission outright.
  • The Settings PoE Metrics dashboard tracks proven, refuted, unproven, and transient counts per detector, plus a computed false-positive rate.
  • Detectors are labeled healthy (under 5% false positives), noisy (5-50%), or DEGRADED (50%+ over at least 20 verified findings).
  • A DEGRADED detector is auto-downranked and skipped on future scans until enough proven evidence brings its rate back down.

Frequently asked questions

What does a PROVEN badge mean in Vexta?
PROVEN means Vexta's proof-of-exploitation engine ran a safe, non-destructive check and confirmed the finding behaves the way it's described, shown as a green badge on the finding row.
Can I submit an UNVERIFIED finding to a bounty platform through Vexta?
Only after ticking an explicit override confirming you manually reproduced it. Vexta gates the submit button on unverified findings until that override is checked.
How does Vexta calculate a detector's false-positive rate?
It divides the number of refuted findings by the number of proven plus refuted findings for that detector, using outcome counters that persist across restarts.
What happens when a detector is marked DEGRADED?
Vexta's orchestrator automatically stops running that detector class on future scans until enough newly proven findings bring its false-positive rate back down.
Why would a PROVEN finding also show an EXPERIMENTAL badge?
EXPERIMENTAL flags detector classes on a curated known-flaky list, such as cache deception or SAML XSW checks, as a reminder to verify manually even when the same finding is also PROVEN.

See which detectors you can trust

Check out Vexta's reliability badges and per-detector false-positive tracking before your next scan.