Finding Deduplication

Prev Next

Overview

SAFE collects security findings (vulnerabilities, misconfigurations, compliance gaps) from multiple scanning tools, Wiz, Tanium, Qualys, CrowdStrike, and others. When findings flow in from these sources, SAFE determines whether a new finding is genuinely new or whether it's a repeat of something already recorded.

This process is called deduplication, deciding when two data points represent the "same" finding and should be merged, versus when they represent distinct issues that should be counted separately.

Finding Types

Type

Label

What It Represents

Vulnerability

va

A known CVE (e.g., CVE-2023-4911) affecting software

Wiz CVE scans, Tanium vulnerability scans

Misconfiguration

ca

A non-CVE security issue — misconfig, EOL software, insecure setting

Wiz cloud configs, Wiz host configs Tanium Comply rules

Issues

finding

A compliance framework control (CIS, ISO, NIST) pass/fail assessment

Wiz issues

Factors That Determine Identity

Every finding's uniqueness in the UI is determined by two primary factors:

  • The Finding Itself — What is the issue? (CVE ID, control name, misconfiguration title)

  • The Software Context — What software is affected? (Application name, version, install location)

If either of these factors differs between two data points, SAFE treats them as separate findings in the UI.

Third factor: The Source (which connector/scanner reported it) — determines how findings are stored in the backend for tracking, but findings from different sources are deduplicated and unified before being shown to the customer.

Three Rules of Finding Identity

Rule 1: Finding Itself

A finding is identified by its core characteristics:

Finding Type

What Makes It Unique

Vulnerability (va)

The CVE ID (e.g., CVE-2023-4911) + its type classification

Misconfiguration (ca)

The finding name + its description + its type classification

Compliance Control (finding)

The control ID + the benchmark/framework it belongs to + the control statement

Key behavior: The finding type label (va, ca, finding) is part of the identity. This means a CVE reported as a "vulnerability" and the same CVE somehow reported as a "misconfiguration" would be treated as two sep

Rule 2: Software Context

Findings can be associated with specific software. When they are, the software identity becomes part of the finding's identity.

  • Software is identified by: Software name (e.g., "OpenSSL") - Version (e.g., "1.1.1k") - Install path (e.g., "/usr/lib/libssl.so") - Username/owner (e.g., "root")

  • Consequence: The same CVE affecting OpenSSL 1.1.1k installed at /usr/lib/ is a different finding instance from the same CVE affecting OpenSSL 1.1.1k installed at /opt/custom/lib/. Different path = different software instance = different finding.

  • When there's no software context: Some findings (device-level misconfigurations, compliance controls) don't have an associated software component. In those cases, the finding is scoped to the device/asset only.

Rule 3: Source

SAFE maintains separate finding instances per connector (source) in its backend for tracking and auditability purposes. However, these are deduplicated and unified by the downstream pipeline before being presented in t in the backend for our sync and re-sync operationshe UI.

  • Consequence: If both Wiz and Tanium report the same vulnerability (for example, CVE-2023-4911) on the same server, SAFE stores each report as a separate backend record—one for each source. However, the UI deduplicates these records and presents them as a single unified finding to the user.

  • Why maintain separate backend instances? Each scanner may provide different metadata, severity assessments, remediation guidance, and observation dates. Keeping per-source records ensures traceability and allows SAFE to intelligently merge the best information from each source into the unified view.

  • What the customer sees: One finding. The downstream pipeline consolidates instances from multiple sources into a single entry, selecting the most relevant metadata (e.g., highest severity, earliest first-observed date).

  • Re-sync behavior: When the same connector re-syncs and reports the same finding again, SAFE recognizes it as the same finding and updates the existing record (idempotent). No duplicate is created.

Deduplication Scenarios

Scenario 1: Same CVE on Different Software Installs

Situation: A server has OpenSSL installed in two locations: - OpenSSL 1.1.1k at /usr/lib/libssl.so (system default) - OpenSSL 1.1.1k at /opt/app/lib/libssl.so (bundled with an application)

Wiz scans the server and reports CVE-2023-0286 on both installations.

What SAFE Does: Records 2 separate vulnerability findings.

#

Finding

Software

Install Path

Source

Result

1

CVE-2023-0286

OpenSSL 1.1.1k

/usr/lib/libssl.so

Wiz

Finding Instance A

2

CVE-2023-0286

OpenSSL 1.1.1k

/opt/app/lib/libssl.so

Wiz

Finding Instance B

Why: Rule 2 applies — different install paths mean different software instances, so each CVE occurrence is tracked independently. This matters because patching one location doesn't necessarily fix the other.

Impact: The finding count reflects the actual remediation effort needed — two separate installations require two separate fixes.

Scenario 2: Same CVE on Same Software from Same Source

Situation: Wiz scans a server and in its vulnerability report, it lists CVE-2021-44228 (Log4Shell) twice against the same software instance — Log4j 2.14.1 at /opt/app/lib/log4j-core-2.14.1.jar, owned by user "appuser".

This can happen when: The scanner reports the same finding in multiple scan passes during one sync - Multiple detection rules trigger on the same underlying issue

What SAFE Does: Records 1 vulnerability finding (deduplicated).

#

Finding

Software

Install Path

User

Source

Result

1

CVE-2021-44228

Log4j 2.14.1

/opt/app/lib/log4j-core-2.14.1.jar

appuser

Wiz

Finding Instance A

2

CVE-2021-44228

Log4j 2.14.1

/opt/app/lib/log4j-core-2.14.1.jar

appuser

Wiz

Deduplicated → merges into A

Why: All three rules match — same finding (CVE ID), same software (name + version + path + user), same source (Wiz). SAFE recognizes these as the same issue and keeps only one record, updating metadata if newer data is available.

Impact: No inflated finding counts from redundant scanner detections.

Scenario 3: Same Finding Reported by Wiz AND Tanium

Situation: A server has OpenSSH 7.4p1 installed. Both Wiz and Tanium scan this server and both report CVE-2023-38408 affecting OpenSSH.

What SAFE Does: Shows 1 deduplicated finding in the UI.

#

Finding

Software

Sources

Shown in UI

1

CVE-2023-38408

OpenSSH 7.4p1

Wiz + Tanium

1 unified finding

Behind the scenes: SAFE maintains separate instances in the backend (one per source) for tracking and auditability. These per-source records preserve each scanner's specific metadata. However, the downstream pipeline deduplicates and unifies them before presenting to the customer.

Layer

What Happens

Backend (ingestion)

2 separate instances stored — one from Wiz, one from Tanium

Downstream pipeline

Instances are deduplicated and consolidated

UI (customer-facing)

1 unified finding displayed

Why maintain separate backend instances: Each scanner may assign different severity (Wiz might call it HIGH, Tanium might call it CRITICAL) - Each scanner tracks its own observation dates (first seen, last seen) - Each scanner provides different remediation guidance and fix versions - If one scanner is removed/disconnected, the finding from the other scanner persists - Allows intelligent merging of the best metadata from all sources

Impact: Adding a second scanner does not double the finding count in the UI. Multiple sources reporting the same finding are consolidated into one entry, providing richer metadata and redundancy without inflating counts.

Scenario 4: Same Compliance Control Under Different Policies

Situation: A Wiz compliance scan reports that a server fails the control "Ensure SSH root login is disabled" (control ID: wc-id-183). This same control appears in multiple compliance frameworks: - ISO 27001:2022 (Policy: ISO/IEC_27001-2022) - CIS AWS Benchmark v1.5 (Policy: CIS_AWS_v1.5) - NIST 800-53 (Policy: NIST_800-53)

What SAFE Does: Shows 1 deduplicated compliance finding in the UI.

#

Control

Policies/Frameworks

Shown in UI

1

"Ensure SSH root login is disabled"

ISO 27001:2022, CIS AWS v1.5, NIST 800-53

1 unified control finding

Behind the scenes: SAFE maintains separate compliance posture entries per policy/framework in the backend for per-framework reporting and tracking. However, these are deduplicated at the control level by the downstream pipeline and presented as a single finding in the UI.

Layer

What Happens

Backend (compliance posture)

3 separate posture entries — one per policy/framework

Downstream pipeline

Deduplicated at the control level

UI (customer-facing)

1 unified compliance finding displayed

Why maintain separate posture entries: Pass/fail status may differ by framework version or interpretation - Compliance reporting per-framework requires granular tracking - Audit trails need per-policy evidence

Impact: The customer sees one compliance finding for "Ensure SSH root login is disabled" regardless of how many frameworks reference it. The finding count reflects unique control gaps, not the number of framework mappings. Per-framework compliance posture details remain available for drill-down and compliance reporting.

Scenario 5: Tanium Compliance as 'ca' vs. Default 'finding' Type

Situation: Both Wiz and Tanium report compliance controls for the same server. However, they classify their findings differently: - Wiz sends compliance controls through the standard compliance pathway → SAFE assigns type finding (the default for compliance) - Tanium explicitly classifies its compliance controls as ca (misconfiguration) in its data — a deliberate design choice because Tanium's compliance checks are closer to configuration assessments than framework controls

What SAFE Does: Even if the control name/ID is identical, these produce different finding identities because the type classification differs.

#

Control

Type Classification

Source

Result

1

"Ensure password complexity"

finding (compliance)

Wiz

Finding Identity X

2

"Ensure password complexity"

ca (misconfiguration)

Tanium

Finding Identity

Why: The finding type is part of what determines a finding's identity (Rule 1). ca and finding produce different internal identifiers even for the same control name. This is intentional — it preserves the connector's assessment of what kind of issue this represents.

Important nuance: These would also be separate due to Rule 3 (different sources), but even if they hypothetically came from the same source, the different type classification alone would keep them separate.

Impact: When comparing Tanium and Wiz findings, the type classification difference means they won't appear as "the same control checked twice." They're treated as fundamentally different kinds of findings — one is a configuration check (ca), the other is a compliance framework assessment (finding).

Scenario 6: Re-syncing Same Data from Same Connector

Situation: Wiz runs a daily vulnerability scan. On Monday it reports CVE-2023-4911 on server-01. On Tuesday it runs again and reports the same CVE-2023-4911 on server-01 with the same software details.

What SAFE Does: The finding count remains 1. The existing record is updated (not duplicated).

Day

Finding

Software

Source

Action

Monday

CVE-2023-4911 on glibc 2.26

/lib64/libc.so.6

Wiz

Created — new finding

Tuesday

CVE-2023-4911 on glibc 2.26

/lib64/libc.so.6

Wiz

Updated — timestamps refreshed

Wednesday (fixed)

CVE-2023-4911 no longer reported

—

Wiz

Marked resolved

Why: All identity factors are identical (same CVE, same software, same source). SAFE uses deterministic identifiers — the same inputs always produce the same internal ID. When a finding with an existing ID arrives, it's treated as a refresh/update.

What gets updated on re-sync: Last observed date - Severity (if the scanner changed its assessment) - Remediation/fix information - Exploit status

What doesn't change: First observed date (preserved from original detection) - Finding count (no inflation)

Impact: Daily or frequent re-syncs do not inflate finding counts. Customers can run connectors as often as needed without worrying about duplicate findings accumulating.

Connector Behavior Summary

How Wiz Maps Findings

Wiz Data Source

SAFE Finding Type

Classification Method

Vulnerability API (CVEs)

Vulnerability (va)

Default — Wiz doesn't set explicit types

Cloud Configuration Issues

Misconfiguration (ca)

Default — Wiz doesn't set explicit types

Compliance Framework Checks

Compliance Control (finding)

Default — Wiz doesn't set explicit types

Wiz relies entirely on SAFE's default classification logic. This is because Wiz's APIs can sometimes return mixed content (e.g., their vulnerability API might return non-CVE issues), so letting SAFE classify based on content is safer.

How Tanium Maps Findings

Tanium Data Source

SAFE Finding Type

Classification Method

Vulnerability Scans (CVEs)

Vulnerability (va)

Explicitly set by Tanium connector

Compliance/CIS Checks

Misconfiguration (ca)

Explicitly set by Tanium connector

Key difference: Tanium explicitly sets its compliance findings as ca (misconfiguration), not the default finding (compliance control). This means Tanium compliance data lives in a different classification space than Wiz compliance data.

Summary Comparison

Aspect

Wiz

Tanium

CVE classification

va (default)

va (explicit)

Misconfiguration classification

ca (default)

—

Compliance classification

finding (default)

ca (explicit override)

Classification method

Relies on SAFE defaults

Sets types explicitly

Impact on dedup

Standard behavior

Compliance findings dedup differently than Wiz's

FAQs

Finding Count Questions

Q: A customer says their vulnerability count seems higher than expected. What could cause this?

Common reasons: 1. Multiple software installs — The same CVE on 3 different install paths = 3 findings 2. Multiple versions — Same app at v1.0 and v2.0, same CVE on both = 2 findings 3. Different users — Same software installed for user A and user B = different instances

Note: Multiple scanners reporting the same CVE does not inflate the count — these are deduplicated in the UI.

Q: Will the same CVE from two different scanners appear once or twice?

Once in the UI. While SAFE maintains separate instances per source in the backend for tracking, the downstream pipeline deduplicates them. The customer sees one unified finding with the combined metadata from all sources.

Q: If a customer adds a second scanner, will their finding count roughly double?

No. Findings that overlap between scanners are deduplicated in the UI. Adding a second scanner primarily adds coverage for assets/findings the first scanner didn't detect, plus provides redundancy and richer metadata for overlapping findings.

Q: Why does the customer see more compliance findings than they have unique controls?

This is uncommon since compliance controls are deduplicated at the control level. If a customer reports seeing more findings than expected, it may be due to controls with slightly different names or IDs across frameworks, which the system treats as distinct controls.


Re-sync and Updates

Q: What happens when we re-sync a connector?

Existing findings are updated, not duplicated. The sync refreshes timestamps, severity, and metadata. Finding count stays the same unless: - New vulnerabilities were discovered since last sync (count goes up) - Previously found vulnerabilities are now remediated (count goes down)

Q: Can a re-sync ever increase the finding count without new actual vulnerabilities?

No, if the underlying data hasn't changed. The same inputs always produce the same internal identifiers. Re-sending identical data is a no-op from a count perspective.

Q: What happens to findings when a connector is disconnected?

Findings from that connector remain in SAFE (they aren't deleted). They just won't receive updates. If the connector is reconnected and re-syncs, it picks up where it left off.


Wiz vs. Tanium Specific

Q: Why does Tanium show compliance findings differently than Wiz?

Tanium classifies its compliance checks as misconfigurations (ca) rather than the standard compliance controls (finding). This is a deliberate mapping by the Tanium connector because Tanium's compliance assessments are configuration-check-based rather than framework-control-based.

This means: - Wiz compliance findings appear as "Compliance Controls" in SAFE - Tanium compliance findings appear as "Misconfigurations" in SAFE - They will never deduplicate against each other (different types + different sources)

Q: If a customer uses both Wiz and Tanium, will their compliance data conflict?

No conflict — they coexist as separate finding types. Wiz compliance shows under the "Compliance Control" category, Tanium compliance shows under "Misconfiguration." They provide complementary views, not competing ones.

Q: A customer sees the same CIS control reported differently by Wiz and Tanium. Is this a bug?

No. This is expected behavior due to how each connector classifies its data. The control may have the same name but is assessed differently by each tool and stored in different categories within SAFE.  

Software and Components

Q: What counts as "different software" for deduplication purposes?

Any difference in these attributes creates a separate software instance: - Name (e.g., "OpenSSL" vs "LibreSSL") - Version (e.g., "1.1.1k" vs "1.1.1n") - Install path (e.g., "/usr/lib/" vs "/opt/custom/lib/") - Owner/user (e.g., "root" vs "appuser")

Q: If software is patched (version changes), what happens to the old finding?

The old finding (against the old version) will no longer be reported in subsequent scans. Once the connector re-syncs and doesn't report that CVE anymore, the finding is marked as resolved. The new version may or may not have its own findings.

Q: Some findings don't seem to have software attached. Is that normal?

Yes. Device-level misconfigurations (like "SSH root login enabled") and many compliance controls are assessed at the device/asset level, not tied to specific software. These are still properly deduplicated — just without the software dimension.  

Deduplication Logic

Q: Is deduplication real-time or batch-based?

Deduplication happens at ingest time — when data arrives from a connector sync. There's also a database-level constraint that prevents duplicates even if logic is bypassed. It's effectively real-time per sync.

Q: How does cross-source deduplication work?

Cross-source deduplication is automatic. While the backend maintains separate instances per source for tracking purposes, the downstream pipeline unifies them. The UI presents a single consolidated finding regardless of how many sources reported it.

Q: What determines if a non-CVE finding (misconfiguration) is "the same" across syncs?

For non-CVE misconfigurations, identity is determined by: - The finding name (e.g., "EC2 instance should use IMDSv2") - The finding description - The finding type (ca) - The source connector

If any of these change (e.g., the scanner updates the description text), it technically creates a new finding identity. In practice, scanners keep names/descriptions stable across syncs.