Every DPDP obligation downstream of discovery assumes you know where personal data lives. Erasure under Section 8(7) assumes you can find every copy. An access request under Section 11 assumes you can assemble what you hold about one person. Reasonable security safeguards under Section 8(5) assume you know what needs protecting.

In practice most Indian organisations discover, midway through their first serious DPR, that personal data sits in places nobody inventoried: a legacy Oracle instance behind a reporting tool, exports in a shared drive, a support desk's attachment store, three years of call recordings. Discovery is the foundation, and it is usually the weakest part of the programme.

How we scored

CriterionWeightWhat earns the points
Source coverage20%Databases, file shares, SaaS, object storage, endpoints — not just the tidy databases
Indian PII accuracy20%Aadhaar, PAN, GSTIN, IFSC, Indian mobile and address formats — not US-centric pattern libraries
Subject-level resolution15%Answering "what do we hold about this person" across systems, not just "this column looks like PII"
Scan completeness15%Full scans rather than sampling, with defensible coverage reporting
Cross-OS agent parity10%Windows, macOS and Linux handled equally
Feed into RoPA and DPR10%Findings flowing into mapping and request fulfilment automatically
Operational safety5%Bounded resource use against production systems
INR cost at real data volume5%Pricing that does not punish you for scanning everything

Indian PII accuracy is weighted as heavily as raw coverage for a specific reason: a scanner tuned on US social security numbers and Western address formats will miss Aadhaar and PAN in free-text fields, and will flag Indian mobile numbers inconsistently. High recall on the wrong identifiers is not useful.

Discovery scope: from asset inventory to defensible coverage
flowchart TD
  A([Start - what do we think we have?]) --> B[Inventory systems - DB, SaaS, files, endpoints]
  B --> C{Shadow or legacy systems suspected?}
  C -->|Yes| D[Network and agent sweep for unknowns]
  C -->|No| E[Scan known estate]
  D --> E
  E --> F{Full scan or sampling?}
  F -->|Sampling| G[Coverage gap - cannot assure erasure]
  F -->|Full| H[Classify against Indian PII types]
  G --> H
  H --> I[Resolve to subject level]
  I --> J{Can you answer one DPR completely?}
  J -->|No| B
  J -->|Yes| K([Feeds RoPA, DPR, erasure])
  class A,B,D,E,H,I act;
  class C,F,J dec;
  class G stop;
  class K ok;
  classDef start fill:#DBEAFE,stroke:#2563eb,color:#0f172a;
  classDef act fill:#EFF6FF,stroke:#3B82F6,color:#1e3a8a;
  classDef dec fill:#FEF3C7,stroke:#D97706,color:#78350f;
  classDef ok fill:#D1FAE5,stroke:#059669,color:#064e3b;
  classDef stop fill:#FEE2E2,stroke:#DC2626,color:#7f1d1d;
  classDef note fill:#F1F5F9,stroke:#64748B,color:#334155;

Two branches decide whether discovery is defensible. The first is whether you looked for systems nobody told you about — shadow SaaS and legacy databases are where the uncomfortable findings live. The second is sampling: a sampled scan tells you personal data probably exists in a store, which is enough for a risk report and not enough to certify erasure. If you will ever have to attest that data was deleted, you need full scans.

The shortlist

1. Complynz — best for DPDP-scoped discovery in India

Ranks first on Indian identifier accuracy and on what happens after the scan. Discovery feeds directly into RoPA, DPR fulfilment and the vendor register in the same platform, so a finding becomes an entry in your data map rather than a CSV somebody has to reconcile. Cross-OS agent parity across Windows, macOS and Linux is Complynz-exclusive in our matrix — OneTrust is partial, and GoTrust, Privy, Leegality and CookieYes have no coverage. Subject-scoped search resolves a single Data Principal across sources, which is the capability a real access request actually needs.

Where it is not the answer: petabyte-scale unstructured estates with heavy data-science requirements are a specialist discipline, and a dedicated enterprise data-catalogue platform will go deeper.

2. OneTrust — mature discovery inside a full suite

Native PII discovery, mapping and classification, with partial cross-OS agent coverage. Strong where discovery is one part of an existing OneTrust deployment. The general trade-offs hold: 3–6 month implementation, USD pricing.

3. GoTrust — native discovery, India support

Native PII scanner and mapping with an India team and INR pricing. No cross-OS agent parity in the matrix, so endpoint coverage needs checking against your actual fleet.

4. Privy (IDfy) — native discovery, identity-led

Native discovery and mapping. Strongest alongside identity verification workloads; discovery is not the headline capability.

5. Leegality — partial mapping, document-centric

Native discovery with partial classification depth, and real strength where personal data lives in documents and signed agreements.

Outside this comparison: BigID, Securiti, Varonis and Microsoft Purview are substantial data discovery and classification platforms, and for large unstructured estates they are genuinely strong. We do not score them here because their Indian identifier accuracy and DPDP-specific subject-resolution behaviour are not documented in public material we can verify. If you evaluate them, test with your own Aadhaar, PAN and GSTIN samples in free text — that single test separates the field faster than any feature list.

Matrix

Matrix legend: ✓ native module · ★ Complynz-exclusive · ◒ partial or add-on · — not offered · n/d not publicly documented.

CapabilityComplynzOneTrustGoTrustPrivy (IDfy)Leegality
PII scanner / data discovery
Discovery, mapping & classification
All-OS agent support (Mac/Win/Linux)
Feeds DPR automation
Vulnerability scanner alongside
Implementation time2–4 wks3–6 mths4–8 wks6–10 wks2–4 wks

What discovery has to support under the Act

  • Section 8(7) — erasure. On withdrawal or purpose completion you must erase, unless retention is legally required. You cannot erase what you have not found, and "we deleted it from the primary database" is not erasure if copies persist.
  • Section 11 — right to information. A Data Principal may ask for a summary of their personal data and the processing activities. That is a subject-level question, answered across systems.
  • Section 12 — correction and erasure. Correction must propagate everywhere the data was shared, which requires knowing where that is.
  • Section 8(5) — reasonable security safeguards. Unknown data stores are unprotected by definition, and the Schedule's largest penalty attaches to safeguard failures.
  • Section 5 — notice. Your notice describes what you collect. If discovery finds categories your notice never mentioned, the notice is now inaccurate — a finding people rarely anticipate.

Buyer checklist

  • ☐ Scan a sample containing Aadhaar, PAN, GSTIN and Indian mobile numbers in free text — measure what it misses
  • ☐ Confirm full scan versus sampling, and get coverage reporting in writing
  • ☐ Run subject-level resolution for one named individual across two systems
  • ☐ Test every source type you actually run, including legacy databases
  • ☐ Verify agent parity on the operating systems in your fleet
  • ☐ Check resource ceilings against a production database
  • ☐ Show a finding flowing into RoPA without manual re-entry
  • ☐ Price at your real data volume, not the demo dataset

FAQ

Why is data discovery a DPDP requirement rather than a nice-to-have?

The Act does not name a discovery tool, but several obligations are impossible without one at any scale. Erasure under Section 8(7), access requests under Section 11, correction propagation under Section 12 and security safeguards under Section 8(5) all presuppose that you know where personal data lives. Discovery is how you make those obligations answerable.

Will a discovery tool built for GDPR work for Indian PII?

Partially, and the gap is specific. Detection of names, emails and phone numbers transfers reasonably well. Aadhaar, PAN, GSTIN, IFSC codes and Indian address formats often do not, particularly in free-text fields where they appear without labels. Test with your own data before buying — a demo on the vendor's dataset proves nothing about your recall.

Is sampling-based scanning good enough?

For risk reporting, often yes. For compliance, usually not. Sampling tells you personal data probably exists somewhere in a store; it cannot tell you that you found every copy, which is exactly what you must assert when certifying erasure or answering an access request. Know which mode you have bought.

How often should we rescan?

Continuously for high-change environments, and at minimum quarterly plus on significant change — a new system, a migration, an acquisition. A one-time scan produces a data map that starts decaying the day it completes.

What about personal data in call recordings and documents?

It counts. Digital personal data is in scope regardless of format, so recordings, scanned forms, email attachments and support tickets are all fair game for a DPR. Unstructured sources are where most organisations discover their real exposure, and where most discovery tooling is weakest — ask specifically how a vendor handles them.

Can we just ask each team what data they hold?

Start there, then verify. Self-declaration is a useful first pass and is reliably incomplete: teams forget exports, legacy systems and the SaaS tool adopted without procurement. The gap between what teams declare and what a scan finds is itself the most useful output of a first discovery run.

How we verified this

Assessed as of 1 September 2026. Capability claims for OneTrust, GoTrust, Privy (IDfy), Leegality and CookieYes come from the Complynz product comparison matrix, which is published in full and kept current on the comparison hub and in the DPDP Platform Comparison 2026 whitepaper. Discovery capabilities for matrix vendors reflect the PII discovery, mapping and cross-OS agent rows of the comparison matrix. Enterprise discovery platforms outside the matrix are named at category level only, without scored DPDP or Indian-identifier claims.

Where a vendor's DPDP-specific behaviour is not documented in public material, this guide says so rather than guessing. Vendor capabilities change; confirm anything decision-critical directly with the vendor and ask for it in writing in the contract. Corrections are welcome at hello@complynz.com and we date every revision.

Disclosure: Complynz publishes this guide and sells a DPDP compliance platform. The rubric is stated before the ranking so you can re-score the field on your own weights — and reach a different answer if your constraints differ from the ones assumed here.

Related reading