Data Discovery & Classification
Privacy-first DSPMEvery piece of personal data, found, understood and under control.
Discovery built for Indian data, from Aadhaar columns to scanned KYC files. Every finding tells you whose data it is, how exposed it is, who owns the fix, and whether it came back.
Trusted by leading enterprises
Scattered data in. An owned inventory out.
Discovery reads your systems where they sit and turns what it finds into facts your teams can act on.
- Databases
- Cloud storage
- File shares
- Laptops and desktops
- Scanned KYC documents
- Backups and exports
- Sensitive data inventory
- Whose data, by population
- Exposure and safeguards per system
- Obligations by section of the Act
- Tasks with named owners
- Questions for what the scan can't see
Measured on every release. Published with the method.
99%
Recall target on Indian identifiers, release by release
95%
Minimum recall every release must pass today
100%
Of planted trap columns caught from content alone
Tested against a seeded Indian estate with trap columns and decoys. A release below the minimum does not ship.
Five steps from unknown to under control
What discovery does for you
Built for Indian data
Aadhaar, PAN, UPI, GSTIN, IFSC, voter ID, passport, driving licence and more, each with its own validation, context and rejection rules. A 12-digit order number stays an order number. Masked, hashed and Base64-encoded copies are caught and labelled for what they are.
Whose data it is
A column of Aadhaar numbers means different things for customers, employees and guarantors. Discovery tags each finding by population, so the right lawful basis follows. Dates of birth under 18 raise a children's-data flag for review. Aadhaar in a marketing tool gets flagged as out of place.
Exposure you can act on
For every system: is the data in plaintext, masked, encrypted or hashed? How many accounts can read it? Is the bucket public? Does a backup hold plaintext copies the live system encrypts? Hashed mobile numbers are marked reversible, because they are.
Safe to point at a bank
Read-only credentials, enforced. Sampling caps, scan windows and a kill switch that stops a scan within a minute. On-premises systems connect through a relay that needs no inbound firewall opening. Scanned documents are read by local OCR and discarded. No client data goes to any external AI.
What the Act asks, and what Discovery shows
Each finding is mapped to the obligation it touches by fixed rules, so two scans of the same system produce the same report.
- s.4, s.5
A lawful ground and a notice for every use
Which identifiers each system holds, for whom, as input to the notice and the lawful-basis register
- s.8(5), Rule 6
Reasonable security safeguards
Plaintext identifiers, reversible hashes, public buckets, wide read access, per system
- s.8(6)
Breach readiness
The exposure surface: where a breach would reach personal data, and how much
- s.8(7)
Erase when the purpose ends
Copies in backups, exports, file shares and laptops, outside the system of record
- s.8(2)
Processors under contract
Personal data found in third-party and SaaS-hosted sources
- s.9
Children's data
Dates of birth implying under-18s, flagged for review
- s.11 to s.14
Data Principal rights
Where one person's records sit across systems, so access and erasure can be answered
- s.16
Cross-border transfers
The hosting region of every source holding personal data
What Discovery hands on
- Data MappingDiscovered systems and data categories pre-load the register, so mapping starts from what exists.
- Data Principal Rights PortalKnows where a person's records sit, so access and erasure requests reach every copy.
- Assessment & Compliance ReportingSafeguard evidence and findings feed control testing, marked by how they were established.
Deploy it your way
Built for regulated Indian enterprises: your data stays where your policies say it must.
- 01
SaaS, hosted in India
Managed by SARC AI on infrastructure in Indian data centres.
- 02
Private cloud
Runs in your own cloud account, under your keys and access policies.
- 03
On-premises
Installed in your data centre. Scanners read in place, and nothing leaves your network.
Connects to
PostgreSQL
MySQL
Oracle Database
Microsoft SQL Server
MongoDB
MariaDB
- IDIBM Db2
SAP HANA
Snowflake
Amazon Redshift
Google BigQuery
Databricks
Teradata
Amazon DynamoDB
Azure SQL
Apache Cassandra
Redis
Elasticsearch
Couchbase
ClickHouse
Apache Hive
Apache Kafka
SQLite
Neo4j
Connectors are enabled during onboarding. Logos belong to their owners and indicate compatibility, not endorsement.
Certified
Frequently asked questions
What is DSPM, and how is Discovery different from DLP?
DSPM tells you what sensitive data you hold, where it sits and how exposed it is. DLP stops data leaving while it moves, by email, USB or upload. Discovery is DSPM built for privacy: it adds whose data it is and which section of the DPDP Act each finding touches. It pairs well with a DLP tool you already run.
Does any of our data leave our environment?
Classification runs where the data sits. Discovery stores findings, counts and fingerprints, and short redacted samples only where you switch them on per source. Laptop sweeps upload findings, never files. No client data, sample or column name is sent to any external AI service.
Will scanning affect our production systems?
Discovery connects with read-only credentials and refuses any credential that can write. It samples instead of reading whole tables, runs inside the windows you set, respects rate limits, and can be stopped within a minute. Every scan is recorded, with who authorised it and what was covered.
How accurate is it, and how do you know?
Every release runs against a seeded Indian test estate with trap columns and decoys, and has to reach 95% recall and 90% precision before it ships. Our target is 99% recall. We publish the method so your team can check it, and we're glad to run the same test on a sample of yours.
Which systems and files can it scan?
PostgreSQL, MySQL and MariaDB, SQL Server, Oracle, MongoDB and Amazon S3 are validated today. It also reads file shares, server folders, Windows and Linux laptops, and PDF, Word, Excel, email and database dump files, with local OCR for scanned Aadhaar and PAN cards. Further connectors are enabled during onboarding.
Does Discovery change or delete data?
No. It's read-only by design, so it can safely run against a bank's live systems. Each finding goes to an owner with the action it needs, such as encrypting a column or removing copies from laptops. Your team makes the change, and the next scan confirms it.
Trusted by leaders across industries
“We knew patient data sat in our hospital information system. We didn't know how much had spread into lab exports, scanned reports and shared drives until DataSurity's assessment showed us. The team understood hospital realities, from paediatric records to staff data, and gave us a plan we could actually run. Implementation is now moving ward by ward, with consent and rights handled in one place.”
Rakesh G
Head - Compliance, KDSG Hospitals (350-bed multispecialty hospital)





