DataSurity

Data Discovery & Classification

Privacy-first DSPM

Every piece of personal data, found, understood and under control.

Discovery built for Indian data, from Aadhaar columns to scanned KYC files. Every finding tells you whose data it is, how exposed it is, who owns the fix, and whether it came back.

Trusted by leading enterprises

  • Indiabulls Securities
  • KDSG Super-Speciality Hospital
  • Modicare
  • Express Inn Hotels & Resorts
  • Econo Broking
  • DAMS
  • Freesia by Express Inn
  • MBL
  • Trident Group
  • Dhani
  • Indiabulls Asset Reconstruction

Scattered data in. An owned inventory out.

Discovery reads your systems where they sit and turns what it finds into facts your teams can act on.

  • Databases
  • Cloud storage
  • File shares
  • Laptops and desktops
  • Scanned KYC documents
  • Backups and exports
DataSurity Discovery
  • Sensitive data inventory
  • Whose data, by population
  • Exposure and safeguards per system
  • Obligations by section of the Act
  • Tasks with named owners
  • Questions for what the scan can't see
Accuracy you can check

Measured on every release. Published with the method.

  • 99%

    Recall target on Indian identifiers, release by release

  • 95%

    Minimum recall every release must pass today

  • 100%

    Of planted trap columns caught from content alone

Tested against a seeded Indian estate with trap columns and decoys. A release below the minimum does not ship.

How it works

Five steps from unknown to under control

Features

What discovery does for you

01

Built for Indian data

Aadhaar, PAN, UPI, GSTIN, IFSC, voter ID, passport, driving licence and more, each with its own validation, context and rejection rules. A 12-digit order number stays an order number. Masked, hashed and Base64-encoded copies are caught and labelled for what they are.

02

Whose data it is

A column of Aadhaar numbers means different things for customers, employees and guarantors. Discovery tags each finding by population, so the right lawful basis follows. Dates of birth under 18 raise a children's-data flag for review. Aadhaar in a marketing tool gets flagged as out of place.

03

Exposure you can act on

For every system: is the data in plaintext, masked, encrypted or hashed? How many accounts can read it? Is the bucket public? Does a backup hold plaintext copies the live system encrypts? Hashed mobile numbers are marked reversible, because they are.

04

Safe to point at a bank

Read-only credentials, enforced. Sampling caps, scan windows and a kill switch that stops a scan within a minute. On-premises systems connect through a relay that needs no inbound firewall opening. Scanned documents are read by local OCR and discarded. No client data goes to any external AI.

What the Act asks, and what Discovery shows

Each finding is mapped to the obligation it touches by fixed rules, so two scans of the same system produce the same report.

Section
  • s.4, s.5

    A lawful ground and a notice for every use

    Which identifiers each system holds, for whom, as input to the notice and the lawful-basis register

  • s.8(5), Rule 6

    Reasonable security safeguards

    Plaintext identifiers, reversible hashes, public buckets, wide read access, per system

  • s.8(6)

    Breach readiness

    The exposure surface: where a breach would reach personal data, and how much

  • s.8(7)

    Erase when the purpose ends

    Copies in backups, exports, file shares and laptops, outside the system of record

  • s.8(2)

    Processors under contract

    Personal data found in third-party and SaaS-hosted sources

  • s.9

    Children's data

    Dates of birth implying under-18s, flagged for review

  • s.11 to s.14

    Data Principal rights

    Where one person's records sit across systems, so access and erasure can be answered

  • s.16

    Cross-border transfers

    The hosting region of every source holding personal data

Deploy it your way

Built for regulated Indian enterprises: your data stays where your policies say it must.

  • 01

    SaaS, hosted in India

    Managed by SARC AI on infrastructure in Indian data centres.

  • 02

    Private cloud

    Runs in your own cloud account, under your keys and access policies.

  • 03

    On-premises

    Installed in your data centre. Scanners read in place, and nothing leaves your network.

Connects to

  • PostgreSQL
  • MySQL
  • Oracle Database
  • Microsoft SQL Server
  • MongoDB
  • MariaDB
  • IDIBM Db2
  • SAP HANA
  • Snowflake
  • Amazon Redshift
  • Google BigQuery
  • Databricks
  • Teradata
  • Amazon DynamoDB
  • Azure SQL
  • Apache Cassandra
  • Redis
  • Elasticsearch
  • Couchbase
  • ClickHouse
  • Apache Hive
  • Apache Kafka
  • SQLite
  • Neo4j

Connectors are enabled during onboarding. Logos belong to their owners and indicate compatibility, not endorsement.

Certified

  • CERTIFIEDISO 27001INFORMATION SECURITY
  • CERTIFIEDISO 27701PRIVACY INFORMATION

Frequently asked questions

What is DSPM, and how is Discovery different from DLP?

DSPM tells you what sensitive data you hold, where it sits and how exposed it is. DLP stops data leaving while it moves, by email, USB or upload. Discovery is DSPM built for privacy: it adds whose data it is and which section of the DPDP Act each finding touches. It pairs well with a DLP tool you already run.

Does any of our data leave our environment?

Classification runs where the data sits. Discovery stores findings, counts and fingerprints, and short redacted samples only where you switch them on per source. Laptop sweeps upload findings, never files. No client data, sample or column name is sent to any external AI service.

Will scanning affect our production systems?

Discovery connects with read-only credentials and refuses any credential that can write. It samples instead of reading whole tables, runs inside the windows you set, respects rate limits, and can be stopped within a minute. Every scan is recorded, with who authorised it and what was covered.

How accurate is it, and how do you know?

Every release runs against a seeded Indian test estate with trap columns and decoys, and has to reach 95% recall and 90% precision before it ships. Our target is 99% recall. We publish the method so your team can check it, and we're glad to run the same test on a sample of yours.

Which systems and files can it scan?

PostgreSQL, MySQL and MariaDB, SQL Server, Oracle, MongoDB and Amazon S3 are validated today. It also reads file shares, server folders, Windows and Linux laptops, and PDF, Word, Excel, email and database dump files, with local OCR for scanned Aadhaar and PAN cards. Further connectors are enabled during onboarding.

Does Discovery change or delete data?

No. It's read-only by design, so it can safely run against a bank's live systems. Each finding goes to an owner with the action it needs, such as encrypting a column or removing copies from laptops. Your team makes the change, and the next scan confirms it.

Trusted by leaders across industries

“We knew patient data sat in our hospital information system. We didn't know how much had spread into lab exports, scanned reports and shared drives until DataSurity's assessment showed us. The team understood hospital realities, from paediatric records to staff data, and gave us a plan we could actually run. Implementation is now moving ward by ward, with consent and rights handled in one place.”

Rakesh G

Head - Compliance, KDSG Hospitals (350-bed multispecialty hospital)

Knowledge resources

Point it at one database. See what's really there.