Redact Sensitive Data
Detect personal information across many categories and mask, hash, or tokenize it — with a report of what was found and where.
Illustrative example. This capability is on our roadmap and is not wired to a live demo yet — request early access.
What it is
Automatic detection and removal of personal data from documents — names, addresses, SSNs, card numbers, dates of birth, and medical identifiers — so a document can be shared, stored, or used for analysis without leaking PII.
Why it matters
Redacting by hand is slow and unreliable, and under regimes like GDPR and HIPAA a single missed identifier is a compliance incident. Teams need it done consistently and provably across every document.
What Xberg does
Xberg's engine detects personal information across many categories and applies the treatment you choose — mask, hash, tokenize, or drop — then returns a report of exactly what it found and where.
- Detects many PII categories, not just a fixed handful.
- Choose the treatment per need: mask, hash, tokenize, or drop.
- An audit report lists every match and its location in the document.
- The rest of the document stays intact and usable.
More use cases
RAG Pipeline Ingestion
Turn a pile of PDFs, Office docs, and HTML into clean, chunked, embedded data for your vector database — in one call.
Document-Reading Agents
Give your AI agents one tool to read any document — 100+ formats, structured output, every framework.
Replace Legacy IDP
Swap brittle, template-based processing for one API that returns schema-mapped JSON — no templates to maintain.