Split Combined Scans
Detect document boundaries inside one big scanned PDF and split it into separate, labelled documents.
What it is
Automatic splitting of a combined scan — one giant PDF holding many unrelated documents — into its separate, logical documents.
Why it's painful
Scan a stack and you get twenty documents in one file. Before you can process anything, you have to find where each one starts and ends — which teams do by hand.
What Xberg does
Xberg detects document boundaries inside the combined file and splits it into separate documents automatically — one messy stack in, clean separate files out.
- Automatic boundary detection inside a combined file.
- Splits into separate, logical documents.
- A first-class option, not a brittle workaround.
- The natural first step before extraction in mailroom and back-office digitization.
More use cases
RAG Pipeline Ingestion
Turn a pile of PDFs, Office docs, and HTML into clean, chunked, embedded data for your vector database — in one call.
Document-Reading Agents
Give your AI agents one tool to read any document — 100+ formats, structured output, every framework.
Replace Legacy IDP
Swap brittle, template-based processing for one API that returns schema-mapped JSON — no templates to maintain.