Open source
All primitives, in 14 languages. Embed in your app or self-host the extraction server. Use any subset of the pipeline.
Get the librariesA self-hostable document intelligence and RAG platform — an integral component in AI meshes. Open-source primitives, one managed backend. 96 file formats, 306 programming languages, RAG built in.
96 formats. 306 code languages. Lossless HTML.
Collections. Hybrid retrieval. Reranking.
Everything teams build on the pipeline — run any of them and see structured output. No account, no setup.
Turn a pile of PDFs, Office docs, and HTML into clean, chunked, embedded data for your vector database — in one call.
Pipeline stages
Input · Documents → chunked, embedded data
Your output will appear here
Choose a sample or upload a document, then run the extraction to inspect the result.
Run to extract a real document through the live API — no account, no setup.
Drop-in integrations for the AI frameworks and tools you already build with.
from langchain_xberg import XbergLoader# Load any document as LangChain Documentsloader = XbergLoader("report.pdf")docs = loader.load()# Drop straight into your vector storedb.add_documents(docs)
Install Xberg in your stack and call the same extraction API — from Python to Rust to the browser.
Python, TypeScript, Rust, Go, Java, Ruby, and more.
Functions, classes, imports, symbols — all parsed.
Run our images or the single-binary CLI.
Free for personal, internal, and commercial use.
All primitives, in 14 languages. Embed in your app or self-host the extraction server. Use any subset of the pipeline.
Get the librariesDirect roadmap input, early access to new pipeline stages, 24-month locked pricing. We work closely with a small cohort across commercial, non-profit, and research teams. Free-tier access available at our discretion in exchange for attribution.
Apply for partnershipProcess documents in milliseconds instead of seconds. Your RAG pipeline moves at the speed of API calls, not extraction bottlenecks. Index millions of documents without waiting weeks for processing to complete.
Effectively process large numbers of documents in bulk. Xberg is built for batch processing, and our cloud infrastructure is designed to scale.
Ultra-fast embeddings via a Rust-native ONNX engine. 4 presets out of the box, extensible to any model. No separate embedding pipeline needed.
Semantic chunking across code, markdown, and plain text. Token reduction, keyword extraction, and rich metadata — structured output ready for any AI pipeline.
Extract functions, classes, imports, and symbols from code files across 306 programming languages. Structured output, ready for semantic chunking and RAG pipelines.
Go beyond extraction. Use vision language models as an OCR backend, extract structured JSON from documents using a schema, and generate embeddings — all via 143 LLM providers, including local models with zero API key configuration.
Textract, Document AI and Azure DI stop at OCR — and send your files to their cloud. Xberg runs the whole pipeline, open source, on your own infrastructure.
| Capability | Xberg | Google Document AI | Amazon Textract | Azure Doc Intelligence |
|---|---|---|---|---|
| Deployment | Self-host or managed | Cloud only (GCP) | Cloud only (AWS) | Cloud only (Azure) |
| Source model | Open source | Proprietary | Proprietary | Proprietary |
| Pipeline scope | Acquire → retrieve | Extract / OCR | Extract / OCR | Extract / OCR |
| Data residency | Stays in your infra | Sent to vendor | Sent to vendor | Sent to vendor |
| Formats | 96 + 306 code langs | Docs / forms | PDF / images | Docs / forms |
| Code intelligence | Built in | — | — | — |
| Pricing | Usage or self-host | Per page | Per page | Per page |
Want the numbers — latency, throughput, accuracy? Every figure is reproducible from our open-source harness.
See the head-to-head benchmarksFeed your vector database with semantically accurate document chunks. Preserve table structure so your AI understands relationships. Bulk-process your knowledge base in hours instead of weeks.
Extract content, metadata, and structure to power smart routing. Automatically categorize incoming documents. Reduce manual sorting and classification workflows.
Extract and structure compliance documents, contracts, and regulatory filings. Preserve table relationships and metadata for audit trails. Support for scanned documents with OCR means nothing falls through the cracks.
See how teams put document intelligence to work. Explore all use cases
Our first design partners are onboarding now. Want your logo here? Apply for partnership
Xberg uses cookies to improve your experience, personalize content, and analyze traffic. You can manage your preferences at any time.