Xberg
The fast, precise document-intelligence engine — for every language.
Point Xberg at anything — a PDF, a scanned image, a spreadsheet, an audio file, a URL, a whole archive, or a source tree — and get back clean text, tables, metadata, and structured data. One engine handles format detection, reading, OCR, and extraction, so you never stitch a pipeline together from a dozen libraries.
100 formats · 120 file extensions · 371 code languages · 15 language bindings · 6 output formats · OCR · transcription · embeddings
The fastest, most precise open-source document and PDF-to-Markdown engine — see the benchmarks.
Install · What you get · Capabilities · CLI · Docs
Xberg is the next iteration of Kreuzberg. Same document-intelligence engine, rebuilt and rebranded under a fresh v1 line.
What you get
Point Xberg at anything — a PDF, a spreadsheet, a scanned image, an audio file, a URL, an archive, a source tree — and get back clean, structured content you can use right away. One core does the format detection, reading, and extraction, so you don't assemble a pipeline yourself. Call it from Rust, Python, Node.js, Go, Java, C#, Ruby, PHP, Elixir, Dart, Swift, Zig, WASM, Kotlin, or C FFI, and run it as a library, CLI tool, REST API, or MCP server.
| Capability | What you get |
|---|---|
| 100 document formats | PDFs, Office, images, HTML, email, e-books, scientific publications, structured data across 120 file extensions — intelligent MIME detection, streaming for multi-GB files. |
| URLs & the web | Point Xberg at an http(s) URL — it fetches and extracts a single document, or crawls and follows links (Auto / Document / Crawl modes via the crawlberg engine). Requires the url-ingestion feature.
|
| Audio & video transcription | Speech-to-text from MP3, M4A, WAV, WebM, and MP4 tracks via Whisper ONNX (tiny → large-v3). Requires the transcription feature.
|
| Archives, traversed | List and recursively extract nested .zip, .tar, .gz, .7z — documents inside documents — guarded by zip-bomb, compression-ratio, and nesting-depth limits. |
| OCR on demand | Tesseract, PaddleOCR, Candle, or VLM backends — fallback chains, confidence scores, language auto-detection, extensible via plugins. |
| Layout & tables | ML layout models (PP-DocLayout-V3, RT-DETR) and table structure (TATR, SLANet) reconstruct reading order and cell grids for clean Markdown. |
| Code intelligence | Functions, classes, imports, symbols, docstrings from 371 programming languages. Syntax-aware chunking for RAG pipelines. |
| Embeddings & search | Local (ONNX) or provider-hosted embeddings (165 providers via liter-llm), sparse and late-interaction, cross-encoder reranking. |
| Enrichment | NER, keyword extraction (YAKE/RAKE), summarization, translation, redaction, page classification, QR detection, language detection, token reduction (TOON). |
| Structured extraction | Schema-driven JSON straight from any document via local (Ollama, LM Studio, vLLM) or hosted LLMs — no prompt engineering. |
| 6 output formats | Plain text, Markdown, Djot, HTML, JSON tree, or Structured (same text as Plain, tagged with a structured metadata label). |
| Runs anywhere | Library, CLI (12 commands), REST API (xberg serve), MCP server, Docker, Helm — no GPU needed. Content-hash caching, parallel batch, per-file timeouts. |
Capabilities marked requires a feature are Cargo feature flags on the core crate (
url-ingestion,transcription,reranker, layout/ORT). Prebuilt language packages and the Docker image bundle the common set; a from-source build enables only what you select.
Installation
Language Packages
pip install xbergSee Python README for full documentation.
npm install @xberg-io/xbergSee Node.js README for full documentation.
cargo add xbergSee Rust README for full documentation.
go get github.com/xberg-io/xberg/packages/go@latest⚠️ The repository root is not a Go module —
go get github.com/xberg-io/xbergwill fail. Always target the/packages/gosubdirectory as shown above.
See Go README for full documentation.
Available on Maven Central as io.xberg:xberg. See Java README for the dependency snippet.
dotnet add package XbergSee C# README for full documentation.
gem install xbergSee Ruby README for full documentation.
composer require xberg-io/xbergSee PHP README for full documentation.
Add {:xberg, "~> 1.0"} to your mix.exs dependencies. See Elixir README for full documentation.
npm install @xberg-io/xberg-wasmSee WebAssembly README for full documentation.
Available on Maven Central as io.xberg:xberg-android. See Kotlin README for the dependency snippet.
Add via Swift Package Manager. See Swift README for full documentation.
dart pub add xbergSee Dart README for full documentation.
Add via zig fetch. See Zig README for full documentation.
Build from source as part of this workspace. See C (FFI) README for full documentation.
CLI & Deployment
brew install xberg-io/tap/xberg12 commands: extract, batch, detect, formats, version, cache (stats/clear/manifest/warm), serve, mcp, api, embed, chunk, completions.
See CLI usage guide for detailed documentation.
docker pull ghcr.io/xberg-io/xberg:latestRun in API, CLI, or MCP modes. See Docker guide for examples.
xberg serve --host 0.0.0.0 --port 8000One POST endpoint handles all formats. Returns JSON or Markdown. Stream large files. See API server guide.
xberg mcp --transport stdio9 tools (extract, extract_batch, detect_mime_type, cache_stats, list_formats, cache_clear, get_version, cache_manifest, cache_warm). 3 prompts (extract_document, extract_with_ocr, semantic_search). 4 resources (formats, models, OCR languages, embedding presets).
Add to Claude Desktop or Cursor:
{
"mcpServers": {
"xberg": { "command": "xberg", "args": ["mcp"] }
}
}AI Coding Assistants
Install the Xberg plugin from xberg-io/xberg. Ships extraction APIs, OCR backends, configuration, and language conventions.
/plugin marketplace add xberg-io/xberg
/plugin install xberg@xberg
/plugins add https://github.com/xberg-io/xberg
Search for xberg and select Install Plugin.
Settings → Plugins → Add from URL → https://github.com/xberg-io/xberg, then select xberg.
gemini extensions install https://github.com/xberg-io/xberg
droid plugin marketplace add https://github.com/xberg-io/xberg
droid plugin install xberg@xberg
copilot plugin marketplace add https://github.com/xberg-io/xberg
copilot plugin install xberg@xberg
Add to opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["@xberg-io/opencode-xberg"]
}Quick Start
Extract text from a document:
use xberg::{extract, ExtractInput, ExtractionConfig};
#[tokio::main]
async fn main() -> xberg::Result<()> {
let config = ExtractionConfig::default();
let output = extract(
ExtractInput::from_uri("document.pdf"),
&config
).await?;
println!("{}", output.results[0].content);
Ok(())
}Common use cases — see Quick start guide for language-specific examples, OCR, batch processing, and API configuration.
Capabilities
Supported File Formats (100 formats · 120 file extensions)
100 formats across 120 file extensions in 8 major categories with intelligent format detection and comprehensive metadata extraction.
Office Documents
| Category | Formats | Capabilities |
|---|---|---|
| Word Processing |
.docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6
|
Full text, tables, images, metadata, styles |
| Spreadsheets |
.xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers
|
Sheet data, formulas, cell metadata, charts |
| Presentations |
.pptx, .pptm, .ppt, .ppsx, .potx, .potm, .pot, .odp, .key
|
Slides, speaker notes, images, metadata |
.pdf |
Text, tables, images, metadata, OCR support | |
| eBooks |
.epub, .fb2
|
Chapters, metadata, embedded resources |
| Database | .dbf |
Table data extraction, field type support |
| Hangul |
.hwp, .hwpx
|
Korean document format, text extraction |
Images (OCR-Enabled)
| Category | Formats | Features |
|---|---|---|
| Raster |
.png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif
|
OCR, table detection, EXIF metadata, dimensions, color space |
| Advanced |
.jp2, .jpx, .jpm, .mj2, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm
|
OCR via pure-Rust JPEG2000 decoder, JBIG2 support, table detection |
| HEIC family |
.heic, .heics, .heif, .avif, .avcs
|
EXIF metadata, optional pixel decoding |
| Vector | .svg |
DOM parsing, embedded text, graphics metadata |
Audio & Video
| Category | Formats | Features |
|---|---|---|
| Audio |
.mp3, .mpga, .m4a, .wav, .webm
|
Whisper transcription |
| Video audio track |
.mp4, .mpeg, .webm
|
Audio-track transcription only |
Web & Data
| Category | Formats | Features |
|---|---|---|
| Markup |
.html, .htm, .xhtml, .xml, .svg
|
DOM parsing, metadata (Open Graph, Twitter Card), link extraction |
| Structured Data |
.json, .yaml, .yml, .toml, .csv, .tsv
|
Schema detection, nested structures, validation |
| Text & Markdown |
.txt, .md, .markdown, .djot, .mdx, .rst, .org, .rtf
|
CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode |
Email & Archives
| Category | Formats | Features |
|---|---|---|
.eml, .msg, .pst
|
Headers, body (HTML/plain), attachments, threading | |
| Archives |
.zip, .tar, .tgz, .gz, .7z
|
File listing, nested archives, metadata, recursive extraction |
Academic & Scientific
| Category | Formats | Features |
|---|---|---|
| Citations |
.bib, .ris, .nbib, .enw
|
Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX |
| Scientific |
.tex, .latex, .typ, .typst, .jats, .ipynb
|
LaTeX, Typst, Jupyter notebooks, PubMed JATS |
| Publishing |
.fb2, .docbook, .dbk, .docbook4, .docbook5, .opml
|
FictionBook, DocBook XML, OPML outlines |
Code Intelligence (371 Languages)
Extract structure from 371 programming languages via tree-sitter:
| Feature | Description |
|---|---|
| Structure Extraction | Functions, classes, methods, structs, interfaces, enums |
| Import/Export Analysis | Module dependencies, re-exports, wildcard imports |
| Symbol Extraction | Variables, constants, type aliases, properties |
| Docstring Parsing | Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats |
| Syntax-Aware Chunking | Split code by semantic boundaries for RAG pipelines |
| Diagnostics | Parse errors with line/column positions |
Powered by tree-sitter-language-pack.
Output Formats (6)
| Format | Use case | Example |
|---|---|---|
| Plain | Raw text, no markup | "Chapter 1\nIntroduction" |
| Markdown | Readable, structured, RAG-friendly | "# Chapter 1\n## Introduction" |
| Djot | Modern lightweight markup | Similar to Markdown but stricter |
| HTML | Styled, browser-ready | <h1>Chapter 1</h1> |
| JSON | Machine-readable tree structure | Hierarchical sections with heading levels |
| Structured | Same plain-text content as Plain, distinguished only by its structured output-format metadata label |
Byte-identical to Plain output |
Deployment Modes
| Mode | Command | Transport | Use case |
|---|---|---|---|
| Library | xberg::extract() |
Async functions | Embed in your application |
| CLI | xberg extract document.pdf |
12 commands | Scripts, batch jobs, CI/CD |
| REST API | xberg serve |
HTTP POST | Microservice, serverless deployment |
| MCP Server | xberg mcp |
stdio or HTTP | Claude, Cursor, IDE agents |
| Docker | docker run ghcr.io/xberg-io/xberg |
All modes | Container deployment |
OCR Backends
- Tesseract — Native C FFI (Linux/macOS/Windows) and WASM (browser)
- PaddleOCR — ONNX Runtime, mobile-optimized models
- Candle — Pure Rust, CPU-only, lightweight
- VLM — GPT-4 Vision, Claude Vision, Gemini Vision, or 165 providers via liter-llm
Fallback chains. Extensible via plugin system.
Embeddings
Local (ONNX Runtime):
- Preset models: fast, balanced (default), quality, multilingual
- Dimensions: 384, 768, 1024
Provider-hosted:
- OpenAI, Anthropic, Google, Hugging Face, Mistral, Cohere, and 165 providers total
- Via liter-llm integration
Reranking:
- Local ONNX rerankers (cross-encoder models)
- Provider-hosted: Cohere Rerank, others
Structured LLM Extraction
Local engines: Ollama, LM Studio, vLLM
Remote: OpenAI, Anthropic, Google, Mistral, Cohere, and 165 providers via liter-llm
Schema validation. Temperature, top-p, frequency penalty tuning.
Enrichment
- NER — GLiNER or LLM-based entity recognition
- Redaction — Mask PII (phone, email, SSN, credit card, addresses)
- Summarization — Document and section summaries via LLM
- Translation — Multi-language via LLM
- Page Classification — Tag document pages (cover, toc, content, etc.)
- QR Code Detection — Extract and decode QR codes from images
- Keyword Extraction — YAKE or RAKE algorithms
- Language Detection — Detect document language
- Layout Detection — RT-DETR + TATR models for document structure
- Table Extraction — Cell-level structure and content
- Token Reduction — TOON wire format (~30–50% fewer tokens than JSON)
CLI Reference
| Command | Subcommands | Purpose |
|---|---|---|
extract |
— | Extract text from a single document (path, URL, or stdin) |
batch |
— | Extract from multiple documents in parallel |
detect |
— | Identify MIME type of a file |
formats |
— | List all supported formats and MIME types |
version |
— | Show Xberg version |
cache |
stats, clear, manifest, warm
|
Manage extraction cache and models |
serve |
— | Start REST API server (default: http://127.0.0.1:8000) |
mcp |
— | Start MCP server (stdio or HTTP transport) |
api |
schema |
Output OpenAPI 3.1 specification |
embed |
— | Generate embeddings for text (local or provider-hosted) |
chunk |
— | Split text into chunks (text, markdown, YAML, or semantic) |
completions |
— | Generate shell completion scripts |
Run xberg --help or xberg <command> --help for detailed options.
Documentation
Full guides, API references for every binding, format reference, and configuration docs live at xberg.io.
- Getting Started
- Quick Start
- Guides
-
API Reference (Rust core) — every binding has its own page under
/reference/ - Format Reference
- Live Demo (browser, WASM)
Built with Xberg
Projects that declare Xberg as a dependency. Xberg was previously published as kreuzberg, and most of these projects declare the package under that name.
| Project | What it is | Stars |
|---|---|---|
| basemind | AI context and content layer for coding agents over one MCP server: code map, document RAG, shared memory and web crawl | |
| delulu | A suite of MCP servers and CLI tools that give your LLM better search and fewer hallucinations | |
| docs-mcp-server | Grounded documentation MCP server, an open-source alternative to Context7, Nia and Ref.Tools | |
| erato | The open-source AI platform | |
| fastmail-cli | CLI and MCP server for Fastmail: email, contacts, masked email, attachments and text extraction | |
| ghfdb-portal | Web portal for the Global Heat Flow Database | |
| hawki-toolkit-file-converter | Prepares and converts PDF files for the HAWKI toolkit | |
| haystack-core-integrations | Integrations that extend Haystack with extra components and document stores | |
| kreuzakt | A search engine for humans and computers, aimed at your most boring documents | |
| lilbee | The whole local AI stack in one executable, with conversational search and cited answers over your files, code and the web | |
| llm-workflow-engine | Power CLI and workflow manager for LLMs | |
| MANSPIDER | Spiders entire networks for files sitting on SMB shares, searching filenames or contents with regex | |
| otoroshi-llm-extension | Connect, secure and manage LLM models behind one OpenAI-compatible API | |
| sift-kg | Turns a collection of documents into a knowledge graph, extracting entities and relationships with an LLM | |
| sirchmunk | Turns raw data into a self-evolving, real-time search and intelligence layer | |
| support-chatbot | Level-1 support chatbot for the Netherlands Red Cross 510 team |
Using Xberg in your project? Open a PR adding it to this list.
Contributing
Contributions are welcome! See CONTRIBUTING.md for guidelines.
Join our Discord community for questions and discussion.
Part of Xberg.dev
Xberg is one of six open-source projects from Kreuzberg, Inc.:
- Xberg — document intelligence: text, tables, metadata from 100 formats with optional OCR.
- Xberg Enterprise — managed extraction API with SDKs, dashboards, and observability.
- crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- html-to-markdown — fast, lossless HTML→Markdown engine.
- liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
- tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
- alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.
License
MIT License (MIT) — see LICENSE for details.