html-to-markdown
Turn messy, real-world HTML into clean Markdown — from the language you already work in.
What and Why?
Feed html-to-markdown the HTML you actually have — unclosed tags, CDATA, custom elements, broken entities, nested tables, mixed encodings — and get back clean CommonMark (or Djot) without losing content. One convert() call does it, and it returns the same result whether you run it from Python, TypeScript, Go, Ruby, Java, or 11 more languages.
You get more than the text: pull page metadata (Open Graph, Twitter, JSON-LD) and structured tables in the same pass, or hook into the conversion to reshape the output. It is fast enough for whole-corpus jobs, and the messy-input handling is automatic — you never choose a parsing strategy or tune anything to get correct output.
Features
| Feature | Description |
|---|---|
| 16 languages, one Rust core | Rust, Python, Node.js, WASM, Java, Go, C#, PHP, Ruby, Elixir, R, Dart, Kotlin (Android), Swift, Zig, and a C ABI |
| Tiered dispatch | Byte scanner → DOM walker → html5ever repair, with byte-equal output across tiers |
| Real-HTML robust | Unclosed tags, CDATA, custom elements, malformed entities, nested tables, mixed encodings — handled without losing content |
| GFM tables | Padded cells, alignment, and pipe escaping |
| Djot output | Set output_format = "djot" to emit Djot instead of Markdown |
| Metadata extraction | Parse <head> into structured metadata (Open Graph, Twitter, JSON-LD, microdata, RDFa, header hierarchy) |
| Inline images | Opt-in mirroring of data URIs and remote image references |
| Visitor API | Feature-gated traversal to transform the converted Markdown AST |
| Configurable preprocessing | Standard, strict, and lenient presets — or build your own |
| Fast | 19–116 MB/s on the Wikipedia/mdream corpus; per-group regression thresholds enforced on every PR |
⭐ Star this repo to show your support — it helps others discover html-to-markdown.
Quick Start
convert() is the single entry point — it returns a structured result with content, warnings, and optional metadata.
Language Packages
cargo add html-to-markdown-rsSee Rust README for full documentation.
pip install html-to-markdownSee Python README for full documentation.
npm install @xberg-io/html-to-markdownSee Node.js README for full documentation.
go get github.com/xberg-io/html-to-markdown/packages/go/v3See Go README for full documentation.
Available on Maven Central as io.xberg:html-to-markdown. See Java README for the dependency snippet and current version.
dotnet add package XbergIo.HtmlToMarkdownSee C# README for full documentation.
gem install html-to-markdownSee Ruby README for full documentation.
This is a native PHP extension (Rust ext-php-rs), so install it with PIE — not composer require:
pie install xberg-io/html-to-markdownSee PHP README for full documentation.
Add {:html_to_markdown, "~> 3.6"} to your mix.exs dependencies. See Elixir README for full documentation.
install.packages("htmltomarkdown", repos = "https://xberg-io.r-universe.dev")See R README for full documentation.
dart pub add h2mSee Dart README for full documentation.
Available on Maven Central as io.xberg:html-to-markdown-android. See Kotlin README for the dependency snippet and current version.
Add via Swift Package Manager. See Swift README for full documentation.
See Zig README for installation and usage.
npm install @xberg-io/html-to-markdown-wasmSee WebAssembly README for full documentation.
Pre-built .so / .dll / .dylib from GitHub Releases. See FFI crate for full documentation.
cargo install html-to-markdown-cli# Or install the prebuilt binary via cargo-binstall:
cargo binstall html-to-markdown-clibrew install xberg-io/tap/html-to-markdownSee CLI usage for full documentation.
AI Coding Assistants
Install the html-to-markdown plugin from xberg-io/html-to-markdown. It ships the html-to-markdown agent skills and works with every major coding agent — expand your harness below.
/plugin marketplace add xberg-io/html-to-markdown
/plugin install html-to-markdown@html-to-markdown
/plugins add https://github.com/xberg-io/html-to-markdown
Then search for html-to-markdown and select Install Plugin.
Settings → Plugins → Add from URL → https://github.com/xberg-io/html-to-markdown, then select html-to-markdown.
gemini extensions install https://github.com/xberg-io/html-to-markdown
droid plugin marketplace add https://github.com/xberg-io/html-to-markdown
droid plugin install html-to-markdown@html-to-markdown
copilot plugin marketplace add https://github.com/xberg-io/html-to-markdown
copilot plugin install html-to-markdown@html-to-markdown
Add the package to opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"plugin": ["@xberg-io/opencode-html-to-markdown"]
}Documentation
Full guides, the convert() API for every binding, tier architecture, the metadata and visitor APIs, and performance benchmarks live at docs.html-to-markdown.xberg.io.
Part of Xberg.io
- Xberg — the open-source content-intelligence engine: text, tables, and metadata from 101 formats (115 file extensions), with OCR, transcription, and code intelligence. MIT.
- Xberg Pro — a complete self-hosted content-intelligence backend in a single container. Commercial.
- Xberg Enterprise — the distributed, governed content-intelligence platform, scaled on Kubernetes with team governance and support. Commercial.
- crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- html-to-markdown — fast, lossless HTML→Markdown engine.
- liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
- tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
- alef — the polyglot binding generator that produces every per-language binding across the 5 polyglot repos.
Contributing
Contributions welcome! See CONTRIBUTING.md for setup instructions and guidelines.
License
MIT License — see LICENSE for details.