Document to Markdown conversion has become a critical foundation for RAG, AI agents, and semantic search. Two open source tools dominate the conversation right now, Marker from Datalab and AnyDoc from Firecrawl. Both are free and both produce clean Markdown, but they were born from very different philosophies.
Marker is a Python pipeline built around a Vision-Language Model or VLM, focused on complex PDFs and maximum accuracy. AnyDoc is a pure Rust library with no AI model at all, built for extreme speed and broad format coverage. So the choice comes down to your document mix.
Quick Summary
| Aspect | Marker | AnyDoc |
|---|---|---|
| Core language | Python | Rust with Node.js, Python, and WASM bindings |
| Approach | Surya VLM plus CPU heuristics | Pure parser with no ML |
| Supported formats | PDF, images, PPTX, DOCX, XLSX, HTML, and EPUB | DOC, DOCX, PPT, PPTX, XLS, XLSX, ODT, ODS, ODP, RTF, EPUB, CSV, and PDF |
| Speed | Around 2.9 pages per second in balanced mode on a B200 GPU | Median around 4.4 ms per document |
| GPU requirement | Recommended for high accuracy mode | Not needed at all |
| OCR for scanned PDFs | Yes through the Surya VLM | No local OCR, text based PDFs only |
| License | Apache 2.0 for code and Open RAIL-M for models | MIT |
Numbers above come from each repository README as of September 2026. Each project runs its own benchmark, so expect some framing bias. Marker is measured with the third party olmocr-bench, while AnyDoc is measured with an internal benchmark that uses an LLM judge on 100 real world documents.
Approach and How They Work
Marker is built on top of Surya, a document VLM served through a local inference server. The pipeline starts by extracting embedded text with pdftext in reading order, then detects page layout, then decides per page whether the embedded text is usable or needs full re-OCR. In balanced mode, layout detection uses the VLM, which is more accurate but needs a GPU. In fast mode, layout uses the lightweight rf-detr detector on CPU, and the VLM is only called for equations, broken blocks, or pages that look scanned. Tables are reconstructed from the PDF text layer with CPU heuristics, with a VLM fallback when confidence is low.
AnyDoc takes the opposite path. Each format gets its own parser that maps into one shared document model, then a single serializer renders everything as GitHub-Flavored Markdown. Tables, headings, footnotes, numbered lists, and escaping therefore behave consistently whether the input was a legacy .doc file or a modern .pptx file. PDFs are handled locally through pdf-inspector with no external service, and OMML equations from Word and PowerPoint plus MathML from OpenDocument and EPUB become LaTeX with $ inline and $$ blocks. The selective approach makes Marker faster than full page VLM systems, but still far heavier than a non-AI parser like AnyDoc.
Accuracy and Output Quality
For difficult PDFs, Marker is built to handle tables, math formulas, multi-column layouts, and scanned documents. On olmocr-bench with 1,403 PDFs, balanced mode scores 76.0 percent overall and 83.5 percent on born-digital PDFs, ahead of MinerU and Docling in an apples to apples pipeline comparison.
AnyDoc is measured differently. Its internal benchmark covers 100 real world documents, with Claude Sonnet 5 judging output against LibreOffice renders. AnyDoc scores 81 and leads on every format against LibreOffice, Unstructured, MarkItDown, Pandoc, Docling, and Mammoth, with per format scores in the 72 to 88 range. The two numbers cannot be compared directly because they measure different things, one heavy scientific PDFs and the other mixed office documents.
Marker goes further for RAG. Beyond Markdown it emits JSON as a full tree with bounding boxes, HTML, and a flat chunks format. Images, tables, LaTeX equations, code blocks, footnotes, and tables of contents come along. The --use_llm flag connects Marker to Gemini, Claude, OpenAI, or local models through Ollama, so cross-page tables get merged, inline math gets cleaned up, and form values get extracted. AnyDoc has no such path. All of its conversion is deterministic parser logic, so broken tables or missing formulas cannot be repaired the way Marker hybrid mode can. For tricky multi-column papers, inline math inside PDFs, and nested forms, AnyDoc falls behind the visual approach of Marker.
Speed and Compute Requirements
AnyDoc is pure Rust with no ML models or external services, so median conversion time is under 5 ms per document. Pipelines that process thousands of office files per day suit this path. Marker on a single B200 host only reaches sustained throughput around 2.9 pages per second in balanced mode and around 7.4 pages per second in fast mode, while no-OCR mode reaches around 23.7 pages per second with no ability to read equations or scanned pages.
The best quality Marker mode ideally needs an NVIDIA GPU with Docker and the NVIDIA Container Toolkit, while on CPU or Apple Silicon the inference server runs through llama-server from llama.cpp, which drops throughput sharply. Minimum requirements are Python 3.10 plus PyTorch, a vLLM or llama.cpp backend, and settings such as SURYA_INFERENCE_URL and SURYA_INFERENCE_BACKEND when pointing at an existing server. AnyDoc instead runs with no GPU and no Python server, Node.js conversion runs on the libuv thread pool without blocking the event loop, and the WASM build converts directly on the user device.
Format Coverage and File Detection
In its internal benchmark, AnyDoc is the only tool to cover 14 out of 14 tested formats, spanning Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF, including legacy variants such as .doc, .ppt, and .xls plus modern variants such as .xlsb and .docm. Format detection reads file content rather than the extension, from the PDF header, the RTF open group, OLE stream names, down to the ZIP package mimetype. Only CSV needs an explicit name because it has no binary marker, so archives with messy naming stay readable.
Marker supports PDF, images, PPTX, DOCX, XLSX, HTML, and EPUB, but development focus remains PDF. For high volume mixed office corpora, coverage is narrower than AnyDoc.
Handling Scanned Documents
Marker includes a complete OCR path through the Surya VLM, so scanned PDFs, document photos, and old archives remain processable. Its scores hold up on the old archive categories in olmocr-bench, while no-OCR mode drops to zero on scanned math.
AnyDoc cannot do this because it only reads the PDF text layer. Scanned PDFs or image-only pages fail with a NeedsOcr error. The --ocr hosted option sends the document to the paid Firecrawl Parse service, so data leaves the local machine. For born-digital archives this limitation hardly matters, but for stacks of scans it decides the choice.
Installation Distribution and Licensing
Standard installation uses pip install marker-pdf, with marker_single for one file and marker for a whole folder. Datalab backs it with a managed API, a batch service claimed to process over 1 billion pages per week, and an on-prem option for privacy sensitive workloads. The managed Chandra model sits one accuracy tier above Marker.
It ships as a CLI through npx @firecrawl/anydoc, as a Rust library, as Node.js and Python bindings, and even as WebAssembly that runs directly in the browser. The online demo converts files locally so files never leave your machine. One command, npx skills add firecrawl/anydoc, teaches coding agents such as Claude Code and Cursor to read any document they encounter, and that is why AnyDoc spread quickly in agent workflows.
Marker code is Apache 2.0 and free for commercial use, but model weights use an Open RAIL-M license that restricts commercial use above 5 million dollars in funding or revenue. Larger commercial teams should read the Datalab pricing page carefully. AnyDoc carries an MIT license with no revenue threshold, unlike the Marker model license.
Which One Should You Choose
Each tool fits a different scenario.
Choose Marker if your situation looks like this.
- Your documents are mostly scanned PDFs, photos, or old archives
- You process scientific papers with math formulas and complex cross-page tables
- You need JSON output with bounding boxes or RAG ready chunks
- You can provide a GPU and tolerate a more complex Python setup for maximum accuracy
Choose AnyDoc if your situation looks like this.
- You process high volumes of office files such as Word, Excel, PowerPoint, ODF, RTF, and CSV
- Your documents are mostly born-digital and text based
- You need very high speed with no GPU server
- You want browser or edge conversion without uploading files
- You need a lightweight pipeline for agents with no AI model dependency
Many teams use both in a complementary way. AnyDoc serves as the fast default path for office files and simple PDFs, while Marker serves as the specialized fallback for scanned or layout heavy PDFs.
Conclusion
For a mixed archive, start with AnyDoc because it is cheap and fast. Route failed or low quality files to Marker. Easy files finish in milliseconds, difficult files still get read.





