Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
MinerU — High-accuracy document parsing engine for LLM · RAG · Agent workflows
Converts PDF · DOCX · PPTX · XLSX · Images · Web pages into structured Markdown / JSON · VLM+OCR dual engine · 109 languagesMCP Server · LangChain / Dify / FastGPT native integration · 10+ domestic AI chip support
🔍 Core Parsing Capabilities
- Native support for
DOCX,PPTX, andXLSXparsing - Formulas → LaTeX · Tables → HTML, accurate layout reconstruction
- Supports scanned docs, handwriting, multi-column layouts, cross-page table merging
- Output follows human reading order with automatic header/footer removal
- VLM + OCR dual engine, 109-language OCR recognition
🔌 Integration
| Use Case | Solution |
|---|---|
| AI Coding Tools | MCP Server — Cursor · Claude Desktop · Windsurf |
| RAG Frameworks | LangChain · LlamaIndex · RAGFlow · RAG-Anything · Flowise · Dify · FastGPT |
| Development | Python / Go / TypeScript SDK · CLI · REST API · Docker |
| No-Code | mineru.net online · Gradio WebUI · Desktop client |
🖥️ Deployment (Private · Fully Offline)
| Inference Backend | Best For |
|---|---|
| pipeline | Fast & stable, no hallucination, runs on CPU or GPU |
| vlm-engine | High accuracy, supports vLLM / LMDeploy / mlx ecosystem |
| hybrid-engine | High accuracy, native text extraction, low hallucination |
Domestic AI chips: Ascend · Cambricon · Enflame · MetaX · Moore Threads · Kunlunxin · Iluvatar · Hygon · Biren · T-Head
from https://github.com/opendatalab/MinerU
( https://opendatalab.github.io/MinerU/quick_start/)
No comments:
Post a Comment