Before you start
- Decide whether embedded text is sufficient or OCR is needed for scanned pages.
- Keep the source PDF so that uncertain passages can be compared later.
PDF conversion
Convert text PDFs and scanned PDFs into Markdown with optional advanced text recognition text recognition.
Start with the original PDF and check the Markdown against difficult pages before publishing or reusing its text.
The PDF and generated files are temporary server files. Default cleanup is 30 minutes and conversion is stopped after eight minutes.
Most server file tools have a default 30-minute retention period. Expired files are removed by scheduled cleanup, so this is not an exact deletion timestamp.
How to Troubleshoot Document Conversion: Formats, Fonts, Encryption, and PagesKnow the limit: Markdown preserves useful text structure, not the exact visual layout of every PDF page.
Turn normal PDFs and scanned PDFs into Markdown files, with optional OCR recognition when the PDF has no usable text layer.
Upload a PDF, choose automatic OCR, forced OCR, or text-layer-only conversion, and MV Tools creates a Markdown file plus TXT and structured JSON downloads. Normal PDFs use embedded text for speed, while scanned PDFs can be rendered and recognized with OCR.
Uploaded PDFs, rendered page images, Markdown files, TXT files, and JSON results are processed temporarily on the server and are cleaned up automatically after the retention window.
Yes. Auto mode falls back to OCR when the PDF does not contain enough extractable text.
No. The first version focuses on readable text, page sections, paragraphs, simple headings, and lists. Complex tables and multi-column layouts may need manual cleanup.
Yes. Choose text-layer-only mode to extract embedded text without rendering pages for OCR.