#OCR, PDFs, and images
PDFs and image-heavy documents are the hardest inputs to convert reliably. They may contain selectable text, scanned page images, embedded screenshots, charts, image-only tables, or mixed content.
#Basic PDF conversion
For simple born-digital PDFs:
make-markdown-library make sources -o library.md --converter autoAuto mode tries MarkItDown first and can fall back to LiteParse if output is empty.
#Scanned or complex PDFs
For scanned, layout-heavy, or table-heavy PDFs:
make-markdown-library make sources -o library.md \
--converter auto \
--liteparse-complexity-checkThis allows the tool to prefer LiteParse when a PDF looks complex or OCR-heavy.
#LiteParse options
make-markdown-library make sources -o library.md \
--converter hybrid \
--liteparse-image-mode placeholder \
--liteparse-ocr-language eng \
--liteparse-dpi 200Supported option flags:
--liteparse-image-mode off|placeholder|markdown|base64
--liteparse-no-links
--liteparse-no-ocr
--liteparse-ocr-language eng
--liteparse-target-pages 1,2,5-8
--liteparse-dpi 150
--liteparse-max-pages 50
--liteparse-password PASSWORD
--liteparse-complexity-check#Images in PDFs and Office files
Image-containing files can mean different things:
| Image type | Desired behaviour |
|---|---|
| Logo/photo | Preserve a note or placeholder. |
| Image containing text | OCR the image where possible. |
| Diagram/chart/screenshot | Preserve visual context or at least record that image handling was needed. |
| Scanned page | Use OCR/layout-aware parsing. |
#What gets recorded
The index records converter options and complexity metadata so you can audit why a document used MarkItDown or LiteParse.
{
"complexity": {
"checked": true,
"complex": true,
"reason": "ocr_required"
},
"converter_options": {
"image_mode": "placeholder",
"ocr_language": "eng",
"dpi": 200
}
}#Windows OCR prerequisites
For scanned PDFs or images containing text, install Tesseract OCR and confirm it with doctor. See Windows prerequisites.
#Practical recommendation
Use MarkItDown for structured Office/HTML/CSV material. Use LiteParse or hybrid routing for scanned PDFs, layout-sensitive PDFs, and image-heavy inputs where OCR or spatial layout matters.