OCR to Markdown: How to Convert Scanned Documents and Images
Updated: September 20, 2026
If you cannot select the words on a page, you may be looking at an image of the document. Turning those pixels into text requires optical character recognition, or OCR. The output can be organized as Markdown with paragraphs, headings, lists and tables.
OCR interprets what it sees. It can mistake a zero for a letter, miss a line or read columns in the wrong order. A readable source and a review of the output are part of the workflow, even when the text looks neatly formatted.
Which Files You Can Convert
On the file2markdown website, you can upload JPG and PNG images for text extraction within your plan's limits. Scanned-PDF OCR requires Pro and processes up to 50 pages per file.
The free allowance is 5 files daily without an account or 15 with a free account, up to 25 MB each. Pro supports 100 MB files and batches of 10. OCR through URL or MCP access requires Pro; those conditions differ from uploading an image directly on the website.
Convert an Image or Scan
- Open the converter.
- Upload a JPG, PNG or scanned PDF supported by your plan.
- Convert it and compare the Markdown with the original image.
- Correct any errors before copying or downloading the text for use.
For PDFs, see the PDF conversion guide. A file may mix text pages and scans. Check that both appear in the output: initial detection uses a sample and does not guarantee recovery of every page in a mixed document.
Prepare a Readable Source
Straighten the page, avoid glare and check that small letters are clear. A blurred photo may not contain enough information to recover every word reliably. Enlarging it later does not restore missing details.
When cropping an image, keep headings, units, notes and table edges. A number without its column header can mean something different. For multi-page documents, maintain the order and check that no pages are missing.
Handwriting, overlapping stamps and low-contrast backgrounds can be difficult to interpret. A plausible sentence is not proof that the image was read correctly.
Review the Markdown
Check names, dates, amounts, minus signs and decimal separators. Pay attention to accented characters in the source language. The interface language does not translate the document.
This table illustrates output structure; it is not an accuracy measurement:
# Example Invoice
| Description | Quantity | Amount |
|---|---|---|
| Service | 2 | EUR 40.00 |
Make sure each value is under the correct header. Keep notes that explain taxes, units and time periods. If a complex table remains unclear, use the original for that part of your work.
OCR and Jev Have Different Jobs
OCR extracts text from images. Quality checks inspect the text after conversion. Every conversion includes basic checks, which can detect some anomalies and flag OCR pages cut short by the model's output limit.
For PDFs, you can enable the optional Jev check. It examines excerpts for possible problems. It does not reread the image, correct the text or compare it against the original. The option does not apply to standalone JPG or PNG uploads.
How the Data Is Processed
OCR uses Claude to interpret images or scanned pages. file2markdown processes files in memory and does not store them. That does not mean everything happens in your browser or that no external provider is involved.
Jev is a separate option, off by default. Enabling it for a PDF sends excerpts of the converted text to TypeSafe. The privacy policy describes the processing.
Frequently Asked Questions
Does every PDF need OCR?
No. A usable text layer can be extracted directly. PDFs containing only images need visual recognition.
Does OCR always preserve tables?
No. It can recover structure, but complex tables, merged cells and blurry sources need comparison with the original.
Can I convert a photo for free?
Yes. You can upload JPG or PNG directly on the website within the free limits. Scanned-PDF OCR and OCR through URL or MCP require Pro.
The Markdown Memo
A fortnightly note for lawyers, researchers, accountants, and anyone else drowning in PDFs, scans, and decks. No spam.