file2markdown

How one associate fed 340 scanned PDFs to AI without burning the token budget

July 2, 2026

How one associate fed 340 scanned PDFs to AI without burning the token budget

The Markdown Memo header

3 min read.

It was 11:47 p.m. on a Tuesday when Marcus Hale, a second year associate at a mid sized litigation firm in Chicago, got the email he had been dreading: 340 scanned PDFs, discovery review by Monday.

His supervising partner, Eleanor Voss, had forwarded a compressed archive: depositions, contracts, correspondence, and internal memos, all part of a commercial dispute. Her note was short. "Need a full summary of relevant communications by Friday. Flag anything mentioning the Delray shipment." Marcus opened the first file, a 47 page scanned fax with a footer photocopied so many times it had gone soft and grey, pasted a passage into the AI assistant the firm had licensed that quarter, and watched the token counter climb faster than he liked. He did the math. At this rate, the archive would burn through his team's entire monthly allocation before he got past the first fifty documents.

He reframed the problem. The PDFs weren't the obstacle, the weight of them was: embedded image layers, scan artifacts, and formatting clutter the model had to wade through before reaching a single useful sentence. Three things changed once he saw it that way.

Step 1: Separate the content from the container.

A scanned PDF isn't text. It's a photograph of text, wrapped in a file format built for printing, not for reading by a model. Markdown strips that down to what's actually there: headings, paragraphs, tables. Nothing else.

Step 2: Convert before you analyze, not while you analyze.

Marcus ran the full batch through file2markdown.ai before touching the AI assistant again. The tool read the scans, pulled the noise out, and handed back clean Markdown, with document titles preserved as headings and columnar data preserved as actual tables instead of scrambled rows of text.

Step 3: Feed the model the clean version, not the raw one.

The difference showed up immediately. Answers came back faster and sharper, and token usage dropped by roughly 60 percent compared to working from the raw PDFs, simply because the model was no longer spending its budget parsing scan noise before it could even start reading. By Thursday afternoon, Marcus had a flagged, cross referenced summary sitting in Eleanor's inbox, sixteen hours ahead of the Friday deadline. Her reply was two words. "Good work." Nobody on the team asked how he'd moved that fast. They assumed he'd worked through the night. He hadn't. He'd just stopped asking the AI to do a job it was never suited for: reading scanned paper directly.

If you've got your own pile of scanned PDFs waiting for an AI assistant to choke on, that's worth fixing before discovery review starts, not during it.

Robin

P.S. Next issue: a researcher feeds an entire literature review into Obsidian and NotebookLM. What works, what breaks, and the one step she almost skipped. Lands in your inbox in two weeks.

What kind of work are you trying to make AI useful on?

The Markdown Memo

A fortnightly note for lawyers, researchers, accountants, and anyone else drowning in PDFs, scans, and decks. No spam.