@llama_index
If you've ever worked in or around legal, you know that discovery is where document parsing really gets stress-tested. Low-resolution scans. Black and white images. Handwritten annotations. Charts buried in reports. Files that are technically PDFs but practically unreadable. And hundreds of thousands of them. Traditional OCR tools struggle with degraded scans, and anything visual (photographs, slide decks, tables) falls through the cracks entirely. That means your search index is noisy, your recall suffers, and relevant documents go unfound. This blog by @tuanacelik walks through how to set up LlamaParse for a legal discovery use-case: handling difficult scans with vision models, surfacing image and chart content, and using custom parsing instructions to guide output for predictable document patterns. The quality of everything downstream depends on what happened at ingestion. Worth getting right. Read the full blog here: https://t.co/MkUWjaJzSm