A worked example, rendered from real sample data. Sign in to run the tool on your own input.
image: invoice-scan-p1.png
bytes: 500
data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAoAAAADICAAAAABdMg/XAAAACXBIWXMAAC4jAAAuIwF4pT92AAABpklEQVR42u3SoREAMAwDsQzR/bcsDywI7hlEgoa+rwtB5QIEiABBgAgQBIgAQYAIEASIAEGACBAEiABBgAgQBIgAQYAIEASIAEGACBAEiABBgAgQBIgAQYAIEAGCABEgCBABggARIMQDPM9cjPtGARoFaBSgAI0CNApQgEYBGgUoQKMAjQIUoFGARgEK0ChAowAFaBSgUYACNArQKEABGgVoFKAAjQI0ClCARgEaBShAowCNAhSgUYBGAQrQKECjAOErASJABAgCRIAgQAQIAkSAIEAECAJEgCBABAgCRIAgQAQIAkSAIEAECAJEgCBABAgCRIAgQAQIAkSACBAEiABBgAgQBIgAQYAIEASIAEGACBAEiABBgAgQBIgAQYAIEASIAEGACBAEiABBgAgQBIgAQYAIEAGCABEgCBABggARIAgQAYIAESAIEAGCABEgCBABggARIAgQAYIAESAIEAGCABEgCBABggARIAJ0AQJEgCBABAgCRIAgQAQIAkSAIEAECAJEgCBABAgCRIAgQAQIAkSAIEAECAJEgCBABAgCRIAgQOIaiOlfHOGllrsAAAAASUVORK5CYII=═══ What this tool did ═══
✗ No text was extracted from this image. There is no OCR engine in this application — not tesseract, not a cloud API, nothing. Any tool that returned you a paragraph of text here would be inventing it.
✓ What it did do is read the image header itself and work out whether the image is good enough to OCR, and with which settings.
═══ Image you attached ═══
File: invoice-scan-p1.png
Size on disk: 500 bytes
Format: PNG
Dimensions: 640 × 200 pixels (0.13 MP)
Pixel format: 8-bit greyscale
Resolution: 300 DPI (declared in the file)
═══ Is this image good enough to OCR? ═══
⚠ 200px on the short edge is small. A full page at 300 DPI is about 2550 × 3300 — this is well under that, so expect character errors in anything below about 12pt.
Implied resolution if this is a full A4 page: about 300 DPI.
✓ That is in the range tesseract is trained for.
ℹ Order that actually matters: greyscale first, then deskew, then a threshold. Sharpening a blurry scan adds ringing and makes OCR worse, not better.
═══ Commands that actually extract the text ═══
tesseract input.png output -l eng --psm 6 --dpi 300 && cat output.txt
tesseract input.png - -l eng+deu --psm 4 # two languages, one column per block
tesseract input.png out hocr # word boxes + per-word confidence
ocrmypdf --force-ocr --deskew --clean in.pdf out.pdf
pdftotext -layout in.pdf - # try this FIRST: a digital PDF needs no OCR at all
magick input.png -colorspace gray -resize 300% -unsharp 0x1 -threshold 60% prepped.png
shortcuts run 'Extra
…
Read an image's real header, get the right tesseract command for it, and clean up the OCR text you paste back — no text is extracted here. Part of the DevTools Surf developer suite. Browse more tools in the Images collection.