Optical character recognition turns a picture of text back into text. A scan, a photograph of a page or a PDF made from scans contains no characters at all — only shapes that a human eye resolves into letters — and until something reads those shapes the file cannot be searched, edited or reflowed.
This scanner takes one file at a time, up to fifty megabytes, as a PDF or as a PNG, JPEG, TIFF, BMP or GIF. You pick the language before scanning, from a list of more than thirty, and for a PDF you can give a start and end page so a long document does not have to be processed in full. The recognition itself runs on the server rather than in your browser, which means the file is uploaded.
The engine is trained on printed type, and that single fact predicts nearly all of your results. Clean typeset pages come back well. Faint photocopies, skewed phone photographs, decorative faces, tables and multi-column layouts come back progressively worse, and handwriting is a different problem that engines of this kind are not built for.
It is uploaded. Recognition happens on our server, which is why large PDFs do not depend on your device's speed. The file is written to disk there and is not deleted automatically once scanning finishes, so avoid uploading anything confidential.
No. Scanning is free and no credits are deducted, so you can run the same page repeatedly at different settings without cost. That makes experimenting with the language selection and page range the sensible approach rather than trying to get it right first time.
Not reliably. The engine is trained to match the consistent letterforms of printed type, and handwriting varies in slant, spacing and shape in ways that break those assumptions. Handwritten pages are best transcribed by hand; the output you would get here would take longer to correct than to retype.