Skip to content

Image to Text (OCR)

Drop an image here, paste it, or choose a fileChoose an image from your device

PNG, JPG, WebP, GIF, BMP, TIFF · never leaves your device

About

Optical Character Recognition (OCR) turns the printed text in an image into text you can copy and edit. This tool runs Tesseract.js - the open-source Tesseract engine and its neural-network (LSTM) text recognizer, compiled to WebAssembly - entirely in your browser. It is built for printed and typed text; handwriting is generally not recognized reliably.

How to use

  1. 1.

    Pick the Language of the text in your image - English is the default

  2. 2.

    Choose an image, drop it onto the page, or paste a screenshot - PNG, JPG, WebP, GIF, BMP, or TIFF

  3. 3.

    Wait while the text is recognized - the Text panel header shows each stage and the progress

  4. 4.

    Copy the extracted text, or download it as a .txt file named after the image

Common uses

  • ▸

    Copying text out of a screenshot, slide or photo where it cannot be selected

  • ▸

    Digitizing printed receipts, invoices, or business cards

  • ▸

    Turning scanned book or magazine pages into editable text

  • ▸

    Reading text in other supported languages, such as Spanish, Chinese or Arabic

Similar tools

Frequently asked questions

What languages are supported?

Ten languages: English, Spanish, French, German, Portuguese, Italian, Chinese (Simplified), Japanese, Korean, and Arabic. One language is used at a time, so for an image that mixes languages, pick the main one. Changing the language re-reads the current image.


How accurate is the text extraction?

Accuracy is best on clear, high-resolution images of printed text with good contrast. Blurry or low-resolution photos, decorative fonts, and skewed or curved text reduce accuracy, and handwriting is generally not recognized reliably. Photos are turned upright using their orientation metadata before recognition.


Is my image sent to a server?

No. Recognition runs entirely in your browser using Tesseract.js, and your image never leaves your device. The only downloads are the OCR engine itself, served from this site, and the data for each language, which comes from the jsDelivr CDN the first time you use that language and is then stored in your browser (in IndexedDB). Neither request includes your image.


What image formats are supported?

PNG, JPEG, WebP, AVIF, GIF (first frame only), and BMP work in all modern browsers. Uncompressed and LZW-compressed TIFF files can be read as well; other TIFF variants and HEIC photos only work in browsers that can open them, such as Safari. Images larger than 3000 pixels on their long side are scaled down before recognition.


Why does processing take a while the first time?

The first run downloads the OCR engine (about 4 MB) and the data for the selected language (up to 3 MB). Both are kept by your browser, so later runs start much faster, and each additional language is downloaded only once. If a download fails, an error with a Retry button appears.