Tesseract icon
Tesseract icon

Tesseract

Tesseract.js is a javascript library that gets words in almost any language out of images.

Tesseract screenshot 1

Cost / License

  • Free
  • Open Source

Application type

Platforms

  • Mac
  • Windows
  • Linux
111likes
0articles
Save

Features

Tesseract News & Activities

Highlights All activities

Recent News

No news, maybe you know any news worth sharing?

Share a News Tip

Recent activities

Tesseract information

Our users have written 2 comments and reviews about Tesseract, and it has gotten 111 likes

Tesseract was added to AlternativeTo by Akasam on and this page was last updated .

Comments and Reviews

   
Top Positive Comment
tylerszabo
5

In terms of OCR this tesseract is fantastic. I compared it to ABBYY 14 and tesseract had fewer errors on dictionary words. While it doesn't offer layout preservation with the OCR (i.e. converting into an editable document that should print similarly) you'll likely make up for that in the reduced time needed to fix OCR errors.

For handling PDFs you'll need to convert them to an image file, first - pdftopng (an Open Source tool that can be found in the Xpdf project)

-3

Update 2025 Now that I use Liberica, Java is fine with me. It seems to me that I've used ocr libraries before, but it's been a while and offhand I'm not sure of a program to use it in. I'd change this to a comment vs rating.

Featured in Lists

Core, Development & Services

A list with 807 apps by AmileyaRyver without a description.
By AmileyaRyver807 appsUpdated

Adobe cloud FOSS Alternatives

What a adobe creative cloud FOSS alternative(including Discontinued Apps and linux)? Well there is not a full suite but here is one by one guide.
By moonstone17 appsUpdated

My FOSS Apps (Linux)

By brickelt963-altto13 appsUpdated

Official Links

Install from the command line

  • winget install -e --id tesseract-ocr.tesseract

Be careful what you paste into a terminal. Check what a command installs before running it.

What is Tesseract?

Tesseract.js is a javascript library that gets words in almost any language out of images.

The Tesseract OCR engine was one of the top 3 engines in the 1995 UNLV Accuracy test. Between 1995 and 2006 it had little work done on it, but it is probably one of the most accurate open source OCR engines available. The source code will read a binary, grey or color image and output text. A tiff reader is built in that will read uncompressed TIFF images, or libtiff can be added to read compressed images. There are language files for many languages, even for text set in Fraktur and blackletter typefaces.