Blog — Computer Vision

Setting up Hindi OCR
using Pytesseract.

Pytesseract, a Python wrapper for Google's Tesseract-OCR engine, is a popular tool for implementing OCR in Python applications. Here's how to set it up to read Hindi text.

Dwayo Team Jul 12, 2022
Prerequisites

What you'll need first.

Before you begin, ensure you have the following prerequisites installed on your system:

  1. Python and pip. Make sure you have Python installed on your system — you can download it from python.org. Pip, the package installer for Python, should also be installed.
  2. Tesseract OCR engine. Install Tesseract on your system. You can download it from the official GitHub repository. Follow the installation instructions for your operating system.
  3. Pytesseract. Install the Pytesseract library using pip:
pip install pytesseract
  1. Pillow (PIL fork). Pillow is a powerful image processing library in Python. Install it using:
pip install pillow

Setting up Hindi language support.

By default, Tesseract supports multiple languages, but we need to specify Hindi for our OCR setup. Follow these steps:

  1. Download Hindi language data. Visit the Tesseract GitHub page for language data and download the Hindi language data file (hin.traineddata). Place the downloaded file in the Tesseract installation directory.
  2. Specify the language in Pytesseract. In your Python script or application, set the language parameter to 'hin' when using Pytesseract. For example:
import pytesseract
from PIL import Image

pytesseract.pytesseract.tesseract_cmd = r'C:\Program Files\Tesseract-OCR\tesseract.exe'  # Set your Tesseract installation path
image_path = 'path/to/your/image.png'
text = pytesseract.image_to_string(Image.open(image_path), lang='hin')
print(text)

Make sure to replace the Tesseract path (tesseract_cmd) with the path where Tesseract is installed on your system.

  1. Run your script. Execute your Python script, and Pytesseract will use Tesseract with Hindi language support to perform OCR on the specified image.
Troubleshooting

Tips if the output looks wrong.

By following these steps, you can set up Hindi OCR using Pytesseract and extract text from images written in Hindi. Experiment with different images and tune the OCR parameters as needed for optimal results.

Where this fits.

Document and vision pipelines like this are part of our Vision AI work — extracting structured, usable text from real-world documents and images, including in Indic scripts.

Get started

Tell us what
you're trying to build.

Book a 30-minute call — we'll help you think through document and vision extraction for your product.

Book a free 30-min call → More articles →