This repository contains code designed to explore the fundamentals of Tesseract, a robust Optical Character Recognition (OCR) tool. The code in the Jupyter notebook was developed concurrently with the "Python in Digital Humanities" YouTube tutorial, showcasing various Tesseract functionalities as demonstrated in the video.
Optical Character Recognition (OCR) is a transformative technology that converts different types of documents, such as scanned paper documents, PDFs, or digital images, into editable and searchable data. Tesseract is one of the most widely used OCR engines, renowned for its accuracy and versatility.
- Document Digitization: Converting printed documents into digital format for efficient storage and retrieval.
- Automated Data Entry: Minimizing manual data entry efforts by automatically extracting text from images.
- Content Searchability: Enabling search functionality within scanned documents or images.
- Text Recognition in Natural Scenes: Extracting text from images captured in natural settings, such as street signs or product labels.
Pre-processing is a crucial step in the OCR pipeline to enhance the quality of input images, thereby facilitating the OCR engine's ability to accurately recognize text. Key pre-processing techniques include:
- Noise Reduction: Eliminating extraneous noise that can hinder text recognition.
- Binarization: Converting images to a binary format to distinguish text from the background.
- Deskewing: Correcting any tilt or rotation in the image to properly align the text.
- Border Removal: Removing borders that may interfere with text recognition.
- Morphological Operations (Dilation and Erosion): Enhancing text features and eliminating minor imperfections.
- Color Inversion: Adjusting the text color from white to black or vice versa, as needed.
Implementing these pre-processing techniques significantly improves the accuracy and reliability of OCR results.
- Matplotlib
- OpenCV
- PyTesseract
- Pillow
- NumPy
This repository includes all the images utilized and a .ipynb file, which encompasses three main sections:
- Opening an Image File Using Pillow
- Pre-Processing Techniques in OpenCV
- Final OCR Implementation in Tesseract
The pre-processing subsection implements the following techniques:
- Opening the Image Using OpenCV and Matplotlib
- Inverting an Image
- Binarization of an Image
- Noise Reduction in an Image
- Morphological Operations (Dilation and Erosion) on an Image
- Deskewing an Image
- Border Removal from an Image
- Adding Borders to an Image
The final OCR subsection examines two types of images:
-
Image with Latin Text:
- This image contains three columns of text. Direct OCR on this image yields suboptimal results as it reads the text left to right. Instead, the text needs to be read column-wise. Various pre-processing techniques were applied to ensure the final OCR output is accurate. The objective was to extract names written in Latin by reading the image column-wise and storing them in a list.
-
Image with a Footnote:
- The task involved extracting the content above a footnote, necessitating the removal of the footnote from the image.
- The entire notebook was developed while following the playlist: Python in Digital Humanities
- GitHub repository for the tutorial: OCR Python Textbook
All images used or generated through pre-processing techniques and OCR are uploaded in this folder.
Please note that none of the images used in this project are proprietary. The sources for these images can be found at the following link: OCR Python Textbook
The paths used in the file are Absolute Paths. You may need to change each path to Relative Path. To do this, replace the /Users/ayaanchoudhury/Desktop/Python Coding/Tesseract in each path to /images