Athena SynCognition TechnologiesAthena SynCognitionTechnologies

Computer Vision / OCR

Bangla Compound Character Recognition

A computer-vision system for recognizing Bangla compound characters (যুক্তাক্ষর) — one of the harder OCR problems in low-resource languages.

Overview

Bangla script includes a large set of compound characters formed by joining multiple glyphs, which standard OCR systems handle poorly. This project builds a recognition model specifically for that challenge.

Problem

Off-the-shelf OCR tools are trained primarily on high-resource languages and simple character sets. Bangla compound characters are visually complex and highly varied, leading to poor recognition accuracy on real-world scanned documents.

Approach

We built and curated a labeled dataset of Bangla compound characters, then trained convolutional neural network architectures with data augmentation to handle handwriting and print variation. The pipeline includes preprocessing (binarization, segmentation) and a classification model tuned specifically for compound glyph shapes.

Architecture

  1. 1Image preprocessing: binarization, noise removal, segmentation
  2. 2Character/glyph segmentation pipeline
  3. 3CNN-based classification model
  4. 4Evaluation harness against held-out labeled samples

Technology

PythonPyTorchOpenCVNumPyJupyter

Results

Classification accuracy
Indicative — see note
Character classes covered
Indicative — see note
Processing time per page
Indicative — see note

This case study uses indicative placeholders pending publication of our verified benchmark results.

Have a similar problem?

Talk to our team — we'll tell you honestly how close this is to what you need.