Computer Vision / OCR
Bangla Compound Character Recognition
A computer-vision system for recognizing Bangla compound characters (যুক্তাক্ষর) — one of the harder OCR problems in low-resource languages.
Overview
Bangla script includes a large set of compound characters formed by joining multiple glyphs, which standard OCR systems handle poorly. This project builds a recognition model specifically for that challenge.
Problem
Off-the-shelf OCR tools are trained primarily on high-resource languages and simple character sets. Bangla compound characters are visually complex and highly varied, leading to poor recognition accuracy on real-world scanned documents.
Approach
We built and curated a labeled dataset of Bangla compound characters, then trained convolutional neural network architectures with data augmentation to handle handwriting and print variation. The pipeline includes preprocessing (binarization, segmentation) and a classification model tuned specifically for compound glyph shapes.
Architecture
- 1Image preprocessing: binarization, noise removal, segmentation
- 2Character/glyph segmentation pipeline
- 3CNN-based classification model
- 4Evaluation harness against held-out labeled samples
Technology
Results
- Classification accuracy
- Indicative — see note
- Character classes covered
- Indicative — see note
- Processing time per page
- Indicative — see note
This case study uses indicative placeholders pending publication of our verified benchmark results.
Have a similar problem?
Talk to our team — we'll tell you honestly how close this is to what you need.