NLP
AI-Generated Text Detection for Bangla
A classifier that distinguishes human-written Bangla text from AI-generated Bangla text — a low-resource-language detection problem.
Overview
As generative AI tools produce increasingly fluent Bangla text, detecting AI-generated content has become important for education, publishing and content-integrity use cases — yet almost all public detection research and tooling targets English.
Problem
Most AI-text detectors are trained and validated on English corpora and transfer poorly to Bangla, a morphologically rich, comparatively low-resource language with far less labeled training data available.
Approach
We assembled a paired corpus of human-written and AI-generated Bangla text across multiple domains, then evaluated classical stylometric features alongside transformer-based classifiers fine-tuned on Bangla language models, focusing on robustness across writing styles and topics.
Architecture
- 1Bangla corpus collection & AI-text generation pipeline for training pairs
- 2Feature extraction (stylometric + embedding-based)
- 3Fine-tuned transformer classification model
- 4Evaluation across multiple text domains and generators
Technology
Results
- Detection accuracy
- Indicative — see note
- False-positive rate
- Indicative — see note
- Domains evaluated
- Indicative — see note
This case study uses indicative placeholders pending publication of our verified benchmark results.
Have a similar problem?
Talk to our team — we'll tell you honestly how close this is to what you need.