Athena SynCognition TechnologiesAthena SynCognitionTechnologies

NLP

AI-Generated Text Detection for Bangla

A classifier that distinguishes human-written Bangla text from AI-generated Bangla text — a low-resource-language detection problem.

Overview

As generative AI tools produce increasingly fluent Bangla text, detecting AI-generated content has become important for education, publishing and content-integrity use cases — yet almost all public detection research and tooling targets English.

Problem

Most AI-text detectors are trained and validated on English corpora and transfer poorly to Bangla, a morphologically rich, comparatively low-resource language with far less labeled training data available.

Approach

We assembled a paired corpus of human-written and AI-generated Bangla text across multiple domains, then evaluated classical stylometric features alongside transformer-based classifiers fine-tuned on Bangla language models, focusing on robustness across writing styles and topics.

Architecture

  1. 1Bangla corpus collection & AI-text generation pipeline for training pairs
  2. 2Feature extraction (stylometric + embedding-based)
  3. 3Fine-tuned transformer classification model
  4. 4Evaluation across multiple text domains and generators

Technology

PythonPyTorchHugging Face Transformersscikit-learnFastAPI

Results

Detection accuracy
Indicative — see note
False-positive rate
Indicative — see note
Domains evaluated
Indicative — see note

This case study uses indicative placeholders pending publication of our verified benchmark results.

Have a similar problem?

Talk to our team — we'll tell you honestly how close this is to what you need.