Inspiration

Urdu is spoken by millions of people, yet it remains comparatively under-resourced in Natural Language Processing. While many toxicity-detection systems can classify an entire sentence as toxic, they often fail to identify exactly which words or phrases are responsible. We wanted to address this gap by building resources and models for Urdu Toxic Span Detection. Our motivation was not only to detect harmful content, but to create research infrastructure that could support future NLP work for an underrepresented language. This led us from text-based toxic span detection to exploring a more ambitious question: can multimodal information, particularly audio, provide additional signals for understanding toxic content?

What it does

MUTEX (Urdu Multimodal Toxic Span Detection) is a research project that identifies the specific toxic spans within Urdu content rather than simply assigning a toxic/non-toxic label. The project includes:

A manually annotated dataset of 14K+ Urdu instances Token-level toxic span detection Support for Urdu's complex linguistic variations Handling of Roman Urdu and code-switched content A multimodal extension exploring text and audio together

The goal is to contribute both a detection approach and a reusable research resource for future Urdu NLP research.

How we built it

We began by creating and manually annotating an Urdu toxic-span dataset. The data required careful preprocessing and annotation because Urdu presents challenges that are less common in well-resourced English NLP benchmarks, including morphological variation, Unicode inconsistencies, Roman Urdu, code-switching, and variations in written expression.

For the text-based system, we explored transformer-based NLP models and sequence-labeling approaches to identify toxic spans at a fine-grained level.

We then extended the research into a multimodal direction by incorporating audio alongside textual information. The broader system explores how speech and acoustic features can complement text when analyzing toxic content.

The project was developed using modern NLP and deep-learning tools, with the dataset and research artifacts made available publicly to support reproducibility and further research.

Challenges we ran into

One of our biggest challenges was the lack of existing resources. Unlike English, we could not simply rely on large, mature datasets and benchmarks specifically designed for Urdu toxic span detection.

Other challenges included:

Defining consistent annotation guidelines Manually identifying exact toxic spans Urdu morphological complexity Unicode normalization Roman Urdu Code-switching between Urdu and English Variation in spelling and informal online language Aligning audio and textual information for multimodal analysis Evaluating a task where the goal is to identify the exact location of toxicity, not just classify an entire sentence

Accomplishments that we're proud of

Accomplishments that we're proud of

We are particularly proud that we did not stop at training a model on an existing benchmark. We contributed to the underlying research infrastructure itself.

Our key accomplishments include:

Creating a 14K+ manually annotated Urdu toxic-span dataset Addressing an NLP problem in a comparatively low-resource language Developing a fine-grained approach to identify exact toxic spans Extending the research from text-only analysis toward multimodal text-and-audio analysis Making the research resources publicly accessible through platforms such as GitHub and preprint repositories Turning the work into an ongoing research trajectory rather than a one-time experiment

What we learned

The biggest lesson was that building an AI model is often only a small part of solving a real research problem.

We learned the importance of:

High-quality data and careful annotation Understanding linguistic and cultural context Designing reproducible research resources Measuring more than just a single accuracy score Persisting through failed experiments and unexpected results Thinking about who is represented, and underrepresented, in modern AI research

Most importantly, we learned that a lack of resources should not always be seen as a reason to avoid a problem. Sometimes, it is an opportunity to create the resources that future researchers will need.

What's next for MUTEX: Urdu Multimodal Toxic Span Detection

Next, we want to continue improving the multimodal approach and investigate how different forms of information can improve toxic-span detection in Urdu.

Our future goals include:

Expanding and refining the dataset Improving multimodal fusion between text and audio Evaluating additional Urdu and multilingual language models Improving robustness for Roman Urdu and code-switched content Exploring additional low-resource South Asian languages Creating more accessible tools and benchmarks for researchers working on Urdu NLP Developing the project into a stronger open research resource for the wider NLP community

Our long-term vision is bigger than one toxicity-detection model: we want to help reduce the resource gap that prevents languages like Urdu from receiving the same level of representation in modern AI research.

Built With

  • arxiv
  • crf-platforms/tools:-google-colab
  • github
  • google-cloud-speech-to-text
  • hugging-face-transformers
  • languages:-python
  • roman-urdu-frameworks/libraries:-pytorch
  • urdu
Share this project:

Updates

Submission history