Inspiration

This project was inspired by the need to preserve and digitize regional dialects that are often overlooked in mainstream speech recognition systems. While most ASR models focus on widely spoken languages, dialects like Chatgaiya remain underrepresented. The goal was to build a dataset and model that could bridge this gap, aligning Bangla ↔ English translations to support both ASR and NLP tasks.

What it does

The Regional Speech Recognition Model: Collects and processes speech samples. Aligns audio with Bangla and English translations for bilingual tasks. Provides a dataset pipeline for ASR training, evaluation, and NLP research. Enables reproducible experiments with containerized environments.

How we built it

Languages: Python (data processing, model training), Bash (automation). Libraries: NumPy, Pandas, Librosa (audio features), PyTorch/TensorFlow (ASR/NLP models), NLTK/SpaCy (text preprocessing). Platforms: Google Colab, Jupyter Notebook (experimentation). Storage: Local file system for raw audio, JSON/CSV for metadata, Google Drive/AWS S3 for backups. Dataset Tools: Custom audio recording pipeline, alignment scripts, preprocessing utilities (noise reduction, normalization, segmentation).

Challenges we ran into

Collecting high-quality audio samples in diverse environments. Designing accurate alignment between Chatgaiya, Bangla, and English translations. Handling noise and variability in regional speech recordings. Ensuring reproducibility across different platforms and contributors.

Accomplishments that we're proud of

Built one of the first structured datasets for Chatgaiya speech. Automated preprocessing and alignment pipelines. Enabled bilingual ASR/NLP research with aligned translations. Created a reproducible, containerized workflow for dataset building.

What we learned

The importance of dialect preservation in speech recognition. How to design robust preprocessing pipelines for noisy audio. That collaboration and reproducibility are key in open-source dataset projects.

What's next for Galactico- Project Management Tool

Expanding dataset size with more diverse Chatgaiya speakers. Training baseline ASR models using the dataset. Adding phonetic and linguistic annotations for deeper NLP tasks. Publishing the dataset for open research and community contributions.

Built With

Share this project:

Updates

Submission history