What You’ll Do
Phase 1: Text Data Curation
Audit and profile open-source datasets (Sangraha, Common Crawl, IndicCorp, etc.) to assess quality, coverage, and noise levels
Design and implement data cleaning pipelines: deduplication, script normalisation, encoding fixes, noise removal, sentence boundary detection
Create and apply metadata tagging schemas labelling text by domain (news, legal, literature, health, etc.), subdomain, language, register, and quality tier
Build validation checklists and quality scorecards to benchmark dataset readiness for model training
Document data provenance, licensing, and processing steps for reproducibility
Phase 2: Speech & Voice Data Preparation
Curate high-quality, phonetically diverse text passages suitable for read-speech recording
Ensure text selection covers domain, prosodic, and phonemic variety required for TTS/ASR model training
Assist in defining metadata standards for audio datasets (speaker demographics, recording conditions, transcription format)
Support the pipeline transition from text corpus to aligned speech dataset