A Low-Resource ASR Adaptation Study at CTNLPR
Center for Tamil Natural Language Processing Research (CTNLPR)
Introduction
Automatic Speech Recognition (ASR) has undergone a fundamental transition from language-specific acoustic models toward large-scale multilingual speech foundation models.
For Indic languages, this transition has been particularly significant.
Systems such as IndicWav2Vec, IndicWhisper, IndicConformer, and newer multilingual speech models have demonstrated that a single pretrained representation can support speech recognition across a large number of languages and linguistic conditions. AI4Bharat’s current ASR work spans these model families and is explicitly motivated by broad linguistic coverage, large-scale speech collection, and adaptation to diverse domains and demographics.
However, there remains an important gap between language coverage and target-domain speech robustness.
A model trained on Indian Tamil speech is not automatically optimized for speech produced by Sri Lankan Tamil speakers.
This motivates our research at the Center for Tamil Natural Language Processing Research (CTNLPR):
Can a large pretrained Indic ASR model be efficiently adapted to Sri Lankan Tamil speech using a relatively small, carefully curated labelled speech corpus?
Rather than developing a new ASR architecture from scratch, our research investigates transfer learning and supervised domain adaptation as a practical pathway toward Sri Lankan Tamil speech recognition.
From Indic Language Coverage to Sri Lankan Tamil Adaptation
Modern Indic ASR systems have already established an important foundation.
For example, the IndicConformer-600M-Multi model is a 600-million-parameter multilingual Conformer-based ASR system supporting all 22 officially recognized Indian languages, including Tamil. It provides both CTC and RNNT decoding pathways.
This is an important starting point for Sri Lankan Tamil.
The pretrained model already contains substantial acoustic and linguistic information associated with Tamil and other Indic languages.
Our hypothesis is therefore
Multilingual Indic Representation + Sri Lankan Tamil Supervision → Sri Lankan Tamil ASR
Instead of relearning speech representations from scratch, we adapt an existing multilingual representation toward the target speech distribution.
The Core Research Problem
Let the pretrained multilingual ASR model be:
where:
- represents an input speech signal,
- represents the pretrained ASR model,
- represents its pretrained parameters.
We construct a Sri Lankan Tamil dataset:
where:
- is Sri Lankan Tamil speech,
- is the corresponding transcription.
The adaptation objective becomes:
The resulting model:
should ideally preserve the general acoustic knowledge learned during multilingual pretraining while becoming more specialized for Sri Lankan Tamil speech.
This makes the problem a low-resource transfer-learning problem, rather than a conventional ASR training problem.
Why Sri Lankan Tamil Needs Adaptation
The existence of Tamil in an Indic ASR model does not mean that the model is optimized for every Tamil speech distribution.
A pretrained model may have been exposed predominantly to speech collected from different:
- geographical populations
- recording environments
- speaker populations
- conversational contexts
- acoustic channels
- lexical distributions
- pronunciation patterns
Therefore, the distribution learned during pretraining can be represented as:
while our target distribution is:
The adaptation problem exists because:
The objective of fine-tuning is therefore to reduce this distributional mismatch.
What We Mean by Sri Lankan Tamil Adaptation
Our current research does not attempt to build separate ASR models for individual Sri Lankan geographical regions.
Instead, we treat Sri Lankan Tamil speech as the target domain.
The research corpus is intended to represent Sri Lankan Tamil speech in a general sense, with variation naturally present within the collected data.
The conceptual data structure is:

The focus is therefore Sri Lankan Tamil adaptation, rather than explicit regional dialect classification.
What Existing Indic ASR Research Teaches Us
Our approach is informed by several existing Indic ASR directions.
IndicWav2Vec
AI4Bharat’s IndicWav2Vec work demonstrates the application of wav2vec-style self-supervised speech representations to Indic languages.
The central principle is important for low-resource settings:
Large quantities of speech representation learning can reduce dependence on fully transcribed speech.
IndicWav2Vec models are available for individual Indic languages, illustrating the effectiveness of transferring pretrained speech representations into language-specific ASR tasks.
This motivates our emphasis on pretrained acoustic representations rather than training from random initialization.
IndicConformer: Multilingual Conformer Adaptation
The IndicConformer family provides another important direction.
The 600M multilingual model uses a Conformer-based hybrid CTC + RNNT architecture, supporting 22 Indian languages.
The underlying architecture combines:
- convolutional acoustic modeling
- self-attention
- contextual sequence modeling
- CTC decoding
- RNNT decoding
This makes it particularly attractive for transfer learning.
The model is also explicitly distributed with resources for loading and training/fine-tuning through the AI4Bharat NeMo ecosystem.
For CTNLPR, this suggests a practical research pathway:

IndicWhisper and Domain Adaptation
Another relevant direction is IndicWhisper.
AI4Bharat’s Vistaar work provides a particularly useful lesson here.
Vistaar combines diverse benchmarks and training datasets spanning multiple languages and domains, with more than 10,700 hours of labelled audio across 12 Indian languages. AI4Bharat reports fine-tuning Whisper models on Vistaar training data and evaluating them across 59 language/domain benchmark combinations.
This demonstrates an important principle:
Pretrained ASR + Target-domain labelled data → Improved domain robustness
For our research, the target domain is not simply “Tamil”.
It is:
Recent Adaptation Evidence
A recent 2026 study on Indic ASR, Vividh-ASR, highlights another important issue: pretrained speech models can exhibit strong bias toward studio or broadcast recordings and degrade when exposed to spontaneous speech. The work investigates fine-tuning strategies and curriculum-based adaptation to address this mismatch.
This is directly relevant to our methodology.
A model’s performance is determined not only by:
but also by:
Therefore, our Sri Lankan Tamil corpus should represent the type of speech for which we ultimately want the model to perform well.
Our Proposed Adaptation Architecture
At CTNLPR, we propose the following general architecture:

The important characteristic is that the architecture itself does not need to be redesigned.
The research contribution is the adaptation methodology and evaluation of transfer to Sri Lankan Tamil speech.
Stage 1 — Establish the Pretrained Baseline
Before performing any adaptation, we need to measure the performance of the original multilingual model.
This produces our zero-shot or direct-transfer baseline:

This experiment answers:
How well does the existing Indic ASR model recognize Sri Lankan Tamil without target-specific training?
This baseline is essential.
Without it, an improvement after fine-tuning cannot be quantified.
Stage 2 — Build the Sri Lankan Tamil Corpus
The next stage is dataset engineering.
Each training example should contain:
where:
and:
A practical manifest can follow:
{
"audio_filepath": "/data/train/sample_001.wav",
"text": "தமிழ் உரை",
"duration": 5.42
}
Modern Indic ASR fine-tuning workflows commonly standardize speech at 16 kHz, and the SraVaani fine-tuning workflow demonstrates JSONL manifests and scalable tarred datasets for larger training collections.
Our corpus preparation should therefore include:
- 16 kHz audio normalization
- mono conversion
- audio integrity validation
- silence/segment handling
- transcript alignment
- Unicode normalization
- whitespace normalization
- consistent transcription conventions
- train/validation/test separation
Transcript Normalization Is a Model Component
A frequently underestimated component of ASR adaptation is transcript normalization.
The target sequence:
is the supervision signal used to update the model.
If equivalent utterances are represented using inconsistent orthographic conventions, the model receives conflicting targets.
Therefore:

The same normalization procedure should subsequently be applied to both the reference and hypothesis during evaluation.
This prevents evaluation errors caused by formatting differences rather than genuine recognition errors.
Stage 3 — Fine-Tuning
Once the dataset is prepared, the pretrained ASR model becomes the initialization point.
Conceptually:
where:
- = pretrained multilingual ASR parameters
- = Sri Lankan Tamil adapted parameters
The training objective is:
The objective is not to destroy the pretrained multilingual representation.
It is to perform a controlled parameter update:
where represents the adaptation induced by Sri Lankan Tamil supervision.
Frozen Encoder vs Full Fine-Tuning
One of the most important experimental decisions is how much of the pretrained model should be updated.
Experiment A — Encoder-Frozen Adaptation

This approach protects pretrained acoustic representations and reduces the number of trainable parameters.
It is particularly attractive when the target dataset is small.
The recent SraVaani fine-tuning guide demonstrates this as a practical low-resource adaptation strategy. It describes a complete language-agnostic workflow using custom speech and transcripts, and demonstrates fine-tuning on a single 15 GB GPU.
Experiment B — Full Fine-Tuning
The second configuration updates the complete model:

This provides greater adaptation capacity but increases:
- GPU memory requirements
- optimization sensitivity
- overfitting risk
- catastrophic-forgetting risk
Therefore, full fine-tuning should be treated as an experimental comparison rather than automatically assuming it is superior.
A Better Strategy: Progressive Adaptation
Based on the broader Indic ASR literature, our preferred experimental methodology is not to immediately commit to one training regime.
Instead:

This produces a scientifically interpretable progression.
We can determine whether the target corpus requires:
- decoder adaptation only,
- partial encoder adaptation,
- or complete end-to-end optimization.
Stage 4 — Data-Efficiency Experiments
The most important research question for a low-resource setting is not simply:
“Can we fine-tune the model?”
The more important question is:
How much labelled Sri Lankan Tamil speech is required to obtain meaningful adaptation?
We therefore propose multiple data regimes:

This allows us to construct a learning curve:
where represents labelled speech hours.
Such an experiment is particularly valuable for low-resource language technology because it directly quantifies annotation efficiency.
Stage 5 — Evaluation
The evaluation framework should contain at least two principal metrics.
Word Error Rate
where:
- = substitutions
- = deletions
- = insertions
- = number of reference words
Character Error Rate
where the unit is characters rather than words.
For Tamil ASR, CER provides a complementary signal because word-level errors can sometimes be influenced by segmentation and orthographic conventions.
Baseline vs Adapted Model
Our central experiment becomes:

Let:
be the baseline error rate and:
be the error rate after Sri Lankan Tamil adaptation.
Then:
and:
This provides a direct quantitative measure of adaptation effectiveness.
Error Analysis
Aggregate WER alone cannot tell us whether the model has actually learned the target speech distribution.
Therefore, we also need qualitative error analysis.
We can categorize residual errors into:

This can help determine whether remaining errors originate primarily from:
- acoustic mismatch,
- lexical mismatch,
- transcript normalization,
- language mixing,
- segmentation,
- or insufficient target-domain data.
Preventing Catastrophic Forgetting
An important consideration is that the pretrained model has learned multilingual knowledge.
Aggressive adaptation on a small Sri Lankan Tamil corpus may produce:
that performs well on the target domain but loses useful general capabilities.
Therefore, we should monitor the trade-off between:
and:
This suggests an additional experiment:

This is especially important if the eventual objective is to create a reusable Tamil ASR model rather than a model specialized exclusively for one dataset.
Why We Are Comparing Multiple Indic ASR Foundations
We do not want the research to depend on the assumption that one pretrained model is inherently optimal.
The broader Indic ASR ecosystem provides several candidate model families.
IndicWav2Vec
Strong self-supervised representation-learning foundation for Indic speech.
IndicWhisper
Whisper-based multilingual adaptation demonstrated through large-scale Indic datasets such as Vistaar.
IndicConformer
Conformer-based hybrid CTC/RNNT architecture with a 600M multilingual checkpoint covering 22 Indian languages, including Tamil.
Large multilingual FastConformer systems
Recent multilingual ASR work demonstrates another route: extensive multilingual pretraining followed by targeted supervised adaptation to a new language or domain. The SraVaani fine-tuning workflow explicitly demonstrates this language/domain adaptation paradigm.
Therefore, our research direction is better represented as:
rather than:
Proposed CTNLPR Experimental Framework
Our broader experimental framework can therefore be organized into four phases.
Phase 1 — Foundation Model Evaluation
Evaluate several suitable pretrained Indic ASR models directly on the same Sri Lankan Tamil test set.

This establishes which pretrained representation transfers best before adaptation.
Phase 2 — Sri Lankan Tamil Adaptation
Take the strongest practical candidate models and fine-tune them using the same Sri Lankan Tamil training corpus.

Phase 3 — Data-Efficiency Analysis
Train with progressively larger subsets:
and compare:
This allows us to quantify the value of additional annotation.
Phase 4 — Adaptation Strategy Analysis
Compare:
Frozen Encoder
vs
Partial Fine-Tuning
vs
Full Fine-Tuning
and evaluate:
- WER
- CER
- convergence
- GPU requirements
- training time
- generalization
- target-domain improvement
The Complete CTNLPR Pipeline
The complete research architecture can be summarized as:

The Research Hypothesis
Our primary hypothesis is:
That is, Sri Lankan Tamil-specific supervised adaptation will reduce recognition errors compared with direct inference from a pretrained Indic ASR model.
A secondary hypothesis is:
However, we do not assume that full fine-tuning will always outperform encoder-frozen adaptation.
H3: Optimal adaptation strategy = f(dataset size, model, domain)
This makes the research experimentally testable rather than prescriptive.
Why This Matters for Sri Lankan Tamil
The broader objective is not simply to obtain another ASR checkpoint.
Sri Lankan Tamil speech technology requires models that can perform effectively on speech actually produced within the target linguistic environment.
Large multilingual models provide an important foundation, but their usefulness for underrepresented speech varieties depends on how effectively they can be adapted.
Our research therefore investigates a scalable paradigm:
Large Multilingual ASR + Small High-Quality Target Corpus → Low-Resource Speech Technology
This is potentially much more practical than attempting to construct a Sri Lankan Tamil ASR foundation model entirely from scratch.
Toward a Sri Lankan Tamil Speech Foundation
The long-term direction of this research extends beyond one ASR model.
A successfully adapted model could become a foundation for:
- speech-to-text systems
- voice interfaces
- conversational agents
- speech-enabled search
- document transcription
- accessibility technologies
- spoken-language information retrieval
- speech analytics
- downstream Tamil NLP systems
More importantly, the resulting methodology can establish a repeatable procedure for adapting large multilingual speech models to other underrepresented Sri Lankan speech varieties and domains.
Conclusion
At the Center for Tamil Natural Language Processing Research (CTNLPR), we are investigating a practical route toward improving Sri Lankan Tamil Automatic Speech Recognition through adaptation of pretrained Indic language ASR models.
The central idea is simple:
Pretrained Indic ASR → Sri Lankan Tamil Fine-Tuning → Adapted ASR
But the research questions are considerably deeper.
We want to determine:
- which Indic ASR foundation transfers most effectively to Sri Lankan Tamil;
- how much labelled speech is required;
- whether frozen or full-model adaptation is more effective;
- how transcript normalization influences performance;
- how much WER/CER improvement can be achieved;
- whether adaptation causes degradation of pretrained capabilities;
- and which error categories remain after adaptation.
The goal is therefore not simply to fine-tune an ASR model.
The goal is to establish an evidence-driven, low-resource adaptation methodology for Sri Lankan Tamil speech recognition.
At CTNLPR, this work represents a step toward building speech technologies that extend beyond broad language coverage toward effective representation of Sri Lankan Tamil speech in real-world conditions.
References and Models Investigated
1. IndicConformer-600M-Multi — AI4Bharat
A 600M-parameter multilingual Conformer-based hybrid CTC/RNNT ASR model covering 22 officially recognized Indian languages, including Tamil. The model is released under the MIT license and the associated AI4Bharat repository provides resources for training and fine-tuning through its NeMo ecosystem.
IndicConformer-600M-Multi — Hugging Face
2. IndicWav2Vec — AI4Bharat
A family of wav2vec-based speech representation models developed for Indic languages, illustrating the use of self-supervised acoustic representation learning for low-resource ASR. AI4Bharat identifies IndicWav2Vec as one of its major ASR model families.
3. IndicWhisper / Vistaar — AI4Bharat
IndicWhisper represents Whisper-based adaptation for Indic languages. The Vistaar project provides diverse training and evaluation datasets across languages and domains and reports fine-tuning Whisper models on these datasets.
4. SraVaani-1.0 — ARTPARK-IISc
A multilingual FastConformer ASR system trained across a large set of Indian languages and dialects. Its associated 2026 fine-tuning guide demonstrates a language-agnostic adaptation workflow using custom audio and transcripts, including encoder-frozen and broader fine-tuning strategies.
5. IndicVoices / IndicASR — AI4Bharat
IndicVoices provides large-scale multilingual speech resources and an IndicASR checkpoint trained using IndicVoices data, demonstrating another large-scale pathway for multilingual Indic speech modeling.
6. Vividh-ASR
Recent work investigating recording-condition bias and adaptation strategies for Indic ASR, highlighting the importance of domain mismatch and careful fine-tuning when transferring pretrained speech models to real-world speech.
CTNLPR Research Direction
Pretrained Indic ASR → Sri Lankan Tamil Adaptation → Low-Resource Evaluation → Robust Sri Lankan Tamil Speech Recognition
The central research objective is therefore:
To develop and systematically evaluate an efficient adaptation methodology for transforming large multilingual Indic ASR models into robust Sri Lankan Tamil speech recognition systems using limited labelled target-language data.