Adapting Indic Language ASR Models for Sri Lankan Tamil Speech

A Low-Resource ASR Adaptation Study at CTNLPR

Center for Tamil Natural Language Processing Research (CTNLPR)

Introduction

Automatic Speech Recognition (ASR) has undergone a fundamental transition from language-specific acoustic models toward large-scale multilingual speech foundation models.

For Indic languages, this transition has been particularly significant.

Systems such as IndicWav2Vec, IndicWhisper, IndicConformer, and newer multilingual speech models have demonstrated that a single pretrained representation can support speech recognition across a large number of languages and linguistic conditions. AI4Bharat’s current ASR work spans these model families and is explicitly motivated by broad linguistic coverage, large-scale speech collection, and adaptation to diverse domains and demographics.

However, there remains an important gap between language coverage and target-domain speech robustness.

A model trained on Indian Tamil speech is not automatically optimized for speech produced by Sri Lankan Tamil speakers.

This motivates our research at the Center for Tamil Natural Language Processing Research (CTNLPR):

Can a large pretrained Indic ASR model be efficiently adapted to Sri Lankan Tamil speech using a relatively small, carefully curated labelled speech corpus?

Rather than developing a new ASR architecture from scratch, our research investigates transfer learning and supervised domain adaptation as a practical pathway toward Sri Lankan Tamil speech recognition.


From Indic Language Coverage to Sri Lankan Tamil Adaptation

Modern Indic ASR systems have already established an important foundation.

For example, the IndicConformer-600M-Multi model is a 600-million-parameter multilingual Conformer-based ASR system supporting all 22 officially recognized Indian languages, including Tamil. It provides both CTC and RNNT decoding pathways.

This is an important starting point for Sri Lankan Tamil.

The pretrained model already contains substantial acoustic and linguistic information associated with Tamil and other Indic languages.

Our hypothesis is therefore

Multilingual Indic Representation + Sri Lankan Tamil Supervision → Sri Lankan Tamil ASR

Instead of relearning speech representations from scratch, we adapt an existing multilingual representation toward the target speech distribution.


The Core Research Problem

Let the pretrained multilingual ASR model be:fθ0(X)f_{\theta_0}(X)

where:

  • XX represents an input speech signal,
  • fθ0f_{\theta_0} represents the pretrained ASR model,
  • θ0\theta_0 represents its pretrained parameters.

We construct a Sri Lankan Tamil dataset:DSLT={(xi,yi)}i=1ND_{SLT} = \{(x_i,y_i)\}_{i=1}^{N}

where:

  • xix_i is Sri Lankan Tamil speech,
  • yiy_i is the corresponding transcription.

The adaptation objective becomes:θ∗=arg⁡min⁡θLASR(DSLT;θ)\theta^{*} = \arg\min_{\theta} \mathcal{L}_{ASR} (D_{SLT};\theta)

The resulting model:fθ∗f_{\theta^{*}}

should ideally preserve the general acoustic knowledge learned during multilingual pretraining while becoming more specialized for Sri Lankan Tamil speech.

This makes the problem a low-resource transfer-learning problem, rather than a conventional ASR training problem.


Why Sri Lankan Tamil Needs Adaptation

The existence of Tamil in an Indic ASR model does not mean that the model is optimized for every Tamil speech distribution.

A pretrained model may have been exposed predominantly to speech collected from different:

  • geographical populations
  • recording environments
  • speaker populations
  • conversational contexts
  • acoustic channels
  • lexical distributions
  • pronunciation patterns

Therefore, the distribution learned during pretraining can be represented as:Ppretrain(X,Y)P_{pretrain}(X,Y)

while our target distribution is:PSLT(X,Y)P_{SLT}(X,Y)

The adaptation problem exists because:Ppretrain(X,Y)≠PSLT(X,Y)P_{pretrain}(X,Y) \neq P_{SLT}(X,Y)

The objective of fine-tuning is therefore to reduce this distributional mismatch.


What We Mean by Sri Lankan Tamil Adaptation

Our current research does not attempt to build separate ASR models for individual Sri Lankan geographical regions.

Instead, we treat Sri Lankan Tamil speech as the target domain.

The research corpus is intended to represent Sri Lankan Tamil speech in a general sense, with variation naturally present within the collected data.

The conceptual data structure is:

The focus is therefore Sri Lankan Tamil adaptation, rather than explicit regional dialect classification.


What Existing Indic ASR Research Teaches Us

Our approach is informed by several existing Indic ASR directions.

IndicWav2Vec

AI4Bharat’s IndicWav2Vec work demonstrates the application of wav2vec-style self-supervised speech representations to Indic languages.

The central principle is important for low-resource settings:

Large quantities of speech representation learning can reduce dependence on fully transcribed speech.

IndicWav2Vec models are available for individual Indic languages, illustrating the effectiveness of transferring pretrained speech representations into language-specific ASR tasks.

This motivates our emphasis on pretrained acoustic representations rather than training from random initialization.


IndicConformer: Multilingual Conformer Adaptation

The IndicConformer family provides another important direction.

The 600M multilingual model uses a Conformer-based hybrid CTC + RNNT architecture, supporting 22 Indian languages.

The underlying architecture combines:

  • convolutional acoustic modeling
  • self-attention
  • contextual sequence modeling
  • CTC decoding
  • RNNT decoding

This makes it particularly attractive for transfer learning.

The model is also explicitly distributed with resources for loading and training/fine-tuning through the AI4Bharat NeMo ecosystem.

For CTNLPR, this suggests a practical research pathway:


IndicWhisper and Domain Adaptation

Another relevant direction is IndicWhisper.

AI4Bharat’s Vistaar work provides a particularly useful lesson here.

Vistaar combines diverse benchmarks and training datasets spanning multiple languages and domains, with more than 10,700 hours of labelled audio across 12 Indian languages. AI4Bharat reports fine-tuning Whisper models on Vistaar training data and evaluating them across 59 language/domain benchmark combinations.

This demonstrates an important principle:

Pretrained ASR + Target-domain labelled data → Improved domain robustness

For our research, the target domain is not simply “Tamil”.

It is:Sri Lankan Tamil Speech\boxed{\text{Sri Lankan Tamil Speech}}


Recent Adaptation Evidence

A recent 2026 study on Indic ASR, Vividh-ASR, highlights another important issue: pretrained speech models can exhibit strong bias toward studio or broadcast recordings and degrade when exposed to spontaneous speech. The work investigates fine-tuning strategies and curriculum-based adaptation to address this mismatch.

This is directly relevant to our methodology.

A model’s performance is determined not only by:

Language\text{Language}

but also by:Language+Domain+Acoustic Conditions+Speaker Variation\text{Language} + \text{Domain} + \text{Acoustic Conditions} + \text{Speaker Variation}

Therefore, our Sri Lankan Tamil corpus should represent the type of speech for which we ultimately want the model to perform well.


Our Proposed Adaptation Architecture

At CTNLPR, we propose the following general architecture:

The important characteristic is that the architecture itself does not need to be redesigned.

The research contribution is the adaptation methodology and evaluation of transfer to Sri Lankan Tamil speech.


Stage 1 — Establish the Pretrained Baseline

Before performing any adaptation, we need to measure the performance of the original multilingual model.

This produces our zero-shot or direct-transfer baseline:

This experiment answers:

How well does the existing Indic ASR model recognize Sri Lankan Tamil without target-specific training?

This baseline is essential.

Without it, an improvement after fine-tuning cannot be quantified.


Stage 2 — Build the Sri Lankan Tamil Corpus

The next stage is dataset engineering.

Each training example should contain:(xi,yi)(x_i,y_i)

where:xi=audiox_i = \text{audio}

and:yi=Sri Lankan Tamil transcriptiony_i = \text{Sri Lankan Tamil transcription}

A practical manifest can follow:

{
  "audio_filepath": "/data/train/sample_001.wav",
  "text": "தமிழ் உரை",
  "duration": 5.42
}

Modern Indic ASR fine-tuning workflows commonly standardize speech at 16 kHz, and the SraVaani fine-tuning workflow demonstrates JSONL manifests and scalable tarred datasets for larger training collections.

Our corpus preparation should therefore include:

  • 16 kHz audio normalization
  • mono conversion
  • audio integrity validation
  • silence/segment handling
  • transcript alignment
  • Unicode normalization
  • whitespace normalization
  • consistent transcription conventions
  • train/validation/test separation

Transcript Normalization Is a Model Component

A frequently underestimated component of ASR adaptation is transcript normalization.

The target sequence:Y=(y1,y2,…,yT)Y = (y_1,y_2,\ldots,y_T)

is the supervision signal used to update the model.

If equivalent utterances are represented using inconsistent orthographic conventions, the model receives conflicting targets.

Therefore:

The same normalization procedure should subsequently be applied to both the reference and hypothesis during evaluation.

This prevents evaluation errors caused by formatting differences rather than genuine recognition errors.


Stage 3 — Fine-Tuning

Once the dataset is prepared, the pretrained ASR model becomes the initialization point.

Conceptually:θ0→θSLT\theta_0 \rightarrow \theta_{SLT}

where:

  • θ0\theta_0 = pretrained multilingual ASR parameters
  • θSLT\theta_{SLT} = Sri Lankan Tamil adapted parameters

The training objective is:LSLT=1N∑i=1NL(fθ(xi),yi)\mathcal{L}_{SLT} = \frac{1}{N} \sum_{i=1}^{N} \mathcal{L} \left( f_{\theta}(x_i), y_i \right)

The objective is not to destroy the pretrained multilingual representation.

It is to perform a controlled parameter update:θSLT=θ0+Δθ\theta_{SLT} = \theta_0 + \Delta\theta

where Δθ\Delta\theta represents the adaptation induced by Sri Lankan Tamil supervision.


Frozen Encoder vs Full Fine-Tuning

One of the most important experimental decisions is how much of the pretrained model should be updated.

Experiment A — Encoder-Frozen Adaptation

This approach protects pretrained acoustic representations and reduces the number of trainable parameters.

It is particularly attractive when the target dataset is small.

The recent SraVaani fine-tuning guide demonstrates this as a practical low-resource adaptation strategy. It describes a complete language-agnostic workflow using custom speech and transcripts, and demonstrates fine-tuning on a single 15 GB GPU.


Experiment B — Full Fine-Tuning

The second configuration updates the complete model:

This provides greater adaptation capacity but increases:

  • GPU memory requirements
  • optimization sensitivity
  • overfitting risk
  • catastrophic-forgetting risk

Therefore, full fine-tuning should be treated as an experimental comparison rather than automatically assuming it is superior.


A Better Strategy: Progressive Adaptation

Based on the broader Indic ASR literature, our preferred experimental methodology is not to immediately commit to one training regime.

Instead:

This produces a scientifically interpretable progression.

We can determine whether the target corpus requires:

  • decoder adaptation only,
  • partial encoder adaptation,
  • or complete end-to-end optimization.

Stage 4 — Data-Efficiency Experiments

The most important research question for a low-resource setting is not simply:

“Can we fine-tune the model?”

The more important question is:

How much labelled Sri Lankan Tamil speech is required to obtain meaningful adaptation?

We therefore propose multiple data regimes:

This allows us to construct a learning curve:WER=f(H)WER = f(H)

where HH represents labelled speech hours.

Such an experiment is particularly valuable for low-resource language technology because it directly quantifies annotation efficiency.


Stage 5 — Evaluation

The evaluation framework should contain at least two principal metrics.

Word Error Rate

WER=S+D+INWER = \frac{S+D+I}{N}

where:

  • SS = substitutions
  • DD = deletions
  • II = insertions
  • NN = number of reference words

Character Error Rate

CER=S+D+INCER = \frac{S+D+I}{N}

where the unit is characters rather than words.

For Tamil ASR, CER provides a complementary signal because word-level errors can sometimes be influenced by segmentation and orthographic conventions.


Baseline vs Adapted Model

Our central experiment becomes:

Let:WERbaseWER_{base}

be the baseline error rate and:WERadaptedWER_{adapted}

be the error rate after Sri Lankan Tamil adaptation.

Then:ΔWER=WERbase−WERadapted\Delta WER = WER_{base} – WER_{adapted}

and:Improvement(%)=WERbase−WERadaptedWERbase×100Improvement(\%) = \frac{ WER_{base}-WER_{adapted} }{ WER_{base} } \times100

This provides a direct quantitative measure of adaptation effectiveness.


Error Analysis

Aggregate WER alone cannot tell us whether the model has actually learned the target speech distribution.

Therefore, we also need qualitative error analysis.

We can categorize residual errors into:

This can help determine whether remaining errors originate primarily from:

  • acoustic mismatch,
  • lexical mismatch,
  • transcript normalization,
  • language mixing,
  • segmentation,
  • or insufficient target-domain data.

Preventing Catastrophic Forgetting

An important consideration is that the pretrained model has learned multilingual knowledge.

Aggressive adaptation on a small Sri Lankan Tamil corpus may produce:θSLT\theta_{SLT}

that performs well on the target domain but loses useful general capabilities.

Therefore, we should monitor the trade-off between:Target Adaptation\text{Target Adaptation}

and:General Representation Preservation\text{General Representation Preservation}

This suggests an additional experiment:

This is especially important if the eventual objective is to create a reusable Tamil ASR model rather than a model specialized exclusively for one dataset.


Why We Are Comparing Multiple Indic ASR Foundations

We do not want the research to depend on the assumption that one pretrained model is inherently optimal.

The broader Indic ASR ecosystem provides several candidate model families.

IndicWav2Vec

Strong self-supervised representation-learning foundation for Indic speech.

IndicWhisper

Whisper-based multilingual adaptation demonstrated through large-scale Indic datasets such as Vistaar.

IndicConformer

Conformer-based hybrid CTC/RNNT architecture with a 600M multilingual checkpoint covering 22 Indian languages, including Tamil.

Large multilingual FastConformer systems

Recent multilingual ASR work demonstrates another route: extensive multilingual pretraining followed by targeted supervised adaptation to a new language or domain. The SraVaani fine-tuning workflow explicitly demonstrates this language/domain adaptation paradigm.

Therefore, our research direction is better represented as:Indic ASR Foundation→Sri Lankan Tamil Adaptation\boxed{ \text{Indic ASR Foundation} \rightarrow \text{Sri Lankan Tamil Adaptation} }

rather than:One Specific Model→Fine-Tuning\text{One Specific Model} \rightarrow \text{Fine-Tuning}


Proposed CTNLPR Experimental Framework

Our broader experimental framework can therefore be organized into four phases.

Phase 1 — Foundation Model Evaluation

Evaluate several suitable pretrained Indic ASR models directly on the same Sri Lankan Tamil test set.

This establishes which pretrained representation transfers best before adaptation.


Phase 2 — Sri Lankan Tamil Adaptation

Take the strongest practical candidate models and fine-tune them using the same Sri Lankan Tamil training corpus.


Phase 3 — Data-Efficiency Analysis

Train with progressively larger subsets:D10⊂D25⊂D50⊂DfullD_{10} \subset D_{25} \subset D_{50} \subset D_{full}

and compare:WER(Di)WER(D_i)

This allows us to quantify the value of additional annotation.


Phase 4 — Adaptation Strategy Analysis

Compare:

Frozen Encoder
       vs
Partial Fine-Tuning
       vs
Full Fine-Tuning

and evaluate:

  • WER
  • CER
  • convergence
  • GPU requirements
  • training time
  • generalization
  • target-domain improvement

The Complete CTNLPR Pipeline

The complete research architecture can be summarized as:


The Research Hypothesis

Our primary hypothesis is:H1:WERadapted<WERbaselineH_1: WER_{adapted} < WER_{baseline}

That is, Sri Lankan Tamil-specific supervised adaptation will reduce recognition errors compared with direct inference from a pretrained Indic ASR model.

A secondary hypothesis is:H2:WERadapted(H)↓as labelled speech hours H↑H_2: WER_{adapted}(H) \downarrow \quad \text{as labelled speech hours } H \uparrow

However, we do not assume that full fine-tuning will always outperform encoder-frozen adaptation.

H3: Optimal adaptation strategy = f(dataset size, model, domain)

This makes the research experimentally testable rather than prescriptive.


Why This Matters for Sri Lankan Tamil

The broader objective is not simply to obtain another ASR checkpoint.

Sri Lankan Tamil speech technology requires models that can perform effectively on speech actually produced within the target linguistic environment.

Large multilingual models provide an important foundation, but their usefulness for underrepresented speech varieties depends on how effectively they can be adapted.

Our research therefore investigates a scalable paradigm:

Large Multilingual ASR + Small High-Quality Target Corpus → Low-Resource Speech Technology

This is potentially much more practical than attempting to construct a Sri Lankan Tamil ASR foundation model entirely from scratch.


Toward a Sri Lankan Tamil Speech Foundation

The long-term direction of this research extends beyond one ASR model.

A successfully adapted model could become a foundation for:

  • speech-to-text systems
  • voice interfaces
  • conversational agents
  • speech-enabled search
  • document transcription
  • accessibility technologies
  • spoken-language information retrieval
  • speech analytics
  • downstream Tamil NLP systems

More importantly, the resulting methodology can establish a repeatable procedure for adapting large multilingual speech models to other underrepresented Sri Lankan speech varieties and domains.


Conclusion

At the Center for Tamil Natural Language Processing Research (CTNLPR), we are investigating a practical route toward improving Sri Lankan Tamil Automatic Speech Recognition through adaptation of pretrained Indic language ASR models.

The central idea is simple:

Pretrained Indic ASR → Sri Lankan Tamil Fine-Tuning → Adapted ASR

But the research questions are considerably deeper.

We want to determine:

  • which Indic ASR foundation transfers most effectively to Sri Lankan Tamil;
  • how much labelled speech is required;
  • whether frozen or full-model adaptation is more effective;
  • how transcript normalization influences performance;
  • how much WER/CER improvement can be achieved;
  • whether adaptation causes degradation of pretrained capabilities;
  • and which error categories remain after adaptation.

The goal is therefore not simply to fine-tune an ASR model.

The goal is to establish an evidence-driven, low-resource adaptation methodology for Sri Lankan Tamil speech recognition.

At CTNLPR, this work represents a step toward building speech technologies that extend beyond broad language coverage toward effective representation of Sri Lankan Tamil speech in real-world conditions.


References and Models Investigated

1. IndicConformer-600M-Multi — AI4Bharat

A 600M-parameter multilingual Conformer-based hybrid CTC/RNNT ASR model covering 22 officially recognized Indian languages, including Tamil. The model is released under the MIT license and the associated AI4Bharat repository provides resources for training and fine-tuning through its NeMo ecosystem.

IndicConformer-600M-Multi — Hugging Face

2. IndicWav2Vec — AI4Bharat

A family of wav2vec-based speech representation models developed for Indic languages, illustrating the use of self-supervised acoustic representation learning for low-resource ASR. AI4Bharat identifies IndicWav2Vec as one of its major ASR model families.

3. IndicWhisper / Vistaar — AI4Bharat

IndicWhisper represents Whisper-based adaptation for Indic languages. The Vistaar project provides diverse training and evaluation datasets across languages and domains and reports fine-tuning Whisper models on these datasets.

4. SraVaani-1.0 — ARTPARK-IISc

A multilingual FastConformer ASR system trained across a large set of Indian languages and dialects. Its associated 2026 fine-tuning guide demonstrates a language-agnostic adaptation workflow using custom audio and transcripts, including encoder-frozen and broader fine-tuning strategies.

SraVaani fine-tuning guide

5. IndicVoices / IndicASR — AI4Bharat

IndicVoices provides large-scale multilingual speech resources and an IndicASR checkpoint trained using IndicVoices data, demonstrating another large-scale pathway for multilingual Indic speech modeling.

6. Vividh-ASR

Recent work investigating recording-condition bias and adaptation strategies for Indic ASR, highlighting the importance of domain mismatch and careful fine-tuning when transferring pretrained speech models to real-world speech.


CTNLPR Research Direction

Pretrained Indic ASR → Sri Lankan Tamil Adaptation → Low-Resource Evaluation → Robust Sri Lankan Tamil Speech Recognition

The central research objective is therefore:

To develop and systematically evaluate an efficient adaptation methodology for transforming large multilingual Indic ASR models into robust Sri Lankan Tamil speech recognition systems using limited labelled target-language data.