Beyond Fine-Tuning: Building a Sri Lankan Tamil Text-to-Speech System

Why Sri Lankan Tamil Speech Synthesis Requires Target-Specific Data and Models

Text-to-Speech (TTS) has advanced rapidly from conventional statistical parametric synthesis to neural architectures capable of generating highly natural, expressive and speaker-controlled speech.

Modern systems such as VITS, VITS2, YourTTS, StyleTTS2, F5-TTS, IndicF5 and newer multilingual speech models have demonstrated impressive capabilities in multilingual synthesis, speaker adaptation and voice cloning.

This naturally raises an important question for Sri Lankan Tamil:

If powerful multilingual TTS models already exist, why not simply fine-tune one of them for Sri Lankan Tamil?

The answer is not that fine-tuning is impossible.

Rather, the fundamental challenge is that language support is not equivalent to target-domain speech modeling.

A model can generate Tamil speech while still failing to accurately represent the phonetic, prosodic, lexical, speaker and acoustic characteristics of Sri Lankan Tamil speech.

This distinction motivates our research direction at the Center for Tamil Natural Language Processing Research (CTNLPR).


1. Tamil Support Does Not Mean Sri Lankan Tamil Support

Modern multilingual TTS models increasingly support Tamil.

For example, IndicF5 supports Tamil as one of 11 Indian languages and was trained using 1,417 hours of speech collected from Rasa, IndicTTS, LIMMITS and IndicVoices-R.

Similarly, the Tamil component of the IndicTTS ecosystem contains approximately 20.33 hours of studio-quality Tamil speech from two native speakers.

These resources are extremely valuable.

However, their target distribution is primarily Indian Tamil / Indian multilingual speech, rather than a dedicated Sri Lankan Tamil speech distribution.

This creates a fundamental domain-adaptation problem:

P Indian Tamil ( X , Y ) ≠ P Sri Lankan Tamil ( X , Y ) P_{\text{Indian Tamil}}(X,Y) \neq P_{\text{Sri Lankan Tamil}}(X,Y)

where:

  • XX represents speech acoustics,
  • YY represents linguistic/textual information.

The difference is not simply vocabulary.

It can emerge from:

  • pronunciation realization,
  • phonetic variation,
  • prosody,
  • rhythm,
  • speaking rate,
  • intonation,
  • lexical preferences,
  • code-switching patterns,
  • speaker demographics,
  • recording conditions,
  • conversational style,
  • and regional linguistic variation.

Therefore:

Tamil is a language label; Sri Lankan Tamil is a target speech distribution.

This distinction is fundamental to our research.


2. The Dataset Is the Real Foundation

One of the strongest lessons from recent Indic TTS research is that data quality and diversity are as important as architecture.

AI4Bharat’s IndicVoices-R work constructed 1,704 hours of high-quality speech from more than 10,000 speakers across 22 Indian languages. The researchers specifically focused on speaker diversity, natural speech and high-quality restoration because these factors are critical for scaling TTS systems.

The dataset is not simply:

audio + transcript

It represents a much richer distribution of:

Text+Speech+Speaker+Prosody+Acoustic Conditions\text{Text} + \text{Speech} + \text{Speaker} + \text{Prosody} + \text{Acoustic Conditions}

This is particularly important for Sri Lankan Tamil.

A dedicated system therefore requires a Sri Lankan Tamil speech corpus designed specifically for speech synthesis, rather than simply recycling an ASR dataset.


3. ASR Data and TTS Data Are Not the Same

This distinction is extremely important.

An ASR dataset primarily asks:Speech→Text\text{Speech} \rightarrow \text{Text}

A TTS dataset requires the inverse mapping:Text→Speech\text{Text} \rightarrow \text{Speech}

For ASR, it is acceptable to collect spontaneous speech from many environments.

For high-quality TTS, however, the dataset needs much stronger control over:

  • transcript accuracy
  • segmentation
  • recording quality
  • speaker consistency
  • text coverage
  • pronunciation
  • prosodic variation
  • alignment
  • and acoustic noise.

IndicVoices-R demonstrates this distinction very clearly: its pipeline takes large-scale ASR-oriented speech and applies enhancement and processing to construct high-quality TTS data.

Therefore, at CTNLPR, the goal should not simply be:

Collect Sri Lankan Tamil audio.

It should be:

Construct a synthesis-oriented Sri Lankan Tamil speech corpus.


4. Why Existing Voice Cloning Is Not Automatically Enough

Modern voice-cloning systems make the problem appear deceptively easy.

For example, XTTS-v2 can clone a speaker using a short reference recording and supports multilingual voice generation. However, its documented language list contains 17 languages and does not include Tamil.

This illustrates an important limitation.

A voice-cloning model has to solve two different problems:

Speaker identity

zspeakerz_{speaker}

and:

Linguistic/acoustic realization

zlanguagez_{language}

A good cloning system needs to combine them:Speech=f(Text,Speaker,Language,Style)Speech = f(Text, Speaker, Language, Style)

A model may successfully reproduce the timbre of a speaker while still producing incorrect:

  • pronunciation,
  • phoneme realization,
  • stress,
  • intonation,
  • rhythm,
  • or language-specific prosody.

Therefore:

Voice cloning is not equivalent to language-specific speech synthesis.


5. Sri Lankan Tamil Voice Cloning Is a Different Research Problem

Consider a Sri Lankan Tamil speaker.

A generic multilingual voice-cloning model may extract:zs=speaker embeddingz_s = \text{speaker embedding}

from the reference recording.

But speaker identity alone does not encode the complete target-language distribution.

The synthesized output still depends on the model’s learned:P(speech∣text,speaker,language)P(\text{speech}|\text{text},\text{speaker},\text{language})

If the model has learned the language primarily from Indian Tamil data, then the speaker embedding may correctly represent the person’s vocal characteristics while the linguistic realization remains biased toward the pretrained distribution.

This can produce a particularly interesting failure mode:

Correct speaker identity + incorrect target-language realization.

That is one of the reasons a dedicated Sri Lankan Tamil corpus remains valuable even in the era of zero-shot voice cloning.


6. Existing Research Actually Supports This Direction

This is not merely a theoretical concern.

The LIMMITS’24 challenge explicitly studied multilingual TTS and voice cloning using multilingual base models, additional multi-speaker corpora and few-shot target-speaker adaptation. The challenge provided 560 hours of studio-quality speech across seven Indian languages and evaluated both naturalness and speaker similarity.

One published LIMMITS system used VITS2, multilingual speaker/language conditioning and BERT-based contextual modeling, followed by few-shot target-speaker adaptation.

This gives us an important research precedent:Multilingual Base Model+Target Data+Speaker Adaptation\text{Multilingual Base Model} + \text{Target Data} + \text{Speaker Adaptation}

can work very well.

But it also demonstrates why target data remains central.


7. VITS Provides an Important Baseline

For a Sri Lankan Tamil TTS research, VITS/VITS2 is still a very strong baseline.

VITS combines:

  • variational inference,
  • normalizing flows,
  • adversarial training,
  • stochastic duration modeling,
  • text-to-speech generation,
  • and neural vocoding

within an end-to-end architecture.

YourTTS demonstrates how the VITS architecture can be extended to multilingual and multi-speaker TTS. It achieved strong zero-shot multi-speaker synthesis and showed that very small amounts of target-speaker data can be useful for adaptation.

This makes VITS/VITS2 particularly attractive for CTNLPR because it provides a well-understood experimental baseline.

A research pipeline could therefore begin with:Sri Lankan Tamil Corpus→VITS/VITS2→Sri Lankan Tamil TTS\text{Sri Lankan Tamil Corpus} \rightarrow \text{VITS/VITS2} \rightarrow \text{Sri Lankan Tamil TTS}

and subsequently compare it against modern pretrained systems.


8. But VITS Should Not Be the Final Answer

A major mistake would be to assume:

“We should use VITS because it is popular.”

The architecture should instead be treated as an experimental variable.

Modern TTS has moved beyond classical VITS-style architectures.


9. StyleTTS2: Explicit Modeling of Speaking Style

StyleTTS2 introduces a fundamentally different approach.

Instead of treating style as a simple deterministic conditioning variable, StyleTTS2 models style as a latent random variable using diffusion.

It also incorporates pretrained speech language models such as WavLM as discriminators and uses differentiable duration modeling for end-to-end training. The original work demonstrated strong naturalness and zero-shot speaker adaptation.

This is highly relevant to Sri Lankan Tamil because a future system should ideally model more than:Text→SpectrogramText \rightarrow Spectrogram

It should capture:Text+Speaker+Style+Prosody→SpeechText + Speaker + Style + Prosody \rightarrow Speech

This opens the possibility of studying:

  • neutral speech,
  • conversational speech,
  • expressive speech,
  • speaking-rate variation,
  • emotional speech,
  • and speaker-specific prosody.

10. F5-TTS Changes the Architecture Landscape

Another major development is F5-TTS.

F5-TTS uses flow matching with a Diffusion Transformer (DiT) and removes several traditional components such as explicit duration modeling, text encoders and phoneme alignment. It operates in a fully non-autoregressive framework and demonstrated strong multilingual zero-shot synthesis.

This is important because it gives us another possible research direction:Text+Reference Speech→Flow-Matching TTS\text{Text} + \text{Reference Speech} \rightarrow \text{Flow-Matching TTS}

rather than building a traditional:Text→Duration→Mel→VocoderText \rightarrow Duration \rightarrow Mel \rightarrow Vocoder

pipeline.


11. IndicF5 Is Particularly Interesting for Our Research

IndicF5 is arguably the most important model to benchmark before deciding to build a completely new system.

It is trained on approximately 1,417 hours of high-quality speech from:

  • Rasa,
  • IndicTTS,
  • LIMMITS,
  • IndicVoices-R,

and supports 11 Indian languages including Tamil.

It also uses reference-prompt audio and its transcript to guide speaker and prosody characteristics during generation.

Therefore, We should not immediately discard IndicF5.

Instead, it should become an experimental baseline:

Experiment 1

Zero-shot IndicF5 → Sri Lankan Tamil

Experiment 2

Fine-tuned IndicF5 → Sri Lankan Tamil

Experiment 3

Sri Lankan Tamil-specific TTS model

Then compare:Ezero−shotE_{zero-shot}EadaptedE_{adapted}EdedicatedE_{dedicated}

This would make the research significantly stronger.


12. Indic-Mio Represents Another New Baseline

The current TTS ecosystem has also moved toward newer multilingual architectures.

For example, Indic-Mio reports support for all 22 scheduled Indian languages plus English, 44 kHz synthesis, sub-0.1 real-time factor and zero-shot voice cloning through speaker embeddings.

That makes it another valuable baseline for Sri Lankan Tamil adaptation experiments.

The research question becomes:

How much target-specific adaptation is required before a large multilingual Indic TTS model adequately represents Sri Lankan Tamil speech?

That is much more scientifically interesting than simply asking whether the model can generate Tamil.


13. The Real Research Gap

The existing ecosystem is already strong.

We have:

  • multilingual TTS,
  • Tamil TTS,
  • zero-shot voice cloning,
  • few-shot speaker adaptation,
  • expressive TTS,
  • flow-matching TTS,
  • large Indian speech corpora.

Yet there remains a significant gap:Sri Lankan Tamil-specific TTS\boxed{ \text{Sri Lankan Tamil-specific TTS} }

The major resources we found are predominantly centered around Indian multilingual speech.

For example, the Tamil IndicTTS dataset contains about 20.33 hours from two Tamil speakers, while IndicVoices-R provides large-scale coverage across Indian languages and speakers.

These are excellent resources—but they are not a dedicated Sri Lankan Tamil TTS corpus.


14. A Dedicated Dataset Is Therefore the First Research Contribution

Before deciding which architecture is best, We should construct a high-quality dataset.

The dataset should ideally contain:

Text diversity

Cover:

  • common vocabulary,
  • rare vocabulary,
  • long-form sentences,
  • short utterances,
  • questions,
  • commands,
  • numbers,
  • dates,
  • abbreviations,
  • names,
  • technical terminology,
  • conversational language,

Phonetic coverage

The corpus should be phonologically and syllabically balanced.

This is important because Rasa specifically reports the importance of syllabically balanced data for expressive TTS.

Speaker diversity

For a multi-speaker system:

  • multiple speakers,
  • different ages,
  • genders,
  • speaking styles,
  • voice characteristics,
  • and recording conditions.

Prosodic diversity

Collect:

  • neutral speech,
  • conversational speech,
  • interrogatives,
  • declaratives,
  • emphasis,
  • different speaking rates,
  • expressive speech.

Rasa demonstrates that even relatively small amounts of expressive data can improve expressive TTS when combined with neutral speech.


15. Text Normalization Becomes a Core Component

A Sri Lankan Tamil TTS system cannot simply depend on raw Unicode text.

A proper frontend needs to handle:

Text → Normalization → Tokenization → Phonological Representation → Acoustic Generation

The frontend must correctly handle:

  • Unicode normalization
  • punctuation
  • numerals
  • abbreviations
  • foreign words
  • English code-switching
  • named entities
  • pronunciation variants.

This becomes especially important if we train a target-specific model.

A model cannot learn a pronunciation distinction if the training representation consistently collapses it.


16. Grapheme-to-Phoneme Modeling May Become Important

A deeper research direction is to investigate whether a Sri Lankan Tamil TTS system should operate directly on characters or use an explicit phonological representation.

We can compare:Grapheme→AcousticModelGrapheme \rightarrow Acoustic Model

against:Grapheme→Phoneme→AcousticModelGrapheme \rightarrow Phoneme \rightarrow Acoustic Model

The second approach provides an opportunity to explicitly model pronunciation.

This could be particularly useful when the written form remains similar while pronunciation differs across speech communities.

Therefore, at CTNLPR we are going to investigate:

character-based TTS vs phoneme-aware TTS.


17. Voice Cloning Should Be Treated as a Separate Evaluation Dimension

A TTS system can have good naturalness without being good at speaker similarity.

Therefore, we should evaluate at least:

Naturalness

MOSMOS

Intelligibility

CER/WERCER/WER

Speaker similarity

SIMspeakerSIM_{speaker}

Prosodic similarity

Human evaluation or suitable prosodic metrics.

This gives:TTS Quality=f(Naturalness,Intelligibility,Speaker Similarity,Prosody)TTS\ Quality = f( Naturalness, Intelligibility, Speaker\ Similarity, Prosody )

A model that produces fluent speech but sounds like the wrong speaker should not be considered a successful voice-cloning system.

18. Our Experimental Design at CTNLPR

Rather than assuming that either fine-tuning an existing multilingual TTS model or training a new Sri Lankan Tamil model from scratch is inherently superior, our research at CTNLPR will treat both approaches as competing experimental hypotheses.

The objective is to establish a controlled experimental framework that allows us to quantify the contribution of:

  • pretrained multilingual knowledge,
  • Sri Lankan Tamil-specific supervision,
  • target-domain adaptation,
  • model architecture,
  • speaker diversity,
  • and dataset scale.

This leads us to a three-level experimental progression:

Zero-Shot → Fine-Tuned → Target-Specific

This progression allows us to distinguish between knowledge transfer, target-domain adaptation, and target-specific model learning.


18.1 Baseline A — Existing Tamil-Capable TTS Models

Our first step at CTNLPR will be to establish strong baselines using existing multilingual TTS systems capable of generating Tamil.

Models such as IndicF5, Indic-Mio, and other suitable multilingual TTS systems will be considered depending on their language coverage, architecture, availability of pretrained checkpoints, licensing, and adaptability.

The purpose of this stage is deliberately simple:

How well can an existing multilingual TTS system generate Sri Lankan Tamil without any target-specific adaptation?

This establishes the zero-shot baseline.

We will evaluate the generated speech using the same evaluation protocol that will later be applied to our adapted and target-specific systems.

This is important because without a zero-shot baseline, it becomes difficult to quantify the actual contribution of Sri Lankan Tamil-specific training.


18.2 Baseline B — Sri Lankan Tamil Adaptation

The next stage of our research will investigate fine-tuning existing multilingual TTS models using our Sri Lankan Tamil speech corpus.

The experimental formulation is:

Pretrained Multilingual TTS + Sri Lankan Tamil Corpus → Adapted Sri Lankan Tamil TTS

Here, the pretrained model provides the initial representation and synthesis capability, while the Sri Lankan Tamil corpus provides target-domain supervision.

The central question is:

How much can the pretrained multilingual model be shifted toward the Sri Lankan Tamil speech distribution through supervised adaptation?

This stage will allow us to measure the value of transfer learning before investing in fully target-specific model training.

The approach is also consistent with existing multilingual TTS research. LIMMITS’24, for example, explicitly studied multilingual multi-speaker TTS and few-shot voice cloning using pretrained multilingual systems and target-speaker adaptation.


18.3 System C — Target-Specific VITS / VITS2

In parallel with adaptation experiments, CTNLPR will investigate target-specific TTS training using established neural TTS architectures.

Our first architecture family will be VITS/VITS2.

Rather than starting from an existing Tamil checkpoint, we can initialize the TTS model using the VITS/VITS2 architecture and train it using our Sri Lankan Tamil corpus.

The conceptual formulation is:

Sri Lankan Tamil Corpus → VITS/VITS2 Training → Sri Lankan Tamil TTS

This provides an important comparison against pretrained-model adaptation.

A particularly relevant precedent is the LIMMITS’24 research ecosystem, where VITS2-based multilingual multi-speaker systems were developed for Indic TTS and voice cloning. One published LIMMITS’24 system used VITS2 together with multilingual and speaker conditioning and subsequently performed few-shot speaker adaptation.

For CTNLPR, however, the objective is different: we want to investigate whether a target-specific Sri Lankan Tamil model can learn characteristics that are not sufficiently captured by adaptation of a general multilingual model.


18.4 System D — Style-Aware TTS with StyleTTS2

The second target-specific architecture we plan to investigate is StyleTTS2.

This is particularly relevant because our objective extends beyond simple intelligibility.

A high-quality TTS system should model:

  • linguistic content,
  • speaker identity,
  • prosody,
  • speaking style,
  • duration,
  • acoustic realization,
  • and natural variation.

StyleTTS2 models style as a latent variable using diffusion and incorporates pretrained speech-language-model representations into adversarial training. Its experiments also demonstrate strong zero-shot speaker adaptation capabilities.

Therefore, at CTNLPR, StyleTTS2 provides an opportunity to investigate a different hypothesis:

Can explicit style modeling improve the naturalness and speaker characteristics of Sri Lankan Tamil synthesis compared with conventional VITS-style generation?

We can investigate both:StyleTTS2 from Scratch\text{StyleTTS2 from Scratch}

and, where a compatible pretrained checkpoint is appropriate:

Pretrained StyleTTS2 + Sri Lankan Tamil Data → Fine-Tuned StyleTTS2

The official StyleTTS2 implementation provides both training and fine-tuning workflows, including fine-tuning a pretrained multispeaker model using limited target-speaker data.


18.5 System E — Modern Flow-Matching TTS

Our architecture investigation will not be restricted to VITS-family models.

We will also consider modern flow-matching TTS architectures, particularly F5-style approaches, as the technology landscape continues to move toward large-scale flow-based speech generation.

This gives us another architectural family for comparison:

Sri Lankan Tamil Corpus → Flow-Matching TTS → Sri Lankan Tamil Speech

The purpose is not to assume that a newer architecture will automatically outperform VITS or StyleTTS2.

Instead, we want to determine whether architectural differences translate into measurable improvements for the specific target domain of Sri Lankan Tamil.


18.6 The CTNLPR Experimental Matrix

The resulting experimental framework can therefore be viewed as:

SystemInitial ModelSri Lankan Tamil DataObjective
AExisting multilingual TTSNoneZero-shot baseline
BExisting multilingual TTSFine-tuningTarget adaptation
CVITS/VITS2Full trainingTarget-specific model
DStyleTTS2Full training / adaptationStyle-aware synthesis
EFlow-matching architectureFull training / adaptationModern generative TTS

This structure allows us to separate model capability from adaptation capability and dataset contribution.


19. Proposed CTNLPR Research Pipeline

Based on these experiments, our planned research pipeline consists of five major phases.

Phase 1 — Sri Lankan Tamil TTS Corpus Construction

Our first priority is to construct a high-quality Sri Lankan Tamil speech-synthesis corpus.

The dataset will pair:Sri Lankan Tamil Text+Sri Lankan Tamil Speech\text{Sri Lankan Tamil Text} + \text{Sri Lankan Tamil Speech}

and will be designed specifically for TTS rather than simply repurposing an ASR corpus.

We will focus on:

  • high-quality recordings,
  • accurate transcription,
  • speaker metadata,
  • text diversity,
  • phonetic coverage,
  • speaker diversity,
  • consistent segmentation,
  • normalization,
  • and suitable train/validation/test partitioning.

The resulting resource will serve as the common dataset across our subsequent experiments.


Phase 2 — Existing Model Benchmarking

Once the corpus is established, we will benchmark suitable existing multilingual TTS models.

The first objective is to answer:

What is the performance of current multilingual TTS models on Sri Lankan Tamil before any target-specific training?

This gives us the zero-shot baseline:Mpretrained→DSLTtestM_{pretrained} \rightarrow D_{SLT}^{test}

The test set will remain isolated from model adaptation.


Phase 3 — Target-Domain Adaptation

The strongest suitable pretrained model will then be adapted using the Sri Lankan Tamil training corpus.Mpretrained+DSLTtrain→MadaptedM_{pretrained} + D_{SLT}^{train} \rightarrow M_{adapted}

The adapted system will be evaluated against exactly the same held-out test set used for the zero-shot baseline.

This gives us a direct measurement of adaptation gain:ΔQ=Q(Madapted)−Q(Mzero−shot)\Delta Q = Q(M_{adapted}) – Q(M_{zero-shot})

where QQ represents the relevant synthesis quality measure.


Phase 4 — Target-Specific Architecture Training

We will then investigate whether a model trained specifically around the Sri Lankan Tamil corpus can provide further improvements.

Candidate architectures include:

  • VITS2
  • StyleTTS2
  • flow-matching TTS architectures

The purpose is not simply to produce another model checkpoint.

Instead, we want to determine whether target-specific architectural training provides measurable benefits beyond transfer learning.


Phase 5 — Controlled Comparative Evaluation

Finally, all systems will be evaluated under the same experimental conditions.

Our evaluation will consider multiple dimensions:

Naturalness

How human-like does the generated speech sound?

Intelligibility

Can listeners understand the synthesized output correctly?

Speaker Similarity

Does generated speech preserve the intended speaker characteristics?

Pronunciation

Are Sri Lankan Tamil linguistic and phonological units realized correctly?

Prosody

Does the generated speech exhibit natural rhythm, duration, stress and intonation?

Robustness

How does the model behave across different speakers, text types and acoustic conditions?

This multidimensional evaluation is important because a lower error rate alone does not guarantee natural or speaker-faithful TTS.

LIMMITS’24, for example, evaluated multilingual voice cloning using both naturalness and speaker-similarity subjective tests, illustrating the importance of evaluating dimensions beyond transcription-related measures.


20. Why Are We Building a Sri Lankan Tamil-Specific Model?

The motivation behind our research is therefore not that existing multilingual TTS models are incapable of generating Tamil.

They clearly demonstrate substantial multilingual capabilities.

Our question is more specific:

Are these pretrained representations sufficient to model Sri Lankan Tamil speech at high fidelity?

A multilingual model may already contain useful Tamil representations, but its learned distribution may not adequately represent the target speech domain.

Conceptually:Ppretrained(X,Y)≠PSriLankanTamil(X,Y)P_{pretrained}(X,Y) \neq P_{SriLankanTamil}(X,Y)

where the mismatch can arise from differences in:

  • speech acoustics,
  • pronunciation,
  • phonetic realization,
  • lexical distribution,
  • prosody,
  • speaking style,
  • speaker populations,
  • code-switching,
  • recording environments,
  • and linguistic context.

Therefore, our research does not reject transfer learning.

Instead, we use transfer learning as one of our experimental baselines.

The scientific question is whether adaptation is sufficient—or whether target-specific training provides additional gains.


21. Research Hypotheses

Based on this experimental design, CTNLPR will investigate several hypotheses.

H₁ — Adaptation Hypothesis

QSLTadapted>QSLTzero−shot\boxed{ Q_{SLT}^{adapted} > Q_{SLT}^{zero-shot} }

We expect Sri Lankan Tamil-specific supervised adaptation to improve target-domain synthesis compared with direct inference from the pretrained model.

However, this will be empirically tested rather than assumed.


H₂ — Target-Specific Training Hypothesis

QSLTdedicated≥QSLTadapted\boxed{ Q_{SLT}^{dedicated} \geq Q_{SLT}^{adapted} }

We will investigate whether a model trained specifically on Sri Lankan Tamil data can match or exceed a fine-tuned multilingual model.

Importantly, we do not assume this hypothesis will necessarily hold.

If fine-tuning performs equally well or better, that itself becomes an important finding.


H₃ — Architecture Hypothesis

QSLTarchitecture1≠QSLTarchitecture2\boxed{ Q_{SLT}^{architecture_1} \neq Q_{SLT}^{architecture_2} }

We will investigate whether VITS2, StyleTTS2 and modern flow-matching approaches exhibit different performance characteristics for Sri Lankan Tamil.

The objective is to identify the architecture–data combination that provides the strongest trade-off between:

  • naturalness,
  • intelligibility,
  • speaker similarity,
  • prosody,
  • computational cost,
  • and data requirements.

22. The Larger CTNLPR Research Question

This leads to a broader question that we believe is more valuable than simply asking whether an existing TTS model can generate Sri Lankan Tamil:

How much Sri Lankan Tamil-specific data and architectural adaptation are required to transform a multilingual TTS foundation into a high-fidelity Sri Lankan Tamil speech synthesis system?

Our research therefore investigates:

Multilingual TTS + Sri Lankan Tamil Data + Target-Specific Adaptation —-* Sri Lankan Tamil Speech Synthesis

while explicitly comparing:Fine-TuningvsTarget-Specific Training\boxed{ \text{Fine-Tuning} \quad vs \quad \text{Target-Specific Training} }

This comparison forms the core of our experimental methodology.


23. The CTNLPR Vision

At CTNLPR, our objective is not simply to release another TTS checkpoint.

We aim to establish a data-centric and architecture-aware methodology for Sri Lankan Tamil speech synthesis.

Through this research, we plan to investigate:

  • How much labelled speech is actually required?
  • Which multilingual TTS model transfers most effectively?
  • How much improvement does target-domain fine-tuning provide?
  • Does a dedicated Sri Lankan Tamil model outperform adaptation?
  • Which architecture provides the best naturalness–data trade-off?
  • How does speaker diversity affect generalization?
  • How much data is required for reliable speaker adaptation and voice cloning?
  • Does explicit style modeling improve Sri Lankan Tamil prosody?
  • Do modern flow-matching architectures provide measurable benefits?
  • How strongly does the text frontend influence pronunciation?
  • What linguistic and acoustic errors remain after adaptation?

These questions transform the work from simply training a TTS model into a systematic low-resource speech technology study.


Conclusion

The rapid development of multilingual TTS has created an important opportunity for under-resourced speech communities.

However, at CTNLPR, we do not want to assume that either fine-tuning an existing model or training a new architecture from scratch is automatically the correct solution.

Instead, we plan to establish the evidence experimentally.

Our methodology will therefore progress from:

Existing multilingual TTS

→ Zero-shot evaluation

→ Sri Lankan Tamil fine-tuning

→ Target-specific VITS/VITS2 training

→ Style-aware TTS investigation

→ Modern flow-matching approaches

→ Controlled comparative evaluation

This allows us to answer a much more meaningful research question:

Can existing multilingual TTS models be sufficiently adapted to Sri Lankan Tamil, or does high-fidelity Sri Lankan Tamil speech synthesis require a dedicated target-specific dataset and model architecture?

Rather than assuming the answer beforehand, CTNLPR will test both pathways under a common experimental framework.

That is the central idea behind our planned research.


Key References

  • Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings — demonstrates the importance of neutral/expressive data composition and syllabic balance for low-resource Indian TTS. (ISCA Archive)
  • IndicVoices-R — 1,704 hours, 10,496 speakers and 22 Indian languages, demonstrating the importance of large-scale speaker and speech diversity for Indian TTS. (GitHub)
  • IndicF5 — multilingual flow-matching TTS supporting Tamil and trained on 1,417 hours of Indian speech. (GitHub)
  • YourTTS — VITS-based multilingual and multi-speaker TTS with zero-shot speaker adaptation and low-resource fine-tuning. (Proceedings of Machine Learning Research)
  • StyleTTS2 — style diffusion and speech-language-model-based adversarial training for high-quality and speaker-adaptive synthesis. (PubMed Central (PMC))
  • F5-TTS — flow-matching + Diffusion Transformer approach for highly natural multilingual TTS. (arXiv)
  • LIMMITS’24 — multilingual Indic TTS and voice-cloning research using multi-speaker data and few-shot speaker adaptation. (IEEE Signal Processing Society)
  • IndicTTS Tamil — approximately 20.33 hours of studio-quality Tamil speech from two speakers, illustrating the existing Indian Tamil TTS data landscape. (Hugging Face)
  • TaLK-Corpus — recent Sri Lankan Tamil speech resource demonstrating the growing availability of Sri Lankan Tamil speech data, although it is primarily an ASR benchmark rather than a purpose-built TTS corpus. (GitHub)

Leave a Reply