Testing AI Voice Quality: How to Measure WER and MOS for Vietnamese TTS

Learn how to measure Vietnamese TTS quality using WER and MOS. A practical guide for developers to ensure accuracy, naturalness, and robust code-switching.

Developing reliable Text-to-Speech (TTS) applications requires more than just generating audio; it demands rigorous verification of naturalness and accuracy. Many projects skip systematic quality testing, resulting in poor user experiences and high rejection rates. This guide outlines a professional workflow for evaluating AI speech, focusing on two critical metrics for the Vietnamese language: Word Error Rate (WER) and Mean Opinion Score (MOS).

Why Specific Measurement Matters

When deploying TTS, "audible" is not enough. You need voices that meet technical standards. Without quantitative metrics, comparing model versions becomes subjective and difficult. By applying standardized measurement methods, engineering and product teams can make data-driven decisions, optimize performance, and minimize errors.

Understanding WER in Vietnamese TTS

Word Error Rate (WER) measures the ratio of errors when comparing the TTS system’s output (transcribed back via Automatic Speech Recognition, or ASR) against the original source text. While WER is traditionally an ASR metric, in TTS, it verifies content accuracy: does the system read the correct words without omissions or additions?

The basic WER formula is:

$$WER = \frac{S + D + I}{N}$$

Where:

  • S: Number of substitutions
  • D: Number of deletions
  • I: Number of insertions
  • N: Total number of words in the original sentence

Challenges with Vietnamese Vietnamese presents unique challenges that make WER measurement more complex than for English. Because Vietnamese does not use spaces between words, word segmentation can introduce bias if not handled correctly. Additionally, compound words and reduplication affect the accuracy of the comparison process.

Practical Tips

  • Normalize Input Data: Ensure both the original text and the recognized transcript are normalized for punctuation and capitalization.
  • Use a Dedicated Test Set: Create a test dataset including sentences with complex grammar, numerical data, and Vietnamese proper nouns for comprehensive evaluation.
  • Automate the Process: Use scripts to run batch tests and calculate WER automatically rather than relying on manual comparison.

Subjective Evaluation with MOS Score

If WER measures accuracy, Mean Opinion Score (MOS) measures the naturalness and intelligibility of the voice. This is the most important metric for evaluating the end-user experience.

MOS is based on human evaluation. Listeners rate TTS-generated audio clips on a scale of 1 to 5:

  • 1: Bad, inaudible.
  • 2: Poor, very uncomfortable to listen to.
  • 3: Fair, acceptable.
  • 4: Good, natural.
  • 5: Excellent, indistinguishable from human speech.

Standard MOS Implementation Process

  1. Sample Selection: Randomly select 100–200 sentences from the test set, ensuring diversity in length and content.
  2. Recruit Evaluators: Engage at least 5–10 independent evaluators. Ideally, they should have good hearing and strong proficiency in Vietnamese.
  3. Evaluation Interface: Use an online survey tool where each evaluator listens to and scores each sentence individually.
  4. Data Processing: Remove samples with significant scoring deviations among evaluators to ensure objectivity.

Notes for Vietnamese Evaluation

  • Tone Sensitivity: Evaluators must pay close attention to the six tones in Vietnamese. A voice with incorrect tonal realization will be rated low, even if the phonemes are clear.
  • Prosody and Rhythm: Machine voices often sound "stiff" in questions or exclamations. Include these sentence types in your test set to capture this nuance.

Comprehensive Testing Workflow

For a holistic view, combine both WER and MOS. Here is a quick comparison:

Metric Goal Method Pros Cons
WER Content Accuracy Automated (ASR + Script) Fast, low cost, objective Does not reflect naturalness
MOS Naturalness/Intelligibility Manual (Human) Reflects real user experience Slow, expensive, subjective

Building a Testing Pipeline

  • Step 1: Run automated WER checks on the entire large dataset to eliminate severe content errors.
  • Step 2: Select samples with low WER for manual MOS evaluation.
  • Step 3: Analyze MOS results by sentence group (short, long, numeric) to identify specific weaknesses.

Why Focus on Vietnamese?

Many multilingual TTS models are optimized for English or Chinese, often resulting in suboptimal Vietnamese quality. Common issues include:

  • Mispronunciation of loanwords.
  • Monotone rhythm lacking emotion.
  • Errors in handling code-switching (English words interspersed in Vietnamese sentences).

A robust TTS system must handle these cases smoothly. Specific AI voice evaluation for Vietnamese helps detect these unique errors and improves model effectiveness.

Optimization with AIVISION

In deploying speech solutions, having a platform that supports measurement and testing is crucial. AIVISION, with experience deploying speech AI for hundreds of enterprises in Vietnam and abroad, provides tools to support TTS and ASR development.

AIVISION’s aiv-tts-S.1.0 is designed specifically for Vietnamese, supporting interspersed English and telephony formats for call centers. Its standard REST and WebSocket APIs allow developers to easily build automated testing pipelines. You can use the API to generate audio samples in bulk and compare them with original text to calculate WER quickly.

To learn more about how AIVISION supports speech AI projects, visit our Pricing page for details or read other in-depth articles on our Blog.

Conclusion

Measuring TTS quality is not optional; it is a mandatory requirement for serious projects. Combining WER to ensure accuracy and MOS to ensure naturalness helps you build high-quality AI voice products tailored to the Vietnamese language.

Start by building a standardized test dataset and applying a rigorous measurement process. If you are looking for a stable and easy-to-integrate Vietnamese TTS solution, Start free with AIVISION today to experience the difference.

Frequently asked questions

How do WER and MOS differ in TTS evaluation?

WER measures content accuracy (whether the correct words are spoken) using automated methods, while MOS measures the naturalness and intelligibility of the voice through subjective human evaluation.

How many people are needed to evaluate a Vietnamese MOS score?

It is recommended to have a minimum of 5–10 independent evaluators to ensure objectivity. A larger number increases reliability, but you must balance this with cost and time constraints.

How should I handle word segmentation errors when measuring WER in Vietnamese?

Use an accurate Vietnamese word segmentation tool to split sentences into distinct words before comparison. Ensure both the original text and the recognized transcript are processed using the same segmentation rules.

Does AIVISION support TTS testing?

Yes, AIVISION provides standard TTS and ASR APIs that allow developers to automate audio generation and speech recognition processes, making it easy to calculate quality metrics like WER.

Is it necessary to evaluate English in Vietnamese TTS?

If your product supports code-switching (mixing English and Vietnamese), you should include samples with English words in your test set to ensure that foreign terms are read naturally and accurately.

Try AIVISION's Vietnamese speech AI

$10 free every day for Speech-to-Text, Text-to-Speech and LLM.

Start free → Contact

By

Dr. Giang Vo — CEO, AIVISION

Dr. Giang Vo is the CEO of AIVISION, leading the development of s2speech — speech recognition, AI voices and a large language model for Vietnamese — and working with companies at home and abroad on their deployments.

0981 419 967 · giang.vo@aivgroups.com

#Vietnamese TTS testing#WER calculation#MOS score#AI voice quality#speech synthesis evaluation#code-switching#AIVISION TTS

Related articles