Handling Numbers, Dates and Currency in Vietnamese Speech-to-Text

Learn how to accurately transcribe and normalize Vietnamese numbers, dates, and currency in speech-to-text. Discover technical strategies and AIVISION’s approach.

In Natural Language Processing (NLP), converting Vietnamese speech into text presents a significant challenge when encountering non-linguistic elements such as quantities, dates, and currency. Users expect a system to do more than just "hear"; it must "understand" and normalize text logically to avoid semantic errors. This article analyzes advanced techniques for handling these cases effectively, enhancing the accuracy and practical value of the output data.

The Unique Challenge of Vietnamese Numbers and Dates

Unlike English, Vietnamese has multiple ways to read numbers and dates depending on context—whether in mathematical, conversational, or commercial transaction settings. For example, the number 1,000,000 can be read as "one million," "one point zero zero zero zero zero," or "one zero zero zero zero zero." If a Speech-to-Text system relies solely on a raw Acoustic Model without an optimized Language Model, the output will be a string of disjointed characters that are difficult to read and unusable in databases.

Furthermore, homonyms in Vietnamese increase the difficulty. For instance, "hai" can mean the number 2, refer to "two" (as in "two children"), or be a misheard version of "hài" (comedy/comfortable). Text normalization requires the system to infer context to select the correct written form.

Effective Techniques for Processing Numbers and Currency

To handle quantitative values effectively, AI engineers typically apply a strategy combining acoustic recognition with linguistic rules.

1. Contextual Parsing

The system must identify the type of object being mentioned.

  • Product Quantity: "Buy 5 shirts" -> "Buy 5 shirts".
  • Currency: "Pay 500 thousand" -> "Pay 500,000 VND".
  • Ratio/Probability: "Rate of 50 percent" -> "Rate of 50%".

2. Currency and Unit Processing

In business meetings or transactions, reading out monetary values is extremely common. An intelligent system recognizes currency keywords (dong, thousand, million, billion) to automatically insert thousand separators and determine the unit.

  • Audio Input: "The project cost is two billion five hundred million dong."
  • Normalized Output: "The project cost is 2,500,000,000 VND."

Without this text normalization step, subsequent financial data analysis becomes impossible because computers cannot interpret "two billion five hundred million" as a numeric value.

Handling Dates: Formatting and Semantics

Dates are among the most error-prone elements in Vietnamese Speech-to-Text due to the diversity of expressions.

Common Reading Styles and Normalization

Audio Reading Recommended Normalized Format Note
"Mùng một tháng hai" (1st of the 2nd month) 01/02 Follows common DD/MM format in VN
"Ngày mười lăm tháng ba" (15th of the 3rd month) 15/03 Clear, low error rate
"Yesterday" / "Tomorrow" [Specific Date] Requires real-time time context
"Q3 2024" (read as words) Q3/2024 Requires 4-digit year processing

A crucial point is handling relative time terms like "yesterday" or "last week." To normalize text accurately, the Speech-to-Text system must be integrated with real-time clock data to convert these relative terms into specific calendar dates.

The Role of Language Models in Reducing Errors

Speech recognition is not just an acoustic problem; it is a linguistic one. The Language Model plays a pivotal role in predicting the next word and selecting the highest-probability word sequence within a specific context.

At AIVISION, we understand this challenge deeply while deploying solutions for hundreds of enterprises in Vietnam and internationally. AIVISION’s Speech-to-Text model is specifically designed for Vietnamese, effectively handling code-switching (mixed English-Vietnamese) and specific entities like monetary values and dates.

According to AIVISION’s internal benchmark, our model achieves an average Word Error Rate (WER) of 11.84%. This difference stems from optimized training data (9,043 hours of curated Vietnamese speech) and the ability to perform automatic text normalization during the recognition process.

Practical Advice for Developers

If you are building or using a Speech-to-Text system, apply the following steps to optimize results:

  • Define specific entities in advance: Clearly identify the types of numbers, currencies, and dates that frequently appear in your domain (e.g., healthcare, finance, hospitality).
  • Use APIs with timestamp support: Word-level timestamps help you debug and verify whether the system correctly identifies the position of numbers.
  • Combine with an LLM for cleaning: After recognition, use a Large Language Model (LLM) to review and normalize text finally. For example, use an LLM to convert "two zero two four" to "2024" if the context is a year.
  • Test accuracy on real-world data: Do not rely solely on general datasets. Test on actual business meetings where monetary values and dates are dense.

AIVISION offers a Speech-to-Text API with real-time streaming via WebSocket and file transcription via REST. The system supports 24 languages, facilitating smooth operation for translation and multilingual note-taking apps. With competitive pricing (70% of major competitors' list prices) and $5 of free usage daily for every account, it is an ideal solution to get started.

Conclusion

Handling numbers, dates, and currency is a decisive factor in the quality of Vietnamese Speech-to-Text systems. A good system does not just listen accurately but also normalizes text intelligently, ensuring output data is ready for subsequent analysis. By applying contextual techniques and using specialized AI platforms like AIVISION, you can overcome these challenges and build voice applications that truly deliver value to businesses.

Experience the difference with leading Vietnamese Speech AI technology. Try it now at Start free to see how AIVISION turns speech into accurate, standardized text.

Frequently asked questions

Why is text normalization important in Speech-to-Text?

Text normalization converts ambiguous phrases like "two thousand" into numeric formats like "2,000" or "2000," enabling computers to easily analyze, store, and process data automatically without further human intervention.

How does AIVISION handle dates in Vietnamese speech?

AIVISION uses a language model trained on rich Vietnamese data to recognize various ways of reading dates and automatically converts them into standard formats (DD/MM/YYYY) based on the conversation context.

How can I reduce errors when recognizing large monetary values?

To reduce errors, use a system capable of handling financial and currency context. AIVISION optimizes its model for specific entities like currency, helping accurately recognize keywords such as "thousand," "million," and "billion" while automatically inserting thousand separators.

Can AIVISION be used for meetings with mixed English and Vietnamese?

Yes, AIVISION Speech-to-Text supports code-switching (mixed Vietnamese and English), ensuring high accuracy when speakers switch languages between sentences, which is particularly useful in multinational corporate environments.

What is the cost of using AIVISION's API?

AIVISION applies a per-token pricing model. Every account receives $5 of free usage daily, allowing you to test service quality before committing to long-term use.

Try AIVISION's Vietnamese speech AI

$10 free every day for Speech-to-Text, Text-to-Speech and LLM.

Start free → Contact

By

Dr. Giang Vo — CEO, AIVISION

Dr. Giang Vo is the CEO of AIVISION, leading the development of s2speech — speech recognition, AI voices and a large language model for Vietnamese — and working with companies at home and abroad on their deployments.

0981 419 967 · giang.vo@aivgroups.com

#Vietnamese speech-to-text#text normalization#number recognition#date formatting#currency transcription#AIVISION#NLP#speech AI

Related articles