How to evaluate enterprise Turkish speech technology: an STT and TTS guide
Enterprise Turkish speech technology is evaluated by testing it on your own call recordings and terminology. For speech recognition you measure error rate, accuracy on critical details and latency; for text-to-speech, naturalness and how numbers, amounts and dates are read. Deployment and data control are reviewed separately.
The reliable way to evaluate enterprise Turkish speech technology is to test it on your own call recordings and your own terminology. For speech-to-text (STT), you measure error rates, latency and how well domain terms are recognized. For text-to-speech (TTS), you measure naturalness and how numbers, amounts, dates and abbreviations are read aloud. On-premise deployability, data control and customization complete the picture. Generic demos and showcase recordings answer none of these questions.
Why Turkish speech technology needs its own evaluation
Turkish is an agglutinative language. A single root can take on dozens of word forms through suffixes. Recognizing "hesap", "hesabımdaki" and "hesaplarımızdan" correctly requires far more than looking a word up in a vocabulary. Suffix relationships and sentence patterns carry meaning too.
Enterprise terminology is the second challenge. Product names, campaign names, regulatory terms and abbreviations are often missing from the training data of a general-purpose model. When the model does not know a word, it writes a similar common word instead, and the meaning of the conversation drifts.
The third challenge is the channel. Telephone audio is typically carried at an 8 kHz sampling rate over a narrow frequency band. Add background noise, mobile network dropouts, regional accents and everyday speaking habits. A model that performs well on studio recordings may not hold up on contact-center audio.
How to evaluate speech-to-text
Start with your own data. Build a test set from real calls that represents different customer profiles, line types and topics. Use human-verified transcripts of those recordings as the reference.
Then look at these measures:
- Word error rate (WER): How far the model output deviates from the reference. Substituted, deleted and inserted words are counted and divided by the number of words in the reference. In Turkish, small suffix differences affect this rate, so it should not be read on its own.
- Entity accuracy: Whether the details a transaction depends on, such as customer numbers, amounts, dates, product names and personal names, are captured correctly. A good overall error rate means little if a single digit is wrong.
- Terminology recognition: Prepare a list of your product, campaign and regulatory terms and track their recognition separately.
- Breakdown by condition: Report results separately for quiet and noisy environments, landline and mobile, different accents and speaking speeds. Averages can hide weak spots.
- Latency: Measure how quickly streaming recognition produces text. On the phone, callers notice silence immediately, so latency is part of the experience.
How to evaluate text-to-speech
Two questions matter most: does the voice sound natural, and does it read correctly?
For naturalness, run listening tests. Several listeners rate recordings generated from your real sentences on intonation, stress, pauses and fluency. Pay particular attention to long sentences and questions.
For pronunciation accuracy, build a checklist from the text the system will actually speak:
- Amounts, currencies and decimal values
- Dates, times and due dates
- Long digit strings such as phone numbers, policy numbers and IBANs
- Abbreviations and terms that must be spelled out
- Brand, product and foreign-origin names
Check whether misread items can be corrected through a pronunciation dictionary or a markup language such as SSML. Measure speech generation latency and consistency across a full call as well.
Deployment, data and customization
Voice recordings and transcripts contain personal data, so where the technology runs is as important as the technical scores. Confirm whether speech recognition and synthesis can run inside your environment, whether any data leaves the approved boundary, and whether there is a mandatory dependency on an external service.
Customization belongs on the list too. Can you add a domain vocabulary? Can the model be adapted to your terminology? Can a brand-specific voice be developed? Integration with your existing telephony and contact-center infrastructure is part of the evaluation. Finally, check whether transcripts and failure reasons can be monitored after go-live, because speech technology improves over time with real traffic.
Evaluation steps
- Define the target call flows and the critical information in each.
- Build a representative test set and reference transcripts from your own recordings.
- Prepare a terminology list and a pronunciation checklist.
- Measure STT and TTS under the same conditions, with results broken down by condition.
- Review the deployment model, data flow and customization options with your security team.
- Validate results on live traffic with a limited-scope pilot.
What CBOT offers
CBOT develops its Speech-to-Text and Text-to-Speech technologies in-house as part of CBOT Speech. They are optimized for Turkish language structure, pronunciation and domain-specific terminology. Streaming speech recognition is built for contact-center and mobile audio, ambient noise, natural speaking patterns and domain terminology. Details are on the Voice AI page.
On the synthesis side, CBOT's neural voice technology produces responses with controllable pronunciation, speed, pitch and pauses. It offers custom voice profiles, brand-specific voice development, SSML support and pronunciation dictionaries. Domain vocabulary, brand names, product names and specialist terminology are supported through custom dictionaries, model adaptation and pronunciation configuration. CBOT adapts its text and speech models with approved terminology, business scenarios and communication standards, as described on the CBOT LLM page.
On deployment, the full voice stack, including Speech-to-Text, Text-to-Speech, supported model inference, RAG, runtime, analytics and integrations, can operate inside the customer's environment. In a fully on-premise architecture, voice recordings and conversations stay in the approved enterprise environment. The options are compared on the deployment options page. CBOT integrates with existing IVR, SIP and MRCP infrastructure.
At Halkbank, CBOT's Automatic Speech Recognition (ASR) and text-to-speech (TTS) technologies are integrated into the bank's IVR system and used in the call center. Read more in the Halkbank customer story. In production, call transcripts, transfers and failure reasons can be monitored with analytics and observability tools.
More questions about speech technology are answered on the FAQ page.
Conclusion
When choosing Turkish speech technology, the results you measure on your own data are what count. Once your test set, terminology list and pronunciation checklist are ready, the comparison rests on objective ground. To plan an evaluation with your own recordings, talk to the CBOT team.