![]() |
| Kakao has advanced its 'Kanana-o' voice generation technology. / Kakao |
Kakao announced on the 4th that it has advanced the voice generation technology of its in-house omni AI model "Kanana-o" to a new level.
According to Kakao's tech blog that day, the upgraded technology allows users to control tone of voice, emotion, and even intonation using only natural language instructions, while a self-developed speech tokenizer also boosted generation speed and efficiency.
While existing voice AI has focused on naturally reading text aloud, the newly improved Kanana-o lets users control the vocal expression itself through their instructions. When a user gives a natural-language request such as "read it very fast," "read it in a low voice," "read it in a sad voice," or "read it in a Gyeongsang dialect," the AI generates speech that precisely reflects speed, volume, pitch, emotion, intonation, and intensity accordingly. It can also handle role-based instructions like "read it like a news anchor," as well as compound instructions combining multiple conditions at once, such as "lower the tone and speak quickly in a sad voice." Notably, even though it was trained primarily on Korean data, it can carry out the same kind of speech instructions smoothly in English as well.
Kanana-o scored 94.50 on the Korean-language "InstructTTSEval" benchmark, which evaluates a model's ability to follow speech instructions. This surpasses GPT-4o-mini-tts (91.10) and is comparable to Google's Gemini-2.5-flash-preview-tts (95.38).
Beyond voice quality, speed and efficiency were also significantly improved. To achieve this, Kakao newly applied its self-developed speech tokenizer, "LM-SPT (LM-aligned SPeech Tokenizer)." LM-SPT compresses speech into fewer tokens for the AI to process, reducing the amount of data the AI needs to handle and enabling faster, more efficient voice generation.
In evaluations comparing Korean and English speech understanding and generation across various voice language models, LM-SPT reportedly outperformed models using the latest global technologies such as Mimi, DualCodec, and CosyVoice2. It also scored highest in expert evaluations that rated the naturalness and speaker similarity of synthesized speech after listening.
Kakao plans to continue developing Kanana-o's voice technology going forward. Through research that processes speech understanding and generation capabilities within a single integrated architecture, the company aims to advance its speech processing technology and focus on delivering a natural, seamless voice experience. It also plans to continue developing more refined voice control technology, including the ability to naturally generate non-verbal expressions such as laughter, sighs, and exclamations.
Noh Byung-seok, performance leader of Kakao's Unified Foundation Model team, said, "This upgrade to Kanana-o's voice technology focused on achieving human-like natural voice generation, as well as the ability to express the tone, emotion, and intonation users want through natural language instructions," adding, "We will apply the Kanana-o model to a variety of services going forward to provide a more natural and convenient AI voice experience."
Kim Young-jin
1
2
3
4
5
6
7