Text to Speech on Gendia: 23 Languages, and Why It's Priced Per Character
How Gendia's text-to-speech works: the five Supertone models, which languages each one covers, the emotion styles, and a pricing model that charges by the character rather than the run.
Most things on Gendia are priced per run. Speech isn't: it's priced by the character. That one difference changes how you should use it, so let's start there.
The pricing, because it's unusual
Roughly one credit per second of finished audio. The platform gets there by counting characters, and spaces don't count:
| Language | Characters per credit |
|---|---|
| Korean | 5 |
| Japanese | 5 |
| English and everything else | 13 |
So "안녕하세요", five characters, is 1 credit. A 50-character Korean line is 10 credits. The same second of English costs the same 1 credit, but fits 13 characters into it, because English packs more letters into a second of speech.
That's not a surcharge on Korean. It's the same price for the same amount of audio; the character counts differ because the languages do.
What this means in practice: speech is the cheapest thing on the platform. A 200-word English narration is roughly 1,000 characters, about 77 credits. Compare that with 125 for a single Nano Banana Pro image.
Five models, and how much of the world each covers
This is the choice that matters most, because it decides which languages you can use at all.
sona_speech_2, 23 languages. The default, and the one to use unless you have a reason not to. English, Korean, Japanese, Arabic, German, French, Hindi, Indonesian, Russian, Vietnamese, Spanish, Portuguese, Italian, Dutch, Polish, Czech, Danish, Greek, Estonian, Finnish, Hungarian, Bulgarian and Romanian.
sona_speech_2_flash: the same 23 languages, tuned for speed. Fewer voice settings: pitch shift, pitch variance, speed and duration only.
sona_speech_1, English, Korean and Japanese. The oldest, and it exposes the most controls, including subharmonic amplitude.
sona_speech_2t, English, Korean and Japanese.
supertonic_api_1, English, Korean, Japanese, Spanish and Portuguese. Speed is the only setting it accepts.
If your language isn't English, Korean or Japanese, sona_speech_2 or its flash variant are your only options.
The 300-character limit
Each request takes at most 300 characters. That's a line or two, not a script.
This is a feature more than a limit. Generating narration line by line means:
- A bad read costs you one line to redo, not the whole piece
- You can give different lines different styles
- Lines drop straight onto the narration track in the timeline editor and can be re-timed individually
Write your script in a document, then bring it across a line at a time.
Styles and voice settings
Style applies an emotion to the read: the available styles depend on the voice, so browse rather than guess.
Voice settings, on models that accept them:
- Speed: every model supports this, and it's the one you'll reach for most
- Pitch shift and pitch variance, how high, and how much it moves. Low variance reads flat and formal; high reads animated
- Duration, for fitting a line into a slot
- Similarity (1–5, sona_speech_1 and 2), how closely the output tracks the original voice
- Text guidance (0–4), how strongly the reading adapts to the meaning of the words
Leave them alone at first. The default read is good, and the fastest way to make synthetic speech sound synthetic is to over-tune it.
Writing for the ear
The script matters more than the settings, and it's the part people skip.
Short sentences. A sentence you'd have to re-read on the page is one a listener can't follow at all.
Punctuation is timing. Commas and full stops become pauses. A missing comma is a missing breath.
Spell out what should be spoken. "3" may be read as "three" or "third". Write the word you want to hear. Same for units, currencies and abbreviations.
Read it aloud yourself first. If you stumble, so will the model. This one habit fixes more bad narration than any setting.
What it's good for
- Narration over video, where it drops onto the timeline's narration track
- Localised versions of a piece you've already made, same script, 23 languages, a few credits each
- Drafts and scratch tracks while you decide whether a section needs a human read
- Accessibility: a spoken version of written material
Try it
Open the audio generator, pick sona_speech_2, and generate one sentence. In English that's about 8 credits. Then generate the same sentence at a different speed and compare: that pair tells you more about the controls than the labels do.
- #AITextToSpeech
- #AIVoiceGenerator
- #KoreanTextToSpeech
- #TTSLanguages
- #AINarration
- #SupertoneTTS
- #TextToSpeechPricing
- #AIVoiceover
- #GendiaTutorial
- #Gendia



