Six Voice Models, One App
Voice Creator Pro runs six text-to-speech models. Each is strongest at a different job, so you pick the one that fits the work instead of forcing one model to do everything.
The Models
What each one is best at
Every model here handles text to speech. What separates them is where each one pulls ahead, and that is usually what should decide which you reach for.
OmniVoice
600+ languagesOne of the fastest models for voice cloning, and by far the broadest: it speaks over 600 languages. It clones a voice from a few seconds of audio and holds that speaker's accent even when they switch language. It also does voice design from fixed attributes rather than a description.
Best for: Localization, multilingual cloning, accent work, and high-volume generation
Qwen3-TTS
13 emotions, 5 intensitiesAssign one of 13 emotions to your text and choose how strongly it lands, so the same voice reads a line happy, furious, or afraid on demand. Voice design is free-form: describe the character you want in your own words and Qwen3 builds it.
Best for: Audiobooks, explainers, and any script where the feeling has to be exact
VoxCPM2
48kHz audioThe only model here that outputs studio-quality 48kHz audio directly, with no upsampler. Voice design works in layers, identity then texture then delivery, the way you would write casting notes, and naming the setting steers the whole performance.
Best for: Narration, documentary, trailers, and anything where audio quality is front and center
DramaBox
Prompt-directed actingYou write the voice and its performance together. Quoted text is spoken, unquoted text is stage direction, so a single take can move through boredom, sarcasm, and despair. You can also hand it a voice and then clone and direct that voice.
Best for: Game dialogue, animation, audio drama, and character work
Also in the app
These run on Cloud and desktop alongside the models above, each with its own strength.
Gepard 1.0
Runs on 2 GB VRAMA lightweight cloning model built to run on low-end hardware, needing only about 2 GB of VRAM, so you can clone a voice on a desktop machine without a powerful GPU. It supports streaming and covers four languages.
Kokoro
Runs on CPUAn 82 million parameter model that generates faster than anything else in the lineup and runs fine without a dedicated GPU. It uses 28 preset voices rather than cloning or designing one, which keeps it simple and quick.
Use Cases
Which model for what you're making
Start from the work instead of the model name. Here is the one to reach for first by job, with a line on why it fits.
Audiobooks and long narration
Reach for Qwen3-TTS
Set the emotion and its intensity passage by passage, so a long read holds its feeling instead of flattening out over chapters.
Dubbing and localization
Reach for OmniVoice
Clone a speaker once and keep their accent as the voice switches language, so a dub sounds like the same person in every market.
Voiceovers, trailers, and documentary
Reach for VoxCPM2
Studio-quality 48kHz audio straight from the model, for work where the sound itself has to hold up on good speakers.
Games, animation, and audio drama
Reach for DramaBox
Direct the performance line by line and carry one voice through different emotions inside a single take.
Voice cloning without a high-end GPU
Reach for Gepard 1.0
Clone a voice on a desktop with only a modest GPU, needing about 2 GB of VRAM, with streaming support for longer runs.
Drafts, scratch tracks, and high volume
Reach for Kokoro
Fast generation that runs without a dedicated GPU, for temp voiceover and batch runs where a custom voice is not the point.
Most jobs have a second and third good fit too. Best TTS Model for Every Use Case walks through the alternates for each.
Side by Side
How the models compare
The specs that actually change which model you should pick. For a walkthrough by use case rather than by spec, see the guides further down.
| Model | Languages | Emotion control† | Voice cloning | Voice design | Best for |
|---|---|---|---|---|---|
| OmniVoice | 600+ | Not explicit, inferred from text | Zero-shot | Structured attributes, ten English accents | Localization, multilingual cloning, accent work, and high-volume generation |
| Qwen3-TTS | 10 | 13 emotions, 5 intensities | Zero-shot | Free-form description | Audiobooks, explainers, and any script where the feeling has to be exact |
| VoxCPM2 | 30, plus nine Chinese dialects | Not explicit, inferred from text | Zero-shot | Free-form description | Narration, documentary, trailers, and anything where audio quality is front and center |
| DramaBox | English only | Directed inline, inside the prompt | Only as part of voice design | From a prompt, or over a voice you provide | Game dialogue, animation, audio drama, and character work |
| Gepard 1.0 | 4 | Not supported | Zero-shot | Not supported | Voice cloning on desktops without a high-end GPU, and streaming |
| Kokoro | English, with 28 preset voices | Not supported | Not supported | Not supported | Fast, straightforward narration where you do not need a custom voice |
† Emotion control means explicitly setting or directing which emotion the voice expresses, rather than the tone a model picks up on its own from your text.
Two Ways To Access
Cloud or desktop, same models
You get the same models either way. The difference is where they run and how you pay for them.
Voice Creator Pro Cloud
Run any model in your browser
- Nothing to install, works from any laptop or phone
- No GPU needed, server GPUs run the models for you
- Free tier that renews monthly, paid plans from $5 a month
- Every model on every tier, including the free one
Desktop App
Run them locally and offline
- Runs every model on your own machine, fully offline
- Unlimited generations, one-time purchase from $54.99
- Your text and audio never leave your device
- Local REST API for your own pipelines
Rather not sign up at all?
For quick, casual text to speech, we also host a set of much lighter open-source models that run free on your own device, with no account and no limits. They are simpler and noticeably less capable than the frontier models above, with no 600+ languages, emotion control, or 48kHz output, so they suit fast everyday use rather than professional or commercial work. Chatterbox Turbo, Supertonic, and Kokoro are all there, and Kokoro also appears in the lineup above.
Try them free, no signupGo Deeper
Guides and comparisons
Written walkthroughs that compare the models against each other, and cover the techniques that apply whichever one you land on.
AI Voice Design Guide: 4 Ways to Create a Voice From Scratch
All four models side by side on how you design a voice, with playable samples of each approach.
Best TTS Model for Every Use Case
Which model to reach for by job, from fast narration to expressive, emotion-controlled speech.
Qwen3-TTS vs OmniVoice vs Chatterbox: Voice Cloning Compared
A head-to-head on cloning quality, speed, and how well each model holds an accent.
Best Offline Voice Cloning Tools
How these models compare when you want to clone a voice on your own machine.
Techniques that apply to any model
FAQ
Common Questions
Six: OmniVoice, Qwen3-TTS, VoxCPM2, DramaBox, Gepard 1.0, and Kokoro. Each is strongest at a different job, from language breadth to emotion control to studio-quality output, and all six run on both Voice Creator Pro Cloud and the desktop app.
Yes. You pick the model per generation, so you can design a base voice with one model and switch to another when you need a different strength. Nothing is locked to a plan: every tier, including the free Cloud tier, has access to every model.
OmniVoice, by a wide margin. It covers over 600 languages, and it holds a cloned speaker's accent even when that voice switches language, which is what makes it the pick for localization. VoxCPM2 covers 30 languages plus nine Chinese dialects, and Qwen3-TTS covers 10.
VoxCPM2. It outputs studio-quality 48kHz audio directly with no upsampler, while the other models output at a lower sample rate. That makes it the one to reach for when audio quality is the point, like narration, documentary, and trailers.
Every model except Kokoro can clone a voice from a short reference clip with no training step. OmniVoice is the pick when the clone needs to speak another language and keep its accent, VoxCPM2 when you want the highest speaker similarity and audio quality, Qwen3-TTS when you want to assign emotions to the cloned voice, and Gepard 1.0 when you're cloning on a machine without a high-end GPU. DramaBox clones a voice too, but only together with voice design, so you clone and direct the voice in one step rather than as two separate actions.
You can use every model on the Voice Creator Pro Cloud free tier, which renews monthly and needs no card. Paid Cloud plans start at $5 a month, and the desktop app is a one-time purchase from $54.99 with unlimited generations.
Yes, with the Voice Creator Pro desktop app. It runs every model on your own machine with no connection needed, unlimited generations, and your text and audio stay on your device. Voice Creator Pro Cloud is server-based and needs a connection.
Yes. Audio you generate with any of these models is yours, on Cloud and desktop and including the free tier, with full commercial rights, no royalties, and no attribution required.
Try every model, free
Every model is on the Voice Creator Pro Cloud free tier, with full commercial rights and no card required. Or get the desktop app and run them offline with no generation limits.