Introducing Song Creator Pro — create music with AI, locally on your device. Try it now →

Six Voice Models, One App

Voice Creator Pro runs six text-to-speech models. Each is strongest at a different job, so you pick the one that fits the work instead of forcing one model to do everything.

Every model on every tier
Browser or offline desktop
Full commercial rights

The Models

What each one is best at

Every model here handles text to speech. What separates them is where each one pulls ahead, and that is usually what should decide which you reach for.

OmniVoice

600+ languages
Voice cloningVoice design

One of the fastest models for voice cloning, and by far the broadest: it speaks over 600 languages. It clones a voice from a few seconds of audio and holds that speaker's accent even when they switch language. It also does voice design from fixed attributes rather than a description.

Best for: Localization, multilingual cloning, accent work, and high-volume generation

Qwen3-TTS

13 emotions, 5 intensities
Voice cloningVoice design

Assign one of 13 emotions to your text and choose how strongly it lands, so the same voice reads a line happy, furious, or afraid on demand. Voice design is free-form: describe the character you want in your own words and Qwen3 builds it.

Best for: Audiobooks, explainers, and any script where the feeling has to be exact

VoxCPM2

48kHz audio
Voice cloningVoice design

The only model here that outputs studio-quality 48kHz audio directly, with no upsampler. Voice design works in layers, identity then texture then delivery, the way you would write casting notes, and naming the setting steers the whole performance.

Best for: Narration, documentary, trailers, and anything where audio quality is front and center

DramaBox

Prompt-directed acting
Voice cloningVoice design

You write the voice and its performance together. Quoted text is spoken, unquoted text is stage direction, so a single take can move through boredom, sarcasm, and despair. You can also hand it a voice and then clone and direct that voice.

Best for: Game dialogue, animation, audio drama, and character work

Also in the app

These run on Cloud and desktop alongside the models above, each with its own strength.

Gepard 1.0

Runs on 2 GB VRAM

A lightweight cloning model built to run on low-end hardware, needing only about 2 GB of VRAM, so you can clone a voice on a desktop machine without a powerful GPU. It supports streaming and covers four languages.

Kokoro

Runs on CPU

An 82 million parameter model that generates faster than anything else in the lineup and runs fine without a dedicated GPU. It uses 28 preset voices rather than cloning or designing one, which keeps it simple and quick.

Side by Side

How the models compare

The specs that actually change which model you should pick. For a walkthrough by use case rather than by spec, see the guides further down.

ModelLanguagesEmotion controlVoice cloningVoice designBest for
OmniVoice600+Not explicit, inferred from textZero-shotStructured attributes, ten English accentsLocalization, multilingual cloning, accent work, and high-volume generation
Qwen3-TTS1013 emotions, 5 intensitiesZero-shotFree-form descriptionAudiobooks, explainers, and any script where the feeling has to be exact
VoxCPM230, plus nine Chinese dialectsNot explicit, inferred from textZero-shotFree-form descriptionNarration, documentary, trailers, and anything where audio quality is front and center
DramaBoxEnglish onlyDirected inline, inside the promptOnly as part of voice designFrom a prompt, or over a voice you provideGame dialogue, animation, audio drama, and character work
Gepard 1.04Not supportedZero-shotNot supportedVoice cloning on desktops without a high-end GPU, and streaming
KokoroEnglish, with 28 preset voicesNot supportedNot supportedNot supportedFast, straightforward narration where you do not need a custom voice

Emotion control means explicitly setting or directing which emotion the voice expresses, rather than the tone a model picks up on its own from your text.

Two Ways To Access

Cloud or desktop, same models

You get the same models either way. The difference is where they run and how you pay for them.

Voice Creator Pro Cloud

Run any model in your browser

  • Nothing to install, works from any laptop or phone
  • No GPU needed, server GPUs run the models for you
  • Free tier that renews monthly, paid plans from $5 a month
  • Every model on every tier, including the free one
Try It Free

Desktop App

Run them locally and offline

  • Runs every model on your own machine, fully offline
  • Unlimited generations, one-time purchase from $54.99
  • Your text and audio never leave your device
  • Local REST API for your own pipelines
Get the Desktop App

Rather not sign up at all?

For quick, casual text to speech, we also host a set of much lighter open-source models that run free on your own device, with no account and no limits. They are simpler and noticeably less capable than the frontier models above, with no 600+ languages, emotion control, or 48kHz output, so they suit fast everyday use rather than professional or commercial work. Chatterbox Turbo, Supertonic, and Kokoro are all there, and Kokoro also appears in the lineup above.

Try them free, no signup

FAQ

Common Questions

Six: OmniVoice, Qwen3-TTS, VoxCPM2, DramaBox, Gepard 1.0, and Kokoro. Each is strongest at a different job, from language breadth to emotion control to studio-quality output, and all six run on both Voice Creator Pro Cloud and the desktop app.

Yes. You pick the model per generation, so you can design a base voice with one model and switch to another when you need a different strength. Nothing is locked to a plan: every tier, including the free Cloud tier, has access to every model.

OmniVoice, by a wide margin. It covers over 600 languages, and it holds a cloned speaker's accent even when that voice switches language, which is what makes it the pick for localization. VoxCPM2 covers 30 languages plus nine Chinese dialects, and Qwen3-TTS covers 10.

VoxCPM2. It outputs studio-quality 48kHz audio directly with no upsampler, while the other models output at a lower sample rate. That makes it the one to reach for when audio quality is the point, like narration, documentary, and trailers.

Every model except Kokoro can clone a voice from a short reference clip with no training step. OmniVoice is the pick when the clone needs to speak another language and keep its accent, VoxCPM2 when you want the highest speaker similarity and audio quality, Qwen3-TTS when you want to assign emotions to the cloned voice, and Gepard 1.0 when you're cloning on a machine without a high-end GPU. DramaBox clones a voice too, but only together with voice design, so you clone and direct the voice in one step rather than as two separate actions.

You can use every model on the Voice Creator Pro Cloud free tier, which renews monthly and needs no card. Paid Cloud plans start at $5 a month, and the desktop app is a one-time purchase from $54.99 with unlimited generations.

Yes, with the Voice Creator Pro desktop app. It runs every model on your own machine with no connection needed, unlimited generations, and your text and audio stay on your device. Voice Creator Pro Cloud is server-based and needs a connection.

Yes. Audio you generate with any of these models is yours, on Cloud and desktop and including the free tier, with full commercial rights, no royalties, and no attribution required.

Try every model, free

Every model is on the Voice Creator Pro Cloud free tier, with full commercial rights and no card required. Or get the desktop app and run them offline with no generation limits.