Voice agent teams keep hitting the same wall. The catalog holds 400 voices and the brief asks for the one that is not in it: a Quebecoise receptionist for a Montreal dealership, a narrator in his sixties with lecture hall authority. Briefs outnumber any catalog, and cloning closes the gap one speaker at a time, each carrying sourcing, consent and a licence. Gradium, the Paris-based voice AI company spun out of the Kyutai research lab, has shipped a different answer. Voice Design reads a written description and returns complete new voices in a few seconds. No reference audio, no speaker, no rights to clear.
Write a Prompt, Get a Brand New Synthetic Voice in Seconds
Gradium Voice Design is live in the Gradium API and in Studio, free on every plan including the free tier, and a kept voice runs on the same streaming Text-to-Speech endpoint as any catalog voice, at the same latency and output formats. The casting brief is the only input the model gets. Gradium’s documentation lists the attributes it responds to, and they read like a casting call: gender, age band, accent or origin, pitch, pace, energy, timbre and resonance, register and manner, and the job the voice is doing. Descriptions run 1 to 500 characters in English, French, Spanish, Portuguese or German. Gradium advises ending with the intended use, because it steers delivery and register rather than only the colour of the voice.
How Does Gradium Voice Design Work?
Voice Design is a model that generates a synthetic voice from a text description alone. One request returns 1 to 5 candidate voices, typically ready in 3 to 5 seconds. They are variations on a single character, so a different character means a different description, not more samples. The description is deliberately non-deterministic: the Gradium team expands the description first, and that expansion varies per request, so the same prompt with a fixed seed still yields a different voice each time. Candidates carry three restrictions that converted voices do not: audition text is capped at 100 characters, they are REST only, and the TTS WebSocket and Speech-to-Speech reject them. Unconverted candidates are deleted after 30 days. Converting is free, clears the expiry, and uses one custom voice slot shared with clones.
The API Flow From Casting Brief to Production Voice
The flow is four calls. POST /voice-generator/generate mints candidate ids with ready: falsecodecodecode. GET /voice-generator/embeddings polls until they flip. Each candidate auditions through the ordinary TTS endpoint, using the candidate id as voice_idcodecodecode. POST /voices/from-embedding promotes the one you keep. Candidates carry three restrictions converted voices do not: audition text is capped at 100 characters, they are REST only, and the TTS WebSocket and Speech-to-Speech reject them. Unconverted candidates are deleted after 30 days. Converting is free, clears the expiry, and uses one custom voice slot shared with clones. The free tier holds 5 custom voices, paid plans hold 1,000.
Why Non-Deterministic Sampling Matters for Voice Design
Sampling is deliberately non-deterministic. Gradium team expands the description first, and that expansion varies per request, so the same prompt with a fixed seedcodecodecode still yields a different voice. This means you cannot regenerate the exact same voice by reusing the same prompt. If you like a candidate, you must convert it to a permanent voice_id. That voice_id is then stored and used on the same streaming Text-to-Speech endpoint as any catalog voice, at the same latency and output formats.
Benchmark Results: How Voice Design Performs on Accent Accuracy
Gradium ran accent prompts through six systems and asked native speakers to pick the closer match, without labels. A model judge (Gemini 3.1 Pro) was then run over the same prompts. 7,627 blind pairwise comparisons across English, French, German, Spanish and Portuguese, September 2026. The model judge produced the same ranking as the human raters. Both evaluations were designed and run by Gradium, so read them as vendor-reported. 50 percent is par. Win rate is wins plus half of ties, over all comparisons. The strongest results came from regional accents that catalogs usually flatten into a single national voice.
What Languages and Attributes Does Voice Design Support?
Voice Design currently supports five languages: English, French, Spanish, Portuguese, and German. The description must be between 1 and 500 characters. The attributes the model responds to include gender, age band, accent or origin, pitch, pace, energy, timbre and resonance, register and manner, and the intended use. Concrete descriptions beat evaluative ones. Gradium advises naming pitch, pace and timbre outright rather than asking for “a great narrator voice”, and putting the intended use at the end.
Shipping Notes: What You Need to Know Before You Deploy
When using the API, only four JSON config keys are accepted. An unrecognised json_config key does not raise an error. The request returns 201 and the candidates simply never become ready, so you must bound the polling loop and treat a timeout as a bad request. SSML is spoken aloud, not parsed. Store the voice_id because it cannot be regenerated. The same prompt with a fixed seed still yields a different voice each time. Candidates have a 100-character audition text limit, are REST only, and expire after 30 days unless converted.
Voice Design represents a fundamental shift in how voice teams approach custom voices. By eliminating the need for reference audio, speaker sourcing, consent, and licensing, it opens the door to rapid prototyping and scalable deployment of bespoke voices for every use case imaginable. The catalog will never be large enough, but the casting brief now is. As the technology matures and language support expands, the line between a written description and a producible voice agent will continue to blur, making the voice agent the most flexible interface yet for human-machine interaction.