On 23 September, Google presented two speech synthesis models that promise to design and direct a voice line by line. For a Maison, a voice becomes an editorial asset whose reference, permitted uses and history must be fixed.
What Google announces
Gemini 3.8 Flash TTS targets creative direction: designing a voice from a description (role, accent, character) in more than one hundred languages and dialects, then directing each line with indications of acting, rhythm and tone. Gemini 3.8 Flash-Lite TTS targets volume, notably dubbing and voice agents. Google cites more than 2,000 ready-made voices, the ability to save created voices to limit drift from one project to the next, and continuous runs of several hours with a stable timbre. Editing a library voice (timbre, pitch, accent) is announced as forthcoming.
The company also claims first place in Hume AI’s voice design ranking, with 71.4 points. The announcement does not detail the method in its text, so the figure remains a vendor score.
A voice that is directed needs a reference
Directing a voice shifts part of the work toward editorial direction. One must be able to say what makes a voice right: the way a Maison’s name is pronounced, a restraint, a silence before a sentence. Adjectives are poorly suited to conveying this. Approved excerpts, listened to and annotated, work better, with the same text serving as a benchmark across versions and languages, and a human listen in each language for proper names, intelligibility and intent. A sentence that is correct on paper can sound wrong when spoken.
Permitted uses: the map is uneven
For replication, Google requires the recording of an oral consent from the person whose voice is reproduced, which must match the reference sample. Thirty seconds of audio is enough, of the user’s own voice or of a voice whose rights they hold. Each excerpt carries the SynthID watermark, and C2PA attestations accompany replicated voices, according to Google.
These signals prove the origin of a file, not the extent of the agreement. A created voice, a replicated voice and a library voice call for different authorizations, and each should be recorded with its agreed uses.
Then comes geography. Google states that replication from its developer tool is not offered in the European Economic Area, which includes France, nor in the United Kingdom, Switzerland or India. This limit changes a French Maison’s reasoning even before quality does. Meanwhile, access for businesses via Gemini Enterprise is announced for later.
Finding each version again
Google promises voices that are “saved and managed.” Nothing in the announcement describes a history: is the former version of a voice kept after modification, can it be restored, is it known which file was generated with which? The ease of generation multiplies these decisions. Every adjustment creates a variant that will need to be located later.
A Maison could keep a simple register: for each voice, its type, its owner, the agreed uses, the date of each version, and the conditions for withdrawing an old one. It is this register, kept by the person who decides, that will make the voice a Maison voice.
Cette publication est également disponible en :
