I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.
So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
thangalin•9m ago
https://www.youtube.com/watch?v=WAeHgE94rVo
Locally hosted, no cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Multicomp•7m ago
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
loremm•2m ago
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion