frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Ask HN: Are there AI models for generating sounds based on a text and reference?

22•onemiketwelve•1d ago
I've been having a hard time finding a solution. Is there really no commercialized model that I can feed in a reference sound and text instruction and get another sound out?

Right now having a multimodal inputs to image or text output is a commodotized, solved problem. IE you can put a prompt for some image and use a reference image to guide the model on what you want. After all, a picture is worth a thousand words right?

Ive also used a text and audio input in order to get a text description or classification out.

I cannot for the life of me find a solution for Audio + text -> Audio

My usecase is that I'm trying to generate new sound effects based on a source that doesn't have that many clean examples and thought this would be just another solved workflow but I'm not finding much. Tried elevenlabs SFX but it's only text->audio and without the reference, it's hard to guide with just text and recreate a sound with any accuracy. I tried the Stable Audio 3 model which is supposed to do what I want but the results were terrible. This modality seems to be stuck where image generation was in 2016. Is there something else or is this just not a common need?

Comments

narrationbox•3h ago
Plenty of models take in text + audio and spits out audio. It's the format of most newer generation accent conversion/voice cloning models.

What's your exact use case?

chr15m•2h ago
Sound effects are completely different to voice, which those models are trained to output.
xg15•2h ago
Maybe a dumb idea, but how good are image models with manipulating spectrogram images? Then a workaround could be: convert input audio to spectrogram -> pass spectrogram + prompt to an image model -> convert modified spectrogram back to audio waveform.
Buttons840•2h ago
I wouldn't call it a dumb idea, but there's soooo much subtlety to sound that wont be visible in any reasonably sized image.
0x20cowboy•1h ago
https://huggingface.co/docs/transformers/model_doc/audio-spe...
narrationbox•1h ago
Could be interesting, though spectrogram voice gen is a step back in terms of advancement. Vocoder models that take in Mel Specs suffer from a lot of issues like hissing due to Griffin Lim issues. Would be curious to see if diffusion models can work on neural codecs directly.

I think Google had one called riffusion (the first version was designed for specs)

chr15m•2h ago
Apparently the AudioX and AudioLDM(2) models do this but I think you've found a genuine gap.
moonu•2h ago
Even though they're technically trained for music, it might be worth testing Suno/Lyria to see if they're able to do this. You might have to isolate it afterwards, but seems like it could be viable
thangalin•1h ago
https://github.com/OpenMOSS/MOSS-TTS

KeenLore is my locally hosted full-cast emotive audiobook generator I'm developing for my novel:

https://www.youtube.com/watch?v=WAeHgE94rVo

Another HN user suggested embedding sound effects. The lead author of MOSS-TTS emailed me that the next version will have a SOTA voice designer. Poke around the voice design space, you'll find text-to-effects here and there. I certainly have my eyes on this space.

ninininino•1h ago
KeenLore is really cool.
jallmann•1h ago
Daydream Music - https://daydream.live

The DEMON realtime engine is open source. The hosted service has a number of integrations with popular audio tools. (I'm on the team)

soundworlds•43m ago
You should look into using AI to generate code that synthesizes sounds. Not exactly what you asked for, but I do think it is an approach worth considering: https://m.youtube.com/watch?v=1-i45X5aj94

Ask HN: Are there AI models for generating sounds based on a text and reference?

22•onemiketwelve•1d ago•12 comments

Ask HN: What do you think about Fractional CFOs

3•mtmosestn•1h ago•3 comments

Ask HN: Why is Ask HN only showing me 14 posts?

40•Gooblebrai•5h ago•37 comments

Ask HN: What Is Your Personal Backup Methods

2•Crontab•1h ago•0 comments

Ask HN: What do you run on a $5 VPS that's worth keeping online 24/7?

11•mariocesar•3h ago•4 comments

Ask HN: What games do you religiously play on your phone browser?

6•saimiam•3h ago•13 comments

Ask HN: Do you still own a printer?

9•nunorbatista•8h ago•14 comments

Ask HN: How to deal with AI "true believer" leadership at work

5•microflash•9h ago•3 comments

Ask HN: How do you describe what your relationship brings you?

2•maxignol•7h ago•3 comments

Ask HN: What would you like to see in a new AI technology release?

2•ikishade•8h ago•2 comments

Ask HN: Vibecoding Follies

4•vegnus•10h ago•2 comments

Ask HN: How come everyone is an LLM expert?

3•delis-thumbs-7e•10h ago•3 comments

ChatGPT self-distances when admitting fault

3•chrisjj•4h ago•1 comments

Tell HN: GitHub refuses to remove cracked copies of my software after a month

54•IvanK_net•8h ago•52 comments

Ask HN: Show me your agentic SDE semantics

2•danielovichdk•12h ago•0 comments

Ask HN: Is observability broken for you?

2•tmach32•14h ago•0 comments

Ask HN: My Apple ID is locked for a week, Apple is a single point of failure

11•akg_67•15h ago•3 comments

Ask HN: Agent access to chat data should require participant consent?

2•maxwellito•17h ago•0 comments

Ask HN: Alternatives to NeurIPS, ICLR and ICML

3•john-titor•19h ago•0 comments

Ask HN: What brings you back to personal AI agents like Instinct and Muse?

3•sdrth•21h ago•1 comments

Orcah Studio: A local-first video agent that can search your videos

8•iliashad•1d ago•5 comments

Ask HN: Why does Astra compact context so frequently vs. Fable?

2•yesitcan•1d ago•0 comments

Ask HN: What do you think about AI generated slides for conferences?

4•Hixon10•1d ago•1 comments

Ask HN: How are AI budgets changing in your company?

5•bobby-cb•1d ago•2 comments

Tell HN: 2026 is the year of the Linux desktop, agent-adjusted

3•dvrp•1d ago•2 comments

Tell HN: Uceprotect is extorting website owners

101•goldenmember•1d ago•59 comments

Our digital privacy is being aggressively disbanded this week

12•Steaglsz•1d ago•1 comments

DOS Game Stunts Port to Linux,Windows,Browser,etc.

4•LowLevelMahn•2d ago•1 comments

Who is cleaning up all the garbage LLMs generate?

6•kbrannigan•2d ago•7 comments

Ask HN: Anyone accepted into OpenAI "Codex for open source"?

4•awb•1d ago•0 comments