Poddie does it e2e for free: it generates transcript using local Whisper model, lets you edit by reading through the transcript and deleting words, auto-trims silence, offers video/audio/captions export with optional caption burn-in. You can also do waveform editing for more precise cuts if desired.
I'm open sourcing it so other creators could try it out and use it for their own workflows.
Happy to hear any feedback!