Uses Rubberband engine pitch shifting with formant preservation to avoid the ogre-like sound of pitching down, and HTDemucs converted to Core ML (NN model) to do live audio separation via sliding audio windows.
I think this is the best tool of its kind (lowest latency, highest quality, realtime). It's invaluable for me now with karaoke, practicing singing / piano along with Spotify songs.