However, I figured out an interesting way to forgo that upfront embedding cost.
algo:
old model/index -> retrieve top-K docs -> score those docs with the new model -> cache/materialize the new embeddings
so instead of rebuilding the entire vector store upfront, the old index keeps getting retrieved from, while the new model reranks those candidates.
This works surprisingly well for some model pairs, (i tested 63 source-> target migrations on h100s, on upto 1M documents).
For example, on a 1M document Natural Questions dataset,
native Qwen3-Embedding-8B: 0.6812 nDCG@10 Qwen3-4B -> Qwen3-8B, K=50: 0.6816 Qwen3-0.6B -> Qwen3-8B, K=50: 0.6638 MiniLM -> Qwen3-8B, K=50: 0.6486
(the hard part is determining k, I held the k constant above to give some sense of migratability).
You can install it with pip
pip install embedflow
and the code is on github
dancemonster35•31m ago
coolArnav•17m ago