This week (October 2026) Google released EmbeddingGemma 2, which puts text, images, audio and video in the same vector space, so I moved the app over to it from the siglip2 model i was using before. The main differences I've seen, apart from video and audio support, are the multilingual capabilities, and the better understanding of long queries, although it still struggles with generic images being too "in the middle" of the embedding space and ranking high on unrelated searches.
I know small SaaS products like this are a tough sell these days. I built it anyway because I grew up learning to code by hand, reading about solo developers with crazy ideas shipping little products that made some side money, and I wanted to go through that whole process myself: build it, deploy it, and put it in front of people.
How it works:
The model runs in your browser, transformers.js, WebGPU or WASM, depending on your device. It does work on mobile, although image upload do take way too much there.
Your files and their vectors are stored on the server so collections sync across devices. I know this could be a privacy concern, but at least for me the ease of use is worth it.
For quantisation im using fp16 on desktop with capable GPUs and q4 on any other device.
The big caveat is that first visit downloads the model (about 170 MB, or about 520 MB on desktops with fast WebGPU), it's cached after that, but you do have to wait a bit.
You can try it for free without an account on the example collections. As a side not, i took the travel photos (the Photography collection, from Venice) and the flower photos myself. There’s a free tier if you want to upload your own files, and paid plans if you need more.
If you have this problem too, I hope this shows that searching your own files using embedding models is genuinely useful, whether you use this or build your own. And if you’d rather not build it, I’d be happy to have you as a user. Any feedback is very welcome.