we have people using it for, routing a ticket, classifying intent, deciding whether something is urgent, or rating severity.
Architecture - 12B Gemma fine-tuned to score options. since it is gemma based it has multimodality built in - When there are >120 labels, we have a 150M retriever that narrows them to a shortlist (recall@16 is 0.9–1.0 on our evals), and the 12B model scores that shortlist. - Context goes up to 49k tokens, with up to 8 images per request.
we have open sourced it,
model: https://huggingface.co/parsecai/tod API: https://parseclab.ai/tod
We are working on getting the RL environments to finetune for custom use cases.
would love to hear your opinions on this.