I saw some examples of using jev-like models to check if a file content is relevant to some query or not as "is this code file related to feature X?". This is already interesting for reducing token usage, but it would be even better if you can have a "mask" over the input tokens that tells which tokens are relevant to the answer, maybe for already giving file line ranges in a tool call result.
Another nice variant is having returned multiple masks one per topic i.e. segmentation but for text.
I think this is already pretty doable with current model architectures and would be very useful, but maybe this needs a model fine-tuned specifically on this task to work well. Does anybody know about the current state of this?