No, because the cache is local to the GPUs and there would be no cost/compute reason to have long-lived caches outside of the 1 hour cache already done by LLM APIs.
sourav_biswas•43m ago
What if the real business isn't caching generic prompts/but caching outputs for a specific vertical where one can know two requests are truly equivalent n not just embedding similar?
minimaxir•47m ago
sourav_biswas•43m ago