I think this is an underrepresented view of what we can do with agents if we have access to the lower levels of the LLM model than the conventional API abstraction.
You can mess with KV-cache to make LLMs more interactive without retraining, this allows Qwen3.5 model to run in a doom environment making actions, while thinking and observing frames, without retraining
puhsu•45m ago
You can mess with KV-cache to make LLMs more interactive without retraining, this allows Qwen3.5 model to run in a doom environment making actions, while thinking and observing frames, without retraining