I know that GPT Astra uses looping instead of reasoning tokens. My specific question is this: reasoning tokens pollute the context and makes compaction kick in quickly. Does this also happen to looped transformers like Astra?
I know that GPT Astra uses looping instead of reasoning tokens. My specific question is this: reasoning tokens pollute the context and makes compaction kick in quickly. Does this also happen to looped transformers like Astra?