e.g. 1.40m would become 0.30s.
do people really pay for these priority plans?
I would send a second request if the first request fails to return the first token within, say, 1 second. Then there's a chance the first request is stalling, which is an infrequent event.
I wonder if higher-availability tiers of LLM providers do a similar thing internally.
eigenblake•1h ago
thousand_nights•22m ago