In my testing I got 150 tokens/sec with a single 5090 RTX.
"The machine immediately taught me that capacity estimates are just admission tickets."
"Useful in production, poison in a kernel comparison."
Please don't publish writing like this, it's exhausting to read. You can edit that stuff out.
The best part of this is the "Five xhigh artifacts from the finished model" section at the bottom, I suggest either moving that up or at least prominently promoting it at the top of the article.
supermatt•55m ago
jermaustin1•27m ago
I know this doesn't exactly match your request, but I'm happy with my 3.6's performance. I've seen people claim they can get 120+tps on a 3090, but I'm not impatient when it comes to streaming text faster than I can read it.
Tostino•22m ago
jermaustin1•9m ago
I might need to finally make the switch away from it, but it is so convenient, especially as a chat interface for system prompt experimentation.
pich•5m ago