We found that they used between 2 and (almost) 6 x the input tokens and were never faster on any task we timed.