That seems to me to be a natural progression, from discrete models to models that are just continuously improved. Maybe we'll end up with different models with different rates of improvement rather than static differences in performance, and methodologies for that improvement will be the thing we care about. Maybe over time, even benchmark tests will be primarily concerned with that kind of efficiency.
I think a huge debate right now is the relative value of the "frontier" models from Western companies at the cutting edge, vs distilled versions of those models that are good enough and exponentially cheaper coming from China. But a paradigm of 'always training' means an always active, always advancing frontier, which is a stronger moat than a one-off model that's more advanced for a few months.
I like this one better
27183•4h ago
In other words: bias. Tons and tons of self-congratulatory, glue sniffing bias.
Mond_•4m ago
> Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something.