The "distillation" argument always seemed pretty weak to me. If one argues that training an AI on copyrighted content is merely analogous to reading a book (and therefore is fully transformative), then it stands to reason that training an AI on other AI outputs would also be fully transformative.
It seems to me you can't have it both ways; either training is a violation of copyright or it isn't, and there's no consistent argument where "distillation" is a violation but other things aren't.
easterncalculus•11m ago
This argument only makes sense if you equate a printed book with a chatbot. They aren't the same thing, and the chatbot's outputs aren't copyrighted.
andsoitis•37m ago
How does distillation actually work? If I wanted to distill a model and I have enough money and compute, what’s the process?
jqpabc123•51m ago
https://www.techopedia.com/trump-administration-openai-train...