"sketchy russian website", how about using some more clear description like: A library for sharing books and articles that should be partly public domain because they were paid for by the public. Only some of the material is copyrighted by authors. However, some of their work is so old that it is not reprinted anyway.
But of course such an explanation would not click.
I also don't see a problem with statements about making people jobless. Imagine if every robotic or automation company advertised like this: Yeah, you'll buy tons of expensive robots and still rely on expensive labor from real people without any efficiency gains.
The second thing underscores my point. They use one line and think they've made a great point because one employee called LibGen sketchy. This site has been around since the 2010s, and it has helped many people do research. It's not just a sketchy website that suddenly appeared and is always doing bad things. I think a more nuanced stance is necessary.
No, you are using one line from the post to discredit them. The release has more than that and it’s not the only communication they made on this matter nor is there any indication it will be the last, it’s just the current one.
I dare say that car manufacturers are aware of their impact on the horse and buggy industry. Calculator manufacturers wrecked the livelihood of mathematicians and accountants.
Technology is in the business of putting people out of work, by inventing better ways of doing things. Or rather, any time you invent a better way of doing things, that's fundamentally going to disrupt all the businesses built around older technologies.
Why is book piracy a better way of doing things?
You twist the argument. Your argument would hold if AI's only use would be to generate booksverbatim it was already trained on. Which is certaintly not the case and huge efforts were made to circumvent this kind of usage.
Clearly these books had value to AI companies but they were too weak and too dishonest to pay for that value. That's not impressive.
> AI agent accidentally publishes OpenAI’s unreleased model weights
The optics are bad. The submission marks a turning page in human history.
Given the political power they have, particularly now with the Trump administration, whatever "optics" exist on this forum seems completely insignificant
the economics of those companies will cause a catastrophic wipe out of jobs across the board.
About LibGen, there might be more discussion - fair. However, the second argument is no real discussion IMO. Why is putting people out of work suddenly a bad thing? Since when do we argue this when talking about automation?
(And then Russia invaded Ukraine, turning any association with .ru things into potential corporate suicide.)
Really has nothing to do with LibGen or with OpenAI. It's about people being easy to manipulate into believing bullshit, which is a reasonable worry, and the Authors Guild is trying to do that exact thing OpenAI was worried about.
The whole "sketchy russian website" bit resolves entirely about being seen as associated or supporting troll farms and Putin.
EDIT: look at it this way: no one is calling Internet Archive "a sketchy US website".
I think they just used a Russian torrent site.
>Microsoft knew about OpenAI’s use of LibGen as early as April 2019
(note: https://z-library.sk/ is prettier/nicer)
Do you have any evidence of them being a lobby organisation (as opposed to OpenAI for example which spends millions of dollars hiring actual lobbyists)
>The group lobbies at the national and state levels on censorship and tax concerns, and it has initiated or supported several major lawsuits in defense of authors' copyrights.
Have you looked at their name?
The more interesting question is IMO if AI training actually falls into one of these cases. You can read a book and also copy it, but you do not do because of the law. However, you have the ability to do so. Is having the ability to do something already forbidden?
The law they broke was pirating the materials, not training per se, even though training is what so many people object to: the judge ruled that actually training a model, when the materials you used were ones you otherwise had lawful access to, was not a breach of law.
IMO, the laws need to change to reflect what tech can now do. This wouldn't be the first time, copyright law has had to shift several times before as new means of reproduction are created.
* the Anthropic one
How? The current system enables the GPL. The GPL protects many open source projects.
No, they were responding to the post defending OpenAI that you wrote. If you meant to communicate something other than “criticism of OpenAI in this context is unwarranted” then it looks like you forgot to do that and wrote something else instead
Why are you turning him into perpetuum mobile in his grave?
Do you really believe Aaron would be arguing against AI companies and for publishing / recording guilds on the grounds of intellectual property claims?
No, it's the tech community that did a sudden about-face, and is now all "friendship ended with free access to information and technologies enabling people; now RIAA is my best friend", and this move is as dumb as that meme (https://imgflip.com/memegenerator/137501417/Friendship-ended).
That's not the point. All the rules and laws are enforced when its you and me but when it's big tech the laws are not enforced the same way.
> friendship ended with free access to information and technologies enabling people; now RIAA is my best friend
Big tech will enable access to free information and will help people reach new heights argument is as dumb as the meme you are referring to.
A: "Information is free"
B: "Information is not free"
C: "Information is free only for the rich and not free for everyone else, giving the rich a material advantage over everyone else that not only entrenches but accelerates wealth inequality and impedes class mobility"
You, or Swartz, are an advocate for A. Why, exactly, do you think that obliges you/Swartz to prefer C over B?
There are several multi billion dollar companies where the founding thesis was “what if we just ignore the law?”
I find the brazenness of saying this while running what's arguably the largest copyright theft operation in human history astonishing. If libgen is "sketchy", then what is OpenAI?
Many people think that it was fair use: training is akin to reading, not copying.
Especially the courts.
The law is the law, there can't be different law for corporations with billions in backing. I don't agree with current copyright laws btw and think they should be changed. However, they probably should have lobbied for that before illegally downloading all this material.
100% of the rulings agree with me.
The piracy is not in question. It is unarguably copyright violation.
But that's not what anyone means in this context. Training is what everyone means.
> The law is the law, there can't be different law for corporations with billions in backing.
I didn't say otherwise. That's a straw man.
No.
Judging by how AI threads look like for the past year, they were absolutely right to be worried.
> largest copyright theft operation in human history
In fact, you're doing exactly that right here.
You're dead fucking right they are doing exactly that. They are saying that what is wrong is wrong.
Instead you seem to be making out that AI companies are some kind of victim that has to "worry" about "manipulation". Meanwhile authors are out of a job right now and not by accident. What's up with that?
[1] https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-s...
Also, defending it on the basis that some books on libgen are public domain is a poor excuse, like claiming people use The Pirate Bay to download Linux ISOs. Even if some of that is true, we all know that use case is not the popular one.
I do see a problem with a company loudly announcing that they are going to make people's lives miserable purely for profit. Leaving aside that it goes against OpenAI's stated mission ("to ensure that artificial general intelligence benefits all of humanity"), the disdain for the lives they are intentionally trying to ruin makes it a problem.
And even if you believe that the transition is inevitable, as it is the case with phasing out combustion engines in cars, anyone reasonable would see that the transition is gradual to give people time to adapt. Instead of doing that, OpenAI is burning cash at astonishing rates, polluting the environment, and killing personal computing with the only aim of being the only ones left atop the ruins. I do see a problem with that.
Obviously the authors should sue them to bankruptcy though.
You’re chastising others for a tone you are yourself employing, and are making monumental assumptions based on a few choice quotes. From the quotes alone you can’t tell if OpenAI thought libgen was sketchy or not.
Also, contrary to what you’re claiming, they were wrong. HN in general seems to approve on libgen when used for its purpose of downloading some books on an individual level. The complaint you’re replying to is about what OpenAI did with the data, it has nothing to do with the website they got it from.
papergirl•3h ago
That's just one of several interesting quotes that have surfaced in documents from the Authors Guild's lawsuit against OpenAI.
swiftcoder•1h ago
optimalsolver•1h ago
lensecat•49m ago
'OpenAI Feared “Optics,” Not the Law – “Dario Amodei, OpenAI’s then-Research Director, responded that ‘as a training set [LibGen is] a bit sketchier.’ [OpenAI researcher Sam] McCandlish explained: ‘I was just worried about optics – i.e. ‘openai uses copyrighted data from sketchy russian website’ showing up on [Hacker News] would be unfortunate.”'
gizajob•26m ago
locknitpicker•45m ago
As a side-note, most orgs already cover the need to STFU in their training material for new hires, particularly how personal comments should not and cannot represent the company. I'm sure this lawsuit will be explicitly mentioned in upcoming versions of this sort training material in multiple orgs.