Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
No gain, all liability. Easier to cut them a check for access to training data and say nothing. Unless legal discovery was performed, the outside world would never know, and the payment records would roll off corporate records through a record retention schedule eventually. Could obfuscate it as a contractor consulting fee ("knowledge management subject matter expert") if you wanted to get tricky, depending on the risk appetite of whomever would receive the funds.
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
You missed Gigapedia (library.nu [1]) , which preceded most of the others.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
Probably the greatest achievement of modern society, a rebirth of the Library of Alexandria. Of course, only private for-business corporations are allowed to steal the world's knowledge, apparently, and only to be able to monetise it. There's something deeply wrong with our civilization that Anna's Archive is punished while companies like OpenAI, Google, Amazon, Anthropic, etc are just ... ignored when they do things like destructively digitise books or pirate things.
I don't understand why there is so much love for Anna's archive here. I feel like I'm getting whiplash because there is so much hate for AI companies training and profiting off of the worlds knowledge without licensing it. And yet, a site that directly facilitates that by taking payments from AI companies is lauded as this amazing and honorable thing. Can someone help me understand what im missing?
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
Obviously, not everyone on Hacker News shares the same opinion.
But, in general, Stewart Brand’s sentiment — expressed during a panel discussion at the first Hackers’ Conference in 1984, when he said that “information wants to be free” — is broadly shared here.
His full quote was: “On the one hand information wants to be expensive, because it’s so valuable… On the other hand, information wants to be free, because the cost of getting it out is getting lower and lower all the time.” Steve Wozniak famously responded: “Information should be free, but your time should not.”
Even among those on HN who believe that copyright serves a societal purpose, many still want copyright terms to be drastically shortened (back).
So I’m skeptical that the hostility toward AI companies is primarily based on their training on, and profiting from, the world’s knowledge without licensing it. I suspect it is rather based on their doing so without making that knowledge accessible in return. In the recent news of their destroying books, it even means reducing the chance that that knowledge will ever become open.
Yes, opening the world’s knowledge can lead to good and bad uses — just like free software can be used for great as well as for nefarious purposes. We’re generally too techno-optimistic to judge technology solely by its worst possible outcomes. But for such a benevolent view to be justified, the potential benefits need to be also clear.
AI companies turn public data into closed commercial products. They are allowed to profit off your copyrighted data, but only they get to profit from their own models.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
I don't hate either particularly, but I care more about getting any book I want whenever, wherever, than I care about some abstract far away moralizing of something of consequence to me only ambiguously.
Exactly. All libraries are worthy of being supported and grown. I make no distinction between a physical dead-tree library and a digital library.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
> If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
I consider libraries, public schools and maybe even public fire departments in the list of things that could never be proposed today if they didn't already exist.
It sounds like AA scraping Spotify was an unnecessarily risky move in terms of legal consequences and the toes they step on (who might be behind the downtimes). And it's not even fully in line with their aim of cataloging all the books in existence: while music could be argued is "information", one's sense of IP is much stronger with pirated music than with pirated books.
Fred and Wilhelmina are in bed together. Fred tosses and turns, unable to sleep. "Fred, what's troubling you?"
"You know Bjarne from the bank? Well, a balloon payment is due tomorrow on my loan, and I don't have the cash flow to pay it."
Wilhelmina thinks for a bit, then reaches for her phone. "Betty? Yes, this is Wilhelmina. Sorry to call so late. Would you tell Bjarne that Fred can't make the loan payment tomorrow? Yes, that's all. Good night."
Fred states at Wilhelmina, aghast. "What did you do THAT for?" She smiles. "Now it's Bjarne's problem. Let him toss and turn, you can go to sleep."
I only wish they allowed browsing journals by year and volume. Libgen allowed that but war in Ukraine broken libgen. It lives but as a much shadier alter egos that does not support all that original libgen supported.
My experience is that many of the books at Anna's Archive were scanned at the Internet Archive, using some compromise exposure setting that renders halftones decently, but puts text on a dingy gray background.
Anna's Archive is what TV told me in the 2000s that the future was. An online database with all published books one click away. Far from the dystopian reality than the corporate internet has become.
Might as well say "Annas Archive OWES ELEVENTY HUNDRED BILLIONTY-TRILLIONTY INFINITY DOLLARS!!!!!111" ala elementary school playground make-believe.
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
Many of us did, and I certainly still do, but there obviously weren't enough of us rebelling to convince them of how reprehensible it is to sue your own fans and sic the guns of the state on them.
People can be predisposed towards something, and then have a tipping point brought by periods of great stress
Jason Arday recently committed suicide due to media pressure, and even when the media was told that he was not mentally well and should please back off, they didn't.
The man was accused of plagiarism, and there was no final verdict on if he did. The institutions that hired him clearly had no issues in vetting him. And yet he committed suicide.
There are absolutely horrible people out there right now who's entire careers are shams, and don't even come close to thinking about it. People are built different, and have different triggers
"We are going to jail you for half or more of your life, restrict any medicines you might be on, treat you in deplorable conditions, constitute you as a slave of the state if we decide so". And there is no parole in federal prisons.
Feds did NOT have to choose to go the route they did. Their pursuance of a victimless "crime" lead to Aaron's only path was 'Exit'.
Sure, he committed suicide. Why? Because his life was already significantly threatened with what amounts to torture by prison.
Wikipedia usually has the most recent domains for sites like this. There's even a browser extension[1] that uses an `.idk` TLD to say "Look up the latest URL from this site's Wikipedia article".
Just make them owe $134 Trillion or whatever the evaluation was years ago suggested by the RIAA for their estimation of 'damages' for music piracy. It's about as meaningful.
> U.S. courts can’t reach domains registered beyond their jurisdiction. That’s likely to increase calls for site-blocking legislation, a measure the industry has long favored and that remains high on the political agenda in the United States.
Eventually US will have its own great firewall like China.
Politics is about powergrab and what we’ve seen is more and more power grab.
It’s likely the billionaires are able to buy the govt goons to pass the laws that allow them to be the gatekeepers.
Anthropic, OpenAI and the model builders massively benefited from the archived information. Distill entirety of archived human knowledge.
when I read title .
"Annas archive owes $340 million absolutely first thing that came to mind is
" How much do all these Ai companies owe for their unauthorised use of books and other media?"
how's Anna make money? I know they offered to share the database for ai training (which I would criticize, that's very shady) but do they have donations or something
it doesn't seem like it would be high revenue at least
Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Why don't Google, OpenAI, Anthropic, Facebook & Co defend Anna's Archive publicly?
Coming out would be a bold move for them.
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
No gain, all liability. Easier to cut them a check for access to training data and say nothing. Unless legal discovery was performed, the outside world would never know, and the payment records would roll off corporate records through a record retention schedule eventually. Could obfuscate it as a contractor consulting fee ("knowledge management subject matter expert") if you wanted to get tricky, depending on the risk appetite of whomever would receive the funds.
(not legal advice!)
Those are all law-abiding organizations, which AA is not.
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
I wouldn’t include YouTube given how only Google is allowed to index it
You missed Gigapedia (library.nu [1]) , which preceded most of the others.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
[1] https://en.wikipedia.org/wiki/Library.nu
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
Incredible website. Really one of the dreams of the internet realised, the ability to access all knowledge at your fingertips.
Probably the greatest achievement of modern society, a rebirth of the Library of Alexandria. Of course, only private for-business corporations are allowed to steal the world's knowledge, apparently, and only to be able to monetise it. There's something deeply wrong with our civilization that Anna's Archive is punished while companies like OpenAI, Google, Amazon, Anthropic, etc are just ... ignored when they do things like destructively digitise books or pirate things.
By "ignored" do you mean everyone on the Internet yelling about it all over all the platforms?
I was blown away when I paid to get an API key for it. The process is intricate and state of the art.
What's the process like? I wasn't aware there was a paid API key.
Didn’t they have a default judgment against them because they didn’t show up to court? Do they even know who runs it?
They railroaded Aaron Swartz (which eventually lead to his suicide) for much, much less.
One side wants to freely share knowledge with all of humanity, the other wants to restrict it to make a buck. I know who I support.
I don't understand why there is so much love for Anna's archive here. I feel like I'm getting whiplash because there is so much hate for AI companies training and profiting off of the worlds knowledge without licensing it. And yet, a site that directly facilitates that by taking payments from AI companies is lauded as this amazing and honorable thing. Can someone help me understand what im missing?
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
Obviously, not everyone on Hacker News shares the same opinion.
But, in general, Stewart Brand’s sentiment — expressed during a panel discussion at the first Hackers’ Conference in 1984, when he said that “information wants to be free” — is broadly shared here.
His full quote was: “On the one hand information wants to be expensive, because it’s so valuable… On the other hand, information wants to be free, because the cost of getting it out is getting lower and lower all the time.” Steve Wozniak famously responded: “Information should be free, but your time should not.”
Even among those on HN who believe that copyright serves a societal purpose, many still want copyright terms to be drastically shortened (back).
So I’m skeptical that the hostility toward AI companies is primarily based on their training on, and profiting from, the world’s knowledge without licensing it. I suspect it is rather based on their doing so without making that knowledge accessible in return. In the recent news of their destroying books, it even means reducing the chance that that knowledge will ever become open.
Yes, opening the world’s knowledge can lead to good and bad uses — just like free software can be used for great as well as for nefarious purposes. We’re generally too techno-optimistic to judge technology solely by its worst possible outcomes. But for such a benevolent view to be justified, the potential benefits need to be also clear.
AI companies turn public data into closed commercial products. They are allowed to profit off your copyrighted data, but only they get to profit from their own models.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
> Can someone help me understand what im missing?
One directly threatens the livelihoods of the commenters here, the other doesn't ;-)
I don't hate either particularly, but I care more about getting any book I want whenever, wherever, than I care about some abstract far away moralizing of something of consequence to me only ambiguously.
Anna's Archive is a gift to humanity
Exactly. if AI gets a pass on stealing all of the world's information and gets a pass why shouldn't we get to enjoy the same benefits?
Exactly. All libraries are worthy of being supported and grown. I make no distinction between a physical dead-tree library and a digital library.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
> If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
I consider libraries, public schools and maybe even public fire departments in the list of things that could never be proposed today if they didn't already exist.
Public libraries pay for their books
It sounds like AA scraping Spotify was an unnecessarily risky move in terms of legal consequences and the toes they step on (who might be behind the downtimes). And it's not even fully in line with their aim of cataloging all the books in existence: while music could be argued is "information", one's sense of IP is much stronger with pirated music than with pirated books.
And Google owes 20 decillion dollars to the Russian government. Who cares either way.
"If you own someone $340, that's your problem. If you owe someone $340,000,000, that's there problem."
Might as well be a trillion dollars as it's never getting paid.
"when you're $100,000 in debt, it's your problem. But when you're $1 million in debt, it's the bank's" - Rosalie Goes Shopping (1989)
An old variation:
Fred and Wilhelmina are in bed together. Fred tosses and turns, unable to sleep. "Fred, what's troubling you?"
"You know Bjarne from the bank? Well, a balloon payment is due tomorrow on my loan, and I don't have the cash flow to pay it."
Wilhelmina thinks for a bit, then reaches for her phone. "Betty? Yes, this is Wilhelmina. Sorry to call so late. Would you tell Bjarne that Fred can't make the loan payment tomorrow? Yes, that's all. Good night."
Fred states at Wilhelmina, aghast. "What did you do THAT for?" She smiles. "Now it's Bjarne's problem. Let him toss and turn, you can go to sleep."
I only wish they allowed browsing journals by year and volume. Libgen allowed that but war in Ukraine broken libgen. It lives but as a much shadier alter egos that does not support all that original libgen supported.
Rembember to seed torrent kids
"owes" seems like a wrong word here
My experience is that many of the books at Anna's Archive were scanned at the Internet Archive, using some compromise exposure setting that renders halftones decently, but puts text on a dingy gray background.
So when are Nvidia, Facebook, and everyone else going to chip in and pay up?
Did you miss that they are huge corporations that the law doesn't apply to?
if this is true. what does this say about "rule of law"? is it all fiction?
I don’t get it. I am not a corp. i can download books from AA (who does the law apply to exactly?)
Anna's Archive is what TV told me in the 2000s that the future was. An online database with all published books one click away. Far from the dystopian reality than the corporate internet has become.
I wonder what would happen if Annas archive announced
In response to the 340M fine, .... "We have set up a LLM and are deleting all our books ? "
Unfortunately it’s blocked in Germany
only the local providers in Germany are forced to block it via DNS - 8.8.8.8 to the rescue
Tor Browser is handy, if simply changing your DNS isn't enough
Might as well say "Annas Archive OWES ELEVENTY HUNDRED BILLIONTY-TRILLIONTY INFINITY DOLLARS!!!!!111" ala elementary school playground make-believe.
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
> elementary school playground make-believe
Oh yeah? Well, my dad works at Anna's Archive.
> I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster.
Why didn't Metallica's fans rebel? That would have stopped it quickly and set a precedent and example for the rest.
Many of us did, and I certainly still do, but there obviously weren't enough of us rebelling to convince them of how reprehensible it is to sue your own fans and sic the guns of the state on them.
[flagged]
[flagged]
People can be predisposed towards something, and then have a tipping point brought by periods of great stress
Jason Arday recently committed suicide due to media pressure, and even when the media was told that he was not mentally well and should please back off, they didn't.
The man was accused of plagiarism, and there was no final verdict on if he did. The institutions that hired him clearly had no issues in vetting him. And yet he committed suicide.
There are absolutely horrible people out there right now who's entire careers are shams, and don't even come close to thinking about it. People are built different, and have different triggers
Sure. That he was facing 35 years in prison for his “crimes“ had nothing to with his suicide.
If only you were able to take your perspective one step farther. Why did he commit suicide?
Why did he commit suicide?
"We are going to jail you for half or more of your life, restrict any medicines you might be on, treat you in deplorable conditions, constitute you as a slave of the state if we decide so". And there is no parole in federal prisons.
Feds did NOT have to choose to go the route they did. Their pursuance of a victimless "crime" lead to Aaron's only path was 'Exit'.
Sure, he committed suicide. Why? Because his life was already significantly threatened with what amounts to torture by prison.
What are current domains? I tried the ones from article but got redirected to spam
Wikipedia usually has the most recent domains for sites like this. There's even a browser extension[1] that uses an `.idk` TLD to say "Look up the latest URL from this site's Wikipedia article".
[1] https://github.com/aaronjanse/dns-over-wikipedia
I made a service specifically to translate Wikipedia entries to DNS with Anna's Archive as the main use case: https://whither.link/
.gl works for me.
.gd works for me
Have they released the Spotify scrape already?
Just wished they didn't sell bulk access for LLM training.
Just make them owe $134 Trillion or whatever the evaluation was years ago suggested by the RIAA for their estimation of 'damages' for music piracy. It's about as meaningful.
Reminder for me to make another donation.
Piracy is morally justified at this point.
I wonder by if the US considers this “supporting a terrorist organization”.
Yeah, truthfully. This type of article is generally a nice nudge to donate.
taking shots against the King?
better not miss!
Thanks for reminding me I need to download some books for the holidays!
> U.S. courts can’t reach domains registered beyond their jurisdiction. That’s likely to increase calls for site-blocking legislation, a measure the industry has long favored and that remains high on the political agenda in the United States.
Eventually US will have its own great firewall like China.
Politics is about powergrab and what we’ve seen is more and more power grab.
It’s likely the billionaires are able to buy the govt goons to pass the laws that allow them to be the gatekeepers.
Anthropic, OpenAI and the model builders massively benefited from the archived information. Distill entirety of archived human knowledge.
"owes"
Maybe the US companies could solve their legal boundary problems by taking their issues to the International Court of Justice... /s
[dead]
it "Owes" just like Anthropic and OpenAI "Own" their models trained on the worlds collective IP.
when I read title . "Annas archive owes $340 million absolutely first thing that came to mind is " How much do all these Ai companies owe for their unauthorised use of books and other media?"
I really don't understand why HackerNews gets hard for Anna's Archive but you never see other illegal download websites praised here.
They're not more or less virtuous than any others, they're in it for the money and you're a fool if you think otherwise.
I'm not against piracy, it's great, but to claim they are moral saints is a joke.
Anyway, bring in the downvotes as I know will happen.
how's Anna make money? I know they offered to share the database for ai training (which I would criticize, that's very shady) but do they have donations or something
it doesn't seem like it would be high revenue at least
I've seen support here for the pirate party and other things like pirate bay.
I praise all pirates and leakers
Maybe you can own (some) things, you can't own information
As a hacker News viewer I'm for other illegal download sites as well.
If you just Google 'free media heck yeah' you'll see a great deal of them!
Can't talk for all of HN, but I for one hereby praise most/all piracy websites. Anna's Archive is great, and Rutracker is also great.
It's not a competition, and being a moral saint isn't a good KPI; being a net positive for society is.
[dead]