top of page
Search

The Strange Economics of AI: Why books are being destroyed to build "Smarter" Machines

  • Writer: Saemi Nadine Jung
    Saemi Nadine Jung
  • 3 days ago
  • 10 min read



Just the other day, someone urgently asked me to come over and see something he had found on social media.


It was a video of a machine slicing through the spines of thousands of books, feeding the pages into scanners, and turning physical books into digital data.


My first reaction was... Who is doing this—and why?


The image reminded me of darker moments in history when books were banned or burned as a means of controlling information. It also brought to mind writers such as Paulo Freire, whose works were suppressed by authoritarian regimes because they challenged existing systems of power.


The answer reveals something deeply unsettling about where artificial intelligence is heading.


According to a recent Fortune report, Dutch booksellers began receiving unusual requests to purchase thousands of books in bulk. Some initially assumed the messages were spam or phishing attempts. But it turns out, they were part of a broader effort to acquire physical books, scan them, and use the resulting digital text as training data for AI systems.


At first glance, it can sound absurd. Why would anyone buy thousands of printed books and tear them apart?


Because books have become one of the most valuable commodities in the AI race.

Especially those written pre-AI times (before late 2022). They represent human-authored knowledge created before AI-generated text became widespread on the web, making them particularly attractive as high-quality training data.





Why AI companies are destroying books




The Search for Better AI Data (books ethics)


In the early days of AI, the internet was the primary source of training material for large language models. Web pages, articles, forums, research papers, and digital archives .. these provided an enormous amount of text.


But as AI-generated content spreads across the web, a new challenge emerged. What happens when machines increasingly learn from content created by other machines?

The quality of training data becomes critical.


AI systems need access to reliable, human-created knowledge. Books are attractive because they containt concentrated human expertise. They are written, edited, reviewed, and published by humans through processes designed to improve accuracy and clarity (at least before the rise of gen AI). A single book may contain years of research, experience, and thought compressed into hundreds of pages.


For AI developers, that makes books extremely valuable.



Why Destroy the Books?


This is where the story becomes controversial.


Industrial-scale scanning works best when books can be processed quickly. Removing the spine allows individual pages to move through high-speed scanners, transforming physical pages into searchable digital text. The process is known as destructive scanning.


Once a book has been scanned, the physical copy may no longer have practical value for the scanning operation. Keeping the original can add storage costs and slow the process.


From an efficiency standpoint, the logic is straightforward. The goal is not the book itself. The goal is to extract the knowledge inside it.


But from a cultural and ethical standpoint, many people find this process deeply unsettlling. It can feel as though something essential is being destroyed.


A book is not just a container of words. It is an artifactm a record of a particular moment in history, a product of human creativity, and often a bridge between generations. We experience books with our hands as much as with our eyes. Their weight, texture, annotations, and even signs of age become part of their story.


Watching thousands of books reduced to loose pages feels strange—even if the information survives digitally.





The Strange Economics of Destroying Something Valuable


The irony is difficult to ignore: in order to preserve human knowledge digitally, the physical objects that have preserved that knowledge for centuries are being destroyed.



This creates one of the strangest contradictions of the AI era. For centuries, books gained value because people wanted to read them, collect them, and preserve them.


Now, some books may gain value because machines want to consume them. A forgotten warehouse full of old academic books could suddenly become a valuable AI resource. An out-of-print title that has little commercial value to readers may become important because it contains information that a model can learn from.


The economics have shifted. The physical object may become disposable while the knowledge inside becomes more valuable than ever.


This is not entirely new. Libraries, archives, and databases have always preserved information by moving it into different formats.

But AI introduces a new scale and a new purpose.


The goal is no longer only preservation. It is transformation.

Turning human knowledge into something machines can process, analyze, and generate from.



More Than Just Copyright (Books)


Most discussions about AI training focus on copyright. Did companies have permission to use these works? Should authors and publishers be compensated?

Those questions matter.


But the book-scanning debate raises another question:

Should every efficient solution automatically be considered the right solution?

If someone legally owns a book, they may have the right to scan it. They may also have the right to dispose of it.


The legal question turned out to be more complicated than the emotional reaction. A federal judge ruled that scanning legally purchased books and using them to train AI models could qualify as fair use because the purpose was considered transformative. But the same protection did not extend to books obtained through piracy. The court effectively separated two questions: what companies do with books they legally acquire, and how they obtain the material in the first place. 


But societies often protect things beyond simple ownership. Libraries preserve books because knowledge can disappear. Archive s protect documents because their importance may not be obvious today. Museums preserve objects because their meaning extends beyond their immediate usefulness.


A destroyed book may continue to exist as digital information, but something about the original object is lost.


The debate is not only about legality.

It is about what we choose to value.



A New Supply Chain for AI


The rise of AI is creating demand for something unexpected: collections of human knowledge that were not tained by machines. Those can be found in:

Secondhand bookstores.

Academic publishers.

Bulk book sellers.

Library collections.

Out-of-print inventories.


These are becoming potential sources of AI training material. The first wave of AI relied heavily on mining the internet. The next wave may increasingly involve mining archives, collections, and physical repositories.


The race for better AI is becoming a race for better data. And some of the best data may be sitting quietly on shelves.



The Bigger Question for AI industry



This story is ultimately not just about books. It is about how society decides what knowledge is worth preserving and what is not. As AI systems become more advanced, companies will continue searching for high-quality sources of human expertise. Books are only one example.


The same questions may eventually apply to newspapers, scientific journals, historical documents, government archives, and private collections.


What should become part of an AI system’s memory?

Who decides?

And what happens to the original sources once their information has been extracted?


Writer and researcher Dan McQuillan captured the strange contradiction with a reference to Abbie Hoffman’s famous counterculture book Steal This Book, joking about a future title: “Destructively Scan This Book.”


The joke works because it captures the absurdity of this moment.

A book, a symbol of preserving ideas, is being sacrificed so that machines can learn from those ideas.




Final Thoughts

The image of books being purchased, sliced apart, scanned, and discarded represents one of the defining tensions of the AI era.


On one side is extraordinary technological progress: machines learning from centuries of human knowledge.


On the other is a growing debate about preservation, ownership, and the meaning of the objects that carry that knowledge.


The future of AI will not only depend on bigger models and faster computers.

It will depend on access to authentic human knowledge.


And increasingly, that knowledge may come from places we never expected—including the bookshelves we once thought were disappearing.




Epilogue




A technology built to absorb human knowledge is creating new questions about preservation, ethics, and what we choose to protect as part of our shared human heritage.


Whether destructive scanning becomes a standard industry practice or faces greater scrutiny remains to be seen.


But one thing is already clear:

The next frontier of AI is not just about building bigger models. It is about finding new sources of authentic human knowledge to teach them—and deciding what happens to the human artifacts that carry that knowledge.


The future of AI will depend not only on bigger models and faster computers, but also on access to authentic human knowledge.



AI 기업들은 왜 책들을 잘라버릴까?


*** 한국 미디어에서는 이 기사가 아직 뉴스화되지 않았습니다. 그래서 한국어로 써봅니다.


얼마 전, 한 지인이 다급하게 저를 불렀습니다. 소셜미디어에서 발견한 영상을 보여주겠다며 보게 된 영상 속에는, 한 기계가 수 백 권의 책등을 잘라내고, 페이지를 자동으로 스캐너에 넣어 처리하는 모습이 담겨있었습니다.


그 모습은 마치 살아있는 누군가의 척추를 잘라내는 것 처럼 보였습니다.


“누가 이런 짓을 벌이고 있는 거지? 그리고 왜 굳이 책을 저렇게 까지 해야 할까?"


이 장면은 마치 역사 속 중요한 책들이 금지되고 심지어 불태워졌던 어두운 순간들을 떠올리게 했습니다. 책을 파괴하는 행위가 단순한 폐기가 아니라 정보와 사상을 통제하는 수단으로 사용되었던 시대 말입니다.

저는 특히 파울루 프레이리(Paulo Freire) 같은 사상가들의 책이 떠올랐습니다. 그의 저서는 담고 있던 사상 때문에 권위주의 정권 아래에서 금지되고 억압받았습니다.


그런데 현재는 전혀 다른 이유로 책이 사라지고 있습니다.


최근 포춘(Fortune)의 보도에 따르면, 네덜란드의 일부 서점들은 수천 권의 책을 대량 구매하겠다는 이례적인 요청을 받았는데, 처음에는 스팸이나 피싱 메시지라고 생각했다고 합니다. 실제로는 AI 기업들이 물리적인 책을 확보하고, 이를 스캔해 디지털 텍스트로 변환한 뒤 자신들의 시스템 학습 데이터로 활용하려는 더 큰 목적이었습니다.


왜 이렇게나 수천 권의 책을 구매한 뒤 잘라버리고 물리적으로 파괴하는 걸까요?


그 이유는 AI 시대 이전에 쓰여진 책들은 더 이상 단순한 책이 아니기 때문입니다. 생성형 AI의 손을 타지 않은 금과 같은 책들이기 때문입니다.




더 나은 AI 데이터를 향한 경쟁



수 년 동안 인터넷은 대규모 언어 모델(LLM)을 훈련시키는 가장 중요한 데이터 공급원이었습니다. 웹페이지, 기사, 포럼, 연구 논문, 디지털 아카이브는 엄청난 양의 텍스트를 제공했습니다.


하지만 AI가 생성한 AI slop 과 같은 컨텐츠들이 웹 곳곳에 빠르게 퍼지면서 새로운 문제가 생기고 있습니다. 점점 더 다른 기계가 만든 컨텐츠를 LLM이 학습하게 된다면 어떻게 될까?

결국 중요한 것은 학습 데이터의 품질입니다.


AI 시스템은 신뢰할 수 있는 인간이 만든 지식에 접근해야 합니다.


책이 매력적인 이유는 책이 압축된 인간 전문성의 결과물이기 때문입니다. 대부분의 책은 저자, 편집자, 검토자, 출판 과정을 거치면서 정확성과 명확성을 높이는 과정을 거칩니다. 한 권의 책에는 수년간의 연구, 경험, 사고가 수백 페이지 안에 담겨 있습니다. AI 개발자들에게 책이 매우 가치 있는 이유입니다.


마찬가지의 이유로 academic texts 논문 등이 AI기업들에게 인기있는 텍스트인 이유 역시 몇 번의 검토, revision 등을 거치며 가장 quality 높은 text 로 분류되기 떄문입니다.


특히 생성형 AI가 대중화되기 이전, 즉 2022년 말 이전에 작성된 책들은 더욱 주목받고 있습니다.

이 책들은 AI가 생성한 콘텐츠가 본격적으로 확산되기 전의 순수한 인간 지식이 담긴 자료이기 때문입니다.



왜 책을 파괴하는가?



여기서부터 논란이 시작됩니다.

대규모 스캔 작업은 책을 빠르게 처리할 수 있을 때 가장 효율적입니다. 책등을 제거하면 페이지를 고속 스캐너에 연속적으로 넣을 수 있고, 물리적인 책은 검색 가능한 디지털 텍스트로 변환됩니다.


이 과정을 파괴적 스캐닝(destructive scanning)이라고 합니다.


책이 한번 스캔되고 나면 물리적인 사본은 더 이상 작업 과정에서 큰 가치가 없을 수 있습니다.

원본을 보관하는 것은 추가적인 저장 비용을 발생시키고 처리 속도를 늦출 수 있기 때문입니다.


효율성 측면에서는 논리가 분명합니다. 목표는 책 자체가 아닙니다. 목표는 책 안에 담긴 지식입니다.


하지만 문화적·윤리적 관점에서 많은 사람들은 이 과정을 불편하게 느낍니다. 마치 책이라는 존재의 중요한 일부를 파괴하는 것처럼 느껴집니다.


책은 단순히 글자를 담는 용기가 아닙니다.

책은 하나의 유물입니다.

특정 시대의 기록이고, 인간 창의성의 흔적이며, 때로는 세대와 세대를 연결하는 매개체입니다.

수천 권의 책이 낱장으로 흩어지는 모습을 보는 것은 이상한 감정을 불러일으킵니다.

비록 정보는 디지털 형태로 살아남더라도 말입니다.




가치 있는 것을 파괴하는 이상한 경제학



이것은 AI 시대가 만들어낸 가장 흥미로운 역설 중 하나입니다. 수백 년 동안 책의 가치는 사람들이 읽고, 소장하고, 보존하려 했기 때문에 생겼습니다. 하지만 이제 일부 책은 기계가 그것을 소비하려 하기 때문에 가치가 생길 수 있습니다. 오랫동안 창고에 방치되어 있던 오래된 학술 서적들이 갑자기 AI 데이터 자원이 되고, 독자들에게는 상업적 가치가 거의 없는 절판 도서가 AI 모델에게는 중요한 학습 자료가 됩니다.


물리적인 객체는 폐기될 수 있지만, 그 안의 지식은 그 어느 때보다 높은 가치를 갖게 되는 것입니다.


물론 이것이 완전히 새로운 현상은 아닙니다. 도서관, 아카이브, 데이터베이스는 항상 정보를 다른 형태로 옮겨 보존해 왔습니다. 하지만 AI는 그 규모와 목적을 바꾸고 있습니다.


목표는 더 이상 단순한 보존이 아닙니다.

변환입니다.

인간의 지식을 기계가 처리하고, 분석하고, 새로운 것을 만들어낼 수 있는 형태로 바꾸는 것입니다.



저작권 그 이상의 문제


AI 학습 데이터를 둘러싼 대부분의 논쟁은 저작권에 집중됩니다. 기업들이 이 자료를 사용할 권한이 있었는가? 저자와 출판사는 보상을 받아야 하는가?


이 질문들은 매우 중요합니다.


하지만 책 스캐닝 논쟁은 또 다른 질문을 던집니다.

효율적인 해결책이라고 해서 항상 올바른 해결책인가?


누군가 합법적으로 책을 소유하고 있다면 그것을 스캔할 권리가 있을 수 있다는 것이 이번 US court ruling 이었습니다. 그리고 그 책들을 심지어 폐기할 권리가 있을 수도 있습니다.

(실제로 법적인 판단은 감정적인 반응보다 훨씬 복잡했습니다. 한 연방법원 판사는 합법적으로 구매한 책을 스캔하고 AI 모델 학습에 사용하는 행위가 변형적 이용(transformative use)에 해당할 수 있으며, 공정 이용(fair use)으로 인정될 가능성이 있다고 판단했습니다. 하지만 불법적으로 확보한 책에는 같은 보호가 적용되지 않았습니다).


결국 법원은 두 가지 문제를 구분했습니다. 기업이 합법적으로 확보한 책을 어떻게 활용하는가. 그리고 그 자료를 처음에 어떻게 얻었는가.


하지만 사회는 단순한 소유권 이상의 가치를 보호합니다.

도서관은 지식이 사라질 수 있기 때문에 책을 보존합니다. 아카이브는 오늘날에는 중요성이 보이지 않는 기록도 미래를 위해 보관합니다. 박물관은 물건이 단순한 실용성을 넘어 의미를 갖기 때문에 보존합니다.


파괴된 책은 디지털 정보로 계속 존재할 수 있습니다.

하지만 원래의 물리적 대상이 가진 무언가는 사라집니다.

이 논쟁은 단순히 합법성에 대한 문제가 아닙니다.

우리가 무엇을 가치 있게 여길 것인가에 대한 문제입니다.



AI를 위한 새로운 공급망



AI의 발전은 예상하지 못했던 수요를 만들고 있습니다. 바로 기계에 의해 만들어지지 않은 인간 지식의 집합체입니다.


그 자료는 다음과 같은 곳에 존재합니다.

  • 중고 서점

  • 학술 출판사

  • 대량 도서 판매 업체

  • 도서관 소장 자료

  • 절판 도서 목록

이들은 모두 AI 학습 데이터의 잠재적 공급원이 되고 있습니다.


Gen AI의 첫 번째 물결은 인터넷을 채굴하는 것이었습니다.

다음 물결은 아카이브, 컬렉션, 물리적 저장소를 탐색하는 것이 될 수 있습니다.

더 나은 AI를 만들기 위한 경쟁은 결국 더 나은 데이터를 찾는 경쟁이 되고 있습니다.

그리고 그 데이터 중 일부는 지금도 조용히 책장 위에 놓여 있습니다.



더 큰 질문



이 이야기는 결국 단순히 책에 대한 이야기가 아닙니다.

사회가 어떤 지식을 보존할 가치가 있다고 판단하는지에 대한 이야기입니다.

AI 시스템이 더 발전할수록 기업들은 계속해서 고품질 인간 지식을 찾게 될 것입니다.

책은 그 중 하나일 뿐입니다.


앞으로 같은 질문은 신문, 과학 저널, 역사 문서, 정부 기록, 개인 소장 자료에도 적용될 수 있습니다.

무엇이 AI의 기억 속에 들어가야 할까요?

누가 그것을 결정할까요?


그리고 정보가 추출된 뒤 원본 자료에는 어떤 일이 일어날까요?

작가이자 연구자인 댄 맥퀼런(Dan McQuillan)은 애비 호프먼(Abbie Hoffman)의 유명한 반문화 서적 Steal This Book을 언급하며 미래의 책 제목을 농담처럼 제시했습니다.

“Destructively Scan This Book.”

“이 책을 파괴적으로 스캔하라.”


그 농담이 흥미로운 이유는 바로 지금 우리가 마주한 역설을 정확히 보여주기 때문입니다.

책이라는, 아이디어를 보존하는 상징이 이제 기계가 그 아이디어를 배우도록 하기 위해 희생되고 있습니다.




에필로그


책이 구매되고, 잘리고, 스캔되고, 버려지는 모습은 AI 시대가 가진 불편한 역설을 상징합니다.

인간의 지식을 흡수하기 위해 만들어진 기술이, 동시에 우리가 무엇을 보존하고 무엇을 보호해야 하는지에 대한 새로운 질문을 던지고 있습니다.


파괴적 스캐닝이 앞으로 산업 표준이 될지, 아니면 더 큰 규제와 감시를 받게 될지는 아직 알 수 없습니다.

하지만 한 가지는 분명합니다.


AI의 다음 단계는 단순히 더 큰 모델을 만드는 것이 아닙니다.

인간이 만들어낸 진짜 지식을 찾고, 그것을 학습시키며, 그 지식을 담아온 인간의 흔적을 어떻게 다룰 것인지 결정하는 것입니다.

 
 
 

Comments


 

© 2026 by The Critical Margin.  

 

bottom of page