Copyright risk at AI companies often starts long before a model produces any answer. The harder questions usually come later, during fundraising diligence, enterprise customer reviews, or a complaint from a rights holder: where did the material come from, who had the authority to hand it over, did the company receive only reading access or also the right to copy, split, train on, and retain it, and which knowledge base or model version did the material enter?
That is the frame of an article by Zhao Xuan, which uses Anthropic’s legal settlement as a way to unpack how courts may treat different steps in the AI data pipeline. On July 20, 2026, a U.S. federal court granted final approval to Anthropic’s $1.5 billion class-action settlement with certain authors, publishers, and other rights holders. The article stresses that the $1.5 billion figure was not a court-imposed infringement damages award, nor was it a universal license price for using books to train AI systems. The more important point, it says, is that the court separated data acquisition, database construction, model training, and later use, and reviewed them one by one.
Why training received support, but data intake did not automatically become lawful
Under the facts of the Anthropic case, the court found that using books for model training served to analyze linguistic structure and statistical relationships rather than offer users a substitute for the original books. On that basis, the training use was treated as transformative and could qualify as fair use.
The court also accepted a narrower form of digitization: buying physical books, scanning them into digital files, and destroying the corresponding physical copies. Where the number of copies did not increase and the files were not distributed outward, that kind of format conversion could also receive support.
The article says the outcome was different when it came to downloading large numbers of ebooks from pirate sites and placing them into a permanent, general-purpose library. The court did not accept the argument that front-end piracy and long-term retention should be swept into fair use merely because the files might later be used for training.
That distinction matters. A transformative training purpose does not solve the legality of the acquisition method itself. At the same time, the piece notes that this does not mean the court had already issued a final merits ruling that Anthropic was liable for pirated copying. Liability and damages were originally still set to be litigated before the case ended through settlement. In that sense, Anthropic’s decision to pay $1.5 billion is described more as a price on class-action exposure, statutory damages risk, trial uncertainty, and appeal uncertainty than as a court-ordered "AI training fee."
The author illustrates the point with a simple comparison: a chef studying the structure and writing of a cookbook is not the same thing as downloading an entire collection of cookbooks from a pirate site and storing them permanently in a private repository. The first question is how a work is studied and used. The second starts earlier: where did the material come from, why could it be copied, and why was it kept for the long term?
Seen that way, data acquisition, ingestion and copying, cleaning and chunking, model training or knowledge-base inclusion, retrieval output, and long-term retention are separate links in one chain. A legal basis for one step does not make the entire chain lawful by default.
Using a third-party model does not move all data risk upstream
Many AI startups do not train foundation models themselves. They call third-party models through APIs, then build industry knowledge bases, retrieval-augmented generation systems, agents, or vertical products on top. That structure may reduce some of the risk tied to training a base model directly, but the article argues it does not transfer a company’s own data-handling responsibility to the upstream model provider.
Companies first need to separate "data" from "works." Facts, numbers, formulas, and general information may not qualify as protected works under copyright law. The same dataset, though, can also contain articles, images, reports, personal information, trade secrets, platform-based data interests, and contractual restrictions.
Because of that, a company cannot decide whether material may enter an AI system by asking only whether copyright exists. It also has to ask what the material contains, how it was obtained, and whether the intended use falls within the actual authorization granted.
The article also rejects the idea that RAG sits outside copyright risk. A report that enters a RAG system is commonly downloaded, converted in format, cleaned, chunked, vectorized, and stored in a knowledge base. Even if the model eventually outputs only a few summary lines, copying may already have occurred in the technical steps that came before, and those steps may exceed the database terms or the customer’s authorization. At the same time, the presence of protected works in a knowledge base does not mean every retrieval output is necessarily infringing. The answer still depends on how copying occurred, what the output contains, and what the authorization permits.
For that reason, "we only do RAG" is not a substitute for data review, and using a third-party model does not relieve a company of liability arising from building its own knowledge base, handling client materials, or optimizing a product over time.
Three points Chinese AI companies often get wrong
If something is not a protected work, that does not mean it is free to use
Even where content falls outside copyright protection, it may still involve personal information, trade secrets, platform rules, database-related interests, and contractual duties. A company cannot conclude that no restriction exists simply because it believes a set of materials carries no copyright.
Publicly accessible or paid-for content is not automatically training-ready
Public accessibility only shows that a user may reach content under certain conditions. It does not automatically permit bulk downloading, copying, chunking, fine-tuning, or long-term storage in a commercial knowledge base. The article compares a paid database account to a reading-room ticket: it may allow searching and reading, but not moving the collection into a company repository.
It cites Article 29 of China’s Copyright Law, which states that where a license agreement does not clearly grant a right, the other party may not exercise that right without the copyright holder’s consent. In other words, having a lawful source and having authorization for the current use are two different matters. The right to read, download, use internally, build a knowledge base, and train a model are not naturally one and the same set of rights.
Falling outside the Interim Measures does not remove other compliance duties
Some internal knowledge bases, in-house enterprise tools, or specific B2B services may not directly fall under the Interim Measures for the Administration of Generative Artificial Intelligence Services. Even so, the business may still be constrained by copyright rules, personal information protection, data security obligations, unfair competition law, client contracts, and platform terms.
Fair use is not a blanket pass
The article says the Anthropic case applies U.S. copyright law and cannot simply be transplanted into China. Article 24 of China’s Copyright Law lists specific fair-use scenarios, and China has not yet created a broad text-and-data-mining exception that generally applies to commercial AI training.
That does not mean all AI training is necessarily infringing. It does mean companies should not jump to a fair-use conclusion on the basis that training is innovative, that the model does not output the original text, or that the use does not substitute for the original work. The real analysis still turns on whether the technical process created copies of protected expression, at which step the copying occurred, whether the model, knowledge base, or final output retains or presents identifiable protected expression, and whether the conduct fits within a specific category recognized by current law.
The piece adds that China’s Supreme People’s Court has already called for prudent handling of new types of cases involving training corpora for large models and AI-generated content infringement. That suggests domestic adjudication standards are still forming. As a result, companies should not build core products and data assets entirely on top of an unsettled fair-use defense.
What to do with old data: stop the inflow first, then deal with the backlog
For startups that have already been operating for some time, the practical question is often not what to do next, but what to do with data that is already inside the system. Many companies cannot stop products immediately, empty a knowledge base, and retrain from scratch. The article says any remediation plan has to balance legal risk, technical cost, and business continuity.
The more realistic order, it argues, is to stop adding high-risk data first and then inventory what is already there. Preserve source records, contracts, and usage logs first, then decide whether the material should be isolated, replaced, or deleted.
Companies can begin with a data list that records the source of the materials, the date of acquisition, the provider, applicable contracts and platform terms, purpose of use, storage location, and the knowledge base, fine-tuning task, or model version the materials entered.
High-risk data should be handled first. The article lists files from pirate sites or files with no explainable origin, data supplied by vendors that cannot show a rights chain, materials used beyond the scope of a client’s authorization, and content obtained by bypassing paywalls, anti-crawling measures, or access controls. Conduct like that may create copyright and contractual risk and, where conditions are met, unfair competition liability as well.
Remediation is not limited to a choice between continued use and immediate deletion. A company can stop new intake, limit access, and isolate suspicious data from the production environment. After that, it can seek supplemental authorization, replace the material with compliant data, rebuild the knowledge base, or fine-tune again. If deletion becomes necessary, the company should be able to locate related files, chunked data, vector records, and backup copies, while keeping a complete internal approval and remediation record.
Those measures do not retroactively legalize past conduct. The article’s point is narrower: they can keep risk from spreading further and help show, in later financing, customer review, or dispute resolution, that the company can identify problems, contain impact, and complete repairs.
For startups with limited resources, a data ledger is the practical starting point
The article does not call for a heavy compliance system from day one. For resource-constrained startups, it says the most useful first step is a data ledger maintained jointly by business teams, technical staff, and legal personnel.
For each batch of data, the minimum record should include who provided it, where it was obtained, which contract or platform terms apply, what scenarios are permitted, what scenarios are prohibited, which product, knowledge base, or model version it entered, how long it may be retained, and how it should exit once a complaint arrives, a license expires, or a partnership ends.
Contracts should also reflect real technical use cases. Rather than relying on broad language such as "the data source is lawful" or "the vendor guarantees non-infringement," the article recommends spelling out whether the material may be downloaded, copied, cleaned, reformatted, chunked, entered into a knowledge base, used for retrieval, fine-tuning, evaluation, or model optimization, along with the relevant product scope, authorized users, time period, and procedures for handling third-party rights claims.
Companies also need to know exactly which knowledge base or model version a dataset entered and whether it can be isolated, replaced, or removed on its own. If data cannot be located, it becomes hard to contain damage once a complaint is made. If it cannot be separated and exited, uncertainty rises in financing, M&A, and enterprise customer reviews.
The broader warning from the Anthropic case
The article ends by saying the most useful lesson for Chinese AI companies is not the isolated phrase that training may qualify as fair use, and not the $1.5 billion figure by itself. It is the court’s decision to examine each stage of data use separately: where the material came from, why it could be copied, for what purpose it entered a repository, and why it was kept over time.
Fair use is not a shield that cures an unlawful source. A lawful source is not the same thing as unlimited rights to train, retrieve, and retain. Calling a third-party model may reduce part of the risk tied to foundation-model training, but it does not solve the responsibility created when a company builds its own knowledge base, handles client material, or refines a product.
For startups, the article says, it is not too late to start. Control today’s new data first, then work through the historical backlog. The earlier a company establishes source records, authorization boundaries, and exit mechanisms, the lower the cost of later remediation, and the easier it becomes to turn data from a latent risk into an asset that can withstand financing and customer scrutiny.

