LearnerBox logo LearnerBox Infosystems LLP

AI Ethics and Governance

Insights and guidelines on responsible AI development, algorithmic fairness, safety regulations, and ethical decision-making.

  • AI copyright lawsuits represent the evolving legal framework for artificial intelligence
    AI Ethics and Governance

    The Critical Wave of AI Copyright Lawsuits Reshaping the Future of Generative Models

    Six Million Pirated Books and a Billion Dollar Question

    On July 20, 2026, a federal judge in San Francisco approved the largest copyright class action settlement in United States history. Anthropic agreed to pay 1.5 billion dollars to a class of authors and publishers, roughly 3,000 dollars for each of an estimated 500,000 works, after admitting it had downloaded as many as seven million pirated books to train its Claude models.

    That single ruling has become the anchor point for understanding the entire wave of AI copyright lawsuits now working through American courts, lawsuits that touch every major AI lab and that will, collectively, determine whether the current generation of large language models was built on a legally sound foundation or a legally precarious one.

    The scale of the problem is genuinely industry wide. Dozens of cases are pending against OpenAI, Google, Microsoft, Meta, Midjourney, and Stability AI, filed by novelists, journalists, musicians, visual artists, and news organisations, all making some version of the same core allegation: that these companies trained their models on copyrighted material without consent, without a licence, and in many cases without even paying for the content in the first place.

    How the Training Actually Happened

    To understand why AI copyright lawsuits have multiplied so quickly, it helps to understand what actually happened during the early training runs of today’s frontier models. Building a capable large language model requires ingesting enormous volumes of text, historically hundreds of billions to trillions of words. Assembling a dataset at that scale through licensed content alone would have been prohibitively slow and expensive in the early 2020s, when the race to build the first genuinely capable chatbots was at its most intense.

    Court filings across multiple AI copyright lawsuits reveal that several major labs took shortcuts. In the case against Anthropic, court records showed the company downloaded books directly from known pirate library websites, essentially the same category of site used for illegal book sharing, and stored them in a centralised internal dataset used to train Claude.

    The court drew a sharp legal distinction that has become central to nearly every subsequent case: training a model on lawfully acquired copyrighted books can plausibly be fair use, because the training process transforms the material into statistical patterns rather than reproducing it. But acquiring the books through piracy in the first place is a separate, independently unlawful act, regardless of what happens to the data afterward.

    The Anthropic Settlement: A Landmark With Limited Reach

    The Anthropic settlement deserves close examination because it will shape negotiating positions across every other pending case. The underlying case, Bartz v. Anthropic, was filed in August 2024 by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson. In June 2025, Judge William Alsup issued a pivotal ruling: training an AI model on lawfully acquired copyrighted books was fair use because the process was sufficiently transformative, but Anthropic’s use of pirated copies to build its library was not protected, and that narrower claim would proceed to trial.

    Rather than face trial, Anthropic settled for 1.5 billion dollars, an amount its own lawyers described as the largest publicly reported copyright recovery in history. The settlement required Anthropic to destroy its pirated dataset entirely. Crucially, legal experts covering the wave of AI copyright lawsuits have been careful to note what the settlement does not do. It does not establish binding legal precedent, because a settled case never reaches an appeals court.

    It does not grant Anthropic a licence for any future training. And it does not resolve the central industry wide question, whether training AI models on copyrighted material is lawful, since that question was never actually decided at trial. As one law professor put it, appeals courts still need to weigh in on the larger question of if and how AI companies can use copyrighted works, and they will have plenty of opportunities ahead, because dozens of similar AI copyright lawsuits remain active against other companies right now.

    Meta: A Parallel Case Still in Progress

    Kadrey v. Meta Platforms, filed by a similar group of authors, follows an almost identical fact pattern to the Anthropic case, and its diverging trajectory illustrates how unpredictable this legal landscape remains. The court granted Meta a partial win on fair use grounds regarding the actual training of its Llama models. But a separate and more damaging allegation survives: that Meta engaged in what is known as seeding during the torrenting process.

    Meaning that in addition to downloading pirated books, Meta’s systems may have simultaneously redistributed those pirated files to other users on the same file sharing network, a potentially more serious violation than simple downloading. That claim remains active in the Northern District of California, and unlike the Anthropic case, Meta has not settled, meaning this particular thread of AI copyright lawsuits is still heading toward further discovery and potentially trial.

    OpenAI and The New York Times: The Fight Over Memorisation

    The most closely watched of all the active AI copyright lawsuits is The New York Times v. OpenAI and Microsoft, filed in December 2023 after nine months of failed licensing negotiations. Unlike the Anthropic and Meta cases, which centre on how training data was acquired, the Times case centres on a different and arguably more consequential legal question: whether ChatGPT can reproduce, or regurgitate, substantial portions of Times articles verbatim when prompted in specific ways.

    The Times alleges that its journalists’ work does not merely inform the model statistically but can, in certain circumstances, be extracted from it nearly word for word, a claim that goes to the heart of whether training itself was transformative or whether the model functions, in part, as an unlicensed distribution mechanism for the underlying copyrighted text.

    The case has become unusually contentious on discovery grounds. A magistrate judge ordered OpenAI to produce twenty million de-identified ChatGPT conversation logs, a demand OpenAI fought vigorously on user privacy grounds before ultimately complying under court order. OpenAI has publicly argued that its use of Times articles is a transformative, non-expressive analytical use protected by fair use, pointing to the Anthropic and Meta rulings on training as precedent in its favour.

    As of mid-2026, the case remains in the Southern District of New York, consolidated with similar suits from other news organisations into a multidistrict litigation, with expert reports completed in late 2025 and summary judgment briefing concluding in April 2026. No trial date has yet been set, but this case, more than any other among the current AI copyright lawsuits, is widely viewed as the one most likely to produce a binding, tested legal precedent, precisely because OpenAI has shown far less inclination to settle than Anthropic did.

    The Broader Pattern: Every Creative Industry Is Suing

    The scope of AI copyright lawsuits extends well beyond books and news journalism. Disney and other major studios have sued Midjourney over image generation trained on copyrighted characters and artwork. Getty Images sued Stability AI over the training of its image generation models. Recording labels have pursued similar claims against AI music generation platforms. As one legal expert tracking the litigation observed, the lawsuits are coming from essentially every sector of human creativity, newspapers, recording labels, movie studios, and more, reflecting a consistent grievance across creative industries that their work was used as raw material for a multi billion dollar commercial product without consent or compensation.

    The Licensing Shift: Settling Before the Courtroom

    One of the more significant developments to emerge from this wave of AI copyright lawsuits is a visible shift in industry behaviour, away from litigation risk and toward proactive licensing. Facing the reputational and financial exposure demonstrated by the Anthropic settlement, several AI companies have begun negotiating licensing agreements with publishers directly rather than waiting to be sued. The New York Times itself, even while actively suing OpenAI, separately reached a multiyear licensing agreement with Amazon in 2025, reportedly worth twenty to twenty five million dollars, for use of its content in Amazon’s AI products, demonstrating that litigation and licensing are not mutually exclusive strategies for the same publisher.

    A Columbia University law professor who studies literary property rights described receiving a steady stream of requests from her own publishers asking her to authorise licensing of her books to AI companies, a pattern that suggests the publishing industry as a whole is moving toward a licensing market for AI training data, driven directly by the legal exposure that the current AI copyright lawsuits have made unmistakably clear.

    What the Outcome Will Determine

    The stakes in this litigation extend well beyond the specific dollar amounts at issue. If courts ultimately rule broadly in favour of AI companies on fair use grounds for lawfully acquired training data, as the preliminary Anthropic and Meta rulings suggest they might, the legal foundation for training future models on publicly available text becomes considerably more secure, provided companies avoid the piracy shortcuts that triggered these specific lawsuits.

    If courts rule against AI companies, particularly on the regurgitation question at the centre of the New York Times case, the entire industry may need to rebuild significant portions of its training pipelines around licensed content, a shift that would fundamentally alter the economics of building frontier AI models and could meaningfully slow the pace of model development across the industry.

    Conclusion

    The current wave of AI copyright lawsuits represents one of the most consequential bodies of litigation in the technology industry’s history, not because any single case will resolve every open question, but because each ruling and each settlement is incrementally shaping the legal architecture within which every future AI model must be built. The Anthropic settlement demonstrated the scale of financial exposure that piracy based training data creates.

    The Meta case demonstrates that even companies who win on the core fair use question can remain exposed on narrower, related claims. And the OpenAI case, still without a trial date but approaching a decisive summary judgment ruling, may ultimately decide whether training itself, done properly and without piracy, is legally sound at all. For publishers, authors, and AI companies alike, the outcome of these AI copyright lawsuits will define the ground rules for how intelligence itself gets built for years to come.

  • Smart car sensors and autonomous vehicle accidents liability
    AI Ethics and Governance

    Who Is Legally Responsible for Autonomous Vehicle Accidents? A Critical Guide to AI, Ethics, and the Law

    Sally Knew the Road. But Who Was Responsible?

    In 1953, Isaac Asimov published a short story called “Sally” in Fantastic magazine. In it, self-driving cars with positronic brains roam a farm for retired automobiles, developing personalities and emotional responses. When a villainous character attempts to exploit them, Sally and the other cars act to protect themselves and the humans they care for. Asimov, as he so often did, had seen something clearly: that autonomous vehicles capable of independent action would inevitably raise questions that went far beyond engineering. Questions about agency, intention, and most pressingly, responsibility.

    Seventy years later, those questions are no longer philosophical. They are legal, regulatory, and deeply urgent. Autonomous vehicle accidents liability has become one of the most contested areas in technology law, as self-driving cars move from test tracks to public roads and courts, insurance companies, and legislators scramble to answer the question Asimov posed in fiction: when an autonomous vehicle causes harm, who is accountable?

    The Scale of the Problem

    Autonomous vehicles are no longer experimental. Waymo, the autonomous driving subsidiary of Alphabet, completed over four million fully driverless trips in 2024 and is currently expanding into new cities including Miami and Tokyo. Tesla’s Full Self-Driving system is active on hundreds of thousands of vehicles on public roads. In 2025, the US National Highway Traffic Safety Administration (NHTSA) reported receiving over 2,500 incident reports involving vehicles with automated driving features, a figure that represents only a fraction of actual incidents due to inconsistent reporting requirements.

    The accidents are real and, in some cases, fatal. In 2023, a Cruise autonomous vehicle in San Francisco struck a pedestrian who had already been hit by another car, then dragged her 20 feet before stopping. California’s Department of Motor Vehicles revoked Cruise’s operating licence. General Motors ultimately shut down the Cruise unit. In 2024, a Waymo robotaxi in Phoenix struck a cyclist who had run a red light. The cyclist was injured but survived. Both incidents raised the same fundamental question that courts and regulators have yet to answer cleanly: who is responsible for autonomous vehicle accidents when the vehicle itself made the decision?

    The Legal Framework in the United States

    Under current US law, autonomous vehicle accidents liability falls into a patchwork of state-level frameworks with no coherent federal standard. This is partly because US traffic law has historically been a state matter, and partly because Congress has repeatedly failed to pass comprehensive autonomous vehicle legislation, most recently when the SELF DRIVE Act stalled in 2021.

    In practice, autonomous vehicle accidents liability in the US currently resolves through three legal theories. The first is product liability: the argument that the autonomous vehicle system was defective and the manufacturer is responsible, in the same way that a tyre manufacturer is liable for a blowout caused by a manufacturing defect.

    The second is negligence, directed either at the manufacturer for deploying an inadequately tested system, or at the human operator, if one was present, for failing to intervene. The third is agency liability, an emerging theory that treats the vehicle’s AI as acting on behalf of its manufacturer, making the manufacturer responsible for the AI’s decisions the way an employer is responsible for an employee’s actions in the course of their work.

    Most autonomous vehicle accident lawsuits in the US have settled before reaching verdict, which has slowed the development of clear case law. Arizona, California, and Texas have enacted their own autonomous vehicle frameworks, but these differ significantly on questions including whether a human must be present in the vehicle, what data must be retained after an incident, and who must report accidents and to whom.

    The NHTSA’s Standing General Order, introduced in 2021 and strengthened in 2023, now requires manufacturers to report all crashes involving automated driving systems within one day if an airbag deployed or a fatality occurred. This has dramatically improved data collection but has not resolved the underlying autonomous vehicle accidents liability question.

    The European Approach: Stricter, Clearer, but Still Evolving

    The European Union has taken a more structured approach to autonomous vehicle accidents liability. The EU’s updated Product Liability Directive, which came into force in December 2024, explicitly covers AI systems and software as “products,” meaning that manufacturers of autonomous vehicle systems can be held liable for damages caused by defects in their AI, including defects that arise from inadequate training data, flawed algorithms, or failure to update the system against known risks.

    Germany moved first among EU member states, passing the Autonomous Driving Act in 2021, which created a legal framework for Level 4 autonomous vehicles (those capable of driving themselves in defined conditions without human intervention) and established that when an autonomous vehicle accidents liability question arises, the vehicle owner’s compulsory insurance covers damages, with the right to pursue the manufacturer if a technical defect is found. Several other EU countries are following Germany’s model.

    The EU AI Act, whose high-risk provisions took full effect in August 2026, classifies autonomous vehicle AI systems as high-risk, requiring conformity assessments, extensive documentation, transparency obligations, and post-market monitoring. For autonomous vehicle accidents liability specifically, the Act requires that high-risk AI systems maintain logs sufficient to trace decisions that led to incidents, which for the first time gives courts access to the AI’s decision record rather than relying solely on witness accounts and physical evidence.

    Ethical Dimensions: The Trolley Problem at 70 Miles Per Hour

    The legal question of autonomous vehicle accidents liability cannot be cleanly separated from the ethical one. Autonomous vehicles must, by design, make split-second decisions that in human drivers arise from instinct and moral intuition. The philosophical thought experiment known as the trolley problem, in which an actor must choose between allowing one harm or actively causing a lesser one, is not abstract for an autonomous vehicle. It is a real design decision encoded in the vehicle’s decision-making algorithm.

    In 2016, a survey published in Science found that while most people agreed autonomous vehicles should be programmed to minimise total casualties, they were less willing to purchase a vehicle programmed to sacrifice its own occupant to save pedestrians. This tension, between what is collectively optimal and what is individually acceptable, has no clean resolution.

    Asimov’s Three Laws of Robotics, which he spent a career demonstrating were insufficient for the complexity of real-world moral situations, come to mind here. Sally’s cars protected their passengers and themselves. A real autonomous vehicle’s priority hierarchy must be set by someone, and whoever sets it is making an ethical choice with legal consequences.

    The MIT Moral Machine experiment, which collected 40 million decisions from participants in 233 countries about autonomous vehicle ethical dilemmas, found significant cultural variation in how people prioritised pedestrians versus passengers, the young versus the old, and law-abiding versus law-breaking road users. There is no universal answer. And yet autonomous vehicle manufacturers must encode one.

    Reducing Autonomous Vehicle Accidents: What Is Actually Working

    Despite the legal and ethical complexity, the safety record of mature autonomous vehicle systems is improving. Waymo published a peer-reviewed study in 2024 showing that its vehicles were involved in significantly fewer injury-causing crashes per mile than human-driven vehicles in comparable environments. The key advances driving this improvement include higher-resolution sensor fusion combining LiDAR, radar, and cameras with redundant processing; improved simulation training using synthetic edge-case scenarios that would be dangerous or impossible to stage in reality; and V2X (vehicle-to-everything) communication, which allows vehicles to share real-time information about hazards, traffic conditions, and each other’s positions.

    Regulators are also improving the frameworks around incident reporting and investigation. The NHTSA’s new autonomous vehicle data portal, launched in 2025, makes incident data publicly available in near real-time, enabling researchers to identify failure patterns and manufacturers to issue targeted system updates far more quickly than the traditional recall process allows.

    Conclusion: Asimov’s Question, Still Unanswered

    Autonomous vehicle accidents liability remains one of the most unresolved legal questions in modern technology law. The US is moving toward resolution through litigation and piecemeal state legislation; the EU is moving toward it through structured regulation. Neither has yet produced a framework that satisfactorily answers the core question: when an AI makes a decision that injures or kills someone, who is responsible?

    Asimov’s Sally knew what she wanted to do and did it. The humans around her were left to reckon with the consequences. In 2026, we are in precisely that position, only the stakes are not fictional, the roads are public, and the answer matters enormously to everyone who shares them.