Did OpenAI Pull Off the Biggest IP Heist of All-Time?
VCR recordings. Book digitization. The use of existing code to build a new operating system.
The legality of some of the most consequential technology and products in the last century have come down to a singular question: Are they protected under fair use, the legal doctrine allowing the use of copyrighted works without a license?
Now, that question has become the central battleground in landmark lawsuits between publishers, studios and recording companies against artificial intelligence companies over large language models. In the near future, these platforms are positioned to break through as the gateway to the internet — and information — to billions of people across the world.
If you ask these AI firms, they’ll say the secret sauce to their technology was innovation. The New York Times and eleven other publishers suing Microsoft and OpenAI over the use of their articles to train AI systems have a different answer: theft.
“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” said a 2023 internal Microsoft document, which noted that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”
The message was uncovered on Thursday in newly unsealed court documents detailing substantial concern within Microsoft and OpenAI over the ways in which they were developing their AI systems at the expense of publishers. They provide perhaps the most revealing indication yet that executives at the companies knew they may been have running afoul of intellectual property laws as they developed their technology.
The ingestion of copyrighted works is described by Microsoft Director of Applied Science Brent Hecht as “an astonishing theft of unprecedented proportions” and potentially “the largest theft of labor in human history,” the documents said.
At this stage of the litigation, OpenAI and Microsoft bear the burden of proving that their conduct constitutes fair use, which turns on a four-factor balancing test. The first prong asks what the copyrighted material was utilized for, including whether the use was commercial and whether it transformed the material for a new purpose.
Here, Supreme Court precedent holding that an “evasive motive” and violations of industry ethical standards undercuts a fair use defense could come into play. Consider an exchange between an OpenAI engineer and president Greg Brockman, who was told of a “hack to get around [the] nytimes paywall” as the company explored ways to scrape the site.
“Ah nice,” Brockman answered.
The response indicates that an executive knew of a technical method to circumvent a publisher’s access restriction and employed it at the company. The publishers argue this was part of broader conduct in which employees were instructed to devise strategies to circumvent paywalls “without detection.”
Contrast Brockman’s exchange with testimony from Microsoft CEO Satya Nadella, who stated “anything that is paywalled should be licensed by anyone who wants to use it,” according to the filing. If he had been made aware that OpenAI had scraped and trained on content that was restricted, he added that he would’ve “invoked” Microsoft’s right to require the Sam Altman-led company to retrain its models.
Another major focal point in the case will be fourth factor of the fair use doctrine, the effect of the use upon the potential market for or value of the copyrighted work. The publishers have argued that the chatbots undermine the market for original news journalism. A core thrust of the Times‘ complaint is that AI systems regurgitate near verbatim excerpts of articles when prompted. These responses, the publisher has said, go far beyond the snippets of texts typically shown with ordinary search results. One example: Bing Chat copied all but two of the first 396 words of its 2023 article “The Secrets Hamas knew about Israel’s Military.” An exhibit shows 100 other situations in which OpenAI’s GPT was trained on and memorized articles from the Times, with word-for-word copying.
“Our products are largely substitutive, period,” said Nick Turley, OpenAI’s head of ChatGPT, Thursday’s filing said.
Since the emergence of generative AI, publishers have seen drastically lower click-through-rates. Internal Microsoft data shows that the Times received significantly fewer clicks when users searched through Bing Chat rather than traditional web search, with rates falling 87% to 93%.
Also at play on this front: Arguments advanced by publishers that existing and future licensing opportunities is being undercut by the use of their copyrighted materials without payment. The publishers have argued that Microsoft and OpenAI can’t credibly claim that no licensing market exists for their content when they’ve reached deals with other publishers.
OpenAI and Microsoft prevailing on its copyright defense would “make a complete mockery of the idea of fair use,” said Hecht, according to the filing.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.