The key question of the trial is whether training neural networks on copyrighted books and news falls under the `fair use` doctrine. This is a macroeconomic Rubicon for the entire IT market. If the court rules that scraping raw data is copyright infringement, the unit economics of LLM development will be destroyed: startups and corporations will have to pay royalties for every terabyte of text fed to a model. Algorithm training will turn from an engineering task into an accounting nightmare, instantly freezing the emergence of new open-source models and leaving only monopolists capable of buying entire publishing houses on the market.
Source: The New York Times / OpenAI / Reuters
CopyrightOpenAILegalFair UseMacroeconomics