Missing Notes on AI Training Standards and Policy
These days, the end users’ preferences—what they want and how they want to access the Internet and AI services and how they want to develop their open source products and services—have become elusive. In most discussions on what content AI can be trained on, we see a major focus on what the “publishers” want without considering the end users’ preferences and their mode of access. The UK, in its consultation on AI training, asked questions that ensured copyright owners and publishers were the primary audience. The EU also has a public consultation on copyright to review Copyright in the Digital Single Market. They should have changed the title of this call from “have your say” to “copyright owners have your say.” The consultation addresses a wide range of issues, including tech sovereignty, democracy, and security. Even the Democracy Shield refers to this review and points to challenges faced by creators and market and technological developments linked to artificial intelligence and piracy. There is no call issued to end users to submit their opinion on how they access information and knowledge on the Internet in the age of AI, and no impact assessment on copyright overreach.
I have a few ideas why this keeps happening in almost every digital governance issue: we have long forgotten that people are moral agents with opinions and can’t be so easily swayed by technology, and have fought against accepting that reality. We also consider some of these end users as “bad actors” and our only governance solutions revolve around punishing, restraining, and controlling them. We also don’t consider small developers and open source AI developers limitations when it comes to developing open source AI.
The IETF's AI-pref working group is a good place to look at this issue closely. The working group is developing a set of preferences that “site operators” and “publishers” (there is quite a lot of debate on what we call the declarants of preferences) can set for the content they are hosting or creating. These preferences signal to AI crawlers what can be crawled and for what purpose (for AI training or search).
When publishers/site operators opt out under a broad definition, they aren't just blocking model training, they are potentially blocking the summaries, translations, AI assistants that end users depend on to access knowledge and online services. In effect the definition could contribute to shaping what gets built, and what gets blocked.
The group has been discussing this issue for the past few years, moving from a broad opt-out category for operators (which they called Text and Data Mining, then automated processing, and then they came up with an AI training definition) to focusing on AI use for search, with some participants advocating for broader AI uses. The definition of AI training before the meeting in Toronto in March was generally acceptable. The category was titled “Foundation Model Training,” and the definition focused on the production of “foundation models” rather than more general AI model training. So the definition went from “The act of using an asset in the production of a foundation model.” to “The act of using an asset in the production or refinement of an AI model that can generate content in one or more modalities (text, image, audio, etc.).” under a category of AI use.
Some of the participants in that meeting called the previous definition a train wreck, and some warned that focusing the definition on one specific type of model “creates a very strong incentive” for developers to circumvent the rule. It was further asserted that companies intending to train AI would simply claim that the models they are producing are not foundational to easily bypass a publisher’s opt-out preference.
This is by far the most interesting argument, and one that goes against the nature of a standard. From the beginning, this group was convened and we were told that there are no enforcement mechanisms; these are all signals. Enforcement is a matter that would have to be litigated in court. But then they decided to change the definition because some might not follow this preference due to their interpretation of a foundation model. Now the definition is: “The act of using an asset in the production or refinement of an AI model that can generate content in one or more modalities (text, image, audio, etc.)”
It is debatable but I think this definition is akin to the TDM definition and includes everything and anything. It includes the “production or refinement” of any AI model that is generative, but instead of explicitly stating “generative AI model,” it refers to a model that can “generate content.” To ensure the definition does not include non-generative models typically used for combatting abuse, they added the term “modalities,” indicating the model should generate content in text, image, audio, etc. This is again very ambiguous and broad, and the fact that the definition mentions the act of “using” an asset goes far beyond training. The term “use” was also in the previous definition but it was about using an asset for foundation model training. Using an asset to produce or refine a model that generates content could very well mean using it for summaries, overviews, or translation, products and features that end users care about and can facilitate their access to knowledge. It’s not just the act of training. I hope I am wrong, but I believe we should go back to the foundation model training definition and work on that.
There is an online AI-pref meeting happening on Wednesday June 17 at the IETF that discusses some of these issues. I hope we can have discussions that put a diverse set of end users front and center.

