Earlier this year, Google's AI chief, Jeff Dean, discussed a concept on a podcast that was largely unknown outside the tech world at the time: distillation.
While discussing the development of Google's AI models, Dean noted that he and his colleagues discovered AI distillation as a way to improve system performance without relying on a single large image recognition model.
"Through distillation—a key technique for making smaller models more powerful—you first need a frontier model, and then you can distill its capabilities into a smaller model," Dean said in February.
Five months later, distillation has suddenly become a hot topic from Silicon Valley to Washington, with tech figures and lawmakers debating whether this practice is becoming a national security threat.
At a high level, distillation refers to using the answers from a chatbot or the output of a sophisticated AI model to train another model. This practice is controversial because, depending on how it is used, it can allow model developers to create competitive products simply by using the outputs of companies that have invested millions or billions of dollars in developing state-of-the-art training techniques.
"It's almost like someone went to the lecture, read the textbook, and did all the hard homework," said Pukar Hamal, founder of the AI security company SecurityPal. "Then another student says, 'Hey, I didn't do that. Can I just copy your homework?'"
Whether due to a post from Krazios or other reasons, the world's largest tech giants united in an unprecedented way on Friday to make a clear statement. Following a week of social media posts, tech giants Nvidia, Microsoft, Meta, Palantir, and over 20 other companies jointly released an open letter, urging policymakers to avoid "premature restrictions" on open-weight AI models, which they argued could "stifle competition or push innovation overseas."
"Distillation, the practice of using one model's output to help train or improve another model, is a widely used technique for model improvement, evolution, and validation," they wrote.
The emergence of distillation has created a dilemma for US policymakers.
Aaron Levie, CEO of Box, was one of the signatories of Friday's open letter. In an interview, Levie stated that to remain competitive, US companies need access to the best technology, regardless of where it is developed.
"The overall trend will be that the more innovation there is—whether from the US or elsewhere—the more AI progress you should expect, and it will generally move towards lower costs and higher efficiency over time," Levie said.
Shashi Bellamkonda, Research Director at Info-Tech Research Group, noted that while much of the current discussion focuses on open-weight AI models, many companies also incorporate distillation techniques when creating their own models. For example, Nvidia used distillation in training its Llama Nemotron series of models, as detailed in its research papers.
"It is a legitimate and highly valuable technique for training smaller, cheaper models based on the output of larger models, and it has been widely used," Bellamkonda said.
However, Anthropic holds a different view, as it observes how its models are used and has a thriving business to protect.
Anthropic, valued at nearly $1 trillion and with plans to go public in the near future, argues that preventing illegal distillation is a matter of national security.
"Anthropic and other US companies have built systems to prevent state and non-state actors from using AI for purposes such as developing biological weapons or conducting malicious cyber activities," the company stated in a February post. Preventing this requires "rapid, coordinated action among industry participants, policymakers, and the global AI community."
OpenAI and Anthropic prohibit distillation in their terms of service. Bellamkonda said they are essentially implying that unauthorized use of their large models represents potential intellectual property theft.
But as AI costs skyrocket, companies will go to great lengths to improve efficiency.
Hamal said that at SecurityPal, he has no objection to using open-weight models, which the company uses to automate security assessments. He said this can save them a significant amount of money.
"We make sure there are no malicious backdoors in the code," Hamal said. "But after we complete our assessment, hosting it on our own infrastructure—why not?"
A major problem for Anthropic and OpenAI in trying to argue the issue of intellectual property theft is that both companies have relied on content from other sources to build their own models and have been sued for it.
Max Pritt, a lawyer at Boies Schiller Flexner LLP who represents book authors in copyright lawsuits against AI companies, said the government is in a similar position.
"At least publicly, this administration has focused its efforts on protecting the intellectual property of tech companies, while largely remaining silent on the unauthorized use of the intellectual property of creators and individuals," Pritt said.
Comments