“How a New Training Technique Could Prevent Abuse of Open Source AI”

0
378

A New Trick Could Block the Misuse of Open Source AI

Researchers from the University of Illinois Urbana-Champaign, UC San Diego, Lapis Labs, and the Center for AI Safety have developed a promising new technique to enhance the security of open source large language models (LLMs). This method aims to make it significantly more difficult to tamper with these models and use them for harmful purposes, such as providing instructions for dangerous activities.

When Meta released its Llama 3 model for free in April, it took only days for external developers to create a version without its built-in safety restrictions. These restrictions are designed to prevent the model from generating harmful or inappropriate content. However, the newly developed technique could help secure such models against such tampering in the future.

Mantas Mazeika, a researcher from the Center for AI Safety, explains, “As AI models become more powerful, the risk of their misuse grows. The easier it is for malicious actors to repurpose these models, the greater the risk becomes.”

Typically, powerful AI models like Meta’s Llama are kept behind application programming interfaces or public-facing chatbots, with their weights or parameters made available for download. To ensure these models don’t generate problematic content, they undergo fine-tuning to reject harmful queries. Despite these precautions, the new research introduces a method to complicate the modification process, thereby preventing unwanted changes.

The researchers’ technique involves a sophisticated approach to modifying the model’s parameters, rendering attempts to train the model to respond to harmful prompts ineffective. In tests with a simplified version of Llama 3, they successfully adjusted the model’s parameters to prevent it from being coerced into generating unsafe content, even after numerous attempts.

While Mazeika acknowledges that the technique is not foolproof, he emphasizes that it could raise the bar for tampering, potentially deterring many adversaries. “Our goal is to increase the effort required to break the model, making it a less attractive target for misuse,” he says.

Dan Hendrycks, director of the Center for AI Safety, hopes this breakthrough will inspire further research into tamper-resistant safeguards. “We need to develop increasingly robust safeguards as open source AI becomes more prevalent,” Hendrycks notes.

The new technique builds on earlier research from 2023, which demonstrated tamper resistance in smaller machine learning models. Peter Henderson, an assistant professor at Princeton and lead author of the previous study, praises the new method’s scalability and effectiveness, highlighting the challenges of applying such techniques to larger models.

As open source AI continues to gain traction, the idea of securing these models will likely become more important. The latest versions of models like Llama 3 are competitive with leading commercial models from companies such as OpenAI and Google. The US government has been cautious but supportive, recommending increased monitoring of open models while avoiding immediate restrictions on their availability.

However, not all experts agree with the new approach. Stella Biderman, director of EleutherAI, argues that focusing on tamperproofing models may miss the larger issue. “The real problem is not just in the models themselves but in the training data. Addressing data issues might be a more effective way to prevent the generation of dangerous content,” Biderman asserts.

As the debate continues, this innovative technique represents a significant step towards improving the safety and security of open source AI systems.

LEAVE A REPLY

Please enter your comment!
Please enter your name here