How to Detect and Clean up Data Contamination in LLMs
Despite the power of large language models (LLMs), they aren’t foolproof. One significant area of concern is data contamination, which can negatively impact the performance of LLMs, leading to inaccurate or unreliable predictions, skewed data or biased results. All these can have serious, real-life implications, especially as LLMs are now being used in an ever-expanding number of areas.
What Is Data Contamination?
Data contamination (or data leakage) in machine learning models occurs when the data used to train the model has been “contaminated” with overlapping data that is later used downstream to test its performance.
This issue commonly arises because, in a standard machine learning setup, machine learning engineers will split their dataset into three sets — one for training, another for testing and the third for validation. However, these separate datasets often need to be “cleaned” before they are used. If there is any overlap between the datasets for training, testing and evaluation, then the model’s performance will be artificially inflated — much like a student who does deceptively well on a test, but only because they were given the answers beforehand.
While cleaning datasets for classic ML models is relatively straightforward, it’s trickier with LLMs because of the size of their datasets, as well as the complexity of the models themselves. Also, the growing lack of transparency around proprietary LLMs like GPT-4 makes it harder to uncover data contamination at its source.
Tools for Identifying Data Contamination
To detect and mitigate data contamination, experts recommend a variety of approaches to tackle the problem, including not using data from the Internet if possible, carefully curating datasets to eliminate the possibility of overlaps and utilizing one of the growing number of tools that are now available.
One of these tools includes detect-pretrain-code-contamination, a publicly available script developed by researchers at University of Washington and Princeton University that helps to detect pre-training code contamination in datasets.
Providing “guided instructions” to black-box LLMs is yet another method to identify data contamination, developed by University of Arizona computer science Ph.D. candidate Shahriar Golchin and associate professor and lum.ai co-founder Mihai Surdeanu. The idea is to detect contamination by “coax[ing] LLMs into outputting their memorized data.” This is done by feeding guided instructions into the model that include the dataset name, partition type, and a random-length initial segment of a reference instance, and then requesting a completion from the LLM. If the LLM’s output is similar to or matches the other segment of the reference, then that instance is considered contaminated.
Another tool being developed by Golchin and Surdeanu is the Data Contamination Quiz (DCQ), which is a simple and effective tool to identify and estimate the amount of contamination that might be present in a dataset. For instance, in evaluating GPT-4 with the DCQ, they found that the actual data contamination level to be at 56 percent, in sharp contrast with OpenAI’s reported contamination rate of only 25 percent.
Beyond these tools, Golchin and Surdeanu offered more advice on how to mitigate data contamination.
“The first potential method that AI experts can employ is to promptly stop using datasets identified as contaminated, whether flagged by our methods or others,” said Golchin and Surdeanu in an email interview. “The second recommendation is to generate private datasets and benchmarks that are not publicly accessible on the internet and utilize these datasets to accurately assess LLM performance. However, questions such as who will be responsible for managing this process, how data should be collected for these datasets and how others can access them are subjects that require further discussion.”
Model Merging
Yet another innovative technique to combat data contamination is model merging, which involves using various methods to combine multiple pre-trained models into one cohesive model, without the need for intensive training or computational resources.
One nifty tool is mergekit, which offers LLM developers a “novel and experimental method for creating sophisticated models at a fraction of the cost, without the need for heavy training and GPU resources.”
Machine Unlearning
Experts are now also looking into machine unlearning algorithms that help machines “forget” the data they’ve learned as yet another method of tackling data contamination, as well as addressing concerns around privacy, security and ethical use of AI.
“We are now working on developing methods to rectify this issue using machine unlearning techniques,” explained Golchin and Surdeanu. “Our focus is on creating ways for LLMs to forget or unlearn specific information that was learned during their pre-training. However, addressing this problem is challenging due to the interconnected nature of LLMs, as altering one aspect can disrupt others. Further, many LLMs facing data contamination issues are closed-source, adding further complexity to the task.”
Transparency Is Key
These challenges highlight the need for more transparency and independent verification when it comes to the LLMs being developed by big tech companies.
In these situations, data contamination analysis is often performed in-house and usually without any external way to cross-validate the results. Without knowing how these powerful models work under the hood, it’s impossible to know if the impressive performance of these powerful LLMs was artificially skewed by contaminated data.
Notably, the scale of the problem becomes even more apparent with recent research suggesting a 1 to 45 percent level of contamination of the underlying datasets behind some of the best-known LLMs — making data contamination a hidden but widespread issue that many who are working in the AI field will ultimately need to grapple with.