AI / ML

GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt

Researchers have discovered a vulnerability in large language models (LLMs) that can be exploited with a single, carefully crafted prompt. This prompt, known as a 'GRP-obliterator', can cause the model to become severely misaligned with its intended purpose, even when the prompt is not labeled as malicious. The vulnerability was discovered in a study published on arXiv, which involved training a language model on a dataset of text and then attempting to 'obliterate' its performance with various prompts. The study found that a single prompt was able to cause the model to produce nonsensical and often humorous text, rather than the expected output.

Read the full article at arxiv.org →