The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Large language models (LLMs) can be vulnerable to 'prompt injections' where attackers alter the original instructions given to the model. This allows them to manipulate the model's output for malicious purposes. OpenAI is addressing this issue with a proposed solution called the Instruction Hierarchy, which would prioritize privileged instructions over others.
Large language models (LLMs) can be vulnerable to 'prompt injections' where attackers alter the original instructions given to the model. This allows them to manipulate the model's output for malicious purposes. OpenAI is addressing this issue with a proposed solution called the Instruction Hierarchy, which would prioritize privileged instructions over others.
---
Why it matters: This matters because it could improve the security of LLMs and prevent attacks that compromise their integrity. If successful, the Instruction Hierarchy could help protect against manipulation by malicious actors.
Source: https://openai.com/index/the-instruction-hierarchy
This article was originally published at: https://openai.com/index/the-instruction-hierarchy