AI

From hard refusals to safe-completions: toward output-centric safety training

OpenAI has developed a new approach to AI safety called 'safe-completions' for its GPT-5 model. This approach aims to improve both the safety and helpfulness of AI responses by moving away from hard refusals, where models simply refuse to answer certain questions. Instead, safe-completions allows the model to provide nuanced and context-dependent answers that can handle dual-use prompts without compromising safety. The goal is to strike a balance between providing accurate in
OpenAI has developed a new approach to AI safety called 'safe-completions' for its GPT-5 model. This approach aims to improve both the safety and helpfulness of AI responses by moving away from hard refusals, where models simply refuse to answer certain questions. Instead, safe-completions allows the model to provide nuanced and context-dependent answers that can handle dual-use prompts without compromising safety. The goal is to strike a balance between providing accurate information and avoiding potential harm. --- Why it matters: This matters because it addresses a critical challenge in AI development: ensuring models can respond safely and responsibly, particularly when faced with ambiguous or sensitive topics. By developing output-centric safety training, researchers can improve the reliability of AI systems and reduce the risk of unintended consequences. Source: https://openai.com/index/gpt-5-safe-completions

This article was originally published at: https://openai.com/index/gpt-5-safe-completions