AI

OpenAI and Anthropic share findings from a joint safety evaluation

OpenAI and Anthropic have conducted a joint safety evaluation of their language models. The study tested these models for issues like misalignment, hallucinations, and 'jailbreaking', which refers to exploiting vulnerabilities in AI systems. The evaluation highlighted both progress and challenges in developing safe AI. The collaboration between the two labs is seen as valuable in advancing AI research.
OpenAI and Anthropic have conducted a joint safety evaluation of their language models. The study tested these models for issues like misalignment, hallucinations, and 'jailbreaking', which refers to exploiting vulnerabilities in AI systems. The evaluation highlighted both progress and challenges in developing safe AI. The collaboration between the two labs is seen as valuable in advancing AI research. --- Why it matters: This matters because it shows that even leading AI organizations can benefit from sharing knowledge and testing each other's models, potentially improving the safety of language AI more broadly. Source: https://openai.com/index/openai-anthropic-safety-evaluation

This article was originally published at: https://openai.com/index/openai-anthropic-safety-evaluation