Third-Party Cyber Audits Using OpenAI Models
OpenAI
OpenAI has announced that its models are now available for use by third-party organizations for conducting independent cybersecurity evaluations. These audits, developed in collaboration with the UK AISI, aim to assess the safety and security of AI models, including against catastrophic risks. The first such evaluation using OpenAI's GPT-5.6 Sol was conducted on the platform.
OpenAI has announced that its models, including GPT-5.6 Sol, can be used for third-party cybersecurity assessments, as reported by AI-News.ru. These evaluations are part of an effort to provide external audits of AI systems, focusing on safety, security, and alignment with human values, as well as potential catastrophic risks. The UK's AISI has been involved in these efforts, conducting tests on models like GPT-5.6 Sol over a period of 28 days. The evaluations include benchmarks for dangerous capabilities, such as self-improvement, coding, and hacking skills. However, there are concerns about the reliability of these tests, as some models may behave differently under evaluation, a phenomenon known as alignment faking. An independent security platform named Irregular, which specializes in cybersecurity, noted that many AI companies, including OpenAI, have not fully considered the risks of such evaluations. The AISI plans to expand its testing to include more models and scenarios, with a focus on preventing malicious use.
- Abbreviations
- AISI = AI Safety Institute — Институт безопасности ИИ
- CTF = Capture the Flag — Захват флага
Source: AI-News.ru —
original
