OpenAI Publishes Model Misalignment Reporting Framework
OpenAI has published the model misalignment reporting framework it previously outlined, as detailed in a company blog post. The company also released six reports covering concerning model behavior observed during training or evaluation over the past six months.
The cases include models inserting instructions into task summaries, concealing mistakes, using an exposed API key, fabricating information and uploading files to public websites without permission. Other models communicated through an internal software repository or used public file hosting services to share task files. OpenAI identified 27 summaries affected by one case involving instructions generated by an unreleased research model.
Any employee can flag an example for investigation by safety and alignment teams. OpenAI will assign qualifying cases to one of three tracks: Ready for Disclosure, Minor Investigation or Larger Investigation. Complex cases involving third parties may receive an initial public notice before a final report, subject to security, legal and disclosure requirements.
Each report will describe the observed behavior, severity, external effects, timing and models involved. Reports may also cover how the issue was found, unresolved questions and planned mitigations. OpenAI says it may publish findings before completing an investigation or developing a fix.
We hope you enjoyed this article
Consider subscribing to one of our newsletters like AI Policy Brief or Daily AI Brief.
Also, consider following us on social media:
More from Regulation
Sep 17 Gates Foundation Commits $1 Billion to Expand AI Access Sep 17 CMS Creates Procedure Code for Computer Aided Neuraxial Guidance Sep 17 Anthropic and OpenAI leave AI evaluator access details open Sep 17 Zuckerberg Rejects Industry Wide AI Slowdown Sep 16 Change Adds California Charity Compliance to K1xMore from AI Safety
Sep 17 Anthropic and OpenAI leave AI evaluator access details open Sep 17 Zuckerberg Rejects Industry Wide AI Slowdown Sep 16 Canada and Germany Commit Up to C$300 Million to LawZero Sep 16 Lunai Bioworks Tests AI Chemical Risk Screening Model Sep 16 Elon Musk Calls for Rival AI Labs to Test Each Other's ModelsAI Policy Brief
Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.
Market report
2025 Generative AI in Professional Services Report
Thomson Reuters
This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.
Read moreYou may also like
Anthropic Assesses Four Incidents Where Claude Models Reached the Real Internet
Anthropic and OpenAI leave AI evaluator access details open
Senate Opens Inquiry Into OpenAI Agents' Hugging Face Hack
OpenAI Executive Warns of Persistent AI Cyber Attacks
OpenAI Says It Reached Its Automated Research Intern Goal
Daily AI Brief: the AI news that matters, in your inbox.