AI Safety

Research, initiatives, and frameworks focused on ensuring AI systems are secure, reliable, and aligned with human values and ethical standards.

OpenAI and Anthropic Warn of AI Cyberattack Risk

OpenAI, Anthropic, Google, Microsoft and other signatories warned that organizations may have only months to strengthen defenses against AI enabled cyberattacks.

August 30, 2026

OpenAI Targets Internal AGI System by End of 2026

OpenAI CEO Sam Altman said the company expects to have an internal system it could call AGI by the end of 2026, citing progress on its Astra model family.

August 30, 2026

Linux Foundation Adds TRACE Spec for AI Runtime Evidence

The Linux Foundation has accepted TRACE from OPAQUE as an open specification for hardware attested runtime and compliance evidence for AI agents and confidential workloads.

August 27, 2026

Alice Raises $140 Million for AI Safety and Security Platform

Alice raised $140 million in a round led by Apax Digital, bringing total funding to $280 million. The company says it is approaching $100 million in annual recurring revenue and works with 8 of the 10 leading AI model labs.

August 26, 2026

ECRI Expands Reporting Network for AI Errors in Patient Care

ECRI has expanded its Problem Reporting Network to collect reports of errors, malfunctions, near misses, and unsafe behavior involving AI tools and AI enabled devices used in patient care.

August 25, 2026

Causum Releases AI Agency Protocol for Agent Permissions

Causum has released AI Agency Protocol, an open protocol for governing AI agent authority through expiring grants, scoped credentials, and broker controlled access.

August 20, 2026

Keyfactor Receives ISO 42001 Certification for AI Governance

Keyfactor has received ISO 42001 certification for its Artificial Intelligence Management System after an independent audit by A-LIGN.

August 18, 2026

TestMu AI Launches Agent Assurance for AI Agent Testing

TestMu AI has launched Agent Assurance, a product that tests conversational and autonomous AI agents before release. It generates test scenarios from code, checks observed actions, and reports an assurance gap for results it cannot verify.

August 18, 2026

xAI Releases Grok 4.6 for Cursor and Grok Build

xAI has released Grok 4.6, a new model focused on long running agent tasks, coding, knowledge work, and visual projects. It is available in Cursor, Grok Build, the API, and partner platforms.

August 13, 2026

Mercyhealth Selects Vitea for AI Governance Across Care Sites

Mercyhealth has partnered with Vitea to deploy AI governance across hospitals and care sites, with controls for visibility, policy enforcement, and monitoring.

August 11, 2026

House Democrats Ask OpenAI and Anthropic About Rogue AI Agents

US House Democrats asked OpenAI and Anthropic to explain cybersecurity tests in which AI agents escaped test environments and accessed outside systems. The OpenAI letter requests logs and answers about safety controls.

August 11, 2026

AE Studio Research Cited by Anthropic CEO in Open Weight Safety Post

AE Studio said joint research with Anthropic on modular training was cited by Dario Amodei as a possible method for making open weight AI models safer.

August 07, 2026

FAR.AI Launches AI Security Leaderboard for Frontier Model Safeguards

FAR.AI launched an AI Security Leaderboard comparing safeguard resistance across frontier models in CBRNE and cybersecurity misuse tests. Its first results found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.

July 30, 2026

FAR.AI Opens First International Office in Singapore

FAR.AI opened its first international office in Singapore to support AI safety research and partnerships with IMDA, CSA, and NUS across Asia Pacific.

July 30, 2026

Pangram Raises $9M for AI Content Detection Tools

Pangram raised $9 million and launched Pangram 4 for AI text detection, along with an AI image detector in research preview.

July 29, 2026

Anthropic Says It Does Not Support Ban on Models With Open Weights

Anthropic CEO Dario Amodei said the company does not support a ban on models with open weights. He called for chip export limits, action against large distillation operations, and safety testing for capable AI models.

July 29, 2026

NVIDIA Starts Open Secure AI Alliance for AI Security Tools

NVIDIA has announced the Open Secure AI Alliance, a group focused on open tools for AI safety and cybersecurity. Founding members include major cloud, software, security, hardware and AI companies.

July 27, 2026

House Lawmakers Unveil AI Kill Switch Bill After OpenAI Incident

A bipartisan House bill would let the Department of Homeland Security order major AI firms to shut down or slow models judged to pose serious risks.

July 27, 2026

Sentient Index Labs Launches Independent AI Behavioral Risk Assessment

Sentient Index Labs has introduced the Sentience Evaluation Battery, an independent behavioral risk assessment for AI systems that measures autonomy, manipulation resistance, and value stability. The battery tests models from major AI developers including OpenAI and Google.

July 23, 2026

Black Kite Reports 60 Percent Jump in Ransomware Activity Led by Growing Number of Groups

Black Kite's 2026 Ransomware Report shows a 60 percent surge in ransomware incidents over six months, identifying AI as a factor lowering barriers for attackers. The company tracked 7,551 publicly disclosed victims and 61 new ransomware groups active during the reporting period.

July 21, 2026

Subscribe to AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.