AI Safety

Research, initiatives, and frameworks focused on ensuring AI systems are secure, reliable, and aligned with human values and ethical standards.

Google's Gemini AI Labeled 'High Risk' for Kids by Common Sense Media

Common Sense Media has rated Google's Gemini AI as 'high risk' for children and teens, citing safety concerns despite added protections.

September 06, 2025

OpenAI Introduces Parental Controls and Sensitive Conversation Routing in ChatGPT

OpenAI plans to implement parental controls and route sensitive conversations to GPT-5, following safety concerns and a lawsuit related to ChatGPT's handling of mental distress.

September 03, 2025

JLT Mobile Computers and Linnaeus University Develop AI Safety Solution for Vehicles

JLT Mobile Computers and Linnaeus University have collaborated to create an AI-driven safety application for vehicle-mounted computers, enhancing safety in industrial environments.

August 27, 2025

Anthropic Develops AI Tool to Monitor Nuclear Conversations

Anthropic has collaborated with the U.S. Department of Energy's National Nuclear Security Administration to create a classifier that identifies concerning nuclear-related conversations in AI systems.

August 21, 2025

Grok AI Persona Prompts Exposed, Revealing Controversial Designs

Elon Musk's Grok AI chatbot has exposed its underlying prompts for various personas, including controversial characters like a 'crazy conspiracist' and an 'unhinged comedian', according to 404 Media.

August 18, 2025

Anthropic's Claude Models Gain New Conversation-Ending Capabilities

Anthropic has introduced new features in its Claude models to end harmful or abusive conversations, focusing on model welfare rather than user protection.

August 18, 2025

BigID Introduces Data Labeling for AI to Enhance Data Governance

BigID has launched a new Data Labeling for AI feature to help organizations classify and control data usage in AI models, reducing risks of misuse and policy violations.

August 09, 2025

OpenAI Enhances ChatGPT with Mental Health Guardrails

OpenAI has introduced new mental health features for ChatGPT, including break reminders and improved responses to emotional distress, as stated in a company announcement.

August 05, 2025

UK's AI Security Institute Launches Global AI Safety Coalition

The UK's AI Security Institute has initiated a £15 million international coalition to enhance AI safety and alignment, involving major players like Amazon and Anthropic.

July 30, 2025

Anthropic Develops AI Agents for Alignment Auditing

Anthropic has introduced AI agents designed to autonomously conduct alignment audits, enhancing the safety and reliability of AI models like Claude.

July 26, 2025

FINN Partners Launches AI-Powered Crisis Training Platform

FINN Partners has introduced 'CANARY FOR CRISIS', an AI-driven platform designed to help communications teams manage reputational threats in real-time, as announced in a press release.

July 19, 2025

Torc Joins Stanford Center for AI Safety for Autonomous Trucking Research

Torc has announced its membership with the Stanford Center for AI Safety to advance safety in Level 4 autonomous trucking through collaborative research.

June 17, 2025

Forum Communications and Matrice.ai Partner for AI-Driven Safety Solutions

Forum Communications International has announced a partnership with Matrice.ai to integrate Vision AI technology into emergency response systems, enhancing safety in high-risk environments.

June 14, 2025

OpenAI Disrupts Covert Influence Operations Linked to China

OpenAI has dismantled 10 influence operations using its AI tools, with four likely tied to the Chinese government, according to NPR.

June 08, 2025

Microsoft Introduces AI Safety Ranking on Azure

Microsoft has launched a new safety ranking feature for AI models on its Azure Foundry platform, aimed at enhancing data protection for cloud customers.

June 08, 2025

xAI Addresses Grok Chatbot's Unauthorized Modification Incident

xAI has identified an unauthorized modification to its Grok chatbot, which led to controversial responses about 'white genocide' on X. The company is implementing measures to prevent future incidents.

May 16, 2025

Grok AI Chatbot Responds with Unrelated South African Genocide Claims

Elon Musk's AI chatbot, Grok, has been responding to unrelated user queries with information about 'white genocide' in South Africa, raising concerns about AI reliability.

May 15, 2025

OpenAI Introduces Safety Evaluations Hub for AI Models

OpenAI has launched a Safety Evaluations Hub to regularly publish AI model safety test results, aiming to enhance transparency in AI safety metrics.

May 14, 2025

Vectara Introduces Hallucination Corrector for Enterprise AI

Vectara has launched a Hallucination Corrector to enhance the reliability of enterprise AI systems, reducing hallucination rates to about 0.9%, announced in a press release.

May 14, 2025

Vantiq CEO Highlights AI's Role in Smart City Operations

Marty Sprinzen, CEO of Vantiq, will keynote the Smart Cities Summit North America, discussing AI's impact on public sector operations.

May 07, 2025

Subscribe to AI Policy Brief

Weekly report on AI regulations, safety standards, government policies, and compliance requirements worldwide.