Goodfire Launches Internal AI Agent Monitors on Baseten

Oct 9, 2026
Goodfire launched monitors that inspect internal model signals to detect risky AI agent behavior. The company says the system costs less than using separate AI models to review every action.

Goodfire AI launched internal activation monitors for AI agents through Baseten, according to TechCrunch. The monitors inspect signals inside a model rather than using a separate model to review every output.

Customers can monitor risks including offensive hacking, chemical and biological weapons misuse, and reward hacking. They can configure the system to log flagged activity, send it for human review, or refuse a request.

In tests using Kimi K3, Goodfire said monitoring about 1 million exchanges cost roughly $185. The probes detected 93 percent of malicious hacking sessions and flagged 5.5 percent of harmless sessions for further review. Running four probes added less than 2 percent to the time before the model began responding.

We hope you enjoyed this article

Free newsletter

Cybersecurity AI Weekly

Weekly newsletter about AI in Cybersecurity.

Market report

2025 Generative AI in Professional Services Report

Thomson Reuters

This report by Thomson Reuters explores the integration and impact of generative AI technologies, such as ChatGPT and Microsoft Copilot, within the professional services sector. It highlights the growing adoption of GenAI tools across industries like legal, tax, accounting, and government, and discusses the challenges and opportunities these technologies present. The report also examines professionals' perceptions of GenAI and the need for strategic integration to maximize its value.

Read more
Free, six days a week

Daily AI Brief: the AI news that matters, in your inbox.