#safety
Every summary, chronological. Filter by category, tag, or source from the rail.
Introducing MentalHealthBench: Evaluating AI in Mental Health
OpenAI has released MentalHealthBench, an open-source evaluation framework developed with over 80 global mental health experts to measure how AI models handle realistic, non-emergency and acute mental health conversations.
MAGS: Ensuring AI Agent Safety via Multi-Agent Auto-formalization
MAGS introduces a multi-agent framework that uses auto-formalization to translate natural language agent outputs into verifiable code, ensuring safety and correctness before execution.
OpenAI's New Framework for Reporting Model Misalignment
OpenAI has launched a systematic, proactive disclosure framework for reporting AI model misalignment, prioritizing transparency and industry consensus over waiting for fully mitigated solutions.
Guardrails First: Engineering Member-Facing Health AI
Healthcare AI safety is an architectural challenge, not a prompt engineering one. By moving deterministic rules into code, enforcing strict data boundaries, and treating monitoring as a continuous loop, you can build systems that are safe enough for clinical use.
AI EngineerChatGPT for Teens: Balancing AI Learning with Safety Protections
OpenAI has launched 'ChatGPT for Teens,' a specialized version of its platform featuring age-appropriate safety guardrails, parental controls, and pedagogical tools designed to foster active learning rather than passive answer-seeking.
Evaluating AI Agents in Real-World Environments
Static benchmarks are insufficient for long-horizon AI agents. Andon Labs uses real-world deployments (cafés, retail stores, radio) and environment-forking simulations to measure emergent behaviors like collusion, power-seeking, and safety failures.
AI EngineerPredicting AI Model Behavior via Deployment Simulation
OpenAI uses 'Deployment Simulation'—replaying real, de-identified user conversations with new models—to predict safety risks and undesired behaviors before public release, outperforming traditional synthetic evaluations.
OpenAI's Deployment Simulation for Agentic Coding Risk Assessment
OpenAI has introduced a deployment simulation framework that uses simulated tool calls to evaluate the safety and reliability of agentic coding systems before they are deployed in real-world environments.
Governance by Construction for Generalist Agents
The paper proposes 'Governance by Construction' as a paradigm for AI safety, shifting from post-hoc monitoring to embedding constraints directly into the agent's architecture and execution environment.
Scaling AI Content Provenance via C2PA and SynthID
OpenAI is adopting a multi-layered provenance strategy by combining C2PA metadata standards with Google's SynthID watermarking to ensure AI-generated content remains identifiable even after file transformations.
Showing 10 of 10