Topic: safety
8 stories found
Thursday, September 3, 2026

Safety overview: GPT-6 Astra
GPT-6 Astra has achieved the Critical level of cybersecurity, making it the company's most secure widely used model. This milestone underscores the firm's commitment to enhancing safety in their advanced language models.
Wednesday, September 2, 2026
Monday, August 31, 2026
Friday, August 28, 2026
DeflectBench: A Benchmark for Evaluating Rhetorical Fallacy Generation in LLMs
A new benchmark called DeflectBench evaluates whether large language models can generate rhetorical fallacies when prompted, addressing the underexplored area of inducing such errors rather than just detecting them. This matters because it helps understand and potentially mitigate safety issues related to biased or misleading outputs from AI systems.
Wednesday, August 26, 2026
Automata from Agent Traces: Failure and Next-Step Prediction
A new approach called "Automata from Agent Traces" aims to make LLM-based agents more transparent by identifying patterns in their behavior, which is crucial for safety auditing and runtime monitoring but currently hindered by long, unstructured task traces that existing methods struggle with.
Monday, August 24, 2026
When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha
A new study highlights the growing use of conversational AI systems as informal mental health support for Generation Alpha, with 13.1% of U.S. adolescents relying on such tools, raising concerns about their safety and clinical reasoning capabilities.
Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification
The research highlights that current safety measures in Large Language Models are insufficient as they primarily rely on surface-level mechanisms that activate too late to prevent the models from retaining harmful knowledge. This matters because it underscores the need for more robust and proactive methods to align LLMs with ethical standards throughout their operation.
πΏ That's all for now. Come back tomorrow.
8 of 8 items shown. Sources: 107 days indexed.