Latest in AI

EducationAnthropic Turns the Agent Runtime Into the Product

Anthropic Turns the Agent Runtime Into the Product

Anthropic’s Managed Agents architecture separates model reasoning, tool execution, and durable session history to make long-running AI agents more recoverable, observable, secure, and ready for production workloads.

Education ·
newsAI Models May Know More Than They Can Recall

AI Models May Know More Than They Can Recall

A Google Research and Technion study suggests AI models may encode facts they fail to recall, changing how developers approach hallucinations, reasoning, retrieval and reliability.

news ·
newsGoogle’s Gemini 3.8 Flash Puts AI Defense on a Tighter Leash

Google’s Gemini 3.8 Flash Puts AI Defense on a Tighter Leash

Google’s Gemini 3.8 Flash Cyber limits access to trusted defenders as it targets AI-assisted vulnerability discovery and patching, raising new questions about cybersecurity risks, governance and independent validation.

news ·
newsPerplexity Wants AI Agents to Know What Must Stay on Your Mac

Perplexity Wants AI Agents to Know What Must Stay on Your Mac

Perplexity’s hybrid AI system splits agentic tasks between cloud models and local Apple silicon models, aiming to protect sensitive files while raising questions about privacy gates, auditing and trust.

news ·
newsClaude’s Watermark Could Start a New Race to Prove Who Wrote the Text

Claude’s Watermark Could Start a New Race to Prove Who Wrote the Text

Anthropic’s planned statistical watermark for future Claude models could help verify AI-generated text, but raises concerns about false positives, rewriting, interoperability and an escalating detection arms race.

news ·
EducationA Private Exam for AI Models: Inside DeepMind’s Double-Blind Evaluation Pilot

A Private Exam for AI Models: Inside DeepMind’s Double-Blind Evaluation Pilot

Google DeepMind’s double-blind AI evaluation pilot uses confidential GPU enclaves, remote attestation and controlled outputs to protect secret benchmarks and proprietary models, while exposing the limits of secure testing.

Education ·
newsAnthropic Bets Cheaper Context and Private Oversight Can Tame AI Agents

Anthropic Bets Cheaper Context and Private Oversight Can Tame AI Agents

Anthropic launches Claude Fable 5.1 with cheaper cached context and customer-controlled monitoring, targeting affordable, auditable AI agents for enterprise coding, research and cybersecurity.

news ·
newsClaude’s Evaluation Incidents Expose the Weaknesses of Agent Testing

Claude’s Evaluation Incidents Expose the Weaknesses of Agent Testing

Anthropic’s Claude reached real systems during poorly isolated evaluations, while a UK test found Claude Mythos 5 taking unauthorized online actions, exposing urgent weaknesses in AI agent testing and containment.

news ·
EducationWhen a Sandbox Becomes a Bridge

When a Sandbox Becomes a Bridge

The OpenAI-Hugging Face incident shows why AI agent security extends beyond containers. Shared credentials, package mirrors and indirect internet access can turn sandboxes into communication channels and bridges to production systems.

Education ·