The latest in news from spAIsee.
Hirundo says it made Alibaba’s Qwen more politically neutral without reducing capability, but independent testing must determine whether the change holds across unfamiliar prompts and real-world evaluations.
Reflection AI’s Beam targets the open-model efficiency race, promising frontier-style reasoning and coding performance with far lower inference costs, though its benchmark and compute claims remain unverified.
Anthropic’s Claude Haiku 5.5 enters the low-cost AI race with a claimed 90% price reduction, challenging OpenAI’s GPT-6 Luna for high-volume enterprise workloads and everyday automation.
Microsoft is turning Windows into an AI workstation, combining local models, cloud services and controlled agents while launching powerful Surface hardware for developers and addressing the security risks of autonomous software.
Tony Fadell explains why Rabbit R1, Humane AI Pin and Limitless Pendant struggled, arguing that AI gadgets must solve clear problems and earn trust before requesting access to users’ private lives.
OpenAI’s GPT-6 rollout pairs a new model with Intelligent UI, testing ChatGPT’s product execution, subscription strategy and ability to turn advanced AI into everyday workflows.
Google has launched SynthID Detector, a public tool that checks images, video and audio for invisible AI watermarks while warning users that negative results do not prove authenticity.
Tab is emerging from stealth with a Messages-based AI assistant that can browse, call businesses, use connected accounts, create files and pause for approval before purchases.
Books By People is launching an “organic literature” certification mark to signal human authorship as AI-generated prose spreads, raising questions about proof, software and trust.
TwelveLabs’ Pegasus 1.6 analyzes first-person work videos to extract actions, objects, hand movements and task boundaries, targeting robotics’ costly shortage of structured training data for physical AI.
Anthropic is bringing Claude into Google Docs, Sheets and Slides, challenging Gemini’s built-in Workspace advantage and turning AI competition into a fight over everyday office workflows, permissions and user habits.
AI safety audits can fail before agents do when they search incomplete evidence, trust model summaries, or miss infrastructure signals. This analysis examines Anthropic’s and OpenAI’s incidents and explains how to build harder-to-fool evaluation pipelines.