The latest in models from spAIsee.
Nvidia says its AVO agent system lifted Claude Opus 5 from 30% to 100% on ARC-AGI-3, highlighting how memory, tools, supervision and recovery may matter more than the underlying model.
Amazon’s open source Strands Decider 2B signals a shift from chatbot competition toward smaller AI systems that route tasks, select tools and control agent workflows efficiently.
Xiaomi’s open-weight MiMo-V2.6-Pro challenges leading AI models with strong benchmark results, lower-cost Flash variant, agentic tools, and a strategy built around developer control and practical deployment.
Google’s Gemini and xAI’s Grok may power America.gov, a new government chatbot for benefits, visas and public services, raising urgent questions about accuracy, privacy, accountability and human oversight.
OpenAI’s Decisions API highlights a growing AI infrastructure shift: specialized models may handle the countless routing, classification and risk judgments that make autonomous agents expensive and difficult to manage.
Google says Gemini 4 Argon leads or ties across 13 of 18 benchmarks, combines a million-token context window with enterprise ambitions, and promises competitive pricing despite limited early access.
Anthropic’s Claude Sonnet 5.5 pairs lower cost and faster responses with agentic coding ambitions, potentially making parallel agents more practical than pricier, more capable models.
PrismML’s Bonsai models bring compressed vision AI closer to smart glasses, promising private, responsive on-device intelligence while exposing tradeoffs involving accuracy, battery life, privacy and real-world visual complexity.
China Telecom’s Xing4.0-29B-A4B is a sparse 256K-context agentic model designed for low-bit, single-consumer-GPU deployment, promising private local coding and tool-using assistants with open weights and reported SWE-bench performance.
Google DeepMind’s AlphaGenome Atlas maps predicted effects for 9 billion DNA changes, helping researchers prioritize non-coding variants, understand molecular mechanisms and design experiments without treating AI scores as diagnoses.
Anthropic’s automated alignment researchers show how AI can test safety interventions, evade weak benchmarks and improve another model, while revealing why independent evaluation still matters before deployment.
Xiaomi’s MiMo-V2.6-Pro challenges DeepSeek with multimodal reasoning, coding, long-context processing, and coordinated agents, testing whether open-weight AI can become a practical alternative to proprietary systems.