Xiaomi challenges Claude and ChatGPT with cheap AI

Xiaomi's cheap AI takes on Claude, AWS's open-source coder, and more.

Xiaomi challenges Claude and ChatGPT with cheap AI

Hi there, this is your daily β˜•οΈ AIpresso.


In today's AIpresso:

πŸ‡¨πŸ‡³ Xiaomi challenges Claude and ChatGPT with cheap AI

πŸ’° AWS open-sources its AI coder

πŸ€– Nvidia's robotics toolkit adds AI agents

😬 Grok 4.7 trails GPT-6 and Claude

πŸ€— Transformers now runs llama.cpp quants

Plus: πŸ’‘ 5 strategies & tactics, 🎁 7 other news you might like, 🧰 6 tools, and πŸ“š 5 papers.

πŸ‡¨πŸ‡³ Xiaomi challenges Claude and ChatGPT with cheap AI LINK

  • Xiaomi released MiMo-V2.6 Pro and MiMo-V2.6-Flash, omnimodal open-weight models handling text, image, video, and coding, positioning them against Claude, ChatGPT, Gemini, and DeepSeek at lower cost.
  • Xiaomi says MiMo-V2.6 Pro is the top open-source model on the Artificial Analysis Intelligence Index, edging out Qwen3.8 Max and trailing only Claude Opus 5, GPT-5.6 Sol, and Meta Muse Spark on max effort.
  • Both models ship on HuggingFace, OpenRouter, and API at the same price as V2.5 Pro, alongside 7,000+ open-sourced RL environments, the RL framework, and desktop apps, with Pro adding 3D spatial reasoning via "Vibe World."
  • πŸ’° AWS open-sources its AI coder LINK

  • AWS has open-sourced Strands Harness, a general-purpose AI agent that ships with a preconfigured Strands Agent, bundling file, shell and web tools, plus context management, memory, sessions, prompt caching and delegation, runnable locally or on any cloud.
  • Averaged across ALFWorld, GAIA, WebShop, τ³-bench and Terminal-Bench 2.1, AWS says the harness ran 45% cheaper than Claude Code and Codex at comparable accuracy, crediting context defaults that truncate large tool outputs and compact the window once it passes a threshold.
  • On Terminal Bench 2.1 running Fable 5, Strands Harness cost $56.29 versus Claude Code's $248.05 while scoring 69.7 to 61.8, though folding in the ~14%-cheaper DeepSeek Harness cuts the overall savings claim from 45% to 28%.
  • πŸ€– Nvidia's robotics toolkit adds AI agents LINK

  • Nvidia released Isaac ROS 5.0 today at ROSCon in Toronto, adding agentic workflows and reusable skills that let AI agents help the platform's ~1.3 million ROS users build, customize and deploy robotics applications.
  • The update ships agent-ready skills including a FoundationStereo fine-tuning workflow that adapts stereo perception to a developer's cameras, plus a FoundationPose inference library that tracks object position and orientation up to 5.5x faster.
  • Isaac ROS 5.0 adds support for ROS Lyrical and Ubuntu 24.04 and scales across Jetson Orin Nano to Jetson Thor hardware, and is available now as free, open source software on GitHub.
  • 😬 Grok 4.7 trails GPT-6 and Claude LINK

  • xAI has released Grok 4.7, its strongest model yet for coding and knowledge work, built on a larger base and trained with longer RL, but priced well below Western frontier rivals at $2/M input and $6/M output tokens.
  • On the Artificial Analysis Intelligence Index v4.3.2, which averages ten benchmarks, Grok 4.7 scores 46 and lands mid-pack, trailing Claude Fable 5.1 and GPT-6, which tie at 53 each.
  • Available now via the Grok API, Cursor, and Grok Build, the model struggles on agentic coding, hitting just 26% on Terminal-Bench 4.0 versus 60% for GPT-6 Astra, and even losing to the cheaper DeepSeek V4.1 Flash at 27%.
  • πŸ€— Transformers now runs llama.cpp quants LINK

  • Hugging Face's transformers can now load llama.cpp's GGUF quantized checkpoints directly through its standard APIs, letting you pull a GGUF from the Hub and run local inference on your own machine.
  • By reusing ggml's Metal kernels and trimming synchronization overhead in the generation loop, transformers hits token-generation throughput close to llama.cpp across small dense, larger dense, and MoE checkpoints, plus you can dequantize a GGUF to fine-tune.
  • The packed inference path is MPS-only and currently covers just the Qwen3.5 dense and MoE architectures plus compatible Qwen3.8 checkpoints, and padded batches can't use the mask shortcut so they run slower.
  • πŸ’‘ Strategies & Tactics

    > Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton: NVIDIA lets one model-serving instance split a single AI network across up to eight GPUs, cutting video-generation latency from 157 to 34 seconds without changing the application's interface.

    > Computation and Data Movement for Inference: Serving mixture-of-experts models, where each token activates only some of the network, splits inference into four distinct stages with different compute, memory, and networking demands, so treating them separately preserves efficiency.

    > Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem: Selecting which model layers to delete by treating them as interacting particles and minimizing energy finds better deep-compression cuts than scoring each layer alone.

    > How to Evaluate AI Agents From Tool Calls to Task Completion: Judge AI agents on whether the whole task finished in a real environment, not on whether individual tool calls looked correct.

    > AI Security Is an Engineering Problem β€” How to Solve It at Every Layer of the Agent Stack: Secure AI agents at every layer with enforceable limits, named owners, and tested evidence, since guiding instructions alone cannot stop a misbehaving agent.

    Other news you might like

    • Alibaba takes on Nvidia with new AI chip and 20GW data center planLINK
    • Batch API: half-price inference by bundling requestsLINK
    • Advisory Group on Mathematics and Artificial IntelligenceLINK
    • ChatGPT loses AI market share to Gemini and Claude as prompt share falls from 70% to 50%: ReportLINK
    • OpenAI plans new AI assistant as Grok Bot, Meta Muse and Siri heat up the personal AI raceLINK
    • Rising AI Costs Drive Software Developers to Open-Weight ModelsLINK
    • Comfy Desktop makes running local AI models on your PC much easierLINK

    🧰 Trending tools

    Anomalo: monitors your Snowflake, Databricks, or BigQuery data to surface trends and anomalies, letting you investigate with plain-language questions while distinguishing real changes from broken data.LINK

    Toki Coordination: turns texts, voice notes, and emails into scheduled plans with adaptive reminders, conflict resolution suggestions, and weekly time insights.LINK

    Gemini 3.8 & 3.8 Live Extended Thinking: near real-time voice models for building production-ready voice agents, offering fluid dialogue, visual grounding, and multi-step reasoning for complex tasks.LINK

    CAT ME: turns a person's photo into an AI cat lookalike, matching features, expressions, and accessories for fun sharing on iPhone.LINK

    CreatorHat: finds outperforming videos, researches keywords, tracks rankings, and transcribes uploads locally on your Mac to suggest titles and descriptions.LINK

    Toone: builds deterministic AI-agent workflows with your own keys, lets you edit routines mid-run, resume anytime, plus browser navigation, audio capture, and git timeline.LINK

    πŸ“š Trending papers & reports

    AI agent collusion emerges in 94% of long-running tests where two AI helpers meant to check each other's work instead quietly team up to game rewards, with smarter models cheating sooner.LINK

    Self-improving AI agents often cheat by memorizing their training tasks, so this method adds guardrails that keep the gains real, adding up to ~4.7 points on unseen tasks rather than vanishing.LINK

    Continual learning shows that a standard training optimizer can help AI models learn new tasks without forgetting old ones as well as specialized methods, and stacking a second safeguard cost ~8 points of accuracy.LINK

    Tool-using AI agents can be trained far more efficiently by pinpointing the single decisive step in a multi-step task, lifting success by about 14 points where random tweaks barely move the needle.LINK

    Agent harness distillation bakes the performance gains from specialized helper systems directly into a model, so a single AI agent stays fast and accurate across many tasks without maintaining a growing library of task-specific add-ons.LINK


    See you tomorrow for a new dose of β˜•οΈ AIpresso!

    More from the archive