Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π€ Grok Bot will route tasks to Claude
β‘ Claude Haiku 5.5 rivals GPT-6
π€ Google Cloud launches Gemini agent
π§ͺ AWS launches AI agent sandbox on Mac
π§ Liquid AI launches open edge models
Plus: π‘ 5 strategies & tactics, π 6 more stories you might like, π§° 6 tools, and π 5 papers.
π€ Grok Bot will route tasks to Claude LINK
- Elon Musk said on October 7 that SpaceXAI's Grok Bot will route customer tasks to external providers, naming Claude Opus 5.5 for language alongside Midjourney for images and Suno for music, with no plan to abandon Grok's own models.
- Routing is task-based with no customer-facing model picker; Cursor manages selection, usage analytics identify the serving model including any failover, and billing follows whichever model handles each request.
- Enterprise teams get a model allowlist that is honored by default, though onboarding warns enforcement is not guaranteed and the announcement supplies no independent test showing better results or full rollout details.
β‘ Claude Haiku 5.5 rivals GPT-6 LINK
- Anthropic shipped Claude Haiku 5.5 on Wednesday at $0.10 per million input tokens and $0.50 per million output for prompts up to 100,000 tokens, matching GPT-6 Luna's rate and cutting Haiku 4.5's short-prompt price by 90%.
- On Artificial Analysis's Intelligence Index it scored 43 at maximum effort, five points above Luna, with Anthropic targeting high-volume support, browser use and subagent retrieval for Opus 5.5 and Sonnet 5.5.
- Above 100,000 tokens every rate multiplies by five and a new tokenizer counts roughly 30% more tokens per text, so Anthropic pegs the real average saving nearer 75%, and it burned about three times Luna's output tokens per task.
π€ Google Cloud launches Gemini agent LINK
- Google Cloud launched Gemini agent, a single universal agent that operates across business apps and enterprise systems from a chat prompt, planning work, using tools, and connecting to systems of record under enterprise policies.
- You give Gemini agent "objectives, not instructions," per CEO Thomas Kurian, and it spins up sub-agents with their own identities and emails, maintaining context across Workspace, Microsoft 365, Slack, and third-party apps while routing between Gemini models and Claude.
- It grounds on a borderless lakehouse querying Amazon S3, Azure Data Lake, Databricks Unity, and Snowflake Polaris without egress fees, with real-time spending caps for cost control, though Google Cloud says other closed and open models will only arrive in the future.
π§ͺ AWS launches AI agent sandbox on Mac LINK
- AWS released Strands Box, an open-source sandbox that uses OS-level isolation plus contextual rule enforcement to keep autonomous agents from taking harmful actions like deleting a production database or racking up costly API calls.
- Box pairs its isolation with the Dogwood Local Engine, which gives the policy engine temporal awareness, so tool calls are checked against prior activity, for example, capping Slack posts to three per ten minutes or limiting Git pushes, with Shell and Python interpreters exposing operations for precise policies.
- The rules enforce deterministically without trusting the agent to comply, but AWS notes developers still decide what access to grant and where human review is needed, and the tool is on GitHub for macOS only, with Linux in development and no Windows release date.
π§ Liquid AI launches open edge models LINK
- Liquid AI released two open-weight d1 decision models on Hugging Face today, built on its Liquid Foundation Models to answer in a single forward pass rather than generating tokens.
- d1-3B tops the Decision Index 0.2.1 among sub-10B models at 48.57, beating Decider 35B-A3B, and averages 82.9 across seven public datasets, while the multimodal d1-omni-600M handles text with image or audio.
- d1-3B answers in under 50 ms on every NVIDIA Jetson device tested, as low as 16 ms on AGX Thor, though d1-omni-600M is an early research release with no reported speed or modality benchmarks yet.
π‘ Strategies & Tactics
> GhostCompact: How Compaction Made OpenAI Inference Nearly Free: A billing bug let OpenAI users route arbitrary AI work through a mispriced context-shrinking feature and pay up to 200 times less.
> Building Spyre as a Native PyTorch Device: IBM makes its Spyre inference chip work through standard PyTorch commands, so developers run models on it with familiar code instead of chip-specific tooling.
> Claude Sonnet 5.5 vs. Opus 5.5: 42% cheaper and perfect on every run: In coding tests, Anthropic's cheaper Claude Sonnet 5.5 beat pricier Opus 5.5 on accuracy and cost, making it the better default for hard coding work.
> Building scalable, agent-friendly APIs for AI applications: Build AI agent services as separate stateless components sharing one database, so conversations keep their context and slow tool calls can't stall regular chat traffic.
> Revamping Skills in Deep Agents: Attach specialized instruction files and their tools together so an AI agent only loads what each task needs, keeping it fast and predictable at scale.
Other news you might like
- GPT-6 and Intelligent UI for everyoneLINK
- Google rolls out improved SynthID AI content detector, now available globallyLINK
- Microsoft brings hybrid AI and agent controls to WindowsLINK
- Report Suggests OpenAI Has Clawed Tons of Market Share Back From Anthropic in 2026LINK
- Kore.βai's AI agent optimization engine Autoloop availableLINK
- At WebexOne, Ciscoβs Jeetu Patel argues the agent era will be won on context, cost and controlLINK
π§° Trending tools
BotBus: connects local coding agents like Codex and Claude Code to your phone, letting you manage tasks, approvals, files, and terminals remotely.LINK
Cekura Bench: automates QA for conversational AI agents, handling pre-production simulation, evaluation, production call monitoring, and CI/CD integration for consistent reliability.LINK
NOVA CLI: generates code from descriptions, auto-fixes runtime errors, refactors files, and handles Git commits all inside your terminal, breaking the fix-and-retry loop.LINK
Tractionwave: analyzes ads and marketing visuals using AI attention prediction, showing heatmaps and ranked hotspots so you can validate designs before spending budget.LINK
judged.systems: routes support tickets through configurable question packs to classify urgency, refund requests, and policy risk, returning scored decisions your thresholds auto-accept or flag for review.LINK
ClawCall: lets ChatGPT, Claude, or Grok place real phone calls for you, navigating phone trees, waiting on hold, and reporting back what was said.LINK
π Trending papers & reports
Training AI agents on the fly gets a cheaper method that keeps policies learning accurately from slightly outdated data, avoiding the costly extra samples or biased shortcuts rival approaches rely on.LINK
AI agent course-correction learns exactly when stepping in actually helps a task-running assistant rather than disrupting it, boosting success by ~7.8 points across tests while beating existing intervention methods by ~2.9 points.LINK
Noised training prompts show that a one line tweak to how diffusion text models learn boosts their puzzle solving accuracy from ~4% to ~25% on hard Sudoku, with no added running cost.LINK
Search instructions in embedding models often get ignored when queries contain distracting text, but retraining with those distractions makes the models follow instructions far more reliably without hurting other tasks.LINK
Fact editing in language models lets you update specific facts a model knows without retraining it or breaking unrelated knowledge, reaching near-perfect success and roughly 3x better accuracy on multi-step reasoning than existing methods.LINK
See you tomorrow for a new dose of βοΈ AIpresso!