β˜•οΈ Meta's open AI Runs on Laptops

Meta's laptop-ready AI, Google's open Raiden chip, and more.

β˜•οΈ Meta's open AI Runs on Laptops

Hi there, this is your daily β˜•οΈ AIpresso.


In today's AIpresso:

πŸ¦™ Meta's open AI Runs on Laptops

πŸ€– Google open-sources its Raiden AI chip

πŸ›‘ OpenAI pauses Astra over cyber risk

πŸ“Œ Claude Code defaults to auto Aug 14

πŸ’¬ Claude Code sessions can now interconnect

Plus: πŸ’‘ 4 strategies & tactics, 🎁 9 other news you might like, 🧰 6 tools, and πŸ“š 5 papers.

πŸ¦™ Meta's open AI Runs on Laptops LINK

  • Meta released Muse Glimmer, a 30B open-weight model that runs agentic tasks locally on a single GPU in a Mac or PC, aimed at developers who want AI on their own hardware instead of the cloud.
  • The model is tuned for tool use, long-horizon tasks, and failure recovery, with multimodal perception built in, and Meta claims competitive agentic and coding performance against larger flagship systems.
  • Muse Glimmer ships under Apache 2.0, and Meta says it plans to follow by releasing weights for Muse Spark 1.2, its most advanced model from the superintelligence team.
  • πŸ€– Google open-sources its Raiden AI chip LINK

  • Google has open-sourced TPU Raiden, an inference optimization library that moves KV-cache data between chips during LLM serving, now live on GitHub under Apache-2.0 as an actively developed project.
  • Raiden occupies the same layer as Nvidia's NIXL, handling KV-cache transfer between prefill and decode instances with modules for direct chip-to-chip moves, cross-VM network transfer, and offloading cache blocks from TPU memory to host RAM.
  • A shared-memory mode keeps cache in DRAM through model-server restarts for routine updates, and vLLM/SGLang integration could cut porting work for TPU users, though Google says it isn't intended for general production use yet.
  • πŸ›‘ OpenAI pauses Astra over cyber risk LINK

  • OpenAI paused internal work on its unreleased Astra model after preliminary evaluations found it could not rule out the model autonomously developing zero-day exploits or executing novel end-to-end cyberattacks from a high-level goal.
  • Astra would be OpenAI's first model to potentially hit the Preparedness Framework's Critical tier, above the High rating given to GPT-5.6-Sol, triggering controls like encrypted weights, sandboxed execution, restricted network access, and monitoring that can interrupt high-risk agentic activity.
  • OpenAI continues benchmarking Astra and plans tests with government agencies and safety orgs, though it has not released the underlying benchmarks, confirmed a final Critical classification, or said which of the two criteria may apply.
  • πŸ“Œ Claude Code defaults to auto Aug 14 LINK

  • Anthropic will make Auto mode the default for all Claude Code Pro, Max, and Team users starting August 14, letting the tool make decisions autonomously instead of surfacing permission prompts on every action.
  • In a study of 1,053 paid testers, Auto mode blocked 89% of dangerous commands versus 13.6% caught by manual human review, and it stopped 800 harmful commands that humans approved while missing only six that humans caught.
  • Auto mode lets models built for long-running work like Claude Opus 5 run uninterrupted for hours on large tasks, and users can revert their default via Shift+Tab, though it stays opt-in for Enterprise users for now.
  • πŸ’¬ Claude Code sessions can now interconnect LINK

  • Claude Code sessions on macOS and Linux can now message each other directly, letting Claude send text summaries or ask questions between terminals instead of you manually copying context around.
  • Communication runs locally when sessions share a machine, and Anthropic supports it both ways, one session can query another and get an answer, or Claude can push a message on its own when a change affects parallel work.
  • Admins can disable the feature via settings, and cross-machine messaging routes through Anthropic's servers with responses only, though it's unavailable on Amazon Bedrock, Google Cloud Agent Platform, or Microsoft Foundry.
  • πŸ’‘ Strategies & Tactics

    > [AINews] Zawinski's Law of MultiAgents: Agents naturally seek to message each other, creating hidden coordination channels that turn multi-agent safety into a central concern for AI labs.

    > I changed one setting in Claude Code, and my token burn dropped by 45%: Lowering Claude Code's reasoning effort to Medium for routine coding cut token usage 45% with no drop in result quality.

    > Self-monitoring doesn't scale (without these 3 countermeasures): Detecting AI models that secretly cover for each other requires fake bad actions, an independent honest reviewer, and scrubbing hidden signals together-removing any one lets sabotage slip through.

    > I vibe coded with Qwen 3.6 on a platform it had never seen, and the model was never the bottleneck: A security-first coding platform keeps API keys out of AI-generated code by injecting them through a proxy, and the model's limited memory, not its skill, proved the real constraint.

    Other news you might like

    • Ultra-High Interactivity on NVIDIA GPUs? - TileRT InferenceXLINK
    • Making Knowledge Distillation Cheap Enough to Run at ScaleLINK
    • xAI's Imagine Image 2.0 lands just behind OpenAI's GPT-Image-2 in Arena benchmarksLINK
    • V4-Flash vs. V4-Pro: DeepSeek promised better and cheaper. It’s true, but not how I expected.LINK
    • Tencent's Team Memory shares AI agent memory across a team, with no governance yet for when it's wrongLINK
    • Why every company wants an AI model router right nowLINK
    • Four AI agents coordinating in real time outperformed Claude Opus 4.8 on enterprise coding tasksLINK
    • TutorMoments: Do AI tutors know when to help and when to hold back?LINK
    • ShadowAI-Watch: Bringing AI Agent Activity Out of the ShadowsLINK

    🧰 Trending tools

    Portfolio Lab: tests AI-generated investment strategies on unseen data and live markets before letting you deploy vetted ones to your brokerage account.LINK

    SecondBrain Note by GenSpark: an AI agent engine that researches topics and generates unbiased Sparkpages, synthesizing trustworthy information to save you time versus SEO-driven search results.LINK

    AI Group Call: lets you state a goal and instantly join a live voice call with six AI participants who discuss it, argue, and pause when you speak, with transcripts and summaries saved for later.LINK

    Paritok: compresses tool outputs, files, and history sent to coding agents, cutting token usage up to 85% for longer local sessions.LINK

    Prime Agent: aggregates and orchestrates global GPU resources into an on-demand compute platform, offering H100s from $1.65/hr and A100s from $0.87/hr.LINK

    oqoqo: runs large-scale eval experiments in realistic environments, letting you build private benchmarks, test agent performance across products, and identify optimal models plus interface friction points.LINK

    πŸ“š Trending papers & reports

    Psychological defense classification generates realistic synthetic examples grounded in clinical definitions to fix scarce, imbalanced therapy-text data, lifting accuracy to 58.26%, a 40.25 point jump over the prior baseline.LINK

    Robot walking training can be sped up by mixing simulator-generated experience with a small amount of imagined, model-predicted experience, cutting training steps needed by about 29% and training time by ~12% without hurting the final walking skill.LINK

    Training data detection can now flag which documents a company's model likely learned from while guaranteeing a strict, provable limit on false accusations, giving copyright and privacy disputes real statistical evidence instead of guesswork.LINK

    Transformer training stability gets a rigorous explanation for where to place a common stabilizing step, letting builders predict and prevent runaway training failures before wasting compute on new model designs.LINK

    Live simulation compression lets scientific supercomputer runs be compressed into a compact model as they happen, matching the quality of compressing the data after the fact, without storing everything first.LINK


    See you tomorrow for a new dose of β˜•οΈ AIpresso!

    More from the archive