Meta open-sources Muse gadget code

Meta open-sources Muse, why AI agents fake tasks, and more.

Meta open-sources Muse gadget code

Hi there, this is your daily โ˜•๏ธ AIpresso.


In today's AIpresso:

๐Ÿ› ๏ธ Meta open-sources Muse gadget code

๐Ÿค– Why AI agents fake task completion

๐Ÿง  Google stops AI agents from memorizing tests

โœ๏ธ Claude Code lets devs rewrite its behavior

๐Ÿ‡ฉ๐Ÿ‡ช Aleph Alpha releases open-weight Kolibri

Plus: ๐Ÿ’ก 5 strategies & tactics, ๐ŸŽ 6 more stories you might like, ๐Ÿงฐ 6 tools, and ๐Ÿ“š 5 papers.

๐Ÿ› ๏ธ Meta open-sources Muse gadget code LINK

  • Meta released Muse Gadgets on Friday, an open-source project that lets developers build their own hardware connecting to Muse, the personal agent that books travel, fills out forms, and shops on a user's behalf.
  • The release includes open-source firmware and a Linux SDK, letting tinkerers wire Muse to a Raspberry Pi or ESP32 board plus displays, buttons, sensors, and actuators, with starter ideas like a color e-ink display or an HDMI TV stick.
  • Meta also built Home Link, a USB-C device that connects Muse to home networks and smart devices, making 5,000 units free to subscribers while supplies last, though Nat Friedman said shipping is still a few weeks out.
  • ๐Ÿค– Why AI agents fake task completion LINK

  • Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades AI agents on the terminal database state and side effects they leave behind rather than their tool calls or final replies, available now on Hugging Face.
  • Across 507 stateful business workflows run 20 times each, 67% of failures still terminated cleanly and reported no tool error, yet executable checks found wrong field values in 78% of them, and Claude Opus 5.5 led at 67.16% pass@1.
  • Consistency diverges sharply from single-attempt skill: Kimi-K3 solves 93.89% of tasks at least once but passes only 13% on all 20 attempts, while Claude Opus 5 clears 48% every time, so pass@20 is the column deployers should watch.
  • ๐Ÿง  Google stops AI agents from memorizing tests LINK

  • Google Cloud AI Research and university partners found that self-improving AI agents, which repeatedly rewrite their own harness based on test feedback, end up memorizing their limited test tasks, boosting training scores while gains on new tasks shrink or vanish.
  • Their fix, called RRSI, caps how many edits an agent can bundle at once and shrinks that budget over time, while a critic rejects proposals that hardcode task names, solutions, or other benchmark-specific tricks.
  • For people building agents, RRSI gained up to 4.7 points on unseen benchmarks and used about 30 percent fewer tokens, showing that self-improvement helps only when repeated feedback becomes lasting, general changes rather than memorized answers.
  • โœ๏ธ Claude Code lets devs rewrite its behavior LINK

  • Anthropic shipped "Mods" for Claude Code, a plugin-based middleware layer that lets developers rewrite how the tool looks and behaves from inside it.
  • Mods are JavaScript or TypeScript functions hooking into events like tool calls, user prompts, and UI rendering, so devs can add custom panels beside the chat, intercept tool calls, or define new commands.
  • They run across the CLI, desktop app, and partly the VS Code extension, with orgs able to whitelist which Mods load, though Mods run with the user's full permissions and aren't sandboxed, so Anthropic warns to install only from trusted sources.
  • ๐Ÿ‡ฉ๐Ÿ‡ช Aleph Alpha releases open-weight Kolibri LINK

  • Aleph Alpha has released Kolibri, an open-weight English-German MoE model under Apache 2.0, downloadable in full from Hugging Face for on-premises deployment in government and regulated industries.
  • The model carries 78.1B total parameters but activates just 3.46B per token via a Mixture-of-Experts design, supports up to million-token contexts, and offers four reasoning settings plus tool calling.
  • Vendor benchmarks show 96.9 on AIME 2025 and strong coding scores, with Aleph Alpha claiming it matches models with up to 4x the active parameters, though these are self-run results using the company's own harnesses and highest reasoning setting.
  • ๐Ÿ’ก Strategies & Tactics

    > Building a High-Performance and Portable vLLM Linear Backend with Helion: Replacing hand-tuned matrix-multiply kernels with one auto-tuned Helion implementation lets vLLM beat its default backends by over 10% throughput on some Nvidia Hopper workloads.

    > Server-Side Code Execution Tools for AI Agents, Compared: Hosted code execution lets AI agents run commands inside a provider's isolated sandbox during an API request, sparing you from securing and patching containers yourself.

    > Failure modes of Claude Code in a supervised but code-blind 60h project: Supervising Claude Code without reading its output breaks down past short tasks because it confidently misexplains bugs and fixes as much as it breaks.

    > Inside-Out AI: Rebuilding Airbnb Behind the Scenes and Across the Guest Experience: Airbnb uses AI internally to speed software development first, then redeploys those same tools to customers, resolving half of support tickets and launching new services faster.

    > We're going to need default hard budget caps on pretty much everything: Cloud services should make hard spending limits the default so a runaway AI-built app can't rack up thousands in surprise charges overnight.

    Other news you might like

    • Google's new Gemini tiers cut free users to its weakest model and lock $5/month subscribers out of ProLINK
    • Open-sourcing AstaBrief, the fast report-generation model in AstaLINK
    • Volantis raises $88M to develop photonic inference systemsLINK
    • MIT's SIFT cuts coding agent eval costsLINK
    • DeepSeek Narrows US AI Benchmark Lead to 3% After September ReleaseLINK
    • NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AILINK

    ๐Ÿงฐ Trending tools

    CoreSpeed: connects apps once and shares memory and accounts across Claude Code, Codex, Cursor, and other MCP agents through a single endpoint with budgets and activity logs.LINK

    Invofox Self Serve: a document parsing API that classifies, validates, and extracts structured data from messy real-world documents, scaling reliably in production workflows.LINK

    devpit: a native desktop app for running Claude Code agents with a visible terminal, card-driven task board, per-call cost tracking, and cross-project orchestration.LINK

    Oogwai Beacon: queries ChatGPT, Gemini, and Claude with buyer questions to show whether AI engines recommend you, who they name instead, and how to improve.LINK

    Marv: hold a hotkey to ask about anything on screen, and it answers aloud while drawing arrows and notes to guide you.LINK

    useagent: provides AI agents with their own cloud computer that use your tools to deliver finished websites, decks, spreadsheets, reports, and pull requests.LINK

    ๐Ÿ“š Trending papers & reports

    Agent training without do-overs teaches AI agents that act in live systems like security sandboxes by learning only from their best past attempts, matching standard methods that need costly repeated trial runs.LINK

    Faster image and video generation automatically picks which saved computation an AI art model can safely reuse to skip costly steps, cutting the work needed while keeping output quality nearly identical to the slower full version.LINK

    Smarter data picking shows that matching how you choose training examples to the learning method used makes AI pretraining more efficient, with one popular method staying the best filter even when paired with newer rivals.LINK

    Fooling image reading AI shows that attacks can slash a model's internal error signal to near zero yet it still describes pictures correctly, because its language side overrules the corrupted visuals, exposing where sabotage attempts fail.LINK

    AI game playtesters let one bot build a browser game while another actually plays it to catch broken interactions, lifting the share of games that work correctly to ~67%, beating standard code generation by ~37 points.LINK


    See you tomorrow for a new dose of โ˜•๏ธ AIpresso!

    More from the archive