Hi there, this is your daily ☕️ AIpresso.
In today's AIpresso:
🤖 AI agents went rogue in UK tests
🚗 Nvidia opens explainable self-driving model
🧮 OpenAI's Astra solves 10 math problems
🐫 Alibaba open-sources its best AI model
🛡️ Mistral launches open AI safety classifier
Plus: 💡 5 strategies & tactics, 🎁 8 other news you might like, 🧰 6 tools, and 📚 5 papers.
🤖 AI agents went rogue in UK tests LINK
🚗 Nvidia opens explainable self-driving model LINK
🧮 OpenAI's Astra solves 10 math problems LINK
🐫 Alibaba open-sources its best AI model LINK
🛡️ Mistral launches open AI safety classifier LINK
💡 Strategies & Tactics
> [AINews] Megakernels are so dead and so back: Nvidia's new Rubin chip design eliminates the launch delays that made painstaking fused megakernels worthwhile, so most inference providers can abandon them.
> PipeNetwork/minimax-h3-mlx: A ported version lets MiniMax's text-and-audio video generator run locally on Apple Silicon Macs, though it needs 115 GB and 45 minutes per clip.
> Nvidia’s NOOA makes an agent one Python class: Nvidia's NOOA framework packs an AI agent's abilities, memory, and prompts into a single Python class, making agent behavior easier to inspect and test.
> Deploy local agents everywhere with LFM2.5-2.6B: A compact 2.6-billion-parameter AI model runs capable agents on phones and laptops, matching models four times larger on tool use and instructions.
> New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging: LLM's new release exposes AI models' reasoning traces and adds provider-run tools like web search, turning the command-line tool into an agent framework.
Other news you might like
- Microsoft Tells Engineers ‘Tokenmaxxing Is Not What We Are Optimizing For’LINK
- Elon Musk says SpaceX will exclusively use Nvidia GPUs 'because they are the best', says optimized Vera Rubin NVL72 will be launched into space next yearLINK
- Samsung announces next-gen 3D memory when all we want is reasonably priced DDR5LINK
- Why An LLM’s Memory Gets Expensive and How to Fix ItLINK
- Nvidia open sources cuFile API, accelerating GPU read/write capability for high-speed storageLINK
- AI helps Microsoft bug hunters chase a record $20M paydayLINK
- The cost of being half-hearted in AI and how to avoid the Solow ParadoxLINK
🧰 Trending tools
Pocketrisu: Self-hosted AI roleplay chat platform you run on your PC or personal server, forked from RisuaiLINK
claude-directory: a curated collection of AI-generated UI experiments made with Claude, featuring landing pages, animations, shaders, and 3D designs.LINK
npm i -g hotcell: self-hostable sandbox SDK for running code securely on any machine, with per-sandbox tokens, egress controls, and no direct exposure of API keys to sandboxed processesLINK
Wispr Flow Notetaker: transcribes meetings with speaker names instead of labels, generating accurate summaries of decisions and next steps for Mac users.LINK
BackEngine MCP: unifies Slack, email, tickets, and CRM data into one continuously updated record, so AI assistants answer from complete context, not fragments.LINK
AdAnt AI: generates and iterates social ad creative for TikTok, Instagram, and YouTube using strategies proven to cut paid acquisition costs by 60%.LINK
📚 Trending papers & reports
Test-time compute strategies for AI reasoning tools vary so much in method and reporting that comparing them fairly requires a shared scoring system, which this framework provides for accurate cross-study benchmarking.LINK
Video research agents that force step-by-step video watching before web searching now answer complex multi-hop video questions with 64.0% accuracy, beating Claude-4.5-Sonnet, GPT-5, and Gemini 2.5 Pro.LINK
Compiler blind spots can be filled by AI that reads surrounding code to spot speed-up opportunities compilers miss, with the top model producing correct fixes ~95% of the time and speeding up code in ~83% of cases.LINK
AI-generated reward signals used to train robots and agents through trial and error can be automatically cleaned up without human labeling, making the training faster and more reliably aligned with the actual goal.LINK
Automated algorithm design uses a research-planning AI system, complete with memory that learns from past failures, to outperform existing automated search methods on classic routing and packing problems, cutting wasted trial-and-error.LINK
See you tomorrow for a new dose of ☕️ AIpresso!