Hi there, this is your daily ☕️ AIpresso.
In today's AIpresso:
🧬 Apple's new AI designs proteins
🤖 GPT-6 Astra pilots a drone
🇨🇳 Chinese AI labs catch up to US
🐛 Research reveals OpenAI agents attacked RubyGems in May
🎮 Nvidia RTX PRO 5500 packs 84GB memory
Plus: 💡 5 strategies & tactics, 🎁 6 other news you might like, 🧰 6 tools, and 📚 5 papers.
🧬 Apple's new AI designs proteins LINK
🤖 GPT-6 Astra pilots a drone LINK
🇨🇳 Chinese AI labs catch up to US LINK
🐛 Research reveals OpenAI agents attacked RubyGems in May LINK
🎮 Nvidia RTX PRO 5500 packs 84GB memory LINK
💡 Strategies & Tactics
> Long Live the Short King: Why 4-hi HBM Wins: Shorter 4-high memory stacks deliver the same bandwidth at lower cost, making them the cheapest way to run AI inference now that huge multi-chip racks provide ample capacity.
> Why an old caching trick is your secret to lower LLM costs: Storing and reusing past AI answers for repeat or near-identical questions cuts model costs sharply, since providers bill each duplicate call separately.
> Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels: Package Helion GPU kernels on Hugging Face's Hub with pre-tuned settings so users load fast, hardware-optimized code without slow local compilation.
> It passed CI. It passed your evals. The customer still got the wrong answer.: Instrument AI features to record what they retrieved and produced, because a technically correct answer from wrong sources passes every standard check.
> Chip Huyen explains how to cut inference costs without new hardware: Cut the recurring cost of running AI models by shrinking their precision, caching repeated prompt text, and batching requests smartly rather than buying more machines.
Other news you might like
- Anthropic's Amodei says China presents 'toughest dilemma' for his proposed AI slowdownLINK
- AI Agents Spending Money Online? New Research Says Not ReallyLINK
- Google, OpenAI and Anthropic Float Idea of AI Standards BodyLINK
- Chinese military researchers and tech giants caught using Claude — US frontier model coded 16 air-defense suppression tools targeting Taiwan, drafted anti-torpedo specs, and fed 151 million training queries to AlibabaLINK
- Iris-mini and Iris-pro are the strongest open-weight search agents in their classLINK
- Long-running AI agents quietly drop compliance rules, and bigger context windows won't fix itLINK
🧰 Trending tools
QApilot MCP for Android: automates mobile app testing with AI agents that create, run, and maintain tests, cutting manual effort and speeding up release cycles.LINK
Widgo: an AI sales rep that answers visitor questions from your docs, identifies companies, scores intent, and auto-books demos on your calendar.LINK
OpenMarket: a multi-agent marketplace where sellers pitch, competitors challenge claims, and truth agents verify evidence, letting products win on merit not marketingLINK
Pascal’s Pager: turns raw webhook JSON into clear iPhone alerts using AI, with private URLs, field masking, grouping, and 30-day payload inspection.LINK
Trancy Air: translates and rewrites text in any app via keyboard shortcuts, with screenshot capture, multi-engine comparison, grammar analysis, and pronunciation scoring for language learners.LINK
tools-for-devops-agent: provides ready-to-use skills, custom agents, and tools that extend AWS DevOps Agent for incident response, root cause analysis, and troubleshooting.LINK
📚 Trending papers & reports
Reasoning-path pruning gets a fix that keeps more of a model's candidate answers alive during problem-solving, improving accuracy or coverage in 14 of 15 tests without any retraining.LINK
Language model slimming figures out the smartest chunks to cut from a big model, halving one 70-billion model's size while scoring nearly 23 points higher on a knowledge test than rival trimming methods.LINK
AI agent scorecards often mislead, since chats rated satisfying failed the customer's actual task ~58% of the time, and the cheap auto-judge misranks nearly-equal agents 31% of the time.LINK
Byte-based language models that read text as raw characters instead of chunks eventually beat conventional models on accuracy given enough compute, matching their performance on one-sixth the training data and about 4% higher accuracy at scale.LINK
Chatbot long-term memory works better when a model treats saved snippets as bookmarks that point back to original conversations rather than standalone facts, giving more accurate answers while using fewer tokens and less delay.LINK
See you tomorrow for a new dose of ☕️ AIpresso!