Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π AI assistants can now make phone calls
π Plugin flaw hits major AI coding agents
π΅οΈ Researchers used Claude to breach OpenAI
π± Android AI benchmark tests long tasks
π€ Claude Code now runs parallel agents
Plus: π‘ 5 strategies & tactics, π 7 other news you might like, π§° 6 tools, and π 5 papers.
π AI assistants can now make phone calls LINK
π Plugin flaw hits major AI coding agents LINK
π΅οΈ Researchers used Claude to breach OpenAI LINK
π± Android AI benchmark tests long tasks LINK
π€ Claude Code now runs parallel agents LINK
π‘ Strategies & Tactics
> Auditing in the age of (good enough) AI: Security auditors used AI agents to build custom analysis tools and machine-checked proofs for an unfamiliar coding language, catching a fund-stealing bug beforehand.
> Two techniques for working with System One models: Handle fast decision-only AI models by layering short-term goals for real-time control and comparing options in tournament rounds to pick the best from many.
> Using LLM-as-a-judge scoring to measure your software factory: Use one AI model to grade past coding-agent sessions on quality and efficiency, giving visibility into performance and a basis for automatic improvement.
> GitHub and Anthropic used their own agents for major Rust rewrites β with very different playbooks: AI coding agents made massive software rewrites affordable, letting GitHub and Anthropic convert huge codebases to Rust in months instead of years.
> I stopped giving Claude Code my entire project, and my Claude usage lasted 4x longer: Point Claude Code straight to the relevant files instead of letting it search your whole project, since exploration burns usage-limit tokens fast.
Other news you might like
- Anthropic Says Claude Leads 26% of Its AI Research, Up From Under 1% in FebruaryLINK
- Anthropic has set up a bio research lab for physical experimentsLINK
- Open-weight models now handle a majority of tokens on Vercelβs AI Gateway. But Anthropic still takes 64% of the spend.LINK
- PrismML hopes its tiny LLM will change how we all use AILINK
- Intel squeezed a 1.58-bit LLM down to 1.485 bits without changing a single weightLINK
- WebKit blog breaks down whatβs new with Safari 27 for developers, including MCP supportLINK
- OpenAI launches Astra for Law, a GPT-6 configuration for legal researchLINK
π§° Trending tools
Toki Coordination: an AI assistant that captures texts, voice notes, and emails into scheduled plans with adaptive reminders, conflict resolution, and weekly insights.LINK
Gemini 3.8 & 3.8 Live Extended Thinking: near real-time voice models for building production-ready voice agents, with fluid dialogue, visual grounding, and multi-step reasoning for complex tasks.LINK
CAT ME: transforms a photo of a person into an AI cat lookalike, matching features, expressions, and accessories for fun sharing with friends.LINK
ABrush: integrates AI models and workflows directly into Photoshop, automating repetitive production tasks while keeping artists in full creative control.LINK
CreatorHat: finds outperforming videos, researches YouTube keywords, tracks rankings, and transcribes uploads locally on your Mac to suggest titles and descriptions.LINK
Toone: build deterministic AI-agent workflows using your own keys, edit routines mid-run, resume anytime, with browser navigation, audio capture, and git timeline included.LINK
π Trending papers & reports
Expert routing gives mixture models, which activate only a few specialist components per task, a smarter way to pick which specialist handles each input, boosting accuracy on both language and image benchmarks.LINK
Vague agent instructions can be fixed by studying the runs an AI already did instead of expensive trial-and-error retesting, cutting cost and reliably beating the original agent across four benchmarks.LINK
A shared data format lets one standard AI model handle plain text, knowledge graphs, and hypergraphs directly, keeping the roles and relationships that get lost when data is flattened into simple word sequences.LINK
Cross-domain graph learning maps very different network datasets onto one shared coordinate system, letting a single model transfer knowledge across wildly varying data and beat existing methods on 14 classification tasks.LINK
Two-armed robot training combines cheap simulation data with real human demonstrations, cutting the need for costly robot practice while succeeding ~63% of the time on unseen scenes, ~54 points above robot-only training.LINK
See you tomorrow for a new dose of βοΈ AIpresso!