Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π¦ New malware uses AI to plan attacks
π° OpenAI cuts GPT-6 API prices in half
π± Qualcomm chip runs 30B AI on phones
π€ Claude Opus 5.5 matches Fable's quality
βοΈ vLLM adds support for more chips
Plus: π‘ 4 strategies & tactics, π 7 other news you might like, π§° 6 tools, and π 5 papers.
π¦ New malware uses AI to plan attacks LINK
π° OpenAI cuts GPT-6 API prices in half LINK
π± Qualcomm chip runs 30B AI on phones LINK
π€ Claude Opus 5.5 matches Fable's quality LINK
βοΈ vLLM adds support for more chips LINK
π‘ Strategies & Tactics
> How Shopify built a continual learning loop with PyTorch and vLLM: Retrain a specialized model daily on real production failures to beat frontier models on quality while cutting serving costs 96%.
> Enabling Private High-Performance Production AI Inference with NVIDIA Confidential Computing: NVIDIA's hardware-encrypted computing runs sensitive AI inference on trusted GPUs while retaining over 96% of normal speed, protecting private prompts and models.
> π¬ An Oscar, Two Asteroids, and the Algorithm in Your sklearn: John Platt on AI for Science: Google's ERA turns scientific problems into scoring tasks that an AI iteratively solves, but scientists must verify results to avoid gaming the metric.
> How OpenAI Built GPT-Live: OpenAI's GPT-Live uses a fast voice model that listens and talks simultaneously while delegating hard questions to a bigger reasoning model, so conversations stay natural without pauses.
Other news you might like
- Meta's Muse personal AI agent tops ChatGPT, Grok and Claude for post-launch downloadsLINK
- China's ByteDance gained access to over 2,000 Nvidia B200 chips through Norway data centerLINK
- Snorkel AI Raises $350 Million at $3.5 Billion Valuation in Series ELINK
- European neocloud Verda raises $189M to build the AI infrastructure of tomorrowLINK
- Text handoffs slow AI models down. C2C lets them communicate through KV caches insteadLINK
- DevDay: New Plans and new platform for building AI appsLINK
- Introducing DigitalOcean Managed Agents: One AI-native stack to power your intelligenceLINK
π§° Trending tools
Ruby UTCP: ruby implementation of the Universal Tool Calling Protocol, letting AI agents call native APIs directly via JSON manifests without MCP wrapper overhead.LINK
Lumiko: records your screen and auto-generates cursor-tracked zoom and pan effects, with browser trimming, secret blurring, webcam bubble, and 4K local exports.LINK
Zella: records and edits screen or camera videos on-device, auto-generating captions, cutting silences, removing filler words, and adding zoom effects offline.LINK
Compute:Arena: community-sourced benchmark database for local AI models, comparing token speeds across hardware, runtimes, and quantisation levels to guide your setup choices.LINK
redcell: an AI red-team platform where autonomous LLM agents run full penetration tests in a Kali container and generate reports.LINK
BiBimba: stores searchable clipboard history and screenshots, using on-device AI to extract text from images, then translate, summarize, and rewrite locally on Mac.LINK
π Trending papers & reports
AI-watching-AI safety checks map the sneaky ways one AI overseer might secretly collude with the AI it's supposed to police, giving companies a clearer test for whether such oversight can actually be trusted before deployment.LINK
Risk-averse AI training shows that teaching a model to play it safe on small bets carries over to enormous ones, lifting cautious choices from 2% to ~70%, a possible safety net if AIs turn out misaligned.LINK
Nuclear escalation in AI agents persists across 13 models even when prompts explicitly warn of nuclear harm, showing that ethical judgment scoring well on isolated dilemmas often fails to steer behavior in complex, high-stakes strategy scenarios.LINK
Audio AI reasoning tends to lose track of what it actually heard as it thinks longer, but a new training method fixes this, roughly doubling perception accuracy to ~64% and boosting overall performance to ~75%.LINK
Sign-based training math pins down exactly when a popular shortcut for teaching models beats the standard method, showing it can improve faster per dollar of compute in noise-heavy settings.LINK
See you tomorrow for a new dose of βοΈ AIpresso!