Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π₯οΈ ChatGPT can now track your Mac activity
π Apple trains its own AI for China
β‘ Google's Gemini 3.7 Flash cuts price 50%
π OpenAI previews GPT-5.6 Sol at 14x speed
π¨π³ China's GLM-5.3 tops Anthropic on coding
Plus: π‘ 5 strategies & tactics, π 7 other news you might like, π§° 6 tools, and π 5 papers.
π₯οΈ ChatGPT can now track your Mac activity LINK
π Apple trains its own AI for China LINK
β‘ Google's Gemini 3.7 Flash cuts price 50% LINK
π OpenAI previews GPT-5.6 Sol at 14x speed LINK
π¨π³ China's GLM-5.3 tops Anthropic on coding LINK
π‘ Strategies & Tactics
> Rubrikβs lessons from one month with Mythos Preview: Rubrik found that Anthropic's Mythos AI uncovers software vulnerabilities faster than humans can fix, forcing it to automate reliable fixes while reserving human judgment for the rest.
> Why Capital One built its multi-agent AI platform around open-weight models: Capital One fine-tunes open-weight AI models with its proprietary data because that customization beats generic models and lifts performance across every use case.
> Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets: Explains how to build an AI agent that records robot demonstrations, trains a control policy by streaming data straight from cloud storage, and redeploys it, so teams avoid re-uploading and re-downloading the same footage each cycle.
> ClickHouse and Hud build a runtime feedback loop for AI-generated software: ClickHouse and Hud link live production data to specific code changes so teams can safely evaluate, verify, and fix AI-generated code.
> Why your AI pipeline costs 10x more after the demo: Cut runaway AI costs by fixing wasteful architecture, caching repeated prompts and summarizing chat history, rather than just switching to a cheaper model.
Other news you might like
- DeepSeek's AI models are about to cost four times moreLINK
- Claude Code now runs daily maintenance on Anthropic's software with a 46 percent merge rateLINK
- FP8 Training on AMD GPUs with TorchTitan and TorchAO: Upstreaming Performance ImprovementsLINK
- Writer introduces new AI model and upgraded harness to contain token costsLINK
- Three Claude agents given conflicting orders sabotaged each other on a shared server, then didn't tell users what they'd doneLINK
- AIβs βmiddle classβ has gotten dramatically better at hackingLINK
- Frontier agents don't comply with standards, even when instructed toLINK
π§° Trending tools
BearDrive: syncs the local folders AI agents work in, automatically versioning and attributing files so teammates access them instantly without moving anything.LINK
BrowserAct Cloud: a browser automation layer for AI agents that bypasses blocks, handles logins and verification, and extracts clean web data at scale.LINK
CodeBurn: tracks AI coding costs by reading local session logs from 40+ tools like Claude Code and Copilot, then finds and fixes token waste automatically.LINK
Hoplite: migrates your local coding agent setup, sessions, MCP servers, dependencies, CLIs, to the cloud, letting you run multiple agents in parallel without laptop crashes or manual configuration.LINK
min.: builds digital profiles of contacts from your emails and meetings, helping you leverage relationship context to get callbacks and close deals faster.LINK
Basedash Tasks: lets you build dashboards and analyze customer data by describing what you need in plain language, skipping manual SQL writing entirely.LINK
π Trending papers & reports
Coding assistants' blind spot for rare programming languages shows they generate poor code by default on Cangjie, but adding basic grammar rules boosts accuracy efficiently, while AI agents score highest yet burn far more compute.LINK
Final model selection for multimodal AI systems is unreliable when candidates score similarly, but cleaning up messy image-text data cuts ranking flip-flops from ~33% to ~11% and nearly doubles evaluation consistency, showing training longer isn't always better.LINK
Multi-turn chatbot jailbreaks can now be trained automatically without any human scripts or outside data, breaking safety filters with an average 80.1% success rate across tested chatbots, exposing a gap current safeguards don't cover.LINK
Body sensor networks get a smarter traffic system that learns which routes to use for medical data, cutting delay, energy use, and network overhead while delivering more packets successfully than existing methods.LINK
Category-by-category safety training tunes chatbots to fix their weakest harm categories individually, instead of one blanket safety score, shrinking the gap between best and worst areas while staying helpful.LINK
See you tomorrow for a new dose of βοΈ AIpresso!