Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π‘οΈ Google unveils Gemini 4
π€ OpenAI builds AI that designs chips
π΅οΈ Chinese AI firm caught probing OpenAI's secret reasoning
π GLM-5.3 nearly matches Claude at hacking
π¨π³ DeepSeek open-sources Huawei chip tools
Plus: π‘ 5 strategies & tactics, π 6 more stories you might like, π§° 6 tools, and π 5 papers.
π‘οΈ Google unveils Gemini 4 LINK
π€ OpenAI builds AI that designs chips LINK
π΅οΈ Chinese AI firm caught probing OpenAI's secret reasoning LINK
π GLM-5.3 nearly matches Claude at hacking LINK
π¨π³ DeepSeek open-sources Huawei chip tools LINK
π‘ Strategies & Tactics
> Confidence Thresholds for Model Escalation Routing: Have a cheap model score its own confidence and escalate only low-scoring answers to a stronger model, cutting cost without shipping confident wrong answers.
> How to Gate Pull Requests on LLM Evals in CI: Block pull requests that change AI prompts from merging when a committed test set's pass rate drops below a threshold measured from repeated runs.
> Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton: Explains how to serve recommendation models faster by compiling them ahead of time and reusing cached user-history computation, cutting latency up to roughly sixfold.
> Tracing Agent Harness Behavior with NVIDIA NeMo Relay: Use NVIDIA NeMo Relay's execution traces alongside pass/fail checks to confirm whether a change to an AI agent genuinely improves results rather than just cutting steps.
> From Upstream Changes to Downstream Confidence: Inside Torch Spyreβs Integration with PyTorch CRCR: IBM's Torch Spyre uses an AI pipeline and a config file to pick and adapt which of PyTorch's many tests matter for custom accelerator hardware, catching regressions without forking test code.
Other news you might like
- OpenAIβs Jev clone could help the frontier lab stop its swarming agentsLINK
- Google figures out how to watermark AI-designed proteinsLINK
- EXCLUSIVE: Broadcom to lend Anthropic up to $42 billion to lease its chips, filing saysLINK
- AI agents inadvertently leak 13,000+ internal screenshots from organizationsLINK
- Scale-up interconnect startup CScale launches with $188M in fundingLINK
- Cohereβs faster query model barely dents retrieval quality in its testsLINK
π§° Trending tools
LUCI Desktop: locally records your screen history and meeting transcripts, letting AI agents like Claude Code and Cursor retrieve forgotten pages and past decisions.LINK
Semos.βai Manager Agents: analyzes your meetings to flag overdue feedback, missed recognition, and avoided conversations, helping managers prioritize what needs attention.LINK
Dots by OpenAI: always-on ChatGPT agents with their own cloud computers and browsers, connecting to 4,000+ apps to autonomously complete tasks while you review results.LINK
Chat.βsh: self-hosted help center with AI search that writes cited answers, lives on your own domain, and exports pages as markdown for LLMs.LINK
America.βgov: ask plain-language questions and get clear answers from U.S. federal agencies in one place, available in English, French, and Spanish.LINK
Dina 4.5: a native macOS app for recording, editing, and captioning video, featuring transcript-based editing, AI captions, and 8K exportsLINK
π Trending papers & reports
Selective self-coaching lets a reasoning model teach itself only at the critical moments in its own work, lifting math and science accuracy by up to ~3 points for roughly 3% extra training cost.LINK
Agent handoffs let one AI pass its already-digested context notes to a different model family instead of re-reading everything, making teamwork up to ~11x faster with no retraining and matching normal quality.LINK
Synthetic textbooks show that organizing AI training material into full, coherently structured books, not just rewritten snippets, lifts model performance by ~1 point, proving how you package data matters as much as what it says.LINK
Smart replanning timing teaches AI agents to decide for themselves how many steps to run before stopping to rethink, boosting task success by ~3 to 16 points while making fewer decisions overall.LINK
Ethical personas in AI training show that teaching a model to be safe on just a few topics reliably makes it behave safely across the board, consistently adopting the broad moral stance it was trained on.LINK
See you tomorrow for a new dose of βοΈ AIpresso!