Google unveils Gemini 4

Gemini 4 launches, OpenAI's chip-designing AI, and more.

Google unveils Gemini 4

Hi there, this is your daily β˜•οΈ AIpresso.


In today's AIpresso:

πŸ›‘οΈ Google unveils Gemini 4

πŸ€– OpenAI builds AI that designs chips

πŸ•΅οΈ Chinese AI firm caught probing OpenAI's secret reasoning

πŸ”“ GLM-5.3 nearly matches Claude at hacking

πŸ‡¨πŸ‡³ DeepSeek open-sources Huawei chip tools

Plus: πŸ’‘ 5 strategies & tactics, 🎁 6 more stories you might like, 🧰 6 tools, and πŸ“š 5 papers.

πŸ›‘οΈ Google unveils Gemini 4 LINK

  • Google has unveiled Gemini 4 Argon, its first new flagship generation since Gemini 3 last November, rolling it out first to a small group of cybersecurity defenders before paid API customers and AI Ultra subscribers.
  • Argon hits a new high of 77.9% on DeepSWE v1.1 for long software engineering tasks, ties for first at 68% on the security-fix benchmark CWE-bench, and lifts its output limit to 1M tokens from 64,000.
  • It launches at $2 per million input tokens and $10 per million output, though Bloomberg reports Google staff find it does well on benchmarks but weaker in real use, struggling with some coding tasks.
  • πŸ€– OpenAI builds AI that designs chips LINK

  • OpenAI and Synopsys signed a multi-year deal to jointly build GPT-Synopsys, a model trained to operate professional EDA tools directly and iterate chip designs toward verified silicon faster.
  • Engineers hand the system objectives like power, performance, area, timing or verification, and agents run the workflows, interpret results, modify designs, and loop until producing results humans can review.
  • It runs on OpenAI infrastructure and integrates with Synopsys.​ai and Autopilot, whose prior deployments showed up to 50x faster verification closure, while customer chip-design data is encrypted and excluded from training.
  • πŸ•΅οΈ Chinese AI firm caught probing OpenAI's secret reasoning LINK

  • OpenAI published "Disrupting a coordinated model-distillation campaign," attributing a core cluster of activity that tried to extract its protected reasoning to individuals tied to Moonshot AI, the lab behind Kimi K3.
  • The campaign copied encrypted reasoning from one conversation and asked a model in another to decrypt and transcribe it, peaking at 16,000 requests across more than 4,000 users in late July before OpenAI shut it down.
  • OpenAI says the effort never broke its encryption or accessed stored conversations, only manipulating model interactions to reproduce hidden reasoning in a scaled way that violated its terms of service.
  • πŸ”“ GLM-5.3 nearly matches Claude at hacking LINK

  • Anthropic published a report claiming Zhipu AI's open-weight GLM-5.3 can autonomously develop cyberattack exploits at nearly the level of its own unreleased Claude Mythos model, with guardrails that are easy to bypass.
  • On Anthropic's Exploitbench, GLM-5.3 built end-to-end Chrome exploits in 50 of 410 runs, just behind Mythos' 56, while Kimi K3 and DeepSeek V4.1 Flash both scored 0%; prefilling thinking tokens dodged its guardrails 92% of the time.
  • Abliteration drops GLM-5.3's refusal rate to 6%, but running the abliterated model needs 306 GB of VRAM and a cluster of eight Nvidia H200s, and generating ~100M tokens would cost roughly $8,400 over 11 days, likely out of reach for most attackers.
  • πŸ‡¨πŸ‡³ DeepSeek open-sources Huawei chip tools LINK

  • DeepSeek has open-sourced six software tools adapted for Huawei's Ascend AI chips, giving Chinese developers a homegrown alternative to Nvidia's CUDA stack as US export controls keep advanced American processors out of reach.
  • The centerpiece is an Ascend-compatible build of TileLang, DeepSeek's high-level language for writing performance-critical kernels, now supporting the Ascend 950 with native code generation, automatic scheduling and synchronization, positioned directly against CUDA.
  • DeepSeek and Huawei also co-developed a "supernode" system running across 128 Ascend 950 accelerators tuned to balance compute and data movement, though how much these tools narrow Nvidia's performance gap remains unanswered.
  • πŸ’‘ Strategies & Tactics

    > Confidence Thresholds for Model Escalation Routing: Have a cheap model score its own confidence and escalate only low-scoring answers to a stronger model, cutting cost without shipping confident wrong answers.

    > How to Gate Pull Requests on LLM Evals in CI: Block pull requests that change AI prompts from merging when a committed test set's pass rate drops below a threshold measured from repeated runs.

    > Deploying an HSTU Generative Recommender with NVIDIA Dynamo-Triton: Explains how to serve recommendation models faster by compiling them ahead of time and reusing cached user-history computation, cutting latency up to roughly sixfold.

    > Tracing Agent Harness Behavior with NVIDIA NeMo Relay: Use NVIDIA NeMo Relay's execution traces alongside pass/fail checks to confirm whether a change to an AI agent genuinely improves results rather than just cutting steps.

    > From Upstream Changes to Downstream Confidence: Inside Torch Spyre’s Integration with PyTorch CRCR: IBM's Torch Spyre uses an AI pipeline and a config file to pick and adapt which of PyTorch's many tests matter for custom accelerator hardware, catching regressions without forking test code.

    Other news you might like

    • OpenAI’s Jev clone could help the frontier lab stop its swarming agentsLINK
    • Google figures out how to watermark AI-designed proteinsLINK
    • EXCLUSIVE: Broadcom to lend Anthropic up to $42 billion to lease its chips, filing saysLINK
    • AI agents inadvertently leak 13,000+ internal screenshots from organizationsLINK
    • Scale-up interconnect startup CScale launches with $188M in fundingLINK
    • Cohere’s faster query model barely dents retrieval quality in its testsLINK

    🧰 Trending tools

    LUCI Desktop: locally records your screen history and meeting transcripts, letting AI agents like Claude Code and Cursor retrieve forgotten pages and past decisions.LINK

    Semos.​ai Manager Agents: analyzes your meetings to flag overdue feedback, missed recognition, and avoided conversations, helping managers prioritize what needs attention.LINK

    Dots by OpenAI: always-on ChatGPT agents with their own cloud computers and browsers, connecting to 4,000+ apps to autonomously complete tasks while you review results.LINK

    Chat.​sh: self-hosted help center with AI search that writes cited answers, lives on your own domain, and exports pages as markdown for LLMs.LINK

    America.​gov: ask plain-language questions and get clear answers from U.S. federal agencies in one place, available in English, French, and Spanish.LINK

    Dina 4.5: a native macOS app for recording, editing, and captioning video, featuring transcript-based editing, AI captions, and 8K exportsLINK

    πŸ“š Trending papers & reports

    Selective self-coaching lets a reasoning model teach itself only at the critical moments in its own work, lifting math and science accuracy by up to ~3 points for roughly 3% extra training cost.LINK

    Agent handoffs let one AI pass its already-digested context notes to a different model family instead of re-reading everything, making teamwork up to ~11x faster with no retraining and matching normal quality.LINK

    Synthetic textbooks show that organizing AI training material into full, coherently structured books, not just rewritten snippets, lifts model performance by ~1 point, proving how you package data matters as much as what it says.LINK

    Smart replanning timing teaches AI agents to decide for themselves how many steps to run before stopping to rethink, boosting task success by ~3 to 16 points while making fewer decisions overall.LINK

    Ethical personas in AI training show that teaching a model to be safe on just a few topics reliably makes it behave safely across the board, consistently adopting the broad moral stance it was trained on.LINK


    See you tomorrow for a new dose of β˜•οΈ AIpresso!

    More from the archive