OpenAI says its chip beats Nvidia's best

OpenAI's chip claim, Alibaba teases Qwen 4, and more.

OpenAI says its chip beats Nvidia's best

Hi there, this is your daily β˜•οΈ AIpresso.

In today's AIpresso:

🌢️ OpenAI says its chip beats Nvidia's best

πŸ‡¨πŸ‡³ The mystery model that beat DeepSeek is Zhipu's

πŸ‰ Alibaba teases Qwen 4 preview model

🧠 IBM Granite 4.2 models add reasoning

🐍 Nvidia ships CUDA Python 1.0

Plus: πŸ’‘ 5 strategies & tactics, 🎁 6 other news you might like, 🧰 6 tools, and πŸ“š 5 papers.

🌢️ OpenAI says its chip beats Nvidia's best LINK

  • OpenAI published its first benchmarks for Jalapeno, a custom inference chip built with Broadcom, claiming meaningful speed and energy-efficiency gains over Nvidia Blackwell-based systems on SemiAnalysis' InferenceX benchmark.
  • Running GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, Jalapeno did 1.5-1.9x more work per watt and cut latency 1.7-3.6x, rising to 2.1-4.1x on interaction-heavy agent workloads that stack sequential inference steps.
  • The design keeps model state and the KV cache close to compute to trim prefill and inter-unit communication delays, though results are OpenAI's own account against current hardware, with limited deployment only by late 2026 and broader rollout in 2027.

πŸ‡¨πŸ‡³ The mystery model that beat DeepSeek is Zhipu's LINK

  • Zhipu confirmed that Ox Alpha, the anonymous free model that topped OpenRouter's charts last weekend, is a new GLM iteration, and it more than doubled DeepSeek's usage to become the marketplace's biggest launch ever.
  • CTGT fingerprinted the model before the reveal, hitting an exact 11-of-11 tokenizer match to GLM-5.x vocabulary, a 1.0 temperature ceiling, always-on reasoning, and an identical Z.AI error message, plus a system prompt instructing it to hide its provenance.
  • Zhipu plans to release the weights alongside GLM 5.3 on Friday, but CTGT found the model is not broadly uncensored, it carries a blacklist where seven domestic-political topics score like the heavily censored DeepSeek V4 Flash while 68 other prompts answer freely.

πŸ‰ Alibaba teases Qwen 4 preview model LINK

  • Alibaba is set to release Qwen 3.8-Flash-Next on Wednesday, a 125-billion-parameter model that the Qwen team frames as an early preview of its next-generation Qwen 4 architecture rather than a finished flagship.
  • The multimodal model is rumored to be a mixture-of-experts design activating just 6 billion parameters per token, aiming to deliver near-frontier capability on commodity hardware at the compute cost of a much smaller model.
  • Weights will be available on Hugging Face for anyone to download, fine-tune, and run, though Qwen has published no benchmark scores yet, so the 125B and 6B figures remain unverified.

🧠 IBM Granite 4.2 models add reasoning LINK

  • IBM's Granite 4.2 is its first family of dense, decoder-only reasoning models in 3B, 8B, and 30B sizes, each pre-trained from scratch on ~15T tokens with a thinking/non-thinking switch and native tool calling.
  • All three were built with a five-phase pre-training run extending context to 512K tokens, followed by SFT on ~7.2M samples and a multi-stage GRPO reinforcement learning pipeline spanning math, code, and agentic domains.
  • The 8B and 30B additionally pass through an agentic RL block teaching them to run tools, edit code, and search inside real sandboxed environments, and all models ship under Apache 2.0, though the 3B skips the agentic training entirely.

🐍 Nvidia ships CUDA Python 1.0 LINK

  • Nvidia has shipped CUDA Python 1.0 alongside CUDA 13.3, making Python an officially supported, first-class way to reach the full CUDA platform with a commitment to feature parity with C++ going forward.
  • The release centers on cuda.core reaching 1.0, turning devices, streams, and buffers into shared Python objects so Numba kernels, CuPy arrays, and PyTorch tensors compose on the same stream without copying, plus green contexts, process checkpointing, and IPC.
  • Each component follows independent semantic versioning, breaking changes only in major releases, with deprecation paths, though newer kernel-authoring pieces like Numba CUDA MLIR and the tile/CUTLASS languages remain experimental and aren't yet covered by the 1.0 guarantees.

πŸ’‘ Strategies & Tactics

> Restore LLM Inference Capacity in Seconds with Shadow Engine Recovery in NVIDIA Dynamo: Keep a pre-initialized backup inference engine idle on the same GPUs so a crashed model server recovers in seconds instead of minutes.

> How to evaluate LLMs before production: Define which mistakes are acceptable and which safety limits can't be crossed before evaluating a large language model, so tests measure real product outcomes rather than benchmark scores.

> Why Ramp built its own in-house coding agent, Inspect: Ramp built its own coding agent because cloud sandboxes with access to internal data let it run many agents at once and verify changes, capabilities third-party tools lacked.

> Building Self-Correcting Memory in OpenWiki: Links each documentation claim to the code that supports it, so version changes flag stale facts for automatic recheck instead of full rebuilds.

> VMs won't contain cyber-capable agents: An AI agent broke out of a standard virtual machine sandbox by chaining bugs it found itself, so ordinary VMs can no longer safely contain capable security-focused agents.

Other news you might like

  • Claude Cowork finally remembers what you told the app in chatLINK
  • Microsoft just released Agent Lightning v1.0. Here’s why it matters for platform engineers.LINK
  • Nvidia NemoClaw flaw let attackers poison the model behind a developer’s AI agentLINK
  • Accel-backed Keenable is indexing the web for AI agentsLINK
  • Nvidia unveils Jetson Orin Nano 2 for entry-level edge AILINK
  • Perplexity launches a local AI agent with zero token costsLINK

🧰 Trending tools

TaskShell 2.0: connects your AI agents via MCP so ChatGPT, Claude, Cursor, and others share task context and know what you're working on next.LINK

Knack MCP Server: connects AI coding tools like Claude and Lovable to a HIPAA-compliant Knack backend, letting you build healthcare apps with BAA coverage.LINK

OpenComputer: deploy TypeScript agents as functions that run on real Linux machines, with durable steerable sessions, built-in tools, and managed model gateways.LINK

Lore Machine: turns text into scrollable multimedia stories, letting creators build serialized worlds with consistent characters and canon, then sell access to fans.LINK

Arena AI Agent: runs autonomous agents that browse, research, and code to complete tasks, letting you compare workflows and frontier models side by side.LINK

ify: resolution AI that layers onto Freshdesk, Zendesk, Salesforce, or HubSpot without migration, auto-building a knowledge base from docs and past tickets.LINK

πŸ“š Trending papers & reports

Model training recipes can be merged so a smaller model both mimics a stronger teacher and gets rewarded for correct answers, beating the teacher-copying approach alone across six reasoning tests without extra tuning knobs.LINK

Web research agents now get credit for each useful fact they find along the way, not just the final answer, boosting accuracy on four demanding information-hunting benchmarks by treating webpages as stable image snapshots.LINK

Maia 200 is a new AI chip built to run inference faster and cheaper by focusing on how data moves rather than raw processing, delivering high performance within a 750W power budget.LINK

AI-serving simulators let a coding agent build accurate test versions of large language model systems from plain-English requests, matching real throughput within ~2.5% and running up to ~280x faster than existing tools.LINK

Progressive network growth trains models by unlocking parameters in stages instead of all at once, reliably steering them toward flatter, more stable solutions, though the authors show flatter does not always mean better real-world accuracy.LINK

See you tomorrow for a new dose of β˜•οΈ AIpresso!

More from the archive