Google's new AI model fits in your pocket

Google's on-device AI, OpenAI's math breakthrough, and more.

Google's new AI model fits in your pocket

Hi there, this is your daily โ˜•๏ธ AIpresso.

In today's AIpresso:

๐Ÿ“ฑ Google's new AI model runs on phones

๐Ÿงฎ OpenAI's math conquest rattles mathematicians

๐ŸŒ Google's Nano Banana 2.1 costs half as much

๐Ÿ‡ช๐Ÿ‡บ Mistral unveils "le Chonk" open AI model

๐Ÿค– Meta, Sierra launch AI agent sign-in standard

Plus: ๐Ÿ’ก 5 strategies & tactics, ๐ŸŽ 5 more stories you might like, ๐Ÿงฐ 6 tools, and ๐Ÿ“š 5 papers.

๐Ÿ“ฑ Google's new AI model runs on phones LINK

  • Google released EmbeddingGemma 2, an open 740M-parameter embedding model that maps text, code, images, audio and video into one space and runs entirely on-device, needing about 191MB for text work on a Pixel 11 Pro.
  • Only 270M parameters load for text, with the vision and audio encoders optional, and it handles 8,192 tokens per modality, four times the first version, or roughly five and a half minutes of audio.
  • Licensed under Apache 2.0, it enables a retrieval pipeline with no network hop, and vectors can be cut to 128 dimensions for up to sixfold storage savings, though the model has had no safety tuning or moderation and performance varies across its many supported languages.

๐Ÿงฎ OpenAI's math conquest rattles mathematicians LINK

  • OpenAI released 722 math manuscripts from an unreleased internal model, with experts claiming it solves many of the field's top open problems, prompting one Anthropic researcher to call it "the most significant moment in mathematical history."
  • The manuscripts, grouped into 372 families drawn from thousands of research problems, used an average of roughly three hours of ChatGPT Pro compute per result, including a claimed quasi-Riemann Hypothesis proof and integer multiplication faster than n log n.
  • OpenAI consulted the Institute for Advanced Study's advisory group on release, and the model stays private, though results are reported by individual commentators and remain unverified, with one critic expecting some not to survive scrutiny.

๐ŸŒ Google's Nano Banana 2.1 costs half as much LINK

  • Google released Nano Banana 2.1, the latest text-to-image and photo-editing model, with a standard 1K image through its API now costing $0.0336, roughly half of Nano Banana 2's $0.067, and it's rolling out across Gemini, Search's AI Mode and developer tools.
  • It scored 1,050 ELO on overall text-to-image preference versus 990 for Nano Banana 2, handles up to 14 reference images while tracking four characters and ten objects, and outputs up to 4K in aspect ratios as wide as 8:1.
  • Developers can tune how long it "thinks" before drawing and ground results in Google Search, while batch jobs cut cost another 50%, though a 4K image still runs $0.0756 against a standard 1K's $0.0336.

๐Ÿ‡ช๐Ÿ‡บ Mistral unveils "le Chonk" open AI model LINK

  • Mistral released a preview of Large 4, a 1T-parameter open-weight model nicknamed "Le Chonk," which the company claims is the strongest open system outside China, available via its API today with weights to follow on October 27.
  • The multimodal model activates just 49B of its trillion parameters per task and leads on agentic coding, scoring 62% on DeepSWE, narrowly ahead of Zhipu's GLM-5.3, while its vision grounding edged out GPT-6 Astra on the DIOR-RSVG test.
  • Once weights ship, banks and governments can run Large 4 on their own machines for daily security scanning that closed US models often price out or refuse, though all current benchmark numbers are preliminary and expected to change before release.

๐Ÿค– Meta, Sierra launch AI agent sign-in standard LINK

  • Meta and Sierra, with partners including Walmart and Stripe, unveiled the Personal Agent Protocol yesterday, a proposed open standard for how personal AI agents authenticate and get permissions when acting on behalf of users at businesses.
  • Sessions run on OAuth: an agent starts on a company's site, may begin as a guest, and gains account access only once the customer signs in and grants read-only or write access, with the business setting what the agent can do.
  • The v0.1 spec is due later this month, but no specification, license, or governing body has been published yet, and OpenAI and Anthropic are absent from the partner list despite Sierra chair Bret Taylor expecting them to join.

๐Ÿ’ก Strategies & Tactics

> Modernizing Table Batched Embeddings with FBTriton: Rewriting Meta's GPU recommendation-lookup code in Triton beats the old CUDA version by a median 1.28x while staying far easier to modify.

> Memory and dreaming: how Devin learns from working with you: Devin's AI coding assistant automatically saves preferences and corrections across sessions, then refines them nightly, so lessons carry forward without manual upkeep.

> Control How Your GPU Shares Work with Green Contexts: Dedicate specific GPU compute units to each workload with NVIDIA's green contexts, preventing a heavy background task from delaying a time-critical one.

> How DOCA GPUNetIO Unifies GPU-Initiated Networking Across the NVIDIA Software Stack: NVIDIA's GPUNetIO lets GPUs drive network transfers directly without the CPU, giving multiple communication libraries one shared codebase to maintain instead of separate overlapping versions.

> One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO: Fine-tuning one base model into math and coding specialists, then pairing each with a generate-and-refine loop, reaches gold-level results at both competitions.

Other news you might like

  • Anthropic opens its most powerful AI models to more security teamsLINK
  • SpaceX seeks $40 billion to buy Nvidia chipsLINK
  • AI computing startup Lambda to raise $4B ahead of planned IPOLINK
  • You can now open Claude directly in Google Docs, Sheets and Slides (and vice versa, open and edit the files in Claude)LINK
  • OpenAI in talks with UAE funds and BlackRock for $30B round at a fixed $1.4T price: ReportLINK

๐Ÿงฐ Trending tools

Pegasus 1.6 by TwelveLabs: multimodal video model that searches, analyzes, and generates text from footage, turning raw video into structured datasets without manual annotation pipelinesLINK

Dots by OpenAI: always-on ChatGPT agents with dedicated cloud computers and browsers that autonomously handle tasks across 4,000+ apps while you review results.LINK

Nano Banana 2.1: searches webpages, images, videos, and more, with specialized tools and filters to help you find exactly what you need quickly.LINK

ParakeetAI 2.0: provides real-time answers during meetings and interviews, helping you respond accurately without pausing to search or second-guess yourself.LINK

Viibeo: turns a prompt, URL, or script into publish-ready faceless shorts with voiceover, B-roll, and timed captions, plus autopilot posting and ad creation.LINK

Thalia: native macOS app for agentic coding with Meta's Muse Code CLI, letting you direct work, inspect diffs, and talk hands-free on Apple Silicon.LINK

๐Ÿ“š Trending papers & reports

Making AI forget gets harder when a model learned by generalizing rather than memorizing, as erasing targeted knowledge from such models does more collateral damage to the data you wanted to keep.LINK

Reasoning transplants let image-understanding AI inherit the problem-solving skills a text model learned, without new training, by copying only the most transferable parts and lifting math-vision scores by ~8.6 points.LINK

Adam tuning picks the one memory setting that controls how much an AI trainer leans on past progress using a quick 200-step warmup, cutting the average performance gap by ~41% across eleven vision and language tasks.LINK

Image generation timing can now be pinpointed feature by feature, measuring exactly when a model fills in broad shapes versus fine details, revealing that different internal setups build pictures in meaningfully different orders.LINK

Outage troubleshooting gets an AI assistant that pinpoints what broke in large online services up to 62% more accurately and 12x cheaper, while explaining its reasoning so operators can actually act on it.LINK

See you tomorrow for a new dose of โ˜•๏ธ AIpresso!

More from the archive