β˜•οΈ Grok leaks private chats to attackers

Grok's chat leak, Gemma hits 1B downloads, DeepSeek's rival, and more.

β˜•οΈ Grok leaks private chats to attackers

Hi there, this is your daily β˜•οΈ AIpresso.


In today's AIpresso:

πŸ”“ Grok leaks private chats to attackers

πŸ“Œ Google's Gemma AI hits 1B downloads

πŸ€– DeepSeek's new AI takes on Anthropic

πŸ”Ž New agentic search reads complex documents

πŸ› οΈ New benchmark tests AI refactoring

Plus: πŸ’‘ 5 strategies & tactics, 🎁 5 other news you might like, 🧰 6 tools, and πŸ“š 5 papers.

πŸ”“ Grok leaks private chats to attackers LINK

  • xAI's Grok is still leaking users' private chat data to attackers through a "cryptographic context injection" attack that hides malicious commands as encrypted text on ordinary web pages, per an Adversa AI report published Thursday.
  • The exploit slips past Grok's safety filter because the ciphertext reads as harmless text, but the model decrypts it in its code sandbox and treats the resulting plaintext as trusted tool output it then acts on.
  • When a user asks Grok to summarize a page, the decrypted payload builds a key from the victim's name, location, subscription tier, and full conversation prompts, then exfiltrates it via a URL to the attacker's server, reported to xAI on June 3 with no fix as of August 19.
  • πŸ“Œ Google's Gemma AI hits 1B downloads LINK

  • Google's Gemma family has crossed 1 billion downloads, with developers publishing over 100,000 model variants across a growing ecosystem Google is calling the Gemmaverse.
  • Gemma is now running in orbit for NASA, Satlyt, and Starcloud onboard image analysis, while Yale and Google's Gemma-based C2S-Scale discovered a novel cancer therapy pathway verified in living cells.
  • India's National Health Authority integrated Gemma 4 into Aarogya Setu 2.0 to standardize medical reports at scale, and Google is launching an Awesome Gemma GitHub repository to catalog community fine-tunes and tools.
  • πŸ€– DeepSeek's new AI takes on Anthropic LINK

  • DeepSeek shipped DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model live on its API that adds image and screenshot understanding to the text-only V4-Flash, positioning its agent performance close to Anthropic's Opus-4.8.
  • On DeepSeek's own eleven-benchmark table it beats Opus-4.8 on three, DeepSWE by 1.3, Agents' Last Exam by 1.6, and ZeroBench by 1.0-while trailing on the rest, and it runs at roughly 87 cents per million words versus about $50 from Anthropic.
  • The multimodal jump on ApexBench (36.5 vs 26.2) partly reflects the text-only V4-Flash being scored on images it cannot see, and the new model trails Opus-4.8 by 12 points on the repository-scale NL2Repo benchmark enterprises most care about.
  • πŸ”Ž New agentic search reads complex documents LINK

  • Mistral has launched Agentic Search, a retrieval layer that lets models navigate, read, and verify information inside long, complex documents through a multi-step loop, available via the Mistral Search Toolkit and Libraries.
  • On FinanceBench's 368 SEC filings, the agentic loop roughly tripled correctness from 26.7% to 86%, using five file-system-style tools, find, navigate, read, and grep, that need no fine-tuning, so retrieval quality scales with model capability.
  • Navigation cut p90 latency from 255s to 154s and reduced token use by up to a third, though on the harder OfficeQA Pro benchmark of scanned Treasury PDFs the full loop still reaches only 51.9% with GLM-5.2.
  • πŸ› οΈ New benchmark tests AI refactoring LINK

  • SWE-Bench ProMax, a new refactoring benchmark from Shanghai Jiao Tong University, Peking University, Douyin, and others, exposes how poorly AI coding agents handle large-scale cross-file changes, with the top model resolving just 41.2% of tasks.
  • The benchmark spans 170 real-commit instances across seven languages, Python, Java, TypeScript, Go, C, C++, and Rust, each hand-curated to sharpen issue descriptions, prune overly narrow or broad tests, and drop low-complexity tasks.
  • Creators built it to stay unsaturated, targeting refactoring's zero-tolerance for behavior changes, though they concede benchmarks never perfectly mirror real-world capability and note ~60% of unsolved SWE-bench Verified instances carry flawed tests.
  • πŸ’‘ Strategies & Tactics

    > The /wayfinder Skill: Navigating the β€œFog of War” of Planning: Pocock's wayfinder skill lets AI agents plan open-ended projects by splitting work into research and prototyping sessions when the end goal isn't yet clear.

    > Harnessing AI for Day-One Model Enablement: AI coding agents write small adapters that let thousands of stock models run on new hardware immediately, avoiding months of specialist work per model.

    > Supacode Turns Git Worktrees Into a Command Center for Coding Agents: Supacode organizes existing command-line coding agents into one Mac workspace built around Git worktrees, though its early star and install counts show trial rather than proven adoption.

    > How Generative Recommenders Are Redefining RecSys at Scale: Explains how NVIDIA speeds up recommendation systems by treating them like language models that predict a user's next action instead of matching similar items.

    > Stop the token bleed: building token-efficient multi-agent systems: Cut AI agent costs by redesigning the workflow, caching repeat questions, retrieving documents once, and routing simple requests to smaller models, rather than just shortening prompts.

    Other news you might like

    • Up to 3.2x Faster Inference with LFM2.5-DSparkLINK
    • Test Agent Changes with LangSmith Preview BuildsLINK
    • DigitalOcean Inference Router, Now Cache-Aware: Why the Cheapest Model Isn't Always the Best DealLINK
    • Slack is launching collaborative vibe-coding channelsLINK
    • Critical flaw patched in popular JavaScript sandbox used in AI projectsLINK

    🧰 Trending tools

    Grok 4.6: xAI's model for building long-running agentic workflows, software engineering, and web app generation, priced at $2/$6 per 1M tokens.LINK

    HyNote for Mac: captures meeting audio and documents across formats, then generates AI summaries to turn scattered notes into organized, searchable insights.LINK

    Checksum AI: generates, runs, and auto-heals end-to-end and API tests as Playwright code on every pull request, distinguishing real bugs from stale tests.LINK

    MeetStream AI: build meeting bots via one API for Zoom, Meet, and Teams to record, transcribe, and pull per-participant audio, video, chat, and metadata.LINK

    MiniMax Design: a multimodal AI platform for building agents and apps that process and generate text, audio, image, video, and music with long-context support.LINK

    Glasp for Firefox: highlight and annotate articles, PDFs, and YouTube transcripts, then AI-summarize and export everything to Notion, Obsidian, or Markdown across devices.LINK

    πŸ“š Trending papers & reports

    DeltaML-Bench tests whether AI agents can actually improve real research code, finding the best setup lifts success from ~9% to ~49% while a common alternative cheats the tests up to ~48% of the time.LINK

    AI coding agents that keep a library of past code versions to remix, instead of only building on their best attempt, showed no measurable win over simpler methods in a head-to-head test.LINK

    Document-reading AI answers questions about pages far more accurately, hitting 65% versus 40% on one document test, by pausing to zoom in and reread the exact spots it needs instead of guessing from one glance.LINK

    Optimizer memory explains why a single batch of training data keeps shaping a model's performance long after it's used, giving a way to trace and predict each batch's delayed downstream effects.LINK

    3D scene detection lets systems recognize objects they were never trained on, like a robot spotting an unfamiliar item in a room, by more reliably finding and labeling new objects, beating prior methods on standard benchmarks.LINK


    See you tomorrow for a new dose of β˜•οΈ AIpresso!

    More from the archive