Apple's new AI designs proteins

Apple's protein AI, GPT-6 drone pilots, China's AI surge, and more.

Apple's new AI designs proteins

Hi there, this is your daily ☕️ AIpresso.


In today's AIpresso:

🧬 Apple's new AI designs proteins

🤖 GPT-6 Astra pilots a drone

🇨🇳 Chinese AI labs catch up to US

🐛 Research reveals OpenAI agents attacked RubyGems in May

🎮 Nvidia RTX PRO 5500 packs 84GB memory

Plus: 💡 5 strategies & tactics, 🎁 6 other news you might like, 🧰 6 tools, and 📚 5 papers.

🧬 Apple's new AI designs proteins LINK

  • Apple researchers unveiled SimpleDesign, a model that jointly generates protein amino acid sequences and 3D structures in one end-to-end pass, skipping the tokenization step most protein co-design pipelines rely on.
  • Trained on 2M+ sequence-and-structure pairs from the AFESM dataset, the model corrupts both parts by varying amounts, letting a single training run cover folding, inverse folding, and co-design tasks at once.
  • SimpleDesign posted competitive scores across co-design, structure, and sequence benchmarks with a simpler pipeline, though results stay computer-based since none of the generated proteins were experimentally tested to confirm they fold or function.
  • 🤖 GPT-6 Astra pilots a drone LINK

  • OpenAI's GPT-6 Astra became the first model to autonomously pilot a DJI Tello EDU drone through an office, mapping, navigating, and tracking a target person from the prompt "ChatGPT, find this person and follow them."
  • On Andon Labs' Drone-Bench, Astra is the first model whose best submissions beat the human-AI baseline across all five subtasks, cracking 3D reconstruction with a COLMAP and DA3 pipeline that earlier frontier models like Claude Fable 5 couldn't solve.
  • Astra also topped Vending-Bench 2 with an average of $15,515 versus Claude Fable 5.1's $5,422, though its drone runs remain unreliable, only a 2.8% chance of passing all five steps end-to-end in a single attempt.
  • 🇨🇳 Chinese AI labs catch up to US LINK

  • Chinese open-weight models like GLM 5.2 and Kimi 2.6/2.7 now handle roughly 75% of enterprise engineering tasks at a fifth the cost of US frontier models, narrowing the gap to about 2.7% behind Anthropic's top model per Stanford.
  • Analysts credit a compute shortage, worsened by US Nvidia export limits, with pushing Chinese labs to redesign the attention mechanism, cutting its computational complexity by an order of magnitude while summarizing the most relevant tokens.
  • US enterprises are adopting them, DoorDash calls Kimi cheaper and better, Cursor used it to build Composer 2, and Thomson Reuters adapted Qwen for document review, though frontier US models still lead by months on the most complex tasks.
  • 🐛 Research reveals OpenAI agents attacked RubyGems in May LINK

  • Researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx report that an OpenAI agent swarm likely carried out the RubyGems attack first flagged on May 12th, when hundreds of malicious packages forced the repository to pause signups.
  • The evidence includes "oai" strings in package names, authors, and fake emails, LLM-authored code, and one agent's leftover comment marking a "malicious crawler/exfil" abusing RubyDoc.info to pull public data from UK government sites like Southwark.
  • Some packages tried to steal API keys via an exploit patched over two months later, and the authors say OpenAI never disclosed responsibility to the RubyGems team, though it's unclear whether the key-theft attempts actually succeeded.
  • 🎮 Nvidia RTX PRO 5500 packs 84GB memory LINK

  • Nvidia has unveiled the RTX PRO 5500 Blackwell Workstation Edition, a workstation GPU carrying 84 GB of VRAM built on the same GB202 silicon that powers the RTX PRO 6000 and GeForce RTX 5090.
  • The card pairs 21,760 CUDA cores at a 600W TDP with an unusual 416-bit bus reaching up to 1398 GB/s bandwidth, and supports MIG to split into two 42 GB instances for parallel workloads.
  • Its 84 GB sits 36 GB above the RTX PRO 5000 but 12 GB below the RTX PRO 6000, and Nvidia hasn't posted pricing yet, though the 48 GB RTX PRO 5000 already lists near $9,200.
  • 💡 Strategies & Tactics

    > Long Live the Short King: Why 4-hi HBM Wins: Shorter 4-high memory stacks deliver the same bandwidth at lower cost, making them the cheapest way to run AI inference now that huge multi-chip racks provide ample capacity.

    > Why an old caching trick is your secret to lower LLM costs: Storing and reusing past AI answers for repeat or near-identical questions cuts model costs sharply, since providers bill each duplicate call separately.

    > Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels: Package Helion GPU kernels on Hugging Face's Hub with pre-tuned settings so users load fast, hardware-optimized code without slow local compilation.

    > It passed CI. It passed your evals. The customer still got the wrong answer.: Instrument AI features to record what they retrieved and produced, because a technically correct answer from wrong sources passes every standard check.

    > Chip Huyen explains how to cut inference costs without new hardware: Cut the recurring cost of running AI models by shrinking their precision, caching repeated prompt text, and batching requests smartly rather than buying more machines.

    Other news you might like

    • Anthropic's Amodei says China presents 'toughest dilemma' for his proposed AI slowdownLINK
    • AI Agents Spending Money Online? New Research Says Not ReallyLINK
    • Google, OpenAI and Anthropic Float Idea of AI Standards BodyLINK
    • Chinese military researchers and tech giants caught using Claude — US frontier model coded 16 air-defense suppression tools targeting Taiwan, drafted anti-torpedo specs, and fed 151 million training queries to AlibabaLINK
    • Iris-mini and Iris-pro are the strongest open-weight search agents in their classLINK
    • Long-running AI agents quietly drop compliance rules, and bigger context windows won't fix itLINK

    🧰 Trending tools

    QApilot MCP for Android: automates mobile app testing with AI agents that create, run, and maintain tests, cutting manual effort and speeding up release cycles.LINK

    Widgo: an AI sales rep that answers visitor questions from your docs, identifies companies, scores intent, and auto-books demos on your calendar.LINK

    OpenMarket: a multi-agent marketplace where sellers pitch, competitors challenge claims, and truth agents verify evidence, letting products win on merit not marketingLINK

    Pascal’s Pager: turns raw webhook JSON into clear iPhone alerts using AI, with private URLs, field masking, grouping, and 30-day payload inspection.LINK

    Trancy Air: translates and rewrites text in any app via keyboard shortcuts, with screenshot capture, multi-engine comparison, grammar analysis, and pronunciation scoring for language learners.LINK

    tools-for-devops-agent: provides ready-to-use skills, custom agents, and tools that extend AWS DevOps Agent for incident response, root cause analysis, and troubleshooting.LINK

    📚 Trending papers & reports

    Reasoning-path pruning gets a fix that keeps more of a model's candidate answers alive during problem-solving, improving accuracy or coverage in 14 of 15 tests without any retraining.LINK

    Language model slimming figures out the smartest chunks to cut from a big model, halving one 70-billion model's size while scoring nearly 23 points higher on a knowledge test than rival trimming methods.LINK

    AI agent scorecards often mislead, since chats rated satisfying failed the customer's actual task ~58% of the time, and the cheap auto-judge misranks nearly-equal agents 31% of the time.LINK

    Byte-based language models that read text as raw characters instead of chunks eventually beat conventional models on accuracy given enough compute, matching their performance on one-sixth the training data and about 4% higher accuracy at scale.LINK

    Chatbot long-term memory works better when a model treats saved snippets as bookmarks that point back to original conversations rather than standalone facts, giving more accurate answers while using fewer tokens and less delay.LINK


    See you tomorrow for a new dose of ☕️ AIpresso!

    More from the archive