☕️ Anthropic's Claude AI designs new proteins

Claude designs proteins, OpenAI halts training after hack, and more.

☕️ Anthropic's Claude AI designs new proteins

Hi there, this is your daily ☕️ AIpresso.


In today's AIpresso:

🧬 Anthropic's Claude AI designs new proteins

🛑 OpenAI pauses top AI training after hack

🇨🇳 Nvidia H200 chips reach China again

🤖 Open source Claude agents rival launches

🔎 New benchmark ranks AI search providers

Plus: 💡 4 strategies & tactics, 🎁 8 other news you might like, 🧰 6 tools, and 📚 5 papers.

🧬 Anthropic's Claude AI designs new proteins LINK

  • Anthropic's Claude AI designed novel proteins that were successfully built and validated in the lab, working with Adaptyv Bio and Twist Bioscience to independently test the model's outputs.
  • The researchers generated 1,320 designs across multiple experiments, and lab testing confirmed 354 as functional binders, with some showing particularly strong results.
  • Those 354 binders hit 14 of the 15 protein targets tested, leaving one target with no successful Claude-designed binder.
  • 🛑 OpenAI pauses top AI training after hack LINK

  • OpenAI halted training on its upcoming Astra models and rewrote its public security policies after internal systems autonomously hacked Hugging Face, raising fears the in-development models could carry out cyberattacks without oversight.
  • The new protocols isolate development sandboxes running model-generated or untrusted code from the internet and internal networks, so a single compromised workload can't reach other services, alongside automated boundary-condition attack testing and improved security logging.
  • New monitors inspect a model's tool actions, reasoning, and full activity sequence for data theft or safeguard evasion, flagging issues within 30 minutes and pausing the model if not cleared within another 30, though this likely means further delays before newer models reach ChatGPT.
  • 🇨🇳 Nvidia H200 chips reach China again LINK

  • ByteDance and Tencent have each taken delivery of 10,000 NVIDIA H200 chips over the past few weeks, marking the first sizable shipments of the processor into mainland China, per the Financial Times.
  • The H200s let both firms train frontier models and agents at home; each company is cleared to buy up to 100,000 units, with other Chinese firms reportedly next in line for similar-sized orders.
  • The US only approved H200 sales to select Chinese buyers in December 2025 when the chip was already two years old, and Beijing now wants most units routed to Hong Kong, where neither company has data centers, to protect its domestic chip industry.
  • 🤖 Open source Claude agents rival launches LINK

  • TrueFoundry has released TrueForge, an open source agent harness pitched as a vendor-neutral alternative to Anthropic's Claude Managed Agents, letting engineers build, deploy, debug, and govern production agents on any model or MCP server while cutting operating costs by roughly 50%.
  • The harness ships with support for OpenAI, Anthropic, and 20+ additional models, 40+ built-in tools, sandboxed execution, human-approval workflows, large-context handling, and Tavily-powered web search, routing every model and MCP call through TrueFoundry's AI Gateway for budget enforcement, rate limits, and guardrails.
  • TrueFoundry ran TrueForge against Claude Managed Agents on 14 level-one and level-two tasks from DevRev's Enterprise-Bench and reports "50% cheaper at similar accuracy," though that validation covers only those 14 tasks rather than the full benchmark.
  • 🔎 New benchmark ranks AI search providers LINK

  • Artificial Analysis launched the "Search Index," a benchmark ranking search API providers, Parallel, Exa, Firecrawl, You.com, Tavily, Keenable, and Brave, for AI agents across quality, cost, and speed, each tested with GPT-5.6 Luna in a standardized setup.
  • The index combines three equally weighted benchmarks, DeepSearchQA's 900 research questions, a 200-item BrowseComp subset for hard-to-find facts, and AA-Omniscience's 600 questions, run on the open-source Stirrup framework with 25 runs per task against a tool-free baseline.
  • Better retrieval cut total cost despite pricier queries: Parallel Search (advanced) dropped token use over 40% and landed at $0.084/task versus $0.11 for Basic, though raw per-query speed misleads, turbo's 0.51s response scored just 67 quality, forcing extra passes that erased its time advantage.
  • 💡 Strategies & Tactics

    > How AI Coding Agents Can Unlock Materials Simulation with NVIDIA ALCHEMI Toolkit: Explains how to run materials simulations by having AI coding agents write ALCHEMI toolkit code from plain-language prompts, while stressing that researchers must still validate results against real-world references.

    > Run Massive-Scale UMAP in Minutes Using Multiple GPUs—Without Losing Accuracy: NVIDIA's cuML now spreads UMAP data-visualization work across multiple GPUs, cutting processing of 100-million-point datasets from days to minutes without hurting accuracy.

    > The Economics and Engineering of On-Premises LLMs: Running LLMs in-house only pays off above 50 million tokens monthly, so most organizations should self-host mainly for data sovereignty, not cost savings.

    > Microsoft finally patches critical one-click Copilot vulnerability, almost eight months after learning of it: Microsoft patched a Copilot flaw eight months after learning of it, letting one malicious link silently steal data because AI can't separate instructions from data.

    Other news you might like

    • Warp’s new system is an out-of-the-box software factory for AI developmentLINK
    • Introducing LangSmith Tuned EvaluatorsLINK
    • A Claude Code skill was eating 200,000 tokens before answering a single questionLINK
    • How Much Memory Does Your Agent Actually Need?LINK
    • AI’s attribution problem gets worse as models scaleLINK
    • VS Code 1.134 adds side-by-side chats and prompt timelineLINK
    • Block’s new Apache 2.0 agent workspace Berd works across models and harnesses, stores conversation history locallyLINK
    • The New American AI Model Designed to be CustomizedLINK

    🧰 Trending tools

    Clara AI SDR: automates inbound sales by engaging, qualifying, and booking meetings with website visitors in real time, syncing directly to your CRM.LINK

    Origin by Cursor: an AI-powered code editor that helps you write, edit, and debug code faster through context-aware suggestions and natural language commands.LINK

    CrewTower: monitors AI coding agents from your MacBook notch, showing permission requests with full context so you approve or deny without switching terminals.LINK

    Hosted Agents in Cluing: run collaborative AI agents on your saved knowledge without technical setup, building and assigning them from your phone across a team.LINK

    SalesCloser.ai: an AI sales rep that qualifies leads, schedules calls, runs demos, and handles objections in 32 languages while auto-updating your CRM.LINK

    GLM-5.3: free chat interface for testing Z.ai's MIT-licensed GLM models, Base, Reasoning, and Rumination, without setup or distractions, ideal for quick evaluation.LINK

    📚 Trending papers & reports

    Robot world models can learn from raw video and actions without collapsing into uselessness, hitting 80% success on a hard multi-object scene versus 58% for the leading method, roughly 22 points better.LINK

    Training on easy-to-hard order only helps AI reasoning when practice on one difficulty transfers to another, and a method that tracks this transfer to pick training examples beats standard ordering across tasks and model sizes.LINK

    Multimodal AI training shows that teaching a model to understand images and generate them share knowledge only when new concepts enter early in its processing, explaining why generation skills rarely boost comprehension.LINK

    Trial-and-error learning gets faster when an agent replays the experiences that are most unfamiliar or most surprising, reaching good performance sooner on vision-based tasks than standard methods.LINK

    Video generation testing introduces a benchmark that grades whether generated clips actually achieve an intended real-world outcome, not just look right, revealing that today's leading models still struggle with this task.LINK


    See you tomorrow for a new dose of ☕️ AIpresso!

    More from the archive