Salesforce, Nvidia launch CRM AI model

Salesforce and Nvidia's CRM AI model, cheaper inference, and more.

Salesforce, Nvidia launch CRM AI model

Hi there, this is your daily β˜•οΈ AIpresso.


In today's AIpresso:

🀝 Salesforce, Nvidia launch CRM AI model

πŸ› AI coding agent fails 60% of tasks

⚑ rack cuts AI inference cost 67x

πŸ–₯️ Claude Code tops rivals at web design

πŸ”§ Tool runs CUDA apps on AMD GPUs

Plus: πŸ’‘ 5 strategies & tactics, 🎁 6 other news you might like, 🧰 6 tools, and πŸ“š 5 papers.

🀝 Salesforce, Nvidia launch CRM AI model LINK

  • Salesforce and Nvidia have trained Koa, a CRM-specialized model built on Nvidia's Nemotron architecture, designed to help Agentforce agents reason through complex multistep workflows across customer relationship data.
  • Koa matched or exceeded leading general-purpose models on CRM tasks with three times fewer errors, trained via supervised fine-tuning, reinforcement learning, and GRPO on synthetic data spanning 14 industries drawn from 27 years of deployment experience.
  • The model runs inside Salesforce's trust boundaries so customer data never leaves, targeting tasks like case routing and lead qualification, though it isn't generally available yet, select customers get pilot access now, with general availability slated for winter.
  • πŸ› AI coding agent fails 60% of tasks LINK

  • Claude Fable 5.1 topped Specific Labs' new Real-SWE benchmark with just a 38.8% success rate, meaning the strongest coding agent tested still failed more than 60% of tasks drawn from real, private production codebases.
  • Real-SWE swaps public GitHub problems for proprietary code licensed from actual companies, a consumer app with 200K+ users and a fintech platform, with solutions spanning a median of 11 files, and tested each model on its native tool: Fable via Claude Code, GPT-6 Astra via Codex CLI at 33.8%, Gemini 3.8 Flash via Gemini CLI at 31.2%.
  • Individual tasks cratered further, with six of 10 below 15%, a tax jurisdiction bug fixed just 3.1% of the time and no model solving the analytics stream reducer across 64 attempts, though it's only 10 tasks, too small to conclude public benchmarks are contamination-inflated.
  • ⚑ rack cuts AI inference cost 67x LINK

  • NVIDIA's Vera Rubin NVL72 rack delivers ~67x the tokens per dollar of TCO versus GB300 on TRTLLM NVFP4 dense at 170 TPS, based on SemiAnalysis's AgentX agentic inference benchmark run on pre-release software.
  • On the DeepSeek V4 Pro agentic workload, Rubin hits ~59.4M tok/s/MW at 100 TPS, a 2.09x edge over the stronger GB300 engine and 7x GB300 SGLang at 150 TPS, while reaching 276 P90 TPS max interactivity versus GB300's 172.
  • At the 60-100 TPS range where providers actually serve this model, Rubin's advantage drops to 1.4x-3x throughput per TCO, and the gap narrows to 2.72x at 200 TPS, though open-source SGLang lets GB300 match Rubin's interactivity.
  • πŸ–₯️ Claude Code tops rivals at web design LINK

  • Claude Code with Opus 5 beat Codex and Google Antigravity in a head-to-head website build, producing the most complete food-delivery site from a single shared prompt with no follow-up instructions.
  • Opus 5 nailed the details that usually need cleanup, consistent typography, spacing, card padding, and image placement, and generated believable mobile mockups plus a live-tracking section where both GPT-5.6 and Gemini 3.8 fell apart.
  • Antigravity's Gemini 3.8 took second on fast, well-presented output, and Codex's GPT-5.6 had the strongest hero section, though Claude Code's hero lacked the floating rating and delivery details that made Codex's version stand out.
  • πŸ”§ Tool runs CUDA apps on AMD GPUs LINK

  • Speedstu's "CUDA-for-AMD-Windows" project lets rigidly CUDA-exclusive workloads run natively on AMD Radeon GPUs in Windows, without virtualization or dual-booting, by bridging the ZLUDA translation layer to AMD's HIP/ROCm SDK.
  • The PowerShell toolkit auto-detects GPU architecture, pins ZLUDA v6-preview.69, and maps the CUDA driver API plus cuBLAS, cuSPARSE, and cuFFT to AMD equivalents, proven by training a 2.2M-parameter PPO network end-to-end on a Radeon RX 9060 XT.
  • The upstream path hit a median 13,278 steps/sec versus 12,876 for a legacy overlay, but cuDNN, TensorRT, and NCCL do not yet resolve, so compatibility is workload-dependent and any tool leaning on cuDNN will fail.
  • πŸ’‘ Strategies & Tactics

    > Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine: Specialized GPU kernels that handle uneven token loads deliver a 10x speedup for mixture-of-experts AI training without sacrificing model quality.

    > One coordinate breaks abliteration on Gemma-3: A single dominant activation channel wrecks Gemma-3's refusal-direction editing, and clipping or masking that outlier restores usable results, showing scale outliers can silently break interpretability methods.

    > Anthropic analyzed 400,000 Claude Code sessions, and it turns out there's a right way to use it: Set the goal and constraints, then let Claude Code choose how to implement it, because delegating rather than micromanaging gets far better results.

    > LLMs as a Judge: How to Know if Your LLM is Healthy: Verify AI agents with automated tests, a second model that scores answers, and human review before launch, since one right reply can't prove reliability.

    > Astra appears to perform belief-propagation-like inference without CoT: Testing showed that OpenAI's GPT-6 Astra solves complex logic puzzles without step-by-step reasoning, seemingly weighing probabilities internally and improving with more filler tokens.

    Other news you might like

    • Nvidia, Palantir, and others restrict advanced AI model usage over privacy concerns, report claims β€” 'paranoia' rising over customer intellectual propertyLINK
    • Trump rejects AI guardrails; for Anthropic, $13.7 billion buys Georgia computeLINK
    • Perplexity Portable Computer Is Now Available on Windows, Powered by NVIDIA RTXLINK
    • Anthropic Deploys Claude to Automate Financial Adviser Prep WorkLINK
    • Chinese AI models dominate OpenRouter’s US token consumption. It can now guarantee that traffic stays entirely in the US.LINK
    • ElevenLabs MCP can now generate voice, music, images & videoLINK

    🧰 Trending tools

    Typewise Nova: an AI agent platform that handles customer service, sales, and billing requests end to end, with human oversight for approvals and handoffs.LINK

    ChatGPT Images 2.5: generates and edits images from text or sketches, preserving reference-photo subjects and rendering natural lighting up to 50% faster.LINK

    Diiverge: turns any image into a branching point-and-click adventure, generating new scenes and short clips as you explore and share saved paths.LINK

    Frigade Assist API: embeds an AI agent and React components that walk users through in-app workflows, cutting onboarding and support effort without custom development.LINK

    Cognition's SWE-2: a coding model that matches top performers on FrontierCode benchmarks while cutting cost and turns, available in Devin Desktop and CLI.LINK

    Perplexity Hybrid Compute: splits tasks between cloud and local models, keeping sensitive files on your Mac while frontier models handle research and drafting.LINK

    πŸ“š Trending papers & reports

    AI agent orchestration adds a decision layer that weighs each action's expected payoff against its cost, cutting wasteful tool calls and delays so agents stay useful without running up unnecessary compute bills.LINK

    Domain-tuned model repair restores the broad skills a specialized model loses during fine-tuning, using ordinary prompts instead of the original training data, and beats existing methods across role-play and medical question-answering tasks.LINK

    Answer-first prompting tells a chatbot to give its answer before explaining why, boosting accuracy in ~88% of tests by an average ~17% while cutting the delay and cost of lengthy reasoning.LINK

    Task-specific coaching notes teach an AI agent to write custom instructions for each new job instead of reusing old examples, lifting task completion from ~73% to ~81% on a software benchmark.LINK

    Minecraft AI agents get a memory that links what they see to why it matters, letting them recover from failures and reorder tasks on the fly, sharply boosting success on long, multi-step missions.LINK


    See you tomorrow for a new dose of β˜•οΈ AIpresso!

    More from the archive