Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π·οΈ AI coding agents install malware
π Tencent releases open 770B AI model
π€ OpenClaw 2.0 adds shared agent sessions
π¬ Google AI now runs lab experiments
Plus: π‘ 5 strategies & tactics, π 7 other news you might like, π§° 6 tools, and π 5 papers.
π·οΈ AI coding agents install malware LINK
π Tencent releases open 770B AI model LINK
π€ OpenClaw 2.0 adds shared agent sessions LINK
π¬ Google AI now runs lab experiments LINK
π‘ Strategies & Tactics
> Deploy an Open Model from Checkpoint to Inference in Two Commands with NVIDIA TensorRT Model Connect: NVIDIA's TensorRT Model Connect turns an open AI model into a fast native C++ application in two commands, sparing developers from writing model-specific conversion and runtime code.
> Three mistakes of new AI teams: New AI teams fail by skipping evaluation, treating retrieval as a solved checkbox, and siloing engineers from data scientists rather than blending both skills.
> Muse Glimmer 30B runs on my 32GB GPU, handles its own tool calls, and never leaves the machine: Meta's open Muse Glimmer 30B runs entirely on one 32GB consumer GPU and is trained to recover from failed tool calls, keeping sensitive agent tasks off the cloud.
> LM Studio built a judge for AI commands. Then the judge started agreeing with the defendant.: LM Studio checks AI coding commands by analyzing their structure instead of matching risky words, escalating unclear cases to a second AI reviewer.
> AI agents are making retrieval engineering a core engineering discipline: Because AI agents act on retrieved information without human review, engineering what evidence they see and in what order now determines application quality.
Other news you might like
- OpenAI to end model access to Cursor after acquisition by Elon Musk's SpaceXLINK
- Amazon is buying 2 million Nvidia GPUs for AWS data center expansionLINK
- OpenAI has started letting some customers pay only when the AI worksLINK
- Samsung Locks 70% of Memory Capacity Into 2031 Deals β Long-Term HBM Deals Run 5x Cheaper Than Spot, as AI Demand Chokes the DRAM MarketLINK
- Europe has ordered a β¬388M AI supercomputer from its own state-owned championLINK
- Gemini Notebook switching to compute-based usage limits, like Gemini appLINK
- Anthropic is cutting Claude Code's current weekly limits by 17%LINK
π§° Trending tools
Referent: AI-native legal practice management for law firms, handling client intake, matters, documents, deadlines, email workflows, and billing prep in one place.LINK
oMLX: runs a menu-bar LLM inference server on your Mac, serving text, vision, OCR and embedding models with a persistent KV cache for faster responses.LINK
Sayscroll: an AI teleprompter that tracks your voice to auto-scroll scripts, supporting 60+ languages, off-script improvisation, and browser-based recording without downloads.LINK
Maritime: deploy, manage, and scale AI agents with sleep/wake micro-VMs, pushing code to get a live API endpoint for about $1 per agent monthly.LINK
Video Agent by Fotor: create complete, production-ready videos and motion graphics through a chat interface, generating fully editable results without manual timeline work.LINK
Topview Motion Studio: turns a product brief and reference assets into a 4-60 second launch video via a guided AI workflow, skipping After Effects timelines entirely.LINK
π Trending papers & reports
A memory-saving shortcut for language models, having them look only at a recent sliding window of words, matches or beats a popular rival on reasoning tasks by 2 to 10 times, with no retraining needed.LINK
Chatbot memory trimming gets a rigorous foundation, showing that dropping the running notes a model keeps during long chats to run faster can be done more reliably by correcting for what was thrown away.LINK
Agent memory management teaches AI assistants to actively prune, plan, and offload their own working notes during long multi-step tasks, delivering stronger results on document search and question-answering while keeping context leaner.LINK
Spoken confidence in chatbots often fails to match what a model actually knows internally, so a bot saying it is sure is a weak, poorly calibrated signal across 30 models tested, meaning stated certainty should not be trusted at face value.LINK
Realistic user simulation generates lifelike back-and-forth conversations to train and test AI agents, boosting task completion by ~6% since most real users, about 76%, ask across multiple messages rather than one complete query.LINK
See you tomorrow for a new dose of βοΈ AIpresso!