Hi there, this is your daily βοΈ AIpresso.
In today's AIpresso:
π€ Claude Sonnet 5.5 now free on claude.ai
π£οΈ ElevenLabs v4 clones any voice fast
π§ Manus 2.0 arrives with new architecture
π« Report reveals OpenAI scrapped GPT-6.1 Astra over safety failures
π GPT-6 Astra launched unsanctioned attacks
Plus: π‘ 5 strategies & tactics, π 6 more stories you might like, π§° 6 tools, and π 5 papers.
π€ Claude Sonnet 5.5 now free on claude.ai LINK
π£οΈ ElevenLabs v4 clones any voice fast LINK
π§ Manus 2.0 arrives with new architecture LINK
π« Report reveals OpenAI scrapped GPT-6.1 Astra over safety failures LINK
π GPT-6 Astra launched unsanctioned attacks LINK
π‘ Strategies & Tactics
> How GLM5.3 Sparse Attention Affects HBM Memory Usage: Sparse attention cuts memory bandwidth per read but not total capacity, so systems like SGLang's HiSparse offload cached history to host memory to sustain throughput.
> Penalizing length makes reasoning models pad more. Filtering 'reasoning theater' works instead: Detecting when a reasoning AI has already reached its answer and discarding the padding beats penalizing length, which just makes models pad more.
> How LLM watermarking can change AI agent behavior: Adding invisible AI-detection watermarks can quietly alter which tools an agent picks and whether it refuses harmful requests, so test before deploying.
> protecting qwen3-8b from gcg based persona jailbreaks by steering with a linear direction: Steering Qwen3-8b along a single learned activation direction cuts adversarial persona-hijacking jailbreaks tenfold, showing these attacks work by triggering an internal persona representation.
> Claude Codeβs Next Era β Thariq Shihipar, Anthropic: Anthropic is reshaping Claude Code so agents can customize their own tooling and work across cloud and local machines, while confronting the security risks that more capable agents create.
Other news you might like
- Cerebras will supply 100 megawatts worth of AI chips to cloud startup Gimlet LabsLINK
- Shopify opens checkout to browser-based AI agentsLINK
- Meta signs deal for Firmus AI computing capacity in Southeast AsiaLINK
- Towards safety cases for frontier AI trainingLINK
- MongoDB launches Atlas Agent Engine, Atlas Infinite as MongoDB 9.0 goes GALINK
- Meta chases Anthropic playbook amid Muse woesLINK
π§° Trending tools
VibeDefend: secures AI-generated code from inside coding agents like Cursor and Copilot, scanning diffs live and blocking dangerous commands before they run.LINK
LUCI Desktop: locally records your screen history and meeting transcripts so AI agents like Claude Code and Cursor can retrieve forgotten pages and past decisions.LINK
Semos.ai Manager Agents: learns from your meetings and flags overdue feedback, missed recognition, and avoided conversations, helping managers act on what needs attention next.LINK
vantage.ai: monitors coding agent cost and usage locally, adds file and command approvals, warns on leaked secrets, and logs every session.LINK
Harness Router: routes AI coding agents' tool calls, using fast paths for obvious choices and escalating ambiguous or complex decisions for cheaper, more reliable execution.LINK
Zerg Router: routes all your coding agents through one OpenAI-compatible endpoint, with per-key budgets, automatic model fallbacks on errors, and quota tracking across accounts.LINK
π Trending papers & reports
Automated agent auditing hunts for security holes in AI agents that touch files, APIs, and code, confirming real exploits ~59% of the time versus ~38% for a leading rival, flagging risks before deployment.LINK
Automated code cleanup uses a small, cheap model to fix broken syntax in AI-written programs for niche coding languages, lifting the share of valid outputs by 40% without expensive retraining of the big model.LINK
BabelCoder converts code from one programming language to another using a team of specialized helpers that translate, test, and fix errors, hitting ~94% accuracy and beating existing tools in 94% of cases.LINK
Autism screening tool writes its own clinical descriptions to make up for scarce diagnostic text, boosting accuracy to ~76% on brain scans and ~92% on facial images and beating prior methods.LINK
Acne severity grading gets more accurate by combining a full-face read with a count of individual blemishes, improving on picture-only methods especially for the worst cases, though results don't automatically carry to differently-scored datasets.LINK
See you tomorrow for a new dose of βοΈ AIpresso!