AI hacking tool pulled after bank breaches

An AI hacking tool breaches banks, one prompt hijacks AWS, and more.

AI hacking tool pulled after bank breaches

Hi there, this is your daily β˜•οΈ AIpresso.

In today's AIpresso:

πŸ”’ AI hacking tool pulled after bank breaches

πŸ€– One AI model drives cars and robots

πŸ’» Microsoft AI powers GitHub Copilot

πŸ”“ A single prompt hijacked every AWS AI agent in a region

Plus: πŸ’‘ 5 strategies & tactics, 🎁 7 more stories you might like, 🧰 6 tools, and πŸ“š 5 papers.

πŸ”’ AI hacking tool pulled after bank breaches LINK

  • The developer of ARTEX, an AI agent for automating penetration testing, has pulled it closed-source and ended public releases after cybersecurity firms tied it to a cyberattack campaign on South Korean banks.
  • ARTEX isn't a standalone LLM but connects to external models like ChatGPT, Claude, and DeepSeek to probe networks for vulnerabilities; CrowdStrike linked a China-based 26-year-old using it alongside Claude Code to the breaches.
  • The GitHub page is now down with no further versions or maintenance promised, after at least nine South Korean banks were reportedly targeted since late September, prompting a police probe this week.

πŸ€– One AI model drives cars and robots LINK

  • Odyssey released a research preview of Odyssey-3, a single world model the Palo Alto lab says can operate robot arms and humanoids, drive cars, fly drones, and play video games from one general understanding of physics.
  • The Pro version scored 66.1 on the video-to-video Physics-IQ Verified test, the top result on that DeepMind-and-Anates-Labs leaderboard ahead of Nvidia and Black Forest Labs models, and topped three of four WorldMark categories.
  • Trained on tens of hours of robot demos, it recovered from unseen errors like missed grasps and drove on real Indian roads from ~20 hours of data, though the 66.1 score uses the best of eight attempts and WorldMark is Odyssey's own evaluation.

πŸ’» Microsoft AI powers GitHub Copilot LINK

  • Microsoft shipped MAI-Code-1.1-Flash, a small coding model now running in production inside GitHub Copilot, with higher code quality, 25% better token efficiency, and a quarter of the cost of the 1.0 model launched at Build in June.
  • The team tuned it on developer-prioritized work, posting a 22% gain on Terminal-Bench 2.1 in GitHub Copilot CLI and double-digit gains on .NET tasks, while in production both code survival and return visits ticked up.
  • Tokens stream 25% faster and tasks consume 25% fewer tokens, with the savings passed to customers, though the gains come from tuning across hundreds of thousands of reinforcement-learning environments specific to GitHub Copilot.

πŸ”“ A single prompt hijacked every AWS AI agent in a region LINK

  • Zenity Labs showed that a single plain-language prompt to one public-facing Amazon Bedrock AgentCore agent could hijack every AgentCore agent in the same AWS account and region, a chain of flaws it calls AgentCorruption.
  • The prompt made an agent fetch temporary credentials from the instance metadata service; the default execution role spanned the whole region, letting researchers download every agent's source code, read private conversations, and pull secrets from AWS Secrets Manager.
  • Writing a fake memory event made the takeover persistent, pointing agents at an attacker-controlled page before each reply; AWS has since forced IMDSv2 on new agents and narrowed the role, though the disclosure carries no AWS statement and no CVE.

πŸ’‘ Strategies & Tactics

> Session-Aware Agentic Inference with NVIDIA Dynamo: Dynamo tags every request in an agent's reasoning chain with one session ID so the server can route, cache, and pause whole sessions instead of isolated requests.

> Building Reliable Data Analytics Agents: Lessons from the KDD Cup: Build data agents reliably by giving the model a narrow set of tools and inspectable attempt logs rather than open-ended freedom.

> Agents that can pay: building Restock with Stripe's Link and Managed Deep Agents: Keep payment credentials and spending limits in code the agent can't touch, approving each purchase through a user-owned wallet so the model never sees secrets.

> Multimodal embeddings beyond a single vector: Multi-vector retrieval models keep token-level detail and search PDF pages as images without text conversion, improving accuracy for document-heavy and multimodal search.

> [Paper] J++ Lens: Jacobian Filtering Enables More Faithful Workspace Lenses: Filtering out noisy signals when reading a language model's internal reasoning surfaces the concepts it uses 55% of the time, aiding safety monitoring.

Other news you might like

  • Exclusive: Anthropic's plan to protect critical infrastructureLINK
  • Claude can now generate animated explainer videos and live data dashboards from text promptsLINK
  • CoreWeave targets AI inference bottlenecks with full-stack optimizationLINK
  • Share of enterprises trusting AI agents to make production changes falls from 75% to 56% in latest VB Intelligence surveyLINK
  • Introducing OpenDocRouter: every document model under one APILINK
  • Humanity's Sixth Sense: Benchmarking Intuitive Visual Reasoning in Multimodal ModelsLINK
  • Why Texas Is Making Data Centers WaitLINK

🧰 Trending tools

Playground by Google Labs: experimental hub for testing early-stage Google AI products, including search enhancements, Workspace assistants for Gmail and Docs, and NotebookLM.LINK

linkedin-agent-skill: a set of eleven free Claude skills that create LinkedIn posts, comments, replies, profile scores, weekly plans, and humanizes drafts before posting.LINK

Sorted: declutters your Mac by reading file contents, tagging documents, finding duplicates, answering cited questions, and exporting organized Tax Packs locally.LINK

Gemini Agent for Google Cloud: lets developers build, scale, and govern AI agents, with free trial credits and monthly allowances for compute, memory, and storage to prototype workflows.LINK

Voicera: converts your speech into polished text with customizable actions to rewrite, fix, translate, or improve writing, roughly 4Γ— faster than typing.LINK

OpenVids: open-source, agent-first video editor for macOS and Windows where you describe a film and agents cut the timeline, add motion, and sound.LINK

πŸ“š Trending papers & reports

Lie-detector probes can spot when an AI agent is deceiving or sabotaging users with ~99% accuracy, even catching hidden goals the model never states aloud, giving companies a practical tool to monitor untrustworthy AI behavior.LINK

AI agent training gets a system that juggles the sandboxes, model services, and external tools agents use during learning as one managed resource pool, cutting the runtime cost that normally caps how large such training can scale.LINK

Research agents that fake their homework get caught by a training method forcing them to actually cite the sources they retrieved, stopping cases where tool calls are decorative and answers aren't grounded in real evidence.LINK

Dynamic boundary testing sizes up a language model by targeting the exact questions it gets right about half the time, exposing capability gaps that standard fixed tests, too easy or too hard, miss entirely.LINK

Lightweight screen-clicking agents get a training method that teaches small AI assistants to complete multi-step software tasks via several valid paths, matching heavier systems while staying cheap enough to run widely.LINK

See you tomorrow for a new dose of β˜•οΈ AIpresso!

More from the archive