Alibaba's new AI makes 30-second videos

Alibaba's 30-second video AI, a 753B model on one GPU, and more.

Alibaba's new AI makes 30-second videos

Hi there, this is your daily β˜•οΈ AIpresso.

In today's AIpresso:

🎬 Alibaba's new AI makes 30-second videos

πŸ–₯️ Run a 753B AI model on one GPU

πŸ—œοΈ 4-bit model beats its full-size original

πŸ”§ Meta built its first AI training chip with networking built in

🏭 Nvidia's Groq 3 AI chip enters production

Plus: πŸ’‘ 4 strategies & tactics, 🎁 8 other news you might like, 🧰 6 tools, and πŸ“š 5 papers.

🎬 Alibaba's new AI makes 30-second videos LINK

  • Alibaba launched Wan3.0 on August 24, 2026, a video generation model that produces clips up to 30 seconds in a single pass at 1080p, double the 15-second ceiling of its predecessor Wan2.7.
  • The model accepts text, images, audio, video, web pages, and documents like PDFs and PowerPoint, generating synchronized facial micro-expressions and multilingual voice output while holding character detail and motion consistency across the clip.
  • Pricing runs $0.05/sec at 480p to $0.20/sec at 1080p ($12/min), undercutting Google's Veo 3.1 at $0.40/sec, though Alibaba closed the weights this release and the quality claims remain unbenchmarked by anyone independent.

πŸ–₯️ Run a 753B AI model on one GPU LINK

  • FreeToken, an edge-native MoE serving engine from UC Berkeley and UT Austin, runs the 753B-parameter GLM-5.2 model on a single workstation GPU by mapping computation across GPU, CPU, memory, and interconnect bandwidth.
  • The engine deploys a 35B model on an 8GB laptop GPU and 284B on a gaming desktop, sustaining 77-83 tokens/sec on Qwen3.6-35B-A3B with an RTX 5090 via bandwidth-adaptive execution and semantic-aware caching.
  • It ships Apache-2.0 on GitHub, as freetoken v0.1.2 on PyPI, and as a one-click Windows/Linux app exposing OpenAI- and Anthropic-compatible endpoints, though the team notes its performance claims still need independent verification.

πŸ—œοΈ 4-bit model beats its full-size original LINK

  • Multiverse Computing's Quantization-Aware Healing recipe took a GPT-OSS 120B model compressed to 60B and quantized to MXFP4, producing a 4-bit model that beats its own bfloat16 source on 7 of 9 benchmarks.
  • Instead of distilling from the recovered checkpoint, QAH distills directly from the original full-size teacher, yielding gains where compression usually hurts most-+7.4 on AA-LCR long-context reasoning and +5.6 on AIME 2025 math.
  • Running at half the teacher's parameters and ~4x less weight memory, the 4-bit model even edges the full-size teacher on LiveCodeBench (66.5 vs. 66.0) and hits peak accuracy ~7x faster than QAT, though it still trails the teacher on the extreme long-context AA-LCR benchmark.

πŸ”§ Meta built its first AI training chip with networking built in LINK

  • Meta has launched MTIA 300, the first chip in its in-house accelerator family built for training recommendation and ranking models, with networking integrated directly on the chip package rather than across a PCIe bus.
  • Two network chiplets pack twelve 800 Gbps RDMA NICs for 1.2 TB/s of I/O, while 16 dedicated message engines offload all collectives, so running GEMMs alongside communication costs under 0.5% compute throughput versus over 20% on GPUs.
  • On a 150B-parameter production model across 40 accelerators, MTIA 300 cut total communication time 3.9x against an equivalent GPU cluster, aided by 216 GB of HBM3E and co-designed with the HCCL library that ties into PyTorch's c10d and torchcomms.

🏭 Nvidia's Groq 3 AI chip enters production LINK

  • Nvidia moved its Groq 3 LPX inference accelerator into full production, extending the Vera Rubin platform to speed token generation for agentic systems that run across thousands of inference steps on large contexts.
  • In Artificial Analysis testing, the chip hit 3,400 output tokens per second on Gemma 4 31B at a 100,000-token context, which Nvidia calls the fastest recorded for that model and four-times the responsiveness of the nearest competitor.
  • Cloud provider Nebius is the first customer, deploying a rack of 256 Groq 3 chips with Vera CPUs and Rubin GPUs via its Token Factory platform, though it only goes live later this year.

πŸ’‘ Strategies & Tactics

> I ran Claude Code against Archon on the same bug six times, and the run that failed told me the most: Wrapping an AI coding agent in a workflow that halts on any unexpected output catches off-script mistakes before they reach your main branch.

> Grok Bot vs. Hermes: Where each draws the security boundary: Grok Bot shares one account and credentials across all its bots while Hermes gives each bot its own profile, so operators must read the security page, not the launch pitch.

> Wire It, Run It, Deploy It: AI Workflows in Gradio: Build multi-step AI apps as a visual node graph in Gradio that runs interactively, exposes each output as an API, and deploys in one command.

> An Old Laptop’s Only M.2 Slot Was Used To Hook Up An RX 7900 XT, Achieving 60 Tokens/s In Qwen3.6 27B & 100K Context Window; External Drive Was Used For The OS: A hobbyist ran a large local AI model fast by plugging a desktop graphics card into a laptop's storage slot, showing cheap hardware can substitute for pricey GPU rigs.

Other news you might like

  • Meta’s Consumer-Focused AI Agent Could Be Weeks From LaunchLINK
  • OpenAI restores 5-hour Codex and Work limits for ChatGPT Plus usersLINK
  • Lambda eyes $12B valuation with $3B pre-IPO raise as AI cloud race heats upLINK
  • Britain Gains First Foreign Access to Ukraine’s Avengers AI LabsLINK
  • Why companies want to own their intelligenceLINK
  • Open-Weight Models Are Gaining Ground in Enterprise AILINK
  • SpaceXAI will deploy standalone Nvidia Vera CPUs for Grok's agentic workloads β€” will use optimized Vera Rubin NVL72 in space with Starmind satelliteLINK
  • JetBrains Junie now runs entirely offline. Can you spare a 64 GB M5 Mac?LINK

🧰 Trending tools

Navigara: connects AI coding performance to your roadmap, tracking costs per item, isolating waste, and routing routine tasks to cheaper models automatically.LINK

Dropstone: a persistent AI runtime that remembers past work, learns from mistakes, and applies skills across code, chat, docs, and spreadsheets.LINK

Lucid Train: generates architecture diagrams from your codebase using local Ollama models, runs fully offline, and feeds diagrams to coding agents as specs.LINK

Claude Watermark: detects and strips hidden characters, invisible spaces, and HTML artifacts from AI-pasted text locally in your browser, no upload needed.LINK

Hosted Agents in Cluing: run collaborative AI agents on your saved knowledge without setup, building and assigning them from your phone across a team.LINK

Vois 2.0: converts scripts, ebooks, and articles into natural speech locally with 63 voices, voice cloning, and editing, no uploads or usage caps.LINK

πŸ“š Trending papers & reports

Task-specific model shrinking maps how far you can compress a large language model for one domain like finance before quality drops, showing specialized skills survive far longer than general knowledge unless you train on step-by-step reasoning.LINK

Video-reasoning training teaches models to reliably pinpoint and track the exact objects and moments that answer a question, using guided hints only during training so accuracy improves without costly extra data or slower analysis at run time.LINK

Long-document reading lets AI blend clues across chunks of a lengthy text instead of jumping between them, boosting multi-step question answering accuracy while keeping memory use low, so models handle huge documents faster and more reliably.LINK

Vision-model jailbreak defenses vary so much by architecture that no single method wins on both safety and usefulness, meaning companies can't reuse one fix across different image-reading AI systems.LINK

Editing image-reading AI lets users fix an AI's mistakes by explaining their reasoning in plain words, and storing that reasoning makes the corrections stick across similar questions far better than prior methods.LINK

See you tomorrow for a new dose of β˜•οΈ AIpresso!

More from the archive