OpenAI watermarks some ChatGPT text

OpenAI watermarks ChatGPT text, Claude aces CPA exam, and more.

OpenAI watermarks some ChatGPT text

Hi there, this is your daily โ˜•๏ธ AIpresso.


In today's AIpresso:

๐Ÿ” OpenAI watermarks some ChatGPT text

๐Ÿงฎ Claude aces test 12 CPAs flunked

๐ŸŒ New AI world model spans many fields

๐Ÿ”Ž New benchmark tests AI code reviewers

๐Ÿค– Reflection unveils its first open model Beam

Plus: ๐Ÿ’ก 5 strategies & tactics, ๐ŸŽ 8 more stories you might like, ๐Ÿงฐ 6 tools, and ๐Ÿ“š 5 papers.

๐Ÿ” OpenAI watermarks some ChatGPT text LINK

  • OpenAI has started a limited rollout of text watermarking and detection to comply with the EU AI Act, which requires generative providers to make AI-generated text identifiable in a machine-readable way.
  • API customers globally can opt into watermarking for select models starting today, though it stays off by default, while invisible watermarks land on eligible ChatGPT and Codex output in the EU over the coming weeks.
  • OpenAI is opening applications for its detector to approved researchers only, cautioning that a watermark shows an OpenAI system touched a passage but not authorship, accuracy, or how much human editing was involved.
  • ๐Ÿงฎ Claude aces test 12 CPAs flunked LINK

  • Claude Opus 5 scored a perfect 100% across all 20 solo attempts on Mercor's real-world accounting tasks, while 12 licensed CPAs averaging 5.4 years of experience managed only about 37%.
  • Drawn from Mercor's APEX-Accounting benchmark, the four month-end close scenarios required searching company files, finding figures, and formatting results; Claude finished each run in under 10 minutes versus 30 minutes to three hours for humans, at roughly 49x lower cost per rubric criterion.
  • The sharper signal is the climb: GPT-4o scored near zero 18 months ago, o3 passed the human average by spring 2025, and GPT-5 hit about 69% before frontier models reached perfect scores, though Mercor stresses the test strips away client communication, judgment, and the longer workflows where model performance still drops.
  • ๐ŸŒ New AI world model spans many fields LINK

  • PhAI Labs, with researchers from Stanford, Oxford and Princeton, released JEPA-Anything, a version of Yann LeCun's Joint-Embedding Predictive Architecture that predicts abstract summaries of future states instead of reconstructing raw data, and tested it across physics, robotics, medicine, weather, fluid dynamics, molecular simulation and biology.
  • The key change: instead of a single prediction module, the model splits the predicted state into several parts, each with its own predictor constrained to capture a different aspect; prediction error fell 35% in a simplified Pong environment and by nearly half on the Burgers fluid equation, and on orbital mechanics it recovered Kepler's third law with an exponent of -1.4991 versus the theoretical -1.5.
  • Code and models are on GitHub and Hugging Face, but the authors caution that cleanly separated parts don't prove the model has learned real cause-and-effect, so how far it can be trusted to guide experiments remains open.
  • ๐Ÿ”Ž New benchmark tests AI code reviewers LINK

  • GitHub released ReviewBench, an offline benchmark that scores AI code review agents on real pull requests, measuring what each system catches, misses, and the precision-recall tradeoffs it makes, with a research preview available today.
  • The corpus holds 219 pull requests from 187 open-source repos across 19 languages, drawn to match the distribution of 103.9M GitHub pull requests, with every finding labeled by severity and category against a multi-source golden set and a Claude Sonnet 5 grader.
  • Six metrics split into grounded scores against known labels and augmented scores crediting newly discovered issues, with senior engineers independently matching ReviewBench's true-positive judgments 96.6% of the time, though augmented recall shifts its denominator per agent, so grounded recall stays the headline cross-system number.
  • ๐Ÿค– Reflection unveils its first open model Beam LINK

  • Reflection has unveiled Beam, its first open-weight model, built to match leading Chinese open models from Deepseek and Qwen on compute-efficient coding, reasoning, and agentic tasks rather than topping raw performance.
  • The mixture-of-experts model activates 23B of 501B parameters per token, matches GLM 5.2 on reasoning using three to four times less compute, and offers a tunable knob to trade thinking time against cost.
  • Weights ship under Apache 2.0 "later this month" alongside a technical report and docs, though Beam is still in final safety testing with only an early version open to select users, and stronger open models like Kimi K3 still beat it on raw performance.
  • ๐Ÿ’ก Strategies & Tactics

    > One MCP server used 18,000 tokens before doing anything. Hereโ€™s the workaround.: Pi keeps MCP (model tool-connection protocol) tools out of the AI's prompt, letting it discover and call them via code to save context space.

    > The AI safety check that runs on a laptop and nearly matched a 35B model: A small laptop-run safety classifier nearly matched a 35-billion-parameter AI model on catching harmful prompts, while responding far faster and cheaper.

    > Claude Code now shows exactly why it burned your tokens, and I finally stopped guessing: Claude Code's cost command now breaks down token usage and cache misses, revealing that leaving sessions idle quietly inflates costs.

    > You can graft SDF changes from base models onto post-trained models: Fine-tune a model's base version on fabricated documents, then transfer that change to the deployed version, preserving the model's grasp on reality.

    > Your local LLM is quietly wasting RAM on context you never use, and one setting gives it back: Lower your local model's context window to match actual usage, freeing several gigabytes of memory that unused capacity otherwise reserves.

    Other news you might like

    • Cohere's North 2 puts AI agents on a budget and gives them a memoryLINK
    • Meta and Microsoft pull back from Claude as Anthropic transforms from partner into competitorLINK
    • Anthropic Subscriptions Offer 5x+ More Value Than OpenAILINK
    • Evolution of the PyTorch Media Processing LandscapeLINK
    • BDH-CQ combines in-context learning with reasoning outside the token streamLINK
    • Cheap AI often costs more; Claude plans buy about five times the usageLINK
    • Clockwork.โ€‹io bags $31M in funding to keep AI inference and training workloads running like โ€ฆ clockworkLINK
    • OpenAI agents tried to hack Wikipedia tools and flooded it with trafficLINK

    ๐Ÿงฐ Trending tools

    Chunk: turns your documents and notes into a searchable knowledge base, delivering fast answers across all your Apple devices.LINK

    WebinarFlow: records your demo once so prospects watch on demand, letting you join live only for real questions while AI handles repeatsLINK

    Speek: a Mac notch voice assistant that dictates text into any app, edits selected content, and runs tasks across your accounts hands-free.LINK

    OnionClaw: gives AI agents full Tor network access and dark web data via a zero-config OpenClaw skill or standalone tool.LINK

    OpenBot: an open-source, local-first runtime that orchestrates multiple AI agents via tagging and channels, routing tasks and managing persistent workflows on your file system.LINK

    Incredible: control your computer with your voice, letting you run tasks and operate your machine hands-free through spoken commands.LINK

    ๐Ÿ“š Trending papers & reports

    Muon's training math now comes with a proof that its popular, speed-tuned optimizer actually reaches any target training accuracy, giving teams theoretical assurance behind a shortcut many already use in practice.LINK

    Low-rank shortcut checks prove that a slimmed-down version of a big calculation is accurate enough using one reusable batch of tests, so verifying many approximations at once does not multiply the checking work.LINK

    Prediction-confidence scores for connected data like social or payment networks often fail to shrink as more data arrives, and a resampling method fixes this so confidence actually improves with evidence.LINK

    Physics-simulation foundation model trains one system across twelve types of physics equations and cuts its internal specialist components in half while matching or improving accuracy, making large-scale scientific simulation cheaper to run and reuse.LINK

    Cordial learning lets separate parties whose data is tangled together train shared AI by exchanging only tiny summaries instead of raw data, reaching optimal accuracy while preserving privacy where standard federated learning fails.LINK


    See you tomorrow for a new dose of โ˜•๏ธ AIpresso!

    More from the archive