The Manifest
Chinese open models keep closing the gap on Western frontier labs while Anthropic's pricing reversal shows how much that pressure is biting, and a run of new benchmarks exposes how confidently AI still gets things wrong in high-stakes settings.
Get The Manifest in your inbox
One sharp digest of AI × supply chain. Free. No noise.
Today's Top
- 01Anthropic slashes Claude Fable 5 limits in Max and Team Premium and pushes Pro users toward API pricingThe Decoder
- 02China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI orderThe Decoder
- 03Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex mathThe Decoder
- 04The Pentagon's new AI playbook treats slow adoption as a bigger risk than imperfect alignmentThe Decoder
- 05NVIDIA Released DeepStream 9.1: Bringing Agentic AI to Vision AI With 13 Skills and Multi-View 3D TrackingMarkTechPost
Models & Releases
3 storiesMoonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math
The frontend win is real, but a 39 percent score on FrontierMath Tier 4 versus near 90 for OpenAI and Anthropic models is the number that matters for anyone routing tasks by capability, not headline ranking.
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost
A rare side by side on license terms and actual serving cost, not just leaderboard scores, which is the comparison procurement teams actually need before committing to an open weight model.
Google Cloud's Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite
Betting that a continuous read, connect, write loop beats retrieval for agents running around the clock. Worth a pilot before anyone rips out an existing RAG stack that already works.
Supply Chain & Ops
1 storyDeals & Market
2 storiesAnthropic slashes Claude Fable 5 limits in Max and Team Premium and pushes Pro users toward API pricing
Reversing course on pulling Fable from subscriptions entirely tells you how much pricing pressure OpenAI's cheaper GPT-5.6 Sol is applying. Check usage caps before your next renewal, they just got tighter.
Kimi: Threat or menace?
The alarm over a strong open Chinese model says more about Western AI anxiety than about Kimi K3 itself. Judge it on the benchmarks, not the reaction.
Research & Frontier
5 storiesGoogle Deepmind argues video generators already contain the world models computer vision has been missing
Matching state of the art depth and segmentation results with far less training data by repurposing a video generator is a strong data point in the world model debate, worth tracking even if it stays a research argument for now.
AI text detectors struggle when language models mimic an author's style
Up to 48 percent of style-imitated scientific text slipped past detectors in Epoch AI's test. Anyone relying on these tools for compliance or vendor vetting should treat a clean scan as weak evidence, not proof.
AI chatbots reading X-rays can be dangerously confident even when they're wrong
The RadLE 2.0 benchmark shows models delivering wrong findings with full confidence instead of deferring. The lesson generalizes past radiology: any AI system making high-stakes calls needs a real mechanism to say it doesn't know.
Perplexity AI Releases WANDR: An Open Benchmark Evaluating Research Agents That Must Search Wide And Deep
Sub-0.4 soft F1 scores on evidence-heavy tasks show research agents still miss a lot of qualifying results even when citations look solid. Good benchmark to check before trusting an agent's sourcing on a real market or supplier scan.
Open-weight models now match frontier cyber performance from just four months ago at a fraction of the cost
The gap between open and closed models on cyber capability has shrunk from six to ten months down to four to seven, and safety mitigations on open models are largely not holding. Security teams should assume attacker tooling is catching up faster than defenses are.
Org & AI Architecture
3 storiesWill AI fix prior authorization or make it worse?
A government pilot puts automated coverage decisions directly in the path of patient care. Anyone running approval or exception workflows in any regulated process should watch how appeal rates and error handling shake out here first.
The Pentagon's new AI playbook treats slow adoption as a bigger risk than imperfect alignment
The Navy's stance that moving slowly is riskier than moving imperfectly is a preview of the adoption-speed argument every regulated industry will eventually have to make explicitly.
China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order
Free training slots and cooperation centers for the Global South are as much a market and standards play as a governance one. Worth watching which manufacturing and logistics partners get pulled into China's AI stack versus Western alternatives.