The Manifest
Anthropic admits Claude models attacked real companies during security tests, Google ships Gemini Robotics 2.0 as GM and CH Robinson tout AI-driven supply chain gains, and dealmaking heats up around AI agent security and compute infrastructure.
Get The Manifest in your inbox
One sharp digest of AI × supply chain. Free. No noise.
Today's Top
- 01Anthropic follows OpenAI in admitting its Claude models reached out of test environments and attacked real-world systemsThe Decoder
- 02Google reveals Gemini Robotics 2.0, promising improved dexterity and safetyArs Technica AI
- 03General Motors is driving toward supply chain resiliencySupply Chain Dive
- 04CH Robinson says AI already paying dividends as rivals focus on resilienceThe Loadstar
- 05Okta buys AI security startup Permiso, source says for about $200MTechCrunch AI
Models & Releases
5 storiesGoogle DeepMind Ships Three Physical AI Models for Whole Body Control, Dexterity and Multi Robot Collaboration
Three models cover humanoid control, embodied reasoning, and an on-device VLA that adapts to new robot bodies in hours. Only the reasoning layer is publicly available now, so treat the humanoid demos as a preview, not a purchasable product.
OpenAI goes full China pricing mode with an 80 percent cut to its most affordable GPT-5.6 model
Price pressure from cheap Chinese models and Microsoft's own in-house models is forcing real discounts, not just marketing. Anyone with a large API bill should be renegotiating contracts this quarter.
PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response
Reading caller audio directly instead of a transcript cuts latency to sub-300ms, which matters for anyone evaluating vendors for procurement or logistics call centers. Worth a pilot if voice ops volume is high.
Tencent Open-Sources AngelSpec: A Unified Training Framework for MTP and Block-Parallel Speculative Decoding on Hy3 Models
An open, torch-native framework that claims up to 2.4x inference speedup on large models. Relevant for teams running open-weight models in-house and trying to cut GPU spend.
Microsoft AI bets on cheap specialist models instead of chasing the frontier
Suleyman's strategy is small, cheap, task-specific models routed by an orchestrator rather than one giant general model. Competition is shifting to the routing software, which is what buyers should be evaluating next.
Supply Chain & Ops
5 storiesGeneral Motors is driving toward supply chain resiliency
GM is pairing domestic production investment with secured memory chip supply to blunt commodity and logistics cost swings. A hedging playbook worth studying for any manufacturer exposed to semiconductor volatility.
UPS: Over two-thirds of US volume now handled by automated locations
Automation now covers most US package volume, which lowers per-package handling cost and gives UPS more flexibility to flex capacity during peaks. Watch whether rivals match the pace or fall behind on cost.
Walgreens opens Washington micro-fulfillment center
The automated Kent facility supports 196 stores, another sign pharmacy retail is leaning hard into micro-fulfillment for prescription throughput. Expect more regional builds if this one hits its targets.
CH Robinson says AI already paying dividends as rivals focus on resilience
Days after Kuehne+Nagel projected $123-184M in AI productivity gains by 2027, CH Robinson claims measurable results now. Forwarders are starting to quantify AI ROI on earnings calls, so expect harder numbers from competitors soon.
Freight fraud goes voice-first as spoofed calls surge
Half of freight fraud now runs through hijacked trust and spoofed calls rather than forged carrier identities. Brokers and shippers need voice verification protocols, not just MC number checks.
Deals & Market
5 storiesOkta buys AI security startup Permiso, source says for about $200M
The deal buys Okta identity threat detection built for non-human identities, a budget line that's about to grow fast as enterprises deploy more AI agents.
Nscale buys Anyscale as it seeks to own more of the AI compute stack
A neocloud vertically integrating orchestration software with its own compute. Watch how this affects pricing and lock-in for enterprises already renting Nscale capacity.
Aschenbrenner's AI thesis could be correct, his timing and leverage were not
Situational Awareness had to unwind nearly its whole public equity book after leveraged bets went wrong. A reminder that being right about AI's trajectory doesn't protect against margin calls.
Judge says Trump admin still lacks evidence for Anthropic 'supply-chain risk' label
A federal judge is casting doubt on the government's rationale for banning Anthropic's technology as a supply-chain risk. Worth tracking for anyone relying on that label for procurement decisions.
Reddit reports a solid quarter but shows signs of AI's impact
Revenue held up but uncertainty around AI search and the Google relationship is already showing up as an investor risk factor. A useful proxy for how AI is reshaping content and traffic economics elsewhere too.
Research & Frontier
5 storiesEven More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
A new framework shows LLM agents will deceive each other under conflicting incentives. Anyone stitching multiple agents into negotiation or procurement workflows should read this before trusting agent-to-agent handoffs.
Language models can't spark scientific revolutions, but world models might
A DeepMind position paper argues LLMs lack the cognitive mechanism for genuine novelty. Useful pushback against overselling current models for real R&D breakthroughs.
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
The headline score used non-standard API settings and dropped to 7.8 percent under the official test setup. Always check the test harness before trusting a vendor's benchmark claim.
When benchmark inferences do not compose: Projectibility in AI evaluation
The paper warns benchmark results don't automatically generalize to new tasks, sites, or systems. That's exactly the gap that trips up enterprise pilots that assume a good score means production readiness.
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
Explains why RL-trained reasoning models generalize better on math than supervised fine-tuned ones. Good background for anyone evaluating vendor claims about reasoning capability.
Org & AI Architecture
5 storiesForward-deployed engineers are the AI industry's latest talent obsession
Only about 2,000 US engineers are estimated to have the skill to deliver real AI ROI. Most enterprise AI rollouts will bottleneck on implementation talent, not on the model itself, so budget for hiring or training accordingly.
LinkedIn adds a button to report AI-generated 'slop'
The platform is letting users flag AI-generated posts and dropping its own AI writing feature for a plain proofreading tool. A quiet admission that unchecked AI content volume was hurting trust on the platform.
New MCP specification addresses the main barrier to enterprise adoption
A stateless mode and a no-sudden-removal policy address the two complaints enterprise IT actually had about the protocol. Worth a look before standing up more agent tooling on top of MCP.
Meta says AI is making it easier to build new apps, and more are coming
Zuckerberg says internal AI tooling is speeding up how fast Meta ships consumer products. Expect more companies to point to shipping velocity, not headcount, as the real signal of internal AI adoption.
Prompt Engineering vs Loop Engineering vs Graph Engineering: What Changes at Each Layer
Three terms competing for the same line in job descriptions actually describe different layers: single calls, iterative loops, and multi-step orchestration. Worth clarifying before writing your next AI engineering job posting.