The Manifest
OpenAI's rogue coding agents and a mounting mathematician revolt expose how little control labs have over their own models, while customs deadlines and tight freight capacity give supply chain buyers real Q4 headaches.
Get The Manifest in your inbox
One sharp digest of AI × supply chain. Free. No noise.
Today's Top
- 01OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could GoogleThe Decoder
- 02OpenAI's feud with mathematicians is only escalatingTechCrunch AI
- 03How hackers used Claude for missiles, drone swarms, and surveillance, while Chinese labs mined it for training dataThe Decoder
- 04CBP: Shippers could lose import privileges if customs info is wrongSupply Chain Dive
- 05Mecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training dataTechCrunch AI
Models & Releases
4 storiesGoogle's new AI model predicts the future from sales data, weather, and discount schedules
TimesFM-3 fills in an entire forecast horizon in one pass instead of chaining daily predictions, and it factors in known future events like promos. Worth a pilot for demand planners tired of compounding error in step-by-step forecasts.
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
An open-weight translation model at this quality level is useful for procurement and logistics teams handling multilingual contracts and customs paperwork, though commercial use still needs a paid license.
Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
Routing tasks to cheaper specialized models instead of running everything on a frontier model is the practical path to cutting agent costs at scale. Watch the price points more than the benchmark scores.
Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
A concrete way to test whether an agent skill actually helps before shipping it, rather than trusting a demo. Any team building internal Claude Code plugins for procurement or planning workflows should adopt this before rollout.
Supply Chain & Ops
5 storiesCBP: Shippers could lose import privileges if customs info is wrong
Starting September 18, bad data on file can void a shipper's right to import at all, not just trigger a fine. Compliance teams should audit customs records now, not after the deadline hits.
Uber Freight warns tight capacity could fuel Q4 freight rate surge
Truckload conditions look stable on the surface, but week-to-week spot buying leaves shippers exposed if capacity tightens fast. Lock in contract rates before peak season rather than riding the spot market.
FedEx launches Shopify app to combat surprise import charges
A checkout-time duty and tax guarantee is a real fix for a chronic ecommerce pain point. Merchants running cross-border Shopify storefronts should test it against their current landed-cost estimates.
Cyber-attack halts operations at Malaysia's Tanjung Pelepas Port
Another major terminal knocked offline by a cyber incident, this time a Maersk joint venture. Port cybersecurity keeps landing in the same headlines as ransomware, and contingency routing plans should assume outages, not just delays.
Lands' End continues backlog recovery from WMS hiccup
A warehouse management system rollout that disrupted Q2 shipments is a reminder that WMS migrations need real cutover buffers, not just go-live confidence. Long-term efficiency gains don't help if the transition period burns customer trust.
Deals & Market
5 storiesMecka AI nears $500M valuation in Sequoia-led deal amid rush for robot training data
Investors are pricing robot training data as the real bottleneck for warehouse and manufacturing automation, not the robots themselves. Ops leaders evaluating robotics vendors should ask who supplies their training data and how.
Kimi-maker Moonshot AI targets $2B in annual revenue
K3 usage is dipping even as Moonshot pushes toward a $2B revenue target, a sign the open-weight Chinese model race is now about monetization, not just benchmark bragging rights. Buyers comparing model vendors should track revenue durability, not just token volume.
Nscale adds former OpenAI exec Fidji Simo to its board ahead of potential IPO
Simo took Instacart public, so her board seat is a clear signal Nscale is preparing for a real IPO process rather than just talking about one. Worth watching as an AI infrastructure IPO bellwether.
Class action lawsuit accuses Anthropic of overselling Claude subscriptions with deceptive usage multipliers
Enterprise buyers negotiating Claude seats should get usage multiplier terms in writing now, before this suit forces a policy change that resets everyone's contracts anyway.
Anthropic's $1.5 billion book settlement descends into chaos as authors and publishers fight over who gets paid
The largest copyright settlement in US history is still unresolved on distribution, a preview of the legal drag that any company training on scraped content should expect to inherit.
Research & Frontier
4 storiesGoogle Research Releases ToolGrad: Answer-First Framework Hits 99.8% Pass Rate for Tool-Use Data Generation
Building the verified tool-call chain first, then writing the query to match it, fixes the low pass rates that plague synthetic agent training data. A small fine-tune on 500 samples matching much bigger baselines is the number to watch.
Can LLMs Engineer Their Own Agent Harness? ByteDance Seed's HarnessDev Says Only 34 of 64 Changes Generalize
Just over half of self-built harness changes actually generalize across tasks, a useful caution against letting agents redesign their own tooling unsupervised. Treat self-modifying harnesses as a research finding, not a production pattern yet.
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
An open post-training recipe for hard reasoning tasks, useful less for math and more as a template for anyone building verification and refinement loops into their own specialist models.
Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge
A framework for testing whether agents stay useful as disruptions pile up over repeated interactions, closer to how real operational workflows actually break than a single pass-fail task benchmark.
Org & AI Architecture
5 storiesOpenAI's feud with mathematicians is only escalating
Twenty-five Fields Medal winners saying labs are eroding the discipline's actual goal is a credibility problem that goes well beyond math. Any function that measures output by solved tickets rather than understanding should pay attention.
An Anthropic researcher's doomsday warning comes at a very interesting time
A resignation warning about racing toward self-improving superintelligence, co-signed by the company's own alignment lead, is not the kind of internal dissent that gets ignored. Expect governance and safety review questions to follow enterprise deals for a while.
Deep learning pioneer Bengio argues the training process itself makes AI dangerous
Bengio's call for independent safety reviews before further training runs is a direct challenge to the current pace of deployment. Buyers signing multi-year AI contracts should ask vendors what independent review they actually submit to.
Ex-Deepmind VP Vinyals says AI self-improvement is coming but won't trigger an intelligence explosion
A useful counterweight to the doom narrative from someone who actually built these systems. His bottlenecks, research taste and reliable evaluation, are the same limits operators run into when trying to automate judgment calls.
ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
A professional got sanctioned for not knowing an LLM can fabricate facts wholesale. The same failure mode applies to anyone using AI output in contracts, compliance filings, or supplier vetting without independent verification.