tandemly.ai
Briefing · JUL 24 2026

July 24, 2026

AI daily briefing

🎯 Top 3 Things to Know

1. The White House accused a Chinese lab of stealing a U.S. model, and the details matter more than the headline. Michael Kratsios, who runs the White House Office of Science and Technology Policy, posted on July 22 that Moonshot AI built its new Kimi K3 model by running large-scale "distillation" against Anthropic's Fable. Distillation means training a smaller model on a larger one's outputs. Kratsios said the normal kind is fine, but that Moonshot ran a covert platform to do it at scale while switching access methods to avoid detection, and that the firm reached banned Nvidia GB300 chips in Thailand. Treasury Secretary Scott Bessent floated sanctions. The friction: Fable only went public on July 1, and several researchers told TechCrunch that timeline is too short to explain how good Kimi K3 is. Worth watching whether Treasury produces evidence or the claim stays an accusation. CyberScoop

2. AMD used its Advancing AI event to claim the fastest AI GPU in the world. On July 23 in San Francisco, CEO Lisa Su launched the Instinct MI400 series, led by the MI455X, which AMD says delivers 34 times the token throughput of last year's MI355X. AMD also shipped its Helios rackscale system and 6th-generation EPYC CPUs, pitching a full-stack answer to Nvidia rather than a single chip. The claim matters because inference cost, not model quality, is what now decides many production deployments. Teams weighing GPU supply should watch independent MLPerf and throughput numbers before pricing the MI455X against Nvidia's rack systems, since vendor throughput figures rarely survive contact with real workloads. AMD

3. A new paper makes context compaction something an agent learns, not a trick bolted on at inference. Long-running agents run out of context window before they finish a task. The usual fix is to summarize the history and keep going. CompactionRL, from a Tsinghua group, folds that summarization into reinforcement learning so the model is trained to produce good compressed states rather than relying on a hand-written heuristic. It matters for anyone building agents that work for hours across many tool calls. Builders hitting context walls on long-horizon agents should test a trained-summary loop against their current truncation approach and measure task completion, not just token count. arXiv

🚀 Frontier Models & Features

Quiet day on new frontier models. The scheduled event to watch is Moonshot's full open-weight release of Kimi K3, its 2.8-trillion-parameter mixture-of-experts model, due July 27. Weights at that scale are roughly 1.4 TB even under MXFP4 quantization, so "open" here does not mean easy to run. VentureBeat

🔬 Research Worth Reading

🏢 Enterprise in the Wild

Quiet day on named production case studies. The infrastructure signal came from AMD's event, where the company said its Helios systems have drawn multi-gigawatt customer commitments, a sign that buyers are contracting compute capacity years ahead of the workloads that will run on it. SiliconANGLE

🛠️ Tooling & Ecosystem

DeepSeek retires its legacy deepseek-chat and deepseek-reasoner API aliases today, July 24, at 15:59 UTC. Applications still pointing at those names must move to deepseek-v4-pro or deepseek-v4-flash or they break. It is a small change with a familiar lesson: pinning to a stable model ID is cheaper than chasing a moving default. DeepSeek V4 guide

⚖️ Policy & Regulation

The Moonshot accusation is as much trade policy as model gossip. Treasury raised the prospect of sanctions and export-control blacklisting, and the specific charge, that Moonshot reached banned Nvidia GB300 servers through Thailand, targets the third-country routing that has become the soft spot in U.S. chip controls. Separately, the EU AI Act's transparency duties and its enforcement powers over general-purpose models take effect August 2, the next hard compliance date for anyone serving models into the EU. TechCrunch

📌 Watch List