[07.05.2026] // 5 min read
Bridgewater fine-tuned Qwen3-235B on expert-labeled financial data and beat every frontier model at 1/14th the cost. The era of renting the biggest generalist may be ending.
[07.05.2026] // 4 min read
A GitHub issue reveals GPT-5.5 Codex responses cluster at exactly 516 reasoning tokens, correlating with task failures. 390K records analyzed, 82% of exact-516 events come from GPT-5.5.
[07.04.2026] // 5 min read
GitHub Models dies July 30. The free playground, API, and model catalog are all gone. Here's what happened, where to migrate, and why this keeps happening.
[07.04.2026] // 6 min read
Microsoft gave 5,000 engineers Claude Code. The annual AI budget lasted four months. Here's the token math that's breaking enterprise budgets.
[07.03.2026] // 5 min read
Rendering code as images before sending it to LLMs can cut token costs by 60%. But the savings come with accuracy tradeoffs and a pricing model that may not last.
[07.03.2026] // 4 min read
claude-real-video uses scene-change detection to extract the right frames from any video, giving LLMs meaningful visual data instead of random snapshots.
[07.02.2026] // 4 min read
Claude Sonnet 5's $2/M intro price hides a 30% tokenizer inflation that makes it cost more per task than Opus 4.8.
[07.02.2026] // 5 min read
Moonshot AI says K2.7 Code cuts reasoning tokens by 30%. Independent testers aren't convinced by the benchmarks. The pricing math still works.
[07.01.2026] // 6 min read
ZAI shipped a desktop app that runs Claude Code with GLM-5.2, an open-weight model that matches Opus on benchmarks at a fraction of the cost.
[07.01.2026] // 4 min read
Lilian Weng traces the full history of scaling law research, showing where Kaplan and Chinchilla disagree and why the exponents matter more than anyone admits.