730B Parameters Still Do Not Make GLM-5.3 Local
GLM-5.3 is open weight, but 141 shards and a million-token ceiling make local deployment an infrastructure decision, not a laptop install.
TOPIC_INDEX
58 published entries in this topic.
GLM-5.3 is open weight, but 141 shards and a million-token ceiling make local deployment an infrastructure decision, not a laptop install.
GLM-5.3-Flash makes hosted agent work look cheap at $0.045 per task, but its 320B total parameters change the local deployment math.
Unsloth Dynamic 3.0 makes 16GB Qwen3.8-27B deployments plausible, but only if you choose the quant for memory headroom and test real coding tasks.
GLM-5.3 looks impressive, but its missing weights and unsettled API change the decision. Here is who should test it now and who should wait.
WorldClaw turns one prompt into editable 3D scenes, but the public evidence still describes a research pipeline rather than a production-ready dependency.
Ling-3.0-flash is open weight, fast, and still a 124B model. Here is what 5.1B active parameters really change for self-hosted agents.
Inkling-Small uses 12B active parameters to approach its 975B sibling on coding and reasoning, but its factuality gap still matters.
DeepSeek V4 Flash 0731 posts an 82.7 on Terminal Bench 2.1 and brings cheap, Codex-ready agent coding to public beta.
Microsoft's 4B Mage-Flow pairs a cheap tokenizer with native-resolution diffusion, hitting 0.59 seconds at 1024² on one A100 while exposing the tradeoffs.
Claude Opus 5 keeps Opus 4.8 pricing, adds adaptive effort controls, and targets cheaper long-running agent work.
Moonshot AI's 2.8T-parameter Kimi K3 is close to the closed frontier, but the real test starts when its promised weights arrive.
Vidu S1 targets live voice-controlled video at 25 FPS, with a reported 42 FPS peak on an RTX 5090.
Google shipped three Gemini models but the promised Pro update remains missing as Gemini 4 pre-training begins. 3.6 Flash delivers 17% fewer tokens and better coding scores.
Alibaba claims Qwen 3.8 is second only to Claude Fable 5. But there are no benchmarks, no model card, and no open weights yet. Here is what we actually know.
Moonshot AI dropped Kimi K3, a 2.8-trillion-parameter open model that matches Claude and GPT on key agentic benchmarks. Open weights arrive July 27.
PrismML's Bonsai 27B compresses a full 27B-class reasoning model to 3.9 GB, running on an iPhone 17 Pro at 11 tok/s. The tradeoffs are real, but so is the trajectory.
A 744B-parameter open-weight model from Beijing lands within a point of Claude Opus 4.8 on agentic benchmarks at one-fifth the cost. The margin mirage is over.
Ploy published every number from moving their production AI agent from Claude Opus to GPT-5.6 Sol. The result: 2.2x faster builds, 27% lower cost, better visual scores, and three hard-won lessons.
xAI's Grok 4.5 was trained on 2 trillion tokens of Cursor developer interactions. It costs $2 per million input tokens, uses 4.2x fewer tokens than Opus 4.8, and trades benchmark leadership for raw cost-per-task efficiency.
Meta's Muse Spark 1.1 API costs $1.25 per million input tokens, undercutting Claude and GPT by 75-80%. The real story isn't the benchmarks.
OpenAI shipped GPT-Live, a full-duplex voice model that listens and speaks simultaneously. The awkward pause is finally dead.
OpenAI's GPT-5.6 Sol benchmarks competitively with Mythos but only 20 government-approved companies can use it. Three tiers, Cerebras at 750 TPS, and a regulatory framework that changes everything.
Claude Sonnet 5's $2/M intro price hides a 30% tokenizer inflation that makes it cost more per task than Opus 4.8.
Moonshot AI says K2.7 Code cuts reasoning tokens by 30%. Independent testers aren't convinced by the benchmarks. The pricing math still works.
Meituan's 1.6T open model matched GPT-5.5 on SWE-bench Pro, trained on 50K domestic cards. MIT license, $0.75/M input tokens.
Zhipu AI's GLM-5.2 is a 753B open-weight MoE model under MIT license that matches Claude Opus on coding benchmarks, with a genuine 1M token context window and $18/month flat-rate pricing.
DeepReinforce's Ornith-1.0 learns to write its own RL scaffolds, and the 35B model beats Qwen 3.5-397B on terminal tasks.
OpenAI shipped GPT-5.6 Sol with a 91.9% coding score. Almost nobody can use it. The US government controls who gets access.
GLM-5.2, a 753B open-weights model from Z.ai, just beat GPT-5.5 on multiple coding benchmarks at one-sixth the cost under an MIT license.
Claude Fable 5 lasted exactly three days before the US government forced Anthropic to pull it. The first model killed by export controls sets a permanent precedent.
Apple announced AFM 3 at WWDC26: a 20B sparse on-device model using Instruction-Following Pruning. Only iPhone 17 Pro/Max/Air can run it. The cloud tier depends on Google's NVIDIA GPUs and Gemini distillation.
Google released DiffusionGemma, a 26B open model that generates text 4x faster by replacing token-by-token decoding with parallel diffusion. The speed is real. The quality gap is too.
Zhipu shipped GLM 5.2 with 1M token context to their coding plan today. No benchmarks. No independent verification. MIT weights coming next week. Here's what we know and what we don't.
Moonshot AI released Kimi K2.7-Code, a 1T-parameter open-source coding model that cuts reasoning tokens by 30% at $0.95/M input. The real story is the economics, not the benchmarks.
Anthropic released Claude Fable 5 and Mythos 5: the same model, split by safety tier. The benchmarks you see everywhere mostly belong to the version you cannot buy.
Ideogram 4.0 is the first open-weight image model to top the DesignArena leaderboard, with 0.97 OCR accuracy and structured JSON prompting that gives designers precise layout control. It runs on a single 24GB GPU.
NVIDIA's Nemotron 3 Ultra is the fastest US open-weight model at 300+ tokens per second, built specifically for long-running agents. 550B parameters, 55B active, hybrid Mamba-Transformer architecture.
NVIDIA released Cosmos 3, an open omnimodel that merges vision reasoning, world generation, and robot action prediction into a single architecture. The five-model stack for physical AI may be obsolete.
MoE architecture and llama.cpp turned a 2020 gaming card into a local AI workstation. Here's the real hardware data.
Google DeepMind released Gemma 4 12B, an encoder-free multimodal model that runs on 16GB laptops and processes text, images, audio, and video through a single transformer. Apache 2.0.
PrismML's Bonsai Image 4B compresses FLUX.2 Klein down to 0.93GB using 1-bit quantization, enabling full local image generation on iPhones, MacBooks, and in-browser via WebGPU. Apache 2.0.
Liquid AI's LFM2.5-8B-A1B brings 128K context to consumer hardware — 1.5B active parameters, 253 tok/s on a laptop, and tool calling that finally feels interactive.
Liquid AI's LFM2.5-8B-A1B packs 128K context and real tool calling into 1.5B active parameters. It runs at 253 tok/s on a laptop. Local agents just became practical.
Tencent's Hy3 preview has quietly become the most-used model on OpenRouter. Here's the tech, the pricing, and what it says about AI in 2026.
Chinese open-weight models now handle over 60% of tokens on OpenRouter, up from 2% a year ago. Xiaomi, Alibaba, DeepSeek, and others are winning on price-performance at 10-20x lower cost for comparable quality. Here is how it happened and what it means for developers.
Cohere released Command A+, a 218B MoE model under Apache 2.0 that runs on two H100s. It's their bid to make sovereign enterprise AI practical , frontier-grade model, no vendor lock-in, deployable on premises or air-gapped.
Google I/O 2026 shifted focus from bigger models to faster ones, launching Gemini 3.5 Flash, Omni for video generation, and Antigravity 2.0 for always-on AI agents.
A Miami startup just claimed 1,000x efficiency gain with a new attention architecture. 13 employees, 11 PhDs, $29M seed, and zero public access. I want it to be real. I also remember Magic.dev.
A Reddit user found Multi-Token Prediction heads hidden inside Gemma 4's model weights. Google said they were 'removed on purpose.' Then Google officially released them anyway. Here's how MTP works and why it matters.
Gemma 4 31B and Qwen3.6 27B both had their refusal mechanisms removed — but with completely different abliteration methods. Forensic analysis shows technique matters more than model.
xAI's Grok 4.3 beta runs a 16-agent architecture at 209 tok/s with up to 2M token context—competitive intelligence at a quarter of Claude's input cost.
IBM's Granite 4.1 8B dense model matches 32B MoE benchmarks while running on consumer hardware. 20x token efficiency vs Qwen.
Google DeepMind's Gemma 4 brings frontier multimodal AI to phones. Apache 2.0 licensed, 140+ languages, thinking mode included.
A 7B model trained with RL achieves DeepSeek-R1-671B level website generation using scaffold-driven generation and cascaded multimodal rewards.
GPT-5.5 released April 23, 2026. 82.7% Terminal-Bench, 84.9% GDPval—beats GPT-5.4, Claude Opus 4.7, Gemini 3.1 Pro. Matches GPT-5.4 latency while delivering higher intelligence. Fewer tokens for same tasks.
DeepSeek V4 just dropped: 1.6T MoE with 1M context, MIT license. Beats GPT-5 and Gemini on LiveCodeBench (93.5%). 27% FLOPs efficiency vs V3.2. Open weights, aggressive pricing.
Alibaba's Qwen3.6-35B-A3B delivers Qwen3.5-27B performance with 3B active parameters. SWE-bench 73.4%, Apache 2.0, runs at 100 t/s on M5 Max.
Inclusion AI's 16B MoE diffusion LLM unifies multimodal understanding and generation, closing the gap with specialized models for the first time.