<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>models — NestFrontier</title><description>Technical AI analysis and research on models.</description><link>https://nestfrontier.com/</link><item><title>730B Parameters Still Do Not Make GLM-5.3 Local</title><link>https://nestfrontier.com/730b-parameters-still-do-not-make-glm-53-local/</link><guid isPermaLink="true">https://nestfrontier.com/730b-parameters-still-do-not-make-glm-53-local/</guid><description>GLM-5.3 is open weight, but 141 shards and a million-token ceiling make local deployment an infrastructure decision, not a laptop install.</description><pubDate>Sat, 29 Aug 2026 00:06:42 GMT</pubDate></item><item><title>A 4.5 Cent Agent Model Still Needs Serious Hardware</title><link>https://nestfrontier.com/a-45-cent-agent-model-still-needs-serious-hardware/</link><guid isPermaLink="true">https://nestfrontier.com/a-45-cent-agent-model-still-needs-serious-hardware/</guid><description>GLM-5.3-Flash makes hosted agent work look cheap at $0.045 per task, but its 320B total parameters change the local deployment math.</description><pubDate>Thu, 27 Aug 2026 00:06:17 GMT</pubDate></item><item><title>16GB Is Enough for Qwen3.8 If You Pick the Right Quant</title><link>https://nestfrontier.com/16gb-is-enough-for-qwen38-if-you-pick-the-right-quant/</link><guid isPermaLink="true">https://nestfrontier.com/16gb-is-enough-for-qwen38-if-you-pick-the-right-quant/</guid><description>Unsloth Dynamic 3.0 makes 16GB Qwen3.8-27B deployments plausible, but only if you choose the quant for memory headroom and test real coding tasks.</description><pubDate>Thu, 20 Aug 2026 05:30:06 GMT</pubDate></item><item><title>The Open-Weight Wait Changes GLM-5.3&apos;s Value</title><link>https://nestfrontier.com/the-open-weight-wait-changes-glm-53s-value/</link><guid isPermaLink="true">https://nestfrontier.com/the-open-weight-wait-changes-glm-53s-value/</guid><description>GLM-5.3 looks impressive, but its missing weights and unsettled API change the decision. Here is who should test it now and who should wait.</description><pubDate>Wed, 19 Aug 2026 00:04:59 GMT</pubDate></item><item><title>The 3D world generator is a pipeline, not a model</title><link>https://nestfrontier.com/the-3d-world-generator-is-a-pipeline-not-a-model/</link><guid isPermaLink="true">https://nestfrontier.com/the-3d-world-generator-is-a-pipeline-not-a-model/</guid><description>WorldClaw turns one prompt into editable 3D scenes, but the public evidence still describes a research pipeline rather than a production-ready dependency.</description><pubDate>Wed, 12 Aug 2026 00:04:05 GMT</pubDate></item><item><title>Open weights do not make Ling 3.0 Flash small</title><link>https://nestfrontier.com/open-weights-do-not-make-ling-30-flash-small/</link><guid isPermaLink="true">https://nestfrontier.com/open-weights-do-not-make-ling-30-flash-small/</guid><description>Ling-3.0-flash is open weight, fast, and still a 124B model. Here is what 5.1B active parameters really change for self-hosted agents.</description><pubDate>Wed, 05 Aug 2026 12:04:06 GMT</pubDate></item><item><title>A Quarter the Size Still Hits 80% on SWE-bench</title><link>https://nestfrontier.com/a-quarter-the-size-still-hits-80percent-on-swe-bench/</link><guid isPermaLink="true">https://nestfrontier.com/a-quarter-the-size-still-hits-80percent-on-swe-bench/</guid><description>Inkling-Small uses 12B active parameters to approach its 975B sibling on coding and reasoning, but its factuality gap still matters.</description><pubDate>Mon, 03 Aug 2026 00:03:35 GMT</pubDate></item><item><title>82.7 Makes Cheap AI Coding Hard to Ignore</title><link>https://nestfrontier.com/827-makes-cheap-ai-coding-hard-to-ignore/</link><guid isPermaLink="true">https://nestfrontier.com/827-makes-cheap-ai-coding-hard-to-ignore/</guid><description>DeepSeek V4 Flash 0731 posts an 82.7 on Terminal Bench 2.1 and brings cheap, Codex-ready agent coding to public beta.</description><pubDate>Fri, 31 Jul 2026 12:03:11 GMT</pubDate></item><item><title>Four Billion Parameters Beat the Big Image Models on Speed</title><link>https://nestfrontier.com/four-billion-parameters-beat-the-big-image-models-on-speed/</link><guid isPermaLink="true">https://nestfrontier.com/four-billion-parameters-beat-the-big-image-models-on-speed/</guid><description>Microsoft&apos;s 4B Mage-Flow pairs a cheap tokenizer with native-resolution diffusion, hitting 0.59 seconds at 1024² on one A100 while exposing the tradeoffs.</description><pubDate>Tue, 28 Jul 2026 00:03:19 GMT</pubDate></item><item><title>Half the price is easy. Token efficiency is the story.</title><link>https://nestfrontier.com/half-the-price-is-easy-token-efficiency-is-the-story/</link><guid isPermaLink="true">https://nestfrontier.com/half-the-price-is-easy-token-efficiency-is-the-story/</guid><description>Claude Opus 5 keeps Opus 4.8 pricing, adds adaptive effort controls, and targets cheaper long-running agent work.</description><pubDate>Sat, 25 Jul 2026 12:03:54 GMT</pubDate></item><item><title>Open weights just reached 2.8 trillion parameters</title><link>https://nestfrontier.com/open-weights-just-reached-28-trillion-parameters/</link><guid isPermaLink="true">https://nestfrontier.com/open-weights-just-reached-28-trillion-parameters/</guid><description>Moonshot AI&apos;s 2.8T-parameter Kimi K3 is close to the closed frontier, but the real test starts when its promised weights arrive.</description><pubDate>Sat, 25 Jul 2026 00:04:22 GMT</pubDate></item><item><title>Video generation now has to listen</title><link>https://nestfrontier.com/video-generation-now-has-to-listen/</link><guid isPermaLink="true">https://nestfrontier.com/video-generation-now-has-to-listen/</guid><description>Vidu S1 targets live voice-controlled video at 25 FPS, with a reported 42 FPS peak on an RTX 5090.</description><pubDate>Fri, 24 Jul 2026 00:03:14 GMT</pubDate></item><item><title>Three Gemini models shipped, but Pro is still missing</title><link>https://nestfrontier.com/three-gemini-models-shipped-but-pro-is-still-missing/</link><guid isPermaLink="true">https://nestfrontier.com/three-gemini-models-shipped-but-pro-is-still-missing/</guid><description>Google shipped three Gemini models but the promised Pro update remains missing as Gemini 4 pre-training begins. 3.6 Flash delivers 17% fewer tokens and better coding scores.</description><pubDate>Wed, 22 Jul 2026 12:15:51 GMT</pubDate></item><item><title>2.4 trillion parameters and no benchmarks to prove it</title><link>https://nestfrontier.com/24-trillion-parameters-and-no-benchmarks-to-prove-it/</link><guid isPermaLink="true">https://nestfrontier.com/24-trillion-parameters-and-no-benchmarks-to-prove-it/</guid><description>Alibaba claims Qwen 3.8 is second only to Claude Fable 5. But there are no benchmarks, no model card, and no open weights yet. Here is what we actually know.</description><pubDate>Mon, 20 Jul 2026 00:28:51 GMT</pubDate></item><item><title>2.8 trillion open-source parameters just changed the AI race</title><link>https://nestfrontier.com/28-trillion-open-source-parameters-just-changed-the-ai-race/</link><guid isPermaLink="true">https://nestfrontier.com/28-trillion-open-source-parameters-just-changed-the-ai-race/</guid><description>Moonshot AI dropped Kimi K3, a 2.8-trillion-parameter open model that matches Claude and GPT on key agentic benchmarks. Open weights arrive July 27.</description><pubDate>Fri, 17 Jul 2026 00:25:34 GMT</pubDate></item><item><title>That 27B model was too big for a phone. Not anymore.</title><link>https://nestfrontier.com/that-27b-model-was-too-big-for-a-phone-not-anymore/</link><guid isPermaLink="true">https://nestfrontier.com/that-27b-model-was-too-big-for-a-phone-not-anymore/</guid><description>PrismML&apos;s Bonsai 27B compresses a full 27B-class reasoning model to 3.9 GB, running on an iPhone 17 Pro at 11 tok/s. The tradeoffs are real, but so is the trajectory.</description><pubDate>Wed, 15 Jul 2026 00:19:10 GMT</pubDate></item><item><title>$4.40 per million tokens just matched the $200 tier</title><link>https://nestfrontier.com/dollar440-per-million-tokens-just-matched-the-dollar200-tier/</link><guid isPermaLink="true">https://nestfrontier.com/dollar440-per-million-tokens-just-matched-the-dollar200-tier/</guid><description>A 744B-parameter open-weight model from Beijing lands within a point of Claude Opus 4.8 on agentic benchmarks at one-fifth the cost. The margin mirage is over.</description><pubDate>Tue, 14 Jul 2026 12:10:23 GMT</pubDate></item><item><title>Production AI build times just got cut in half</title><link>https://nestfrontier.com/production-ai-build-times-just-got-cut-in-half/</link><guid isPermaLink="true">https://nestfrontier.com/production-ai-build-times-just-got-cut-in-half/</guid><description>Ploy published every number from moving their production AI agent from Claude Opus to GPT-5.6 Sol. The result: 2.2x faster builds, 27% lower cost, better visual scores, and three hard-won lessons.</description><pubDate>Mon, 13 Jul 2026 12:23:09 GMT</pubDate></item><item><title>A coding model trained on 2 trillion interactions just undercut everyone</title><link>https://nestfrontier.com/a-coding-model-trained-on-2-trillion-interactions-just-undercut-everyone/</link><guid isPermaLink="true">https://nestfrontier.com/a-coding-model-trained-on-2-trillion-interactions-just-undercut-everyone/</guid><description>xAI&apos;s Grok 4.5 was trained on 2 trillion tokens of Cursor developer interactions. It costs $2 per million input tokens, uses 4.2x fewer tokens than Opus 4.8, and trades benchmark leadership for raw cost-per-task efficiency.</description><pubDate>Fri, 10 Jul 2026 08:06:44 GMT</pubDate></item><item><title>That $1.25 API price just forced Anthropic and OpenAI to respond</title><link>https://nestfrontier.com/that-dollar125-api-price-just-forced-anthropic-and-openai-to-respond/</link><guid isPermaLink="true">https://nestfrontier.com/that-dollar125-api-price-just-forced-anthropic-and-openai-to-respond/</guid><description>Meta&apos;s Muse Spark 1.1 API costs $1.25 per million input tokens, undercutting Claude and GPT by 75-80%. The real story isn&apos;t the benchmarks.</description><pubDate>Thu, 09 Jul 2026 20:06:43 GMT</pubDate></item><item><title>The AI voice interrupt problem just got fixed</title><link>https://nestfrontier.com/the-ai-voice-interrupt-problem-just-got-fixed/</link><guid isPermaLink="true">https://nestfrontier.com/the-ai-voice-interrupt-problem-just-got-fixed/</guid><description>OpenAI shipped GPT-Live, a full-duplex voice model that listens and speaks simultaneously. The awkward pause is finally dead.</description><pubDate>Wed, 08 Jul 2026 20:07:21 GMT</pubDate></item><item><title>Only 20 companies got access to GPT-5.6. The government decided.</title><link>https://nestfrontier.com/only-20-companies-got-access-to-gpt-56-the-government-decided/</link><guid isPermaLink="true">https://nestfrontier.com/only-20-companies-got-access-to-gpt-56-the-government-decided/</guid><description>OpenAI&apos;s GPT-5.6 Sol benchmarks competitively with Mythos but only 20 government-approved companies can use it. Three tiers, Cerebras at 750 TPS, and a regulatory framework that changes everything.</description><pubDate>Mon, 06 Jul 2026 08:07:47 GMT</pubDate></item><item><title>That $2/M intro price is hiding a 30% tokenizer tax</title><link>https://nestfrontier.com/that-dollar2m-intro-price-is-hiding-a-30percent-tokenizer-tax/</link><guid isPermaLink="true">https://nestfrontier.com/that-dollar2m-intro-price-is-hiding-a-30percent-tokenizer-tax/</guid><description>Claude Sonnet 5&apos;s $2/M intro price hides a 30% tokenizer inflation that makes it cost more per task than Opus 4.8.</description><pubDate>Thu, 02 Jul 2026 20:06:11 GMT</pubDate></item><item><title>Stop overthinking: a coding model just cut reasoning tokens by 30%</title><link>https://nestfrontier.com/stop-overthinking-a-coding-model-just-cut-reasoning-tokens-by-30percent/</link><guid isPermaLink="true">https://nestfrontier.com/stop-overthinking-a-coding-model-just-cut-reasoning-tokens-by-30percent/</guid><description>Moonshot AI says K2.7 Code cuts reasoning tokens by 30%. Independent testers aren&apos;t convinced by the benchmarks. The pricing math still works.</description><pubDate>Thu, 02 Jul 2026 08:06:34 GMT</pubDate></item><item><title>Meituan&apos;s open model matched GPT-5.5. They deliver food for a living.</title><link>https://nestfrontier.com/meituans-open-model-matched-gpt-55-they-deliver-food-for-a-living/</link><guid isPermaLink="true">https://nestfrontier.com/meituans-open-model-matched-gpt-55-they-deliver-food-for-a-living/</guid><description>Meituan&apos;s 1.6T open model matched GPT-5.5 on SWE-bench Pro, trained on 50K domestic cards. MIT license, $0.75/M input tokens.</description><pubDate>Tue, 30 Jun 2026 08:03:49 GMT</pubDate></item><item><title>A 753B open-source model just matched Claude on coding</title><link>https://nestfrontier.com/a-753b-open-source-model-just-matched-claude-on-coding/</link><guid isPermaLink="true">https://nestfrontier.com/a-753b-open-source-model-just-matched-claude-on-coding/</guid><description>Zhipu AI&apos;s GLM-5.2 is a 753B open-weight MoE model under MIT license that matches Claude Opus on coding benchmarks, with a genuine 1M token context window and $18/month flat-rate pricing.</description><pubDate>Sat, 27 Jun 2026 20:04:36 GMT</pubDate></item><item><title>Your coding agent&apos;s harness is the bottleneck it can&apos;t see</title><link>https://nestfrontier.com/your-coding-agents-harness-is-the-bottleneck-it-cant-see/</link><guid isPermaLink="true">https://nestfrontier.com/your-coding-agents-harness-is-the-bottleneck-it-cant-see/</guid><description>DeepReinforce&apos;s Ornith-1.0 learns to write its own RL scaffolds, and the 35B model beats Qwen 3.5-397B on terminal tasks.</description><pubDate>Sat, 27 Jun 2026 09:35:23 GMT</pubDate></item><item><title>91.9% coding score. The government decides who uses it.</title><link>https://nestfrontier.com/919percent-coding-score-the-government-decides-who-uses-it/</link><guid isPermaLink="true">https://nestfrontier.com/919percent-coding-score-the-government-decides-who-uses-it/</guid><description>OpenAI shipped GPT-5.6 Sol with a 91.9% coding score. Almost nobody can use it. The US government controls who gets access.</description><pubDate>Sat, 27 Jun 2026 08:05:31 GMT</pubDate></item><item><title>That expensive API subscription just lost to a free download</title><link>https://nestfrontier.com/that-expensive-api-subscription-just-lost-to-a-free-download/</link><guid isPermaLink="true">https://nestfrontier.com/that-expensive-api-subscription-just-lost-to-a-free-download/</guid><description>GLM-5.2, a 753B open-weights model from Z.ai, just beat GPT-5.5 on multiple coding benchmarks at one-sixth the cost under an MIT license.</description><pubDate>Sat, 20 Jun 2026 08:06:05 GMT</pubDate></item><item><title>That 72-Hour AI Model Got Banned by the Government</title><link>https://nestfrontier.com/that-72-hour-ai-model-got-banned-by-the-government/</link><guid isPermaLink="true">https://nestfrontier.com/that-72-hour-ai-model-got-banned-by-the-government/</guid><description>Claude Fable 5 lasted exactly three days before the US government forced Anthropic to pull it. The first model killed by export controls sets a permanent precedent.</description><pubDate>Fri, 19 Jun 2026 20:03:22 GMT</pubDate></item><item><title>Your iPhone can&apos;t run Apple&apos;s new 20B AI model</title><link>https://nestfrontier.com/your-iphone-cant-run-apples-new-20b-ai-model/</link><guid isPermaLink="true">https://nestfrontier.com/your-iphone-cant-run-apples-new-20b-ai-model/</guid><description>Apple announced AFM 3 at WWDC26: a 20B sparse on-device model using Instruction-Following Pruning. Only iPhone 17 Pro/Max/Air can run it. The cloud tier depends on Google&apos;s NVIDIA GPUs and Gemini distillation.</description><pubDate>Mon, 15 Jun 2026 20:03:58 GMT</pubDate></item><item><title>One token at a time is a bottleneck. Google broke it.</title><link>https://nestfrontier.com/one-token-at-a-time-is-a-bottleneck-google-broke-it/</link><guid isPermaLink="true">https://nestfrontier.com/one-token-at-a-time-is-a-bottleneck-google-broke-it/</guid><description>Google released DiffusionGemma, a 26B open model that generates text 4x faster by replacing token-by-token decoding with parallel diffusion. The speed is real. The quality gap is too.</description><pubDate>Sun, 14 Jun 2026 08:05:16 GMT</pubDate></item><item><title>Zhipu just shipped 1M context to coding. Where are the benchmarks?</title><link>https://nestfrontier.com/zhipu-just-shipped-1m-context-to-coding-where-are-the-benchmarks/</link><guid isPermaLink="true">https://nestfrontier.com/zhipu-just-shipped-1m-context-to-coding-where-are-the-benchmarks/</guid><description>Zhipu shipped GLM 5.2 with 1M token context to their coding plan today. No benchmarks. No independent verification. MIT weights coming next week. Here&apos;s what we know and what we don&apos;t.</description><pubDate>Sat, 13 Jun 2026 20:03:30 GMT</pubDate></item><item><title>Your coding agent wastes tokens thinking. This one doesn&apos;t.</title><link>https://nestfrontier.com/your-coding-agent-wastes-tokens-thinking-this-one-doesnt/</link><guid isPermaLink="true">https://nestfrontier.com/your-coding-agent-wastes-tokens-thinking-this-one-doesnt/</guid><description>Moonshot AI released Kimi K2.7-Code, a 1T-parameter open-source coding model that cuts reasoning tokens by 30% at $0.95/M input. The real story is the economics, not the benchmarks.</description><pubDate>Fri, 12 Jun 2026 20:05:46 GMT</pubDate></item><item><title>Anthropic&apos;s strongest model has two versions. You get the weaker one.</title><link>https://nestfrontier.com/anthropics-strongest-model-has-two-versions-you-get-the-weaker-one/</link><guid isPermaLink="true">https://nestfrontier.com/anthropics-strongest-model-has-two-versions-you-get-the-weaker-one/</guid><description>Anthropic released Claude Fable 5 and Mythos 5: the same model, split by safety tier. The benchmarks you see everywhere mostly belong to the version you cannot buy.</description><pubDate>Tue, 09 Jun 2026 20:03:05 GMT</pubDate></item><item><title>Open-weight image models finally stopped losing to closed ones</title><link>https://nestfrontier.com/open-weight-image-models-finally-stopped-losing-to-closed-ones/</link><guid isPermaLink="true">https://nestfrontier.com/open-weight-image-models-finally-stopped-losing-to-closed-ones/</guid><description>Ideogram 4.0 is the first open-weight image model to top the DesignArena leaderboard, with 0.97 OCR accuracy and structured JSON prompting that gives designers precise layout control. It runs on a single 24GB GPU.</description><pubDate>Mon, 08 Jun 2026 20:04:01 GMT</pubDate></item><item><title>US Open-Weight AI Finally Has a Fast One</title><link>https://nestfrontier.com/us-open-weight-ai-finally-has-a-fast-one/</link><guid isPermaLink="true">https://nestfrontier.com/us-open-weight-ai-finally-has-a-fast-one/</guid><description>NVIDIA&apos;s Nemotron 3 Ultra is the fastest US open-weight model at 300+ tokens per second, built specifically for long-running agents. 550B parameters, 55B active, hybrid Mamba-Transformer architecture.</description><pubDate>Mon, 08 Jun 2026 08:06:42 GMT</pubDate></item><item><title>NVIDIA collapsed 5 AI models into one. Robots just got cheaper.</title><link>https://nestfrontier.com/nvidia-collapsed-5-ai-models-into-one-robots-just-got-cheaper/</link><guid isPermaLink="true">https://nestfrontier.com/nvidia-collapsed-5-ai-models-into-one-robots-just-got-cheaper/</guid><description>NVIDIA released Cosmos 3, an open omnimodel that merges vision reasoning, world generation, and robot action prediction into a single architecture. The five-model stack for physical AI may be obsolete.</description><pubDate>Fri, 05 Jun 2026 20:04:19 GMT</pubDate></item><item><title>Your 3070 runs a 35B model at 22 tokens per second</title><link>https://nestfrontier.com/your-3070-runs-a-35b-model-at-22-tokens-per-second/</link><guid isPermaLink="true">https://nestfrontier.com/your-3070-runs-a-35b-model-at-22-tokens-per-second/</guid><description>MoE architecture and llama.cpp turned a 2020 gaming card into a local AI workstation. Here&apos;s the real hardware data.</description><pubDate>Fri, 05 Jun 2026 04:05:15 GMT</pubDate></item><item><title>Your laptop runs multimodal AI now. Nobody told you.</title><link>https://nestfrontier.com/your-laptop-runs-multimodal-ai-now-nobody-told-you/</link><guid isPermaLink="true">https://nestfrontier.com/your-laptop-runs-multimodal-ai-now-nobody-told-you/</guid><description>Google DeepMind released Gemma 4 12B, an encoder-free multimodal model that runs on 16GB laptops and processes text, images, audio, and video through a single transformer. Apache 2.0.</description><pubDate>Thu, 04 Jun 2026 20:06:29 GMT</pubDate></item><item><title>Your phone can now run FLUX image generation locally</title><link>https://nestfrontier.com/your-phone-can-now-run-flux-image-generation-locally/</link><guid isPermaLink="true">https://nestfrontier.com/your-phone-can-now-run-flux-image-generation-locally/</guid><description>PrismML&apos;s Bonsai Image 4B compresses FLUX.2 Klein down to 0.93GB using 1-bit quantization, enabling full local image generation on iPhones, MacBooks, and in-browser via WebGPU. Apache 2.0.</description><pubDate>Sun, 31 May 2026 20:06:14 GMT</pubDate></item><item><title>Your laptop just got 128K context. At 253 tokens per second.</title><link>https://nestfrontier.com/your-laptop-just-got-128k-context-at-253-tokens-per-second/</link><guid isPermaLink="true">https://nestfrontier.com/your-laptop-just-got-128k-context-at-253-tokens-per-second/</guid><description>Liquid AI&apos;s LFM2.5-8B-A1B brings 128K context to consumer hardware — 1.5B active parameters, 253 tok/s on a laptop, and tool calling that finally feels interactive.</description><pubDate>Sat, 30 May 2026 08:11:14 GMT</pubDate></item><item><title>On-device AI agents no longer cap out at 32K context</title><link>https://nestfrontier.com/on-device-ai-agents-no-longer-cap-out-at-32k-context/</link><guid isPermaLink="true">https://nestfrontier.com/on-device-ai-agents-no-longer-cap-out-at-32k-context/</guid><description>Liquid AI&apos;s LFM2.5-8B-A1B packs 128K context and real tool calling into 1.5B active parameters. It runs at 253 tok/s on a laptop. Local agents just became practical.</description><pubDate>Fri, 29 May 2026 20:08:50 GMT</pubDate></item><item><title>Nobody was watching this model until it crushed OpenRouter</title><link>https://nestfrontier.com/nobody-was-watching-this-model-until-it-crushed-openrouter/</link><guid isPermaLink="true">https://nestfrontier.com/nobody-was-watching-this-model-until-it-crushed-openrouter/</guid><description>Tencent&apos;s Hy3 preview has quietly become the most-used model on OpenRouter. Here&apos;s the tech, the pricing, and what it says about AI in 2026.</description><pubDate>Fri, 29 May 2026 08:05:30 GMT</pubDate></item><item><title>Chinese open-weight models are eating Silicon Valley&apos;s lunch</title><link>https://nestfrontier.com/chinese-open-weight-models-are-eating-silicon-valleys-lunch/</link><guid isPermaLink="true">https://nestfrontier.com/chinese-open-weight-models-are-eating-silicon-valleys-lunch/</guid><description>Chinese open-weight models now handle over 60% of tokens on OpenRouter, up from 2% a year ago. Xiaomi, Alibaba, DeepSeek, and others are winning on price-performance at 10-20x lower cost for comparable quality. Here is how it happened and what it means for developers.</description><pubDate>Mon, 25 May 2026 20:05:01 GMT</pubDate></item><item><title>Enterprise-grade AI that actually runs on premises, not just in theory</title><link>https://nestfrontier.com/enterprise-grade-ai-that-actually-runs-on-premises-not-just-in-theory/</link><guid isPermaLink="true">https://nestfrontier.com/enterprise-grade-ai-that-actually-runs-on-premises-not-just-in-theory/</guid><description>Cohere released Command A+, a 218B MoE model under Apache 2.0 that runs on two H100s. It&apos;s their bid to make sovereign enterprise AI practical ,  frontier-grade model, no vendor lock-in, deployable on premises or air-gapped.</description><pubDate>Fri, 22 May 2026 08:05:28 GMT</pubDate></item><item><title>Google Abandoned the Bigger Model Race for Speed</title><link>https://nestfrontier.com/google-abandoned-the-bigger-model-race-for-speed/</link><guid isPermaLink="true">https://nestfrontier.com/google-abandoned-the-bigger-model-race-for-speed/</guid><description>Google I/O 2026 shifted focus from bigger models to faster ones, launching Gemini 3.5 Flash, Omni for video generation, and Antigravity 2.0 for always-on AI agents.</description><pubDate>Wed, 20 May 2026 08:04:38 GMT</pubDate></item><item><title>Your $2,600 RULER Run Just Cost $8</title><link>https://nestfrontier.com/your-dollar2600-ruler-run-just-cost-dollar8/</link><guid isPermaLink="true">https://nestfrontier.com/your-dollar2600-ruler-run-just-cost-dollar8/</guid><description>A Miami startup just claimed 1,000x efficiency gain with a new attention architecture. 13 employees, 11 PhDs, $29M seed, and zero public access. I want it to be real. I also remember Magic.dev.</description><pubDate>Sun, 10 May 2026 14:37:09 GMT</pubDate></item><item><title>Google Hid a 3x Speedup in Gemma 4. The Community Found It in Three Days.</title><link>https://nestfrontier.com/google-hid-a-3x-speedup-in-gemma-4-the-community-found-it-in-three-days/</link><guid isPermaLink="true">https://nestfrontier.com/google-hid-a-3x-speedup-in-gemma-4-the-community-found-it-in-three-days/</guid><description>A Reddit user found Multi-Token Prediction heads hidden inside Gemma 4&apos;s model weights. Google said they were &apos;removed on purpose.&apos; Then Google officially released them anyway. Here&apos;s how MTP works and why it matters.</description><pubDate>Wed, 06 May 2026 01:26:31 GMT</pubDate></item><item><title>Your uncensoring technique matters more than the model you modify</title><link>https://nestfrontier.com/your-uncensoring-technique-matters-more-than-the-model-you-modify/</link><guid isPermaLink="true">https://nestfrontier.com/your-uncensoring-technique-matters-more-than-the-model-you-modify/</guid><description>Gemma 4 31B and Qwen3.6 27B both had their refusal mechanisms removed — but with completely different abliteration methods. Forensic analysis shows technique matters more than model.</description><pubDate>Mon, 04 May 2026 00:52:30 GMT</pubDate></item><item><title>Grok 4.3: xAI&apos;s Heavy Multi-Agent Engine</title><link>https://nestfrontier.com/grok-43-xais-heavy-multi-agent-engine/</link><guid isPermaLink="true">https://nestfrontier.com/grok-43-xais-heavy-multi-agent-engine/</guid><description>xAI&apos;s Grok 4.3 beta runs a 16-agent architecture at 209 tok/s with up to 2M token context—competitive intelligence at a quarter of Claude&apos;s input cost.</description><pubDate>Fri, 01 May 2026 21:37:55 GMT</pubDate></item><item><title>IBM Granite 4.1: The 8B Dense Model That&apos;s Outperforming 32B MoEs</title><link>https://nestfrontier.com/ibm-granite-41-the-8b-dense-model-thats-outperforming-32b-moes/</link><guid isPermaLink="true">https://nestfrontier.com/ibm-granite-41-the-8b-dense-model-thats-outperforming-32b-moes/</guid><description>IBM&apos;s Granite 4.1 8B dense model matches 32B MoE benchmarks while running on consumer hardware. 20x token efficiency vs Qwen.</description><pubDate>Thu, 30 Apr 2026 12:08:24 GMT</pubDate></item><item><title>Gemma 4: Google&apos;s Open Multimodal That Runs on Your Phone</title><link>https://nestfrontier.com/gemma-4-googles-open-multimodal-that-runs-on-your-phone/</link><guid isPermaLink="true">https://nestfrontier.com/gemma-4-googles-open-multimodal-that-runs-on-your-phone/</guid><description>Google DeepMind&apos;s Gemma 4 brings frontier multimodal AI to phones. Apache 2.0 licensed, 140+ languages, thinking mode included.</description><pubDate>Mon, 27 Apr 2026 07:52:53 GMT</pubDate></item><item><title>WebGen-R1: 7B Model Rivals DeepSeek-R1-671B for Website Generation</title><link>https://nestfrontier.com/webgen-r1-7b-model-rivals-deepseek-r1-671b-for-website-generation/</link><guid isPermaLink="true">https://nestfrontier.com/webgen-r1-7b-model-rivals-deepseek-r1-671b-for-website-generation/</guid><description>A 7B model trained with RL achieves DeepSeek-R1-671B level website generation using scaffold-driven generation and cascaded multimodal rewards.</description><pubDate>Sat, 25 Apr 2026 18:16:01 GMT</pubDate></item><item><title>GPT-5.5: OpenAI&apos;s Smartest Model for Real Work</title><link>https://nestfrontier.com/gpt-55-openais-smartest-model-for-real-work/</link><guid isPermaLink="true">https://nestfrontier.com/gpt-55-openais-smartest-model-for-real-work/</guid><description>GPT-5.5 released April 23, 2026. 82.7% Terminal-Bench, 84.9% GDPval—beats GPT-5.4, Claude Opus 4.7, Gemini 3.1 Pro. Matches GPT-5.4 latency while delivering higher intelligence. Fewer tokens for same tasks.</description><pubDate>Fri, 24 Apr 2026 12:33:12 GMT</pubDate></item><item><title>DeepSeek V4: Million-Token Context Intelligence</title><link>https://nestfrontier.com/deepseek-v4-million-token-context-intelligence/</link><guid isPermaLink="true">https://nestfrontier.com/deepseek-v4-million-token-context-intelligence/</guid><description>DeepSeek V4 just dropped: 1.6T MoE with 1M context, MIT license. Beats GPT-5 and Gemini on LiveCodeBench (93.5%). 27% FLOPs efficiency vs V3.2. Open weights, aggressive pricing.</description><pubDate>Fri, 24 Apr 2026 06:41:00 GMT</pubDate></item><item><title>Qwen3.6-35B-A3B: 9x Efficiency With 35B MoE</title><link>https://nestfrontier.com/qwen36-35b-a3b-9x-efficiency-with-35b-moe/</link><guid isPermaLink="true">https://nestfrontier.com/qwen36-35b-a3b-9x-efficiency-with-35b-moe/</guid><description>Alibaba&apos;s Qwen3.6-35B-A3B delivers Qwen3.5-27B performance with 3B active parameters. SWE-bench 73.4%, Apache 2.0, runs at 100 t/s on M5 Max.</description><pubDate>Fri, 24 Apr 2026 01:16:57 GMT</pubDate></item><item><title>LLaDA2.0-Uni: First Diffusion LLM to Close Gap with Specialists</title><link>https://nestfrontier.com/llada20-uni-first-diffusion-llm-to-close-gap-with-specialists/</link><guid isPermaLink="true">https://nestfrontier.com/llada20-uni-first-diffusion-llm-to-close-gap-with-specialists/</guid><description>Inclusion AI&apos;s 16B MoE diffusion LLM unifies multimodal understanding and generation, closing the gap with specialized models for the first time.</description><pubDate>Thu, 23 Apr 2026 10:00:49 GMT</pubDate></item></channel></rss>