<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>research — NestFrontier</title><description>Technical AI analysis and research on research.</description><link>https://nestfrontier.com/</link><item><title>A 35B Agent Team Still Needs a Hard Stop</title><link>https://nestfrontier.com/a-35b-agent-team-still-needs-a-hard-stop/</link><guid isPermaLink="true">https://nestfrontier.com/a-35b-agent-team-still-needs-a-hard-stop/</guid><description>Apodex 1.1 shows when a 35B agent team earns its overhead, and when it only creates more plausible debris. Here is the safe local test path.</description><pubDate>Wed, 26 Aug 2026 00:05:17 GMT</pubDate></item><item><title>Always-On Agents Need a Budget and a Sandbox</title><link>https://nestfrontier.com/always-on-agents-need-a-budget-and-a-sandbox/</link><guid isPermaLink="true">https://nestfrontier.com/always-on-agents-need-a-budget-and-a-sandbox/</guid><description>Headlong makes AI agents think continuously. Here is the safe way to test one without an uncapped bill, exposed host, or shared-memory surprise.</description><pubDate>Tue, 25 Aug 2026 12:44:13 GMT</pubDate></item><item><title>The $8 AI Paper Pipeline Still Needs a Human</title><link>https://nestfrontier.com/the-dollar8-ai-paper-pipeline-still-needs-a-human/</link><guid isPermaLink="true">https://nestfrontier.com/the-dollar8-ai-paper-pipeline-still-needs-a-human/</guid><description>Spark-to-Paper turns a research idea into a checked draft for $8.1, but its gates catch format failures, not scientific truth.</description><pubDate>Sun, 16 Aug 2026 00:03:08 GMT</pubDate></item><item><title>Long videos still drown VLMs in irrelevant frames</title><link>https://nestfrontier.com/long-videos-still-drown-vlms-in-irrelevant-frames/</link><guid isPermaLink="true">https://nestfrontier.com/long-videos-still-drown-vlms-in-irrelevant-frames/</guid><description>ReToken adds one learned token to visual-language models and improves long-context image and video retrieval without retraining the whole backbone.</description><pubDate>Sat, 01 Aug 2026 12:03:02 GMT</pubDate></item><item><title>Production oncall is still a human job at 10% accuracy</title><link>https://nestfrontier.com/production-oncall-is-still-a-human-job-at-10percent-accuracy/</link><guid isPermaLink="true">https://nestfrontier.com/production-oncall-is-still-a-human-job-at-10percent-accuracy/</guid><description>ORCA-bench puts coding agents in a noisy six-day microservice testbed. The best hard-task RCA accuracy was just 10%.</description><pubDate>Sat, 01 Aug 2026 00:03:01 GMT</pubDate></item><item><title>32 Hz Robotics Just Lost Its Billion Parameter Habit</title><link>https://nestfrontier.com/32-hz-robotics-just-lost-its-billion-parameter-habit/</link><guid isPermaLink="true">https://nestfrontier.com/32-hz-robotics-just-lost-its-billion-parameter-habit/</guid><description>TurboVLA removes the LLM from the robot control loop, reaching 32 Hz and 97.7% LIBERO success with 0.2B parameters and 0.9 GB VRAM.</description><pubDate>Fri, 31 Jul 2026 00:03:09 GMT</pubDate></item><item><title>Long context just stopped being a memory problem</title><link>https://nestfrontier.com/long-context-just-stopped-being-a-memory-problem/</link><guid isPermaLink="true">https://nestfrontier.com/long-context-just-stopped-being-a-memory-problem/</guid><description>HiLS-Attention learns which distant chunks deserve attention, reaching million-token extrapolation without dense attention costs.</description><pubDate>Thu, 30 Jul 2026 00:03:07 GMT</pubDate></item><item><title>Small Models Stop Digging Their Own Reasoning Holes</title><link>https://nestfrontier.com/small-models-stop-digging-their-own-reasoning-holes/</link><guid isPermaLink="true">https://nestfrontier.com/small-models-stop-digging-their-own-reasoning-holes/</guid><description>Relay-OPD lets a teacher briefly rescue a small model when its reasoning goes off course, improving math accuracy while cutting training trajectories by more than half.</description><pubDate>Wed, 29 Jul 2026 12:03:07 GMT</pubDate></item><item><title>Models Still Answer Nonsense. BullshitBench Shows Which Ones Don&apos;t</title><link>https://nestfrontier.com/models-still-answer-nonsense-bullshitbench-shows-which-ones-dont/</link><guid isPermaLink="true">https://nestfrontier.com/models-still-answer-nonsense-bullshitbench-shows-which-ones-dont/</guid><description>BullshitBench v2 tests whether AI models reject plausible nonsense instead of confidently answering it. The results are uncomfortable.</description><pubDate>Wed, 29 Jul 2026 00:03:02 GMT</pubDate></item><item><title>Agents Still Fail Half Their Real World Jobs</title><link>https://nestfrontier.com/agents-still-fail-half-their-real-world-jobs/</link><guid isPermaLink="true">https://nestfrontier.com/agents-still-fail-half-their-real-world-jobs/</guid><description>A new 400-task benchmark finds leading agents pass less than half of realistic jobs, with memory, vision, and cross-app coordination still breaking most systems.</description><pubDate>Tue, 28 Jul 2026 12:03:08 GMT</pubDate></item><item><title>Training Agents Still Breaks at the Harness Layer</title><link>https://nestfrontier.com/training-agents-still-breaks-at-the-harness-layer/</link><guid isPermaLink="true">https://nestfrontier.com/training-agents-still-breaks-at-the-harness-layer/</guid><description>OpenForgeRL trains agents inside the real harnesses they use in deployment, posting strong tool-use and GUI results while exposing why error recovery remains hard.</description><pubDate>Mon, 27 Jul 2026 00:03:55 GMT</pubDate></item><item><title>Russian prompts expose a hidden coding agent gap</title><link>https://nestfrontier.com/russian-prompts-expose-a-hidden-coding-agent-gap/</link><guid isPermaLink="true">https://nestfrontier.com/russian-prompts-expose-a-hidden-coding-agent-gap/</guid><description>RuBench tests coding agents on fresh repository fixes written in native Russian, then audits what each product actually read and ran.</description><pubDate>Sun, 26 Jul 2026 00:03:48 GMT</pubDate></item><item><title>Deep Research Got Better When AREX Learned to Doubt Itself</title><link>https://nestfrontier.com/deep-research-got-better-when-arex-learned-to-doubt-itself/</link><guid isPermaLink="true">https://nestfrontier.com/deep-research-got-better-when-arex-learned-to-doubt-itself/</guid><description>AREX improves deep research by verifying partial answers, preserving useful evidence, and sending unresolved claims into a targeted second pass.</description><pubDate>Fri, 24 Jul 2026 12:03:49 GMT</pubDate></item><item><title>Drop 39% of agent tokens by reading the model&apos;s hidden state</title><link>https://nestfrontier.com/drop-39percent-of-agent-tokens-by-reading-the-models-hidden-state/</link><guid isPermaLink="true">https://nestfrontier.com/drop-39percent-of-agent-tokens-by-reading-the-models-hidden-state/</guid><description>ByteDance researchers show the keep-or-prune signal for coding agent context is already inside the model&apos;s own hidden states. SWE-Pruner Pro prunes tool outputs without a separate scoring model.</description><pubDate>Tue, 21 Jul 2026 12:17:14 GMT</pubDate></item><item><title>Gold at IMO and IPhO with just 2B active parameters</title><link>https://nestfrontier.com/gold-at-imo-and-ipho-with-just-2b-active-parameters/</link><guid isPermaLink="true">https://nestfrontier.com/gold-at-imo-and-ipho-with-just-2b-active-parameters/</guid><description>Loopie-20B-A2B, a looped Transformer with only 2B active parameters, won gold at IMO 2025 and IPhO 2025 by reusing layers instead of stacking more, challenging the assumption that frontier reasoning requires massive scale.</description><pubDate>Tue, 21 Jul 2026 00:29:38 GMT</pubDate></item><item><title>An 87-year-old math problem just fell to Claude Fable 5</title><link>https://nestfrontier.com/an-87-year-old-math-problem-just-fell-to-claude-fable-5/</link><guid isPermaLink="true">https://nestfrontier.com/an-87-year-old-math-problem-just-fell-to-claude-fable-5/</guid><description>Claude Fable 5 produced a hand-checkable counterexample to the 87-year-old Jacobian conjecture. The polynomial map fits in one tweet and can be verified in minutes with Wolfram Alpha.</description><pubDate>Mon, 20 Jul 2026 12:17:24 GMT</pubDate></item><item><title>A 10-page prompt closed a 30-year math gap in 148 minutes</title><link>https://nestfrontier.com/a-10-page-prompt-closed-a-30-year-math-gap-in-148-minutes/</link><guid isPermaLink="true">https://nestfrontier.com/a-10-page-prompt-closed-a-30-year-math-gap-in-148-minutes/</guid><description>A Berkeley professor spent a year on an open optimization problem. GPT-5.6 solved it in 148 minutes with a 10-page prompt. The proof is Lean-verified.</description><pubDate>Sun, 19 Jul 2026 00:22:26 GMT</pubDate></item><item><title>975 billion parameters and none of them are locked</title><link>https://nestfrontier.com/975-billion-parameters-and-none-of-them-are-locked/</link><guid isPermaLink="true">https://nestfrontier.com/975-billion-parameters-and-none-of-them-are-locked/</guid><description>Thinking Machines Lab released Inkling: a 975B open-weight MoE model with Apache 2.0 license, controllable thinking effort, and native multimodality across text, images, and audio.</description><pubDate>Thu, 16 Jul 2026 00:23:57 GMT</pubDate></item><item><title>64 Agents Just Proved What Mathematicians Couldn&apos;t in 50 Years</title><link>https://nestfrontier.com/64-agents-just-proved-what-mathematicians-couldnt-in-50-years/</link><guid isPermaLink="true">https://nestfrontier.com/64-agents-just-proved-what-mathematicians-couldnt-in-50-years/</guid><description>GPT-5.6 Sol Ultra proved a 50-year-old graph theory conjecture using 64 parallel subagents in under an hour. The model became publicly available the same day.</description><pubDate>Fri, 10 Jul 2026 20:04:43 GMT</pubDate></item><item><title>5B parameters play Rocket League without a physics engine</title><link>https://nestfrontier.com/5b-parameters-play-rocket-league-without-a-physics-engine/</link><guid isPermaLink="true">https://nestfrontier.com/5b-parameters-play-rocket-league-without-a-physics-engine/</guid><description>MIRA runs four-player Rocket League at 20 FPS using a 5B diffusion transformer with no physics engine, trained entirely on bot gameplay. Open-source.</description><pubDate>Thu, 09 Jul 2026 08:05:24 GMT</pubDate></item><item><title>Real quantum hardware just recovered a 10-bit encryption key</title><link>https://nestfrontier.com/real-quantum-hardware-just-recovered-a-10-bit-encryption-key/</link><guid isPermaLink="true">https://nestfrontier.com/real-quantum-hardware-just-recovered-a-10-bit-encryption-key/</guid><description>Researchers recovered a 10-bit crypto key on real IBM quantum hardware, doubling the previous record. Your AES-128 is still safe.</description><pubDate>Tue, 07 Jul 2026 20:06:10 GMT</pubDate></item><item><title>84.7% accuracy at 1/14th cost. Bridgewater just proved fine-tuning beats frontier</title><link>https://nestfrontier.com/847percent-accuracy-at-114th-cost-bridgewater-just-proved-fine-tuning-beats-frontier/</link><guid isPermaLink="true">https://nestfrontier.com/847percent-accuracy-at-114th-cost-bridgewater-just-proved-fine-tuning-beats-frontier/</guid><description>Bridgewater fine-tuned Qwen3-235B on expert-labeled financial data and beat every frontier model at 1/14th the cost. The era of renting the biggest generalist may be ending.</description><pubDate>Sun, 05 Jul 2026 20:05:55 GMT</pubDate></item><item><title>When does scaling actually stop working?</title><link>https://nestfrontier.com/when-does-scaling-actually-stop-working/</link><guid isPermaLink="true">https://nestfrontier.com/when-does-scaling-actually-stop-working/</guid><description>Lilian Weng traces the full history of scaling law research, showing where Kaplan and Chinchilla disagree and why the exponents matter more than anyone admits.</description><pubDate>Wed, 01 Jul 2026 08:06:29 GMT</pubDate></item><item><title>A $475M startup just proved oscillators can draw</title><link>https://nestfrontier.com/a-dollar475m-startup-just-proved-oscillators-can-draw/</link><guid isPermaLink="true">https://nestfrontier.com/a-dollar475m-startup-just-proved-oscillators-can-draw/</guid><description>Unconventional AI released Un-0, an image generator running on coupled oscillators instead of neural networks. FID 6.74 on ImageNet 64x64. The hardware to make it 1000x more efficient does not exist yet.</description><pubDate>Fri, 26 Jun 2026 08:06:35 GMT</pubDate></item><item><title>Open image models finally got a training pipeline that works</title><link>https://nestfrontier.com/open-image-models-finally-got-a-training-pipeline-that-works/</link><guid isPermaLink="true">https://nestfrontier.com/open-image-models-finally-got-a-training-pipeline-that-works/</guid><description>Krea 2 drops open weights for a 12B image model with a two-checkpoint workflow: train on Raw, generate fast on Turbo. Zero synthetic training data, 2-second inference.</description><pubDate>Wed, 24 Jun 2026 20:04:13 GMT</pubDate></item><item><title>Simulated agent training now beats the real thing</title><link>https://nestfrontier.com/simulated-agent-training-now-beats-the-real-thing/</link><guid isPermaLink="true">https://nestfrontier.com/simulated-agent-training-now-beats-the-real-thing/</guid><description>Alibaba&apos;s Qwen-AgentWorld is the first language model built to simulate agent environments, and training in its fake worlds beats training in real ones.</description><pubDate>Wed, 24 Jun 2026 10:37:04 GMT</pubDate></item><item><title>That 3B model matched Claude Opus at math. Benchmarks broken?</title><link>https://nestfrontier.com/that-3b-model-matched-claude-opus-at-math-benchmarks-broken/</link><guid isPermaLink="true">https://nestfrontier.com/that-3b-model-matched-claude-opus-at-math-benchmarks-broken/</guid><description>A 3B parameter model from Sina Weibo claims to match Claude Opus 4.5 on math benchmarks. The AI community is split on whether this is a breakthrough or benchmark gaming.</description><pubDate>Tue, 23 Jun 2026 20:05:12 GMT</pubDate></item><item><title>Your 11.9B inpainting model just got outperformed by something 50x smaller</title><link>https://nestfrontier.com/your-119b-inpainting-model-just-got-outperformed-by-something-50x-smaller/</link><guid isPermaLink="true">https://nestfrontier.com/your-119b-inpainting-model-just-got-outperformed-by-something-50x-smaller/</guid><description>A new 0.22B parameter model matches or beats 10B-level inpainting models like FLUX and SD3.5 using less than 2% of the parameters and 15x faster inference.</description><pubDate>Mon, 22 Jun 2026 20:03:09 GMT</pubDate></item><item><title>59% SWE-Bench score from a model costing $0.30 per million tokens</title><link>https://nestfrontier.com/59percent-swe-bench-score-from-a-model-costing-dollar030-per-million-tokens/</link><guid isPermaLink="true">https://nestfrontier.com/59percent-swe-bench-score-from-a-model-costing-dollar030-per-million-tokens/</guid><description>MiniMax M3 is an open-weight model scoring 59% on SWE-Bench Pro at $0.30/M tokens. The architecture is novel, the price is aggressive, and the weights are real. But vendor-run benchmarks and Chinese jurisdiction concerns complicate the story.</description><pubDate>Sun, 21 Jun 2026 08:05:07 GMT</pubDate></item><item><title>Robots need three brains to think and move</title><link>https://nestfrontier.com/robots-need-three-brains-to-think-and-move/</link><guid isPermaLink="true">https://nestfrontier.com/robots-need-three-brains-to-think-and-move/</guid><description>Alibaba&apos;s Tongyi Lab launched three specialized robotics foundation models that cover navigation, manipulation, and world prediction. The complete stack runs on NVIDIA Jetson Thor and is already in enterprise pilot testing.</description><pubDate>Wed, 17 Jun 2026 20:05:41 GMT</pubDate></item><item><title>The AI Boom Just Outpaced Every Safeguard We Built</title><link>https://nestfrontier.com/the-ai-boom-just-outpaced-every-safeguard-we-built/</link><guid isPermaLink="true">https://nestfrontier.com/the-ai-boom-just-outpaced-every-safeguard-we-built/</guid><description>Stanford&apos;s 2026 AI Index: $581.7B invested, a 2.7% U.S.-China gap, junior dev hiring down 20%, and transparency scores falling. The systems can&apos;t keep up.</description><pubDate>Wed, 10 Jun 2026 20:05:29 GMT</pubDate></item><item><title>Your LLM&apos;s KV Cache Is Eating 75% of Your GPU Memory</title><link>https://nestfrontier.com/your-llms-kv-cache-is-eating-75percent-of-your-gpu-memory/</link><guid isPermaLink="true">https://nestfrontier.com/your-llms-kv-cache-is-eating-75percent-of-your-gpu-memory/</guid><description>A new lossless compression technique uses a smaller model&apos;s predictions to shrink LLM KV caches by 4x on top of existing methods, with zero quality loss and no retraining required.</description><pubDate>Sun, 07 Jun 2026 20:04:21 GMT</pubDate></item><item><title>Anthropic warns AI could soon build itself without human help</title><link>https://nestfrontier.com/anthropic-warns-ai-could-soon-build-itself-without-human-help/</link><guid isPermaLink="true">https://nestfrontier.com/anthropic-warns-ai-could-soon-build-itself-without-human-help/</guid><description>Anthropic admitted 80% of its own code is now written by Claude. The company says recursive self-improvement is approaching and wants a global pause.</description><pubDate>Sat, 06 Jun 2026 08:05:04 GMT</pubDate></item><item><title>AI Broke Berkeley CS With 35% Student Failure Rates</title><link>https://nestfrontier.com/ai-broke-berkeley-cs-with-35percent-student-failure-rates/</link><guid isPermaLink="true">https://nestfrontier.com/ai-broke-berkeley-cs-with-35percent-student-failure-rates/</guid><description>Berkeley CS failure rates hit 35% in Spring 2026 as AI cheating and dwindling math skills converge. Professors name LLM overreliance as the primary driver.</description><pubDate>Thu, 04 Jun 2026 08:09:12 GMT</pubDate></item><item><title>MiniMax M3: Open-Weight Model With 1M Context and Frontier Coding</title><link>https://nestfrontier.com/minimax-m3-open-weight-model-with-1m-context-and-frontier-coding/</link><guid isPermaLink="true">https://nestfrontier.com/minimax-m3-open-weight-model-with-1m-context-and-frontier-coding/</guid><description>MiniMax M3 is the first open-weight model to combine a 1M-token context window, native multimodality, and frontier coding — beating GPT-5.5 on SWE-Bench Pro while cutting compute by 20x.</description><pubDate>Wed, 03 Jun 2026 12:24:25 GMT</pubDate></item><item><title>Law Professors Picked AI Answers Over Their Own Colleagues</title><link>https://nestfrontier.com/law-professors-picked-ai-answers-over-their-own-colleagues/</link><guid isPermaLink="true">https://nestfrontier.com/law-professors-picked-ai-answers-over-their-own-colleagues/</guid><description>A blind study of nearly 3,000 comparisons found law professors preferred AI-generated answers to student questions 75% of the time — and flagged human answers as harmful 3x more often.</description><pubDate>Wed, 03 Jun 2026 08:09:10 GMT</pubDate></item><item><title>Your LLM works harder after a short nap</title><link>https://nestfrontier.com/your-llm-works-harder-after-a-short-nap/</link><guid isPermaLink="true">https://nestfrontier.com/your-llm-works-harder-after-a-short-nap/</guid><description>A new paper proposes letting LLMs &apos;sleep&apos; between processing chunks, consolidating context into fast weights and clearing the cache. The results on multi-hop reasoning are 3x better than baseline.</description><pubDate>Wed, 27 May 2026 20:03:54 GMT</pubDate></item><item><title>Vision models finally broke free from the token-by-token cage</title><link>https://nestfrontier.com/vision-models-finally-broke-free-from-the-token-by-token-cage/</link><guid isPermaLink="true">https://nestfrontier.com/vision-models-finally-broke-free-from-the-token-by-token-cage/</guid><description>NVIDIA&apos;s LocateAnything uses Parallel Box Decoding to make vision-language models 10x faster at visual grounding while improving accuracy on LVIS and COCO. The 3B model hits 12.7 boxes per second and handles dense detection, GUI, and document grounding from a single checkpoint.</description><pubDate>Wed, 27 May 2026 08:06:56 GMT</pubDate></item><item><title>Long Context Still Breaks Memory. Gated DeltaNet-2 Fixes It</title><link>https://nestfrontier.com/your-llm-uses-cheaper-attention-nvidia-just-made-it-better/</link><guid isPermaLink="true">https://nestfrontier.com/your-llm-uses-cheaper-attention-nvidia-just-made-it-better/</guid><description>Gated DeltaNet-2 gives linear attention separate erase and write gates, targeting the memory interference that breaks long-context retrieval.</description><pubDate>Sat, 23 May 2026 20:04:11 GMT</pubDate></item><item><title>Scaling Laws Need Fewer Questions Than Anyone Thought</title><link>https://nestfrontier.com/stanford-just-made-llm-scaling-laws-99percent-cheaper/</link><guid isPermaLink="true">https://nestfrontier.com/stanford-just-made-llm-scaling-laws-99percent-cheaper/</guid><description>Item Response Scaling Laws estimate model scaling with about 50 questions per benchmark after calibration. That cuts evaluation cost, not model-training cost.</description><pubDate>Fri, 22 May 2026 20:03:41 GMT</pubDate></item><item><title>An 80-Year-Old Math Conjecture Just Fell to AI</title><link>https://nestfrontier.com/an-80-year-old-math-conjecture-just-fell-to-ai/</link><guid isPermaLink="true">https://nestfrontier.com/an-80-year-old-math-conjecture-just-fell-to-ai/</guid><description>An OpenAI model has disproved an 80-year-old mathematical conjecture by Paul Erdős, marking the first time AI has autonomously solved a prominent open problem in mathematics.</description><pubDate>Thu, 21 May 2026 20:05:05 GMT</pubDate></item><item><title>A Benchmark Caught AI Faking Answers to Broken Problems</title><link>https://nestfrontier.com/a-benchmark-caught-ai-faking-answers-to-broken-problems/</link><guid isPermaLink="true">https://nestfrontier.com/a-benchmark-caught-ai-faking-answers-to-broken-problems/</guid><description>A consortium of 64 mathematicians built Soohak, a 439-problem benchmark revealing that frontier AI models confidently produce wrong answers for unsolvable math problems. No model exceeded 50% on the refusal subset.</description><pubDate>Wed, 20 May 2026 20:11:08 GMT</pubDate></item><item><title>CoT Boosts Agent Performance 10x But RL Still Wins Planning</title><link>https://nestfrontier.com/cot-boosts-agent-performance-10x-but-rl-still-wins-planning/</link><guid isPermaLink="true">https://nestfrontier.com/cot-boosts-agent-performance-10x-but-rl-still-wins-planning/</guid><description>Agentick puts RL agents, LLMs, VLMs, and hybrid systems on equal ground across 37 tasks. GPT-5 mini leads overall but PPO dominates planning. Chain-of-thought multiplies LLM performance by up to 10x.</description><pubDate>Tue, 19 May 2026 14:15:13 GMT</pubDate></item><item><title>Your LLM generates one word at a time. This one doesn&apos;t.</title><link>https://nestfrontier.com/your-llm-generates-one-word-at-a-time-this-one-doesnt/</link><guid isPermaLink="true">https://nestfrontier.com/your-llm-generates-one-word-at-a-time-this-one-doesnt/</guid><description>ByteDance&apos;s Cola DLM is a 2.3B parameter language model that ditches autoregressive token prediction for continuous latent diffusion. It beats matched autoregressive baselines on reasoning benchmarks and suggests a path beyond the left-to-right token parade that&apos;s dominated NLP for a decade.</description><pubDate>Sun, 10 May 2026 22:15:36 GMT</pubDate></item><item><title>ExoActor: When Video Generation Becomes Robot Imagination</title><link>https://nestfrontier.com/exoactor-when-video-generation-becomes-robot-imagination/</link><guid isPermaLink="true">https://nestfrontier.com/exoactor-when-video-generation-becomes-robot-imagination/</guid><description>BAAI&apos;s ExoActor framework uses video generation models as a robot&apos;s imagination — generating third-person videos of task execution and translating them into physical humanoid robot behaviors on Unitree G1 hardware.</description><pubDate>Fri, 01 May 2026 21:23:54 GMT</pubDate></item><item><title>Tuna-2: Meta&apos;s Pixel Embeddings Beat Vision Encoders</title><link>https://nestfrontier.com/tuna-2-metas-pixel-embeddings-beat-vision-encoders/</link><guid isPermaLink="true">https://nestfrontier.com/tuna-2-metas-pixel-embeddings-beat-vision-encoders/</guid><description>Meta&apos;s Tuna-2 proves pretrained vision encoders are unnecessary. Direct pixel embeddings achieve SOTA on OCR, counting, and perception benchmarks—no CLIP, no VAE.</description><pubDate>Wed, 29 Apr 2026 12:11:17 GMT</pubDate></item><item><title>Talkie: The 13B Language Model That Thinks It&apos;s 1930</title><link>https://nestfrontier.com/talkie-the-13b-language-model-that-thinks-its-1930/</link><guid isPermaLink="true">https://nestfrontier.com/talkie-the-13b-language-model-that-thinks-its-1930/</guid><description>Alec Radford and team built a 13B LM trained only on pre-1931 data. It writes Python without ever seeing computers, and enables clean experiments on LLM generalization vs memorization.</description><pubDate>Tue, 28 Apr 2026 12:13:13 GMT</pubDate></item></channel></rss>