LPSR: Training-Free Error Correction That Beats 70B Models with 8B
LPSR is a training-free method that improves 8B model MATH-500 scores from 28.8% to 44.0%, beating a 70B model with 8.75x fewer parameters.
NESTFRONTIER / ARCHIVE
Technical analysis, model releases, research, and deployment notes.
LPSR is a training-free method that improves 8B model MATH-500 scores from 28.8% to 44.0%, beating a 70B model with 8.75x fewer parameters.
MiniMax M2.7 is the first AI model to autonomously optimize its own behavioral scaffolding, achieving 30% improvement over 100+ iterations without weight changes.
GLM-5.1 by Zhipu AI became the first open-source model to beat GPT-5.4 on SWE-Bench Pro. MIT licensed, 754B MoE, 40B active parameters, trained on Huawei chips. The real game-changer is the MIT license enabling enterprise self-hosting.
Qwen 3.6 Max Preview claims six benchmark #1s in agentic coding. +9.9 SkillsBench, +6.3 SciCode over Plus. preserve_thinking feature for multi-turn agents. Proprietary, API-only.
Kimi K2.6 is the first open-source model competitive with GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. 1T params, 32B active, leads agentic benchmarks, matches frontier models on coding.
Humanity’s Last Exam has 2,500 expert-built questions, yet frontier scores keep rising. The real warning is benchmark saturation, not a single leaderboard number.
SXSW 2026 XR Audience Award winner Fabula Rasa proves AI NPCs can do more than fetch quests. Nine characters, zero dialogue trees, fully improvised VR theater.
YC S25 startup built a Sims-style 3D game where AI agents negotiate, dance, and show emergent behaviors. Now pivoting to world model research.
LingBot-Map from Ant Group achieves real-time 3D reconstruction at 20 FPS while outperforming offline methods on benchmarks. Apache 2.0 licensed, 2.6k+ GitHub stars.
Honor's Lightning robot beat the human half-marathon world record by 6+ minutes in Beijing, completing 21.1km in 50:26 autonomously—just one year after robots finished an hour behind humans.