A Quarter the Size Still Hits 80% on SWE-bench
Inkling-Small uses 12B active parameters to approach its 975B sibling on coding and reasoning, but its factuality gap still matters.
NESTFRONTIER / ARCHIVE
Technical analysis, model releases, research, and deployment notes.
Inkling-Small uses 12B active parameters to approach its 975B sibling on coding and reasoning, but its factuality gap still matters.
Seedance 2.5 doubles native clip length and expands multimodal references, but its real test is whether continuity survives difficult scenes.
Kimi K3 is a 2.8T open-weight model with frontier scores, a 1.56 TB checkpoint, and a license that changes the commercial math.
ReToken adds one learned token to visual-language models and improves long-context image and video retrieval without retraining the whole backbone.
ORCA-bench puts coding agents in a noisy six-day microservice testbed. The best hard-task RCA accuracy was just 10%.
DeepSeek V4 Flash 0731 posts an 82.7 on Terminal Bench 2.1 and brings cheap, Codex-ready agent coding to public beta.
TurboVLA removes the LLM from the robot control loop, reaching 32 Hz and 97.7% LIBERO success with 0.2B parameters and 0.9 GB VRAM.
Vera tested four production agents with executable safety cases and found a 93.9% average attack success rate under multi-channel attacks.
HiLS-Attention learns which distant chunks deserve attention, reaching million-token extrapolation without dense attention costs.
Relay-OPD lets a teacher briefly rescue a small model when its reasoning goes off course, improving math accuracy while cutting training trajectories by more than half.