Models Still Answer Nonsense. BullshitBench Shows Which Ones Don't
BullshitBench v2 tests whether AI models reject plausible nonsense instead of confidently answering it. The results are uncomfortable.
NESTFRONTIER / ARCHIVE
Technical analysis, model releases, research, and deployment notes.
BullshitBench v2 tests whether AI models reject plausible nonsense instead of confidently answering it. The results are uncomfortable.
A new 400-task benchmark finds leading agents pass less than half of realistic jobs, with memory, vision, and cross-app coordination still breaking most systems.
Microsoft's 4B Mage-Flow pairs a cheap tokenizer with native-resolution diffusion, hitting 0.59 seconds at 1024² on one A100 while exposing the tradeoffs.
Microsoft's Resource2Skill turns tutorials, code, and visual references into executable agent skills, scoring 11.9 points above agents without the wiki.
OpenForgeRL trains agents inside the real harnesses they use in deployment, posting strong tool-use and GUI results while exposing why error recovery remains hard.
Grok 4.5 pairs frontier coding claims with $2 input pricing, 80 TPS, and an open agent harness. The cost story is more convincing than the leaderboard story.
RuBench tests coding agents on fresh repository fixes written in native Russian, then audits what each product actually read and ran.
Claude Opus 5 keeps Opus 4.8 pricing, adds adaptive effort controls, and targets cheaper long-running agent work.
Moonshot AI's 2.8T-parameter Kimi K3 is close to the closed frontier, but the real test starts when its promised weights arrive.
AREX improves deep research by verifying partial answers, preserving useful evidence, and sending unresolved claims into a targeted second pass.