A 35B Agent Team Still Needs a Hard Stop
Apodex 1.1 shows when a 35B agent team earns its overhead, and when it only creates more plausible debris. Here is the safe local test path.
TOPIC_INDEX
47 published entries in this topic.
Apodex 1.1 shows when a 35B agent team earns its overhead, and when it only creates more plausible debris. Here is the safe local test path.
Headlong makes AI agents think continuously. Here is the safe way to test one without an uncapped bill, exposed host, or shared-memory surprise.
Spark-to-Paper turns a research idea into a checked draft for $8.1, but its gates catch format failures, not scientific truth.
ReToken adds one learned token to visual-language models and improves long-context image and video retrieval without retraining the whole backbone.
ORCA-bench puts coding agents in a noisy six-day microservice testbed. The best hard-task RCA accuracy was just 10%.
TurboVLA removes the LLM from the robot control loop, reaching 32 Hz and 97.7% LIBERO success with 0.2B parameters and 0.9 GB VRAM.
HiLS-Attention learns which distant chunks deserve attention, reaching million-token extrapolation without dense attention costs.
Relay-OPD lets a teacher briefly rescue a small model when its reasoning goes off course, improving math accuracy while cutting training trajectories by more than half.
BullshitBench v2 tests whether AI models reject plausible nonsense instead of confidently answering it. The results are uncomfortable.
A new 400-task benchmark finds leading agents pass less than half of realistic jobs, with memory, vision, and cross-app coordination still breaking most systems.
OpenForgeRL trains agents inside the real harnesses they use in deployment, posting strong tool-use and GUI results while exposing why error recovery remains hard.
RuBench tests coding agents on fresh repository fixes written in native Russian, then audits what each product actually read and ran.
AREX improves deep research by verifying partial answers, preserving useful evidence, and sending unresolved claims into a targeted second pass.
ByteDance researchers show the keep-or-prune signal for coding agent context is already inside the model's own hidden states. SWE-Pruner Pro prunes tool outputs without a separate scoring model.
Loopie-20B-A2B, a looped Transformer with only 2B active parameters, won gold at IMO 2025 and IPhO 2025 by reusing layers instead of stacking more, challenging the assumption that frontier reasoning requires massive scale.
Claude Fable 5 produced a hand-checkable counterexample to the 87-year-old Jacobian conjecture. The polynomial map fits in one tweet and can be verified in minutes with Wolfram Alpha.
A Berkeley professor spent a year on an open optimization problem. GPT-5.6 solved it in 148 minutes with a 10-page prompt. The proof is Lean-verified.
Thinking Machines Lab released Inkling: a 975B open-weight MoE model with Apache 2.0 license, controllable thinking effort, and native multimodality across text, images, and audio.
GPT-5.6 Sol Ultra proved a 50-year-old graph theory conjecture using 64 parallel subagents in under an hour. The model became publicly available the same day.
MIRA runs four-player Rocket League at 20 FPS using a 5B diffusion transformer with no physics engine, trained entirely on bot gameplay. Open-source.
Researchers recovered a 10-bit crypto key on real IBM quantum hardware, doubling the previous record. Your AES-128 is still safe.
Bridgewater fine-tuned Qwen3-235B on expert-labeled financial data and beat every frontier model at 1/14th the cost. The era of renting the biggest generalist may be ending.
Lilian Weng traces the full history of scaling law research, showing where Kaplan and Chinchilla disagree and why the exponents matter more than anyone admits.
Unconventional AI released Un-0, an image generator running on coupled oscillators instead of neural networks. FID 6.74 on ImageNet 64x64. The hardware to make it 1000x more efficient does not exist yet.
Krea 2 drops open weights for a 12B image model with a two-checkpoint workflow: train on Raw, generate fast on Turbo. Zero synthetic training data, 2-second inference.
Alibaba's Qwen-AgentWorld is the first language model built to simulate agent environments, and training in its fake worlds beats training in real ones.
A 3B parameter model from Sina Weibo claims to match Claude Opus 4.5 on math benchmarks. The AI community is split on whether this is a breakthrough or benchmark gaming.
A new 0.22B parameter model matches or beats 10B-level inpainting models like FLUX and SD3.5 using less than 2% of the parameters and 15x faster inference.
MiniMax M3 is an open-weight model scoring 59% on SWE-Bench Pro at $0.30/M tokens. The architecture is novel, the price is aggressive, and the weights are real. But vendor-run benchmarks and Chinese jurisdiction concerns complicate the story.
Alibaba's Tongyi Lab launched three specialized robotics foundation models that cover navigation, manipulation, and world prediction. The complete stack runs on NVIDIA Jetson Thor and is already in enterprise pilot testing.
Stanford's 2026 AI Index: $581.7B invested, a 2.7% U.S.-China gap, junior dev hiring down 20%, and transparency scores falling. The systems can't keep up.
A new lossless compression technique uses a smaller model's predictions to shrink LLM KV caches by 4x on top of existing methods, with zero quality loss and no retraining required.
Anthropic admitted 80% of its own code is now written by Claude. The company says recursive self-improvement is approaching and wants a global pause.
Berkeley CS failure rates hit 35% in Spring 2026 as AI cheating and dwindling math skills converge. Professors name LLM overreliance as the primary driver.
MiniMax M3 is the first open-weight model to combine a 1M-token context window, native multimodality, and frontier coding — beating GPT-5.5 on SWE-Bench Pro while cutting compute by 20x.
A blind study of nearly 3,000 comparisons found law professors preferred AI-generated answers to student questions 75% of the time — and flagged human answers as harmful 3x more often.
A new paper proposes letting LLMs 'sleep' between processing chunks, consolidating context into fast weights and clearing the cache. The results on multi-hop reasoning are 3x better than baseline.
NVIDIA's LocateAnything uses Parallel Box Decoding to make vision-language models 10x faster at visual grounding while improving accuracy on LVIS and COCO. The 3B model hits 12.7 boxes per second and handles dense detection, GUI, and document grounding from a single checkpoint.
Gated DeltaNet-2 gives linear attention separate erase and write gates, targeting the memory interference that breaks long-context retrieval.
Item Response Scaling Laws estimate model scaling with about 50 questions per benchmark after calibration. That cuts evaluation cost, not model-training cost.
An OpenAI model has disproved an 80-year-old mathematical conjecture by Paul Erdős, marking the first time AI has autonomously solved a prominent open problem in mathematics.
A consortium of 64 mathematicians built Soohak, a 439-problem benchmark revealing that frontier AI models confidently produce wrong answers for unsolvable math problems. No model exceeded 50% on the refusal subset.
Agentick puts RL agents, LLMs, VLMs, and hybrid systems on equal ground across 37 tasks. GPT-5 mini leads overall but PPO dominates planning. Chain-of-thought multiplies LLM performance by up to 10x.
ByteDance's Cola DLM is a 2.3B parameter language model that ditches autoregressive token prediction for continuous latent diffusion. It beats matched autoregressive baselines on reasoning benchmarks and suggests a path beyond the left-to-right token parade that's dominated NLP for a decade.
BAAI's ExoActor framework uses video generation models as a robot's imagination — generating third-person videos of task execution and translating them into physical humanoid robot behaviors on Unitree G1 hardware.
Meta's Tuna-2 proves pretrained vision encoders are unnecessary. Direct pixel embeddings achieve SOTA on OCR, counting, and perception benchmarks—no CLIP, no VAE.
Alec Radford and team built a 13B LM trained only on pre-1931 data. It writes Python without ever seeing computers, and enables clean experiments on LLM generalization vs memorization.