Eight GPUs Is the Real Price of Open Hy4
Tencent Hy4 preview is open weight, but its official FP8 recipe needs eight GPUs. Here is how to decide between self-hosting and hosted inference.
Technical analysis and architectural deep-dives that separate signal from hype.
Tencent Hy4 preview is open weight, but its official FP8 recipe needs eight GPUs. Here is how to decide between self-hosting and hosted inference.
A practical Meetily setup guide for private local transcription, GPU detection, Ollama summaries, and the traps that make “local” less private than it sounds.
GLM-5.3 is open weight, but 141 shards and a million-token ceiling make local deployment an infrastructure decision, not a laptop install.
A skeptical guide to testing an open AI gateway, controlling agent spend, and checking caching, telemetry, credentials, and fallback behavior before switching.
Gemini 3.5 Transcribe splits live voice from recorded audio. Here is how to choose the API without confusing WER with product quality.
A practical local harness routes simple coding work to cheaper models, keeps artifacts auditable, and reserves frontier context for ambiguity.
GLM-5.3-Flash makes hosted agent work look cheap at $0.045 per task, but its 320B total parameters change the local deployment math.
A practical RAG ladder: start with SQLite FTS5, measure misses, then add embeddings only when semantic search earns its complexity.
Apodex 1.1 shows when a 35B agent team earns its overhead, and when it only creates more plausible debris. Here is the safe local test path.