
Google’s Gemma 4 adds Multi‑Token Prediction (MTP) drafters: small auxiliary models that propose short token chunks in parallel while the full Gemma model verifies them in a single forward pass. That speculative‑decoding pattern amortizes the memory‑bound KV/attention work and can yield up to ~3× tokens/sec on common stacks (LiteRT, vLLM, Hugging Face, MLX) without changing the target model’s final logits or reasoning. It plugs into existing runtimes and has predictable limits around MoE routing and tight VRAM budgets.
Read story →The feed is electric about new robotics foundation models (MolmoAct2 and related releases). There's real excitement — people celebrate open weights and sim-to-real progress — but an equal dose of skepticism: many caution this isn't a 'ChatGPT moment' for robotics yet and highlight gaps between impressive demos and reliable, deployable systems. The tension is between open-source enthusiasm (shared weights, reproducibility) and pragmatic doubts about safety, generalization, and whether today’s models truly deliver robust real-world behavior.
Two overlapping conversations: headline-grabbing capital commitments (Anthropic -> Google Cloud / TPU reports) and high-level takes about compute becoming a tradable, financial asset. People are awed at the scale — and uneasy about the implications. The mood mixes FOMO (this is where the winners will be made) with concern about vendor lock-in, cloud concentration, and how these mega-bets shape who controls AI performance and access.
A clear narrative thread: work is shifting from static chat interfaces to autonomous agents controlling UIs, remote machines, and workflows. Enthusiasm is high — agents promise to automate complex, multi-step work — but posts about layoffs, reorganizations, and product pivots show concrete disruption. The debate centers on whether agents will augment teams or replace roles, how to integrate them safely, and the organizational changes (one-person product teams, new tooling) needed to ship agent-first products.