Today in AI hardware

2026-07-20 · AI Native

The Real Story This Week Isn't About Chips—It's About Inference Economics

PagedWeight's dynamic quantization for MoE serving and the llama.cpp release spree signal the actual arms race: making models cheaper to run, not just bigger to train. TSMC's AI capex flex gets the headlines, but the real margin compression is happening in inference optimization—where PagedWeight's quality-aware approach matters because every percentage point of efficiency directly hits operator costs. Skip the Anthropic IPO speculation and the Taiwan cheerleading; watch whether open-weight inference frameworks can commoditize serving faster than proprietary APIs can lock in customers. That's where the hardware money actually flows next.