InferScale is the real infrastructure story here—personalized LLM serving at scale is where the actual GPU bottleneck lives, and native KV injection addresses the most painful friction point in production deployments. The llama.cpp releases are table stakes (optimization is continuous, not news), and most of the funding noise is either defensive positioning (Visa layoffs signal margin pressure, not AI opportunity) or noise (OceanBase chasing AI funding is desperation, not differentiation). The academic papers are fine but scattered—only the seizure detection work has immediate clinical relevance; the rest are benchmark/dataset contributions that matter to maybe 50 researchers. Watch the GPU utilization math on personalized serving; ignore the stock market rotation theater and most enterprise AI funding announcements until someone actually ships measurable inference cost reductions.