Around 22 billion tokens per minute are currently processed across Google’s APIs and developer products. With 1 token equaling roughly 0.75 words, that is the equivalent of processing the entire English Wikipedia more than three and a half times over every single minute. Scaling infrastructure to sustain that velocity requires rethinking systems from first principles—co-designing hardware and software, optimizing network fabrics, and maximizing goodput across TPUs, GPUs, and CPUs to eliminate idle capacity. For enterprise leaders scaling their own AI workloads, these architectural lessons provide a practical roadmap. Watch this Six Five Summit conversation between Mark Lohmeyer and The Futurum Group’s Daniel Newman breaking down the new infrastructure paradigm → https://goo.gle/4d73bBv
Framing token throughput in terms of reading English Wikipedia 3.5x every minute really puts enterprise AI adoption velocity into perspective. Great discussion between Mark Lohmeyer and Daniel Newman on how full-stack infrastructure optimization is becoming the true backbone of scalable AI deployment
Entire English Encyclopaedia 3.5 times in a minute is insane 😲, growth in infrastructure to maintain this level or even beyond will contribute to a mountains of electronic waste ?
22 billion tokens per minute that’s a lot
Processing 22 billion tokens per minute is a mind-boggling scale. The point about co-designing hardware and software to maximize 'goodput' across TPUs, GPUs, and CPUs is critical—at that volume, minimizing idle compute capacity isn't just an efficiency metric, it's a fundamental cost and performance requirement