Around 22 billion tokens per minute are currently processed across Google’s APIs and developer products. With 1 token equaling roughly 0.75 words, that is the equivalent of processing the entire English Wikipedia more than three and a half times over every single minute. Scaling infrastructure to sustain that velocity requires rethinking systems from first principles—co-designing hardware and software, optimizing network fabrics, and maximizing goodput across TPUs, GPUs, and CPUs to eliminate idle capacity. For enterprise leaders scaling their own AI workloads, these architectural lessons provide a practical roadmap. Watch this Six Five Summit conversation between Mark Lohmeyer and The Futurum Group’s Daniel Newman breaking down the new infrastructure paradigm → https://goo.gle/4d73bBv

Processing 22 billion tokens per minute is a mind-boggling scale. The point about co-designing hardware and software to maximize 'goodput' across TPUs, GPUs, and CPUs is critical—at that volume, minimizing idle compute capacity isn't just an efficiency metric, it's a fundamental cost and performance requirement

Framing token throughput in terms of reading English Wikipedia 3.5x every minute really puts enterprise AI adoption velocity into perspective. Great discussion between Mark Lohmeyer and Daniel Newman on how full-stack infrastructure optimization is becoming the true backbone of scalable AI deployment

Entire English Encyclopaedia 3.5 times in a minute is insane 😲, growth in infrastructure to maintain this level or even beyond will contribute to a mountains of electronic waste ?

Like
Reply

22 billion tokens per minute that’s a lot

Like
Reply
See more comments

To view or add a comment, sign in

Explore content categories