Sign in to view Amin’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Amin’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Mountain View, California, United States
Sign in to view Amin’s full profile
Amin can introduce you to 10+ people at Google
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
42K followers
500+ connections
Sign in to view Amin’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Amin
Amin can introduce you to 10+ people at Google
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Amin
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Amin’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
42K followers
-
Amin Vahdat shared thisWe’re taking our TPUs to low-Earth orbit! Project Suncatcher is a massive systems engineering challenge; from vacuum cooling to radiation resilience. We have come a long way from our first-ever TPUs. Incredibly proud of the teams pushing the absolute limits of scale, resilience, and power efficiency of AI compute. A true moonshot milestone.Amin Vahdat shared thisInspired by our history of moonshots, from quantum computing to autonomous driving, Project Suncatcher is exploring how we could one day build scalable ML compute systems in space. In low-Earth orbit, satellites can access near-constant sunlight, generating up to 8x more solar power than on Earth. Eventually, linked constellations of satellites could handle significant AI workloads while in orbit. Turning that vision into reality starts with a fundamental question: Can our AI hardware actually survive and operate in space? After years of research, Project Suncatcher is scheduled to embark on its first test in orbit. Built in partnership with Planet, we’re launching a prototype satellite aboard the Transporter-18 rideshare mission with SpaceX to test the performance of our TPUs in space. From the intense vibration and g-forces of launch, to radiation exposure, to the unique challenges of cooling chips in a vacuum, it’s safe to say there’s a lot that can go wrong! It reminds me of seeing the first Waymo leave the parking lot in Mountain View, or visiting our Quantum lab to see our early quantum chips cool down to near absolute zero. Every moonshot starts with a milestone like this. Check out our new four-part video series detailing the science and engineering hurdles the team overcame to get to this moment: https://lnkd.in/gWVSeeqa
-
Amin Vahdat shared thisThis is fantastic news! Google Cloud has been named a Leader in The Forrester Wave™: Public Cloud Platforms, Q3 2026, receiving the highest overall score of any vendor evaluated. We earned the highest possible scores in 23 out of 30 criteria, including Vision, Innovation, AI Development, Database, Analytics, and Containers. This validation is the direct result of decades of foundational research and full-stack co-design under one roof. It is also further testament that fragmented infrastructure offerings are not meeting customer needs. From frontier labs training tomorrow's breakthrough models to global enterprises managing complex business logic or innovators building autonomous workflows, Google Cloud provides the unmatched scale, performance, and foundation to support the entire lifecycle of present and future computing demands. Read the complete analysis with commentary from Brad Calder and Mark Lohmeyer.Amin Vahdat shared thisExciting news — Google Cloud was named a Leader in The Forrester Wave™: Public Cloud Platforms (Q3 2026), earning the highest score in the Current Offering category! Even better, we received the highest combined score of any vendor and the top score possible in 23 of the 30 criteria evaluated — including Vision, Innovation, AI Development, GKE / Kubernetes, Databases, and Security. As teams shift from chatbots to autonomous agents, the big takeaway is simple: you can’t run next-gen AI on fragmented infrastructure. You need a platform co-designed from the ground up — from silicon and systems to models and orchestration. What I believe is driving our leadership position: 1. Full-stack co-design: Decades of co-designing our silicon, systems, and software stacks allow customers to scale AI workloads with real enterprise efficiency. 2. Ready for the agentic era: We're continuing to evolve GKE so organizations can run autonomous agents and core workloads on a highly scalable, proven platform. 3. Grounded in live data: Our Agentic Data Cloud bridges operational databases and analytics, turning enterprise data into a real-time reasoning engine. A massive thank you to our incredible engineering, product, and infrastructure teams for their relentless innovation — and to our customers and partners who push us forward every day. Check out the full report: https://lnkd.in/g4MTV3Tv #GoogleCloud #AIInfrastructure #GKE #Kubernetes #GenerativeAI #CloudComputing
-
Amin Vahdat shared thisThe most promising aspect of AI is how it is accelerating scientific progress. James Manyika's blog below is well worth a read as it highlights four recent examples of AI- enabled scientific advances: - The predicted impact of all 9 billion possible single letter genetic changes. - A 50% improvement in weather accuracy prediction. - A prediction engine designed to track and predict global crises early. - Techniques to mitigate the climate impacts of aviation and. For me, the best analogy to this moment remains the Industrial Revolution. The invention of the steam engines for the first time enabled us to transform potential energy to kinetic energy with unprecedented control and scale, making physical labor dramatically more efficient. Today, in this Era of Intelligence, we are transforming energy into intelligence and insight, enabling us to advance the state of scientific understanding at an unprecedented rate. Link to Blog: https://lnkd.in/e9p6X97g
-
Amin Vahdat shared thisI have been reflecting on an absolutely incredible journey to Taiwan. The vibrant energy, the pace of innovation, and the amazing people I met left me more optimistic about our ability to meet the moment in delivering infrastructure at unprecedented scale, velocity, reliability, and efficiency. Here are some of my top takeaways along with a photo montage to capture it all*. 1. Beyond the breathtaking scenery and dynamic cities, the shared ambition for the future of technology. Thank you, Semicon, for the opportunity to deliver the conference keynote this year and to strike a dialog around the technology and collaboration needed to deliver the compute capacity supporting model and services breakthroughs. 2. I was moved by the profound and shared sense of responsibility; local companies, Googlers, and members of the press. There is a collective commitment to support the massive demand for AI and infrastructure in both a responsible and sustainable manner. It is inspiring to see an entire ecosystem work to prioritize rapid technical advancement while being mindful of the myriad costs to do so. 3. The spirit and unity of Google Taiwan made for a special experience. It is always a privilege to witness a team culture that builds so cohesively together. As a long-time Googler, it’s what makes Google special. The passion and mutual support across different teams showed that the best technology is built by those who genuinely care for one another. Google Taiwan also just celebrated its 20th anniversary, and it was wonderful to be part of the celebration. 4. The overriding impression of this trip is the genuine warmth of the people, matched by the technical breakthroughs they are driving. Working in direct collaboration with Google daily, the teams are driving breakthrough work around automation, hardware reliability, and the very limits of physics in chip design and manufacturing. Thank you to everyone in Taiwan who made this visit so memorable. We are building something extraordinary together. * Perhaps most striking is how much everyone loves taking photos, especially with the universal 👍🏻 signifying the joy and optimism that permeated so many of the discussions.
-
Amin Vahdat shared thisWhen we converted a former paper mill in Hamina into a data center 15 years ago, it marked a turning point in how we thought about sustainable infrastructure. This week, we announced a €13 billion investment across Finland, expanding beyond Hamina into Kajaani, Muhos, and Vaala to build the foundation for the next generation of AI workloads. Delivering on this scale in terms of compute capability and power requires holistic systems design across: Co-designing with the grid: To support grid resilience, we are pairing this footprint with a 22-year agreement supporting the Loviisa nuclear plant, new onshore wind capacity, and a 94 MW battery storage system to stabilize local power during cold, windless winter peaks. Circular energy in practice: In Hamina, we continue to route data center waste heat directly into the local municipal district heating network to heat homes and schools with greater efficiency. Ecosystem restoration: We continue to protect the local watersheds where we build our data centers through local wetland regeneration and native forest restoration. All of the above are critical because scaling AI compute only matters if it unlocks human capabilities in a responsible manner. We will be supporting more than 37,000 jobs across Finland during this construction, with thousands of permanent engineering and operations roles to follow. We are also directing €31M to local communities across Hamina, Kajaani, Muhos, and Vaala to fund AI upskilling for 4,400+ workers and vocational training for students starting their careers in infrastructure. A massive thank you to our partners across Finland, Fingrid, Fortum, and the extraordinary Google teams making this all a reality. Read more about our commitment in the blog post by Bikash Koley. https://lnkd.in/e7BGbAiYGoogle deepens its commitment to Finland with a €13 billion investment in AI infrastructureGoogle deepens its commitment to Finland with a €13 billion investment in AI infrastructure
-
Amin Vahdat shared thisI am in Taiwan this week for an opportunity to speak at the Semicon 2026 conference and to celebrate our Google Taiwan office turning 20, which is home to our largest hardware engineering hub outside of the United States. To support our continued growth and meet the AI opportunity ahead, we are expanding our presence in Taipei’s Shilin District, increasing the office space footprint of our AI Infrastructure team by more than 60 percent. This expansion underscores our commitment to strengthening our partnership and engagement in Taiwan. Over the past two decades, Google's teams in Taiwan have supported our biggest platform shifts from mobile and ChromeOS to the cloud, and now AI. Today, our Taipei-based engineering hub builds on this 20-yr history by developing the physical systems needed for this work. This latest expansion is critical as we stand together at the precipice of the Age of Intelligence. Advancing intelligence is both a shared journey and a collective responsibility that requires global and industry-wide collaboration. The team in Taiwan is extremely well positioned to push our work forward along these dimensions. Happy 20th to Google Taiwan! Read our Google Taiwan blog in traditional Chinese, which also includes my recent op-ed for Business Today. For our English speakers, the Translate feature on Google Chrome is very helpful. https://lnkd.in/e4vezP3X
-
Amin Vahdat shared thisThank you, Sridhar Lakshmanamurthy and Norm Jouppi, for your presentation at Hot Chips 2026 showcasing our eighth-generation TPUs. It was a fantastic representation of our TPU8 capabilities, as well as the collaborative journey and learnings that brought us here. Nicely done!Amin Vahdat shared thisIt was truly an honor to have the opportunity to share the stage again with the legend, Norm Jouppi, at Hot Chips 2026. We discussed the trade-offs involved in building two TPU systems this year: 8i for inference and 8t for training. A huge shout-out and thank you to the entire TPU team for your creativity, tenacity and dedication. 👏 . You rock! 🎉 🚀
-
Amin Vahdat shared thisIt’s back-to-school season, and as a former professor, this time of year makes me nostalgic for the energy of a new semester and the deep, energizing debates we would have in the classroom. It is actually why I still love standing at a lectern* for formal presentations; it brings me right back to those teaching days. For me, the best part of academia was watching students unlock their potential and see them wrestle with complex systems to push the boundaries of their understanding and, eventually, scientific knowledge. I have always believed that the true promise of AI goes far beyond the models or infrastructure we build. At its core, AI innovation should act as a mind multiplier for human capability, a tool that clears away the friction of studying so students can focus on learning, critical thinking, and deep comprehension. This is why I am excited that Google is offering Gemini for free this academic year to students. It is incredibly rewarding to see how our products are helping students navigate their academic year. Full details in the link below. Wishing the very best to all students and educators as they kick off this new school year. https://lnkd.in/eSxPYqU6 * Many, including me until a few years ago, often confuse “podium” for “lectern”. The podium is the raised box upon which a speaker stands, though I am still grappling with this new reality.Amin Vahdat shared thisBack-to-school season is here. 🎒 To help students prepare for the upcoming academic year, we're offering student plans for 12 months at no cost, and study tools to make learning easier and more organized. Here are some of the Gemini tools available to help you make the most of this school year: • Student plans for one year, free of charge: Starting today, eligible college students in the U.S. can claim one year of Google AI Pro for free, unlocking 4x higher usage limits in Gemini, access to Gemini Spark, Gemini in Google apps like Gmail and Google Docs, 5 TB of storage, and more. For eligible college students in over 140 countries, we’re offering one year of Google AI Plus, free of charge. This offer unlocks access to Gemini Omni, 2x higher usage limits in Gemini and 400GB of storage. We’re also offering a bundle of Google AI Pro and a YouTube Premium individual subscription for up to 70% off for students in over 80 countries (including the U.S.) looking for ad-free music and video streaming. • New student hub: Students will have a dedicated hub with all of Gemini's tools and offers to stay organized, start a study notebook, create flashcards, take a practice quiz and more. As we roll out new learning tools, you’ll find them in the student hub. • Study notebooks: With study notebooks, Gemini can create a diagnostic quiz from your uploaded materials to identify your unique knowledge gaps. You'll receive custom lessons and quizzes, and be able to track your progress. • Interactive visualizations for visual learning: With interactive visualizations, Gemini can generate functional 3D simulations to help you better understand the topic you’re studying. For example, ask Gemini about a topic that could be explained visually, like “show me how DNA works in 3D,” and you’ll be able to interactively rotate and zoom into a 3D DNA structure. • Gemini Live: Today, we’re bringing Deep Research into Gemini Live so you can launch comprehensive, multi-step research reports and chat with Gemini about them. Ask Gemini to research a topic, then feel free to close the chat, lock your screen, or keep chatting about other things. Gemini works in the background and sends a notification when your report is ready. You can chat through the results, ask follow-up questions, or refine the details. By seamlessly switching from typing to talking, Gemini Live is the ultimate conversational study partner. Learn more and sign up for your free student plan for one year → https://goo.gle/4hDnhX4
-
Amin Vahdat shared thisWe continue to advance our water stewardship commitments to responsibly manage vital water resources where we build and operate data centers. Today, we’re proud to announce an additional $60 million in funding for new water stewardship projects in Arizona, Indiana, Ohio, Oklahoma, and Virginia. This builds on our original $17 million commitment across seven U.S. states, reinforcing our promise to responsibly manage vital water resources in the communities where we build and operate data centers. In addition, we are reaffirming our commitments across five dimensions by: 1. Replenishing more water than we consume at our sites by 2030 2. Modernizing water and wastewater infrastructure for our neighbors 3. Protecting at-risk watersheds with air-cooled solutions 4. Reporting our annual water use transparently 5. Pursuing alternative and reclaimed solutions to protect water resources See our blog post below for full details. https://lnkd.in/gvpF3MWnGoogle’s water stewardship commitments for local communitiesGoogle’s water stewardship commitments for local communities
-
Amin Vahdat liked thisAmin Vahdat liked thisAGI-pilled and want to use AI to change how chips get built? In AI2, we build TPUs: the world’s most cost-effective, high-performance ML accelerators. By pairing hardware design with Google’s breakthrough agentic tools like Teamwork (https://lnkd.in/e76VUBrw), we use state-of-the-art AI to build the next generation of chips. There’s no place like Google: from frontier model development, agents, ML frameworks, profilers, compilers, silicon architecture and data centers all operating as One Team. Coolest chips, planet-scale impact, with the best teammates *everrr*. Hiring for in-person US positions with experience in either of these: (1) Architecture pathfinding, VLIW, understanding the full process of exploring and adding new features. (2) perf sim: modeling LLMs and other workloads on hypothetical HW to guide the roadmap. (3) design verification: System Verilog, UVM, DFT. Bonus points if you build your own AI tools and workflows. Amin Vahdat introducing TPU 8t and 8i 🔥 : https://lnkd.in/enQ4hfWX Interested? DM me with the most impressive project you’ve ever done.Google TPU 8t and TPU 8i: Purpose-built for the Agentic EraGoogle TPU 8t and TPU 8i: Purpose-built for the Agentic Era
-
Amin Vahdat liked thisAnother big step for UCP! Building on what we shared at I/O, UCP is officially expanding to travel with the release of the initial draft spec of UCP for Lodging. Huge thanks to our Lodging Technical Council co-members—including Amadeus, Booking.com, Expedia, Hilton, Marriott Hotels, and Trip.com — for developing this open standard with us to simplify booking lodging across AI surfaces. Looking forward to where we take this next! Details here: https://lnkd.in/gXWCAScaAmin Vahdat liked thisAI is transforming how travelers discover and plan trips – and digital booking needs open, collaborative standards to match. Proud to share that today marks the release of the draft specifications for the Universal Commerce Protocol (UCP) for Lodging! Co-developed with our fellow Lodging Tech Council members (including Amadeus, Booking.com, Expedia, Hilton, Marriott Hotels, and Trip.com), UCP simplifies hotel transactions across AI surfaces while keeping partners in direct control of their customer relationships. We recently introduced conversational hotel booking through AI Mode in Search, and going forward, we’ll work closely with the industry to power this experience using UCP and further refine the specs. The draft is now open for public feedback, and we’d love your input: https://lnkd.in/ehq9j9Dt
-
Amin Vahdat liked thisAmin Vahdat liked thisAs a fellow Mancunian with great memories of my childhood in Sale, I was delighted to join UK Prime Minister Andy Burnham to discuss how technology – responsibly executed – can enable ‘good growth in every postcode’. For the last 20 years, the UK has been home to some of Google’s most groundbreaking work – from our Nobel Prize-winning work pioneered by Google DeepMind, to AI-assisted mammography work with the NHS; and last year, we announced a deepening of Google’s partnership with the UK with a £5B investment over 2 years. As Google invests, our commitment is to deliver responsibly for the communities and countries where we are growing – adding clean energy capacity, strengthening energy resilience, advancing economic growth, creating jobs, and supporting community development and skills. I first met then-Mayor of Manchester Burnham a decade ago at the launch of one of Google’s community-focused digital skilling programs in Manchester; since then, I’m proud that Google has expanded these programs to train more than 1M Brits. As Google’s footprint in the UK continues to grow, we look forward to working closely with the Prime Minister and his team to support investment and growth across the United Kingdom.
-
Amin Vahdat liked thisAmin Vahdat liked thisIn this Google Cloud blog post, Andrés Lagar-Cavilla and I provide a brief overview of how we have been using agentic AI in Google's AI and Infrastructure team to secure the hundreds of millions of lines of Google's code. The key innovation is *pervasive vulnerability scanning* directly embedded into the software development lifecycle. In other words, we *continually* scan *every* code change that happens in our infrastructure. Read the blog to find out how we built and deployed this, and lessons you can take for your security journey. Special call out to Stella Voutsina, Yulong Zhang, Nick Galloway. https://lnkd.in/gQJXsJu6Using AI agents to secure Google infrastructure | Google Cloud BlogUsing AI agents to secure Google infrastructure | Google Cloud Blog
-
Amin Vahdat liked thisAmin Vahdat liked thisCan an autonomous AI independently advance state-of-the-art research published by top human scientists? In our latest paper from Google Cloud & AI Research, we show the answer is yes. I’m thrilled to introduce ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI. Most automated research agents today run into two major bottlenecks: they overfit to single scalar metrics on narrow benchmarks, and they lack the self-correcting empirical rigor of real scientists. They don't systematically run ablations to isolate what actually works, nor do they stress-test their ideas against peer review. We designed ScientistTwo to execute the entire scientific lifecycle end-to-end: - Diagnoses human SOTA limitations rather than guessing blindly. - Formulates and screens novel hypotheses across multi-dataset benchmarks using a subset-to-full-set strategy. - Conducts automated ablation studies to isolate exact causal mechanisms of performance gains. - Simulates closed-loop peer review & rebuttals, where a dedicated Rebuttal Agent writes code and runs new experiments to directly address reviewer critiques. - Drafts full publication-ready manuscripts accompanied by verified, reproducible codebases. To measure it against the highest scientific standards, we benchmarked ScientistTwo across 107 competitive papers accepted at top-tier venues (ICLR, ICML, NeurIPS): 📊 The Key Findings: - 80.4% Success Rate: Advanced 86 out of 107 target problems. - +25.2% Mean Relative Gain: Consistently outperformed human state-of-the-art baselines across LLMs, optimization, RL, and time series. - Passed Venue Acceptance Standards: Surpassed average review scores of human-accepted papers at ICLR 2026 and NeurIPS 2025 under automated AI reviewers (ScholarPeer & Stanford Agentic Reviewer). - Zero Hallucinations: Passed 100% of our Chain-of-Evidence integrity audit—clean method-code alignment, zero specification violations, and verified citations. - Compounding Discovery: When fed its own newly discovered solution as a baseline, ScientistTwo iteratively found subsequent SOTA improvements across multiple generations. Autonomous AI is moving beyond assisted coding to pioneering new knowledge. 📄 Paper: https://lnkd.in/gXjiGpDE 🌐 Website: https://lnkd.in/gJsTPDje Incredible collaboration with Jaehyun Nam, Jinsung Yoon, Yanzhou Pan, Yubo Wang, Rui Meng, Partha Ranganathan, Tomas Pfister! #ArtificialIntelligence #MachineLearning #AutonomousAgents #AIScience #DeepLearning #ResearchInnovation #AI #GoogleCloud #NeurIPS #ICLR #ICML cc Amin Vahdat, Burak Gokturk
-
Amin Vahdat liked thisI have been reflecting on an absolutely incredible journey to Taiwan. The vibrant energy, the pace of innovation, and the amazing people I met left me more optimistic about our ability to meet the moment in delivering infrastructure at unprecedented scale, velocity, reliability, and efficiency. Here are some of my top takeaways along with a photo montage to capture it all*. 1. Beyond the breathtaking scenery and dynamic cities, the shared ambition for the future of technology. Thank you, Semicon, for the opportunity to deliver the conference keynote this year and to strike a dialog around the technology and collaboration needed to deliver the compute capacity supporting model and services breakthroughs. 2. I was moved by the profound and shared sense of responsibility; local companies, Googlers, and members of the press. There is a collective commitment to support the massive demand for AI and infrastructure in both a responsible and sustainable manner. It is inspiring to see an entire ecosystem work to prioritize rapid technical advancement while being mindful of the myriad costs to do so. 3. The spirit and unity of Google Taiwan made for a special experience. It is always a privilege to witness a team culture that builds so cohesively together. As a long-time Googler, it’s what makes Google special. The passion and mutual support across different teams showed that the best technology is built by those who genuinely care for one another. Google Taiwan also just celebrated its 20th anniversary, and it was wonderful to be part of the celebration. 4. The overriding impression of this trip is the genuine warmth of the people, matched by the technical breakthroughs they are driving. Working in direct collaboration with Google daily, the teams are driving breakthrough work around automation, hardware reliability, and the very limits of physics in chip design and manufacturing. Thank you to everyone in Taiwan who made this visit so memorable. We are building something extraordinary together. * Perhaps most striking is how much everyone loves taking photos, especially with the universal 👍🏻 signifying the joy and optimism that permeated so many of the discussions.
Experience & Education
-
Google
***** ************* ** ************** *** ***
-
********** ** *********** ********
*** undefined undefined
-
********** ** *********** ********
** undefined
View Amin’s full experience
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Publications
-
Pip: Detecting the Unexpected in Distributed Systems
Proceedings of NSDI
Bugs in distributed systems are often hard to find. Many bugs reflect discrepancies between a system's behavior and the programmer's assumptions about that behavior. We present Pip, an infrastructure for comparing actual behavior and expected behavior to expose structural errors and performance problems in distributed systems. Pip allows programmers to express, in a declarative language, expectations about the system's communications structure, timing, and resource consumption. Pip includes…
Bugs in distributed systems are often hard to find. Many bugs reflect discrepancies between a system's behavior and the programmer's assumptions about that behavior. We present Pip, an infrastructure for comparing actual behavior and expected behavior to expose structural errors and performance problems in distributed systems. Pip allows programmers to express, in a declarative language, expectations about the system's communications structure, timing, and resource consumption. Pip includes system instrumentation and annotation tools to log actual system behavior, and visualization and query tools for exploring expected and unexpected behavior. Pip allows a developer to quickly understand and debug both familiar and unfamiliar systems.
We applied Pip to several applications, including FAB, SplitStream, Bullet, and RanSub. We generated most of the instrumentation for all four applications automatically. We found the needed expectations easy to write, starting in each case with automatically generated expectations. Pip found unexpected behavior in each application, and helped to isolate the causes of poor performance and incorrect behavior.Other authorsSee publication -
WAP5: Black-Box Performance Debugging for Wide-Area Systems
Proceedings of WWW
Wide-area distributed applications are challenging to debug, optimize, and maintain. We present Wide-Area Project 5 (WAP5), which aims to make these tasks easier by exposing the causal structure of communication within an application and by exposing delays that imply bottlenecks. These bottlenecks might not otherwise be obvious, with or without the application's source code. Previous research projects have presented algorithms to reconstruct application structure and the corresponding timing…
Wide-area distributed applications are challenging to debug, optimize, and maintain. We present Wide-Area Project 5 (WAP5), which aims to make these tasks easier by exposing the causal structure of communication within an application and by exposing delays that imply bottlenecks. These bottlenecks might not otherwise be obvious, with or without the application's source code. Previous research projects have presented algorithms to reconstruct application structure and the corresponding timing information from black-box message traces of local-area systems. In this paper we present (1) a new algorithm for reconstructing application structure in both local- and wide-area distributed systems, (2) an infrastructure for gathering application traces in PlanetLab, and (3) our experiences tracing and analyzing three systems: CoDeeN and Coral, two content-distribution networks in PlanetLab; and Slurpee, an enterprise-scale incident-monitoring system.
Other authorsSee publication
View Amin’s full profile
-
See who you know in common
-
Get introduced
-
Contact Amin directly
Other similar profiles
Explore more posts
-
Bahareh Banijamali
NVIDIA • 7K followers
🏢 Deploy #AI inference at data center scale. NVIDIA Dynamo is now available across major cloud providers to enable efficient multi-node inference on Kubernetes in the cloud, including: ☁️ Amazon Web Services (AWS) ☁️ Google Cloud ☁️ Microsoft Azure ☁️ Oracle Cloud Infrastructure (OCI) And it’s already delivering results: Baseten is seeing faster, more cost-effective complex model serving. Dynamo integrations for #Kubernetes simplify AI inference at data center scale, boosting performance and efficiency for complex AI reasoning and MoE models. ✨ New NVIDIA Grove API in Dynamo streamlines coordinating inference components in Kubernetes, improving predictability and resource utilization. 🔗 Read the blog post to learn more.
15
-
Stanley Thuita
Zone01 Kisumu • 30K followers
🚀 $100 BILLION AI CHIP OPPORTUNITY INCOMING Broadcom CEO Hock Tan just dropped a prediction that should make every tech leader pay attention: AI chip revenue could exceed $100 billion by 2027. Let that sink in. 💡 Here's what's actually happening: The semiconductor industry is experiencing an unprecedented boom. Broadcom posted record Q4 revenues of $18 billion (up 28%), but here's the kicker—their AI chip business surged 74% year-over-year. 74%. That's not just growth. That's a fundamental shift in how the world builds AI infrastructure. 🔥 Why this matters: Every AI model—whether it's training GPT-level systems or running inference at scale—needs specialized chips. Broadcom is betting big on 3D stacked chip technology, with plans to sell at least 1 million units by 2027. Think about that: 1 million advanced chips designed specifically for AI workloads. That's not a niche market anymore. That's the backbone of the next decade of computing. 📈 The real story: This isn't just about Broadcom's success. It's about the entire AI infrastructure race. Every cloud provider, every AI startup, every enterprise building AI systems needs these chips. The demand is insatiable. The companies that control the silicon control the future of AI. ❓ Here's my question for you: As AI becomes more compute-intensive, do you think chip manufacturers will be the real winners of this AI boom—even more than the AI software companies? Share your thoughts below. 👇
1
-
Robert Williamson
Arm • 3K followers
Warner Bros. Discovery (WBD) has achieved ~60% cost savings and significant latency improvements by moving its ML-inference workloads to AWS Graviton-based infrastructure. As someone working at Arm and partnering with cloud hyperscalers like Amazon Web Services (AWS), this story hits home for three reasons: 1) Performance and cost align: WBD achieved improvements in p99 latency ranging from ~7% up to ~60%, while also reducing operating cost by roughly 60%. 2) Arm-based architecture scaling: The #Graviton family (built on Arm architecture) is optimized for vector/matrix operations (Neon, SVE, MMLA) making it a strong fit for #ML #inference workloads. 3) Real-world validation: This isn’t a niche test—WBD is delivering recommendations to 125+ million users across 100+ countries, using this infrastructure change to drive both user experience and business efficiency. For engineering leaders evaluating cloud strategy: this case underscores that moving to Arm-based cloud instances isn’t just a “green” or “nice to have” story—it can materially shift cost, performance and scale. https://lnkd.in/gmkTCcC9
43
2 Comments -
Eliuth Triana
NVIDIA • 7K followers
At the upcoming GTC, you can learn how system-level innovations are accelerating LLM inference. A good example is this new AWS blog on P-EAGLE, introducing GPU parallel speculative decoding in vLLM to generate and verify multiple tokens in parallel. Why this matters? Inference performance is increasingly a systems problem. Techniques like this improve tokens/sec per GPU, helping reduce latency and cost for production GenAI workloads. Good read for anyone working on LLM serving and inference optimization. Huang Xin Florian Saupe Jaime Campos Salas Benjamin Chislett Zhenghang (Max) Xu Faradawn Yang Omri Almog Jiahong Liu AWS blog: https://lnkd.in/g-GjMgvX�
43
5 Comments -
Saurabh Gayen
1K followers
AI infrastructure is starting to look like a single, extended memory hierarchy. I took a pass at mapping this into a unified “memory pyramid,” spanning from on-die SRAMs to network-attached storage, and capturing key AI trends across LPUs, GPUs, rail-optimized networks, and KVCache-optimized inference fabrics. Any other key trends I should consider including? More here (full article): https://lnkd.in/gsMcNn9W
35
-
Keith Strier
Special Competitive Studies… • 31K followers
AMD and IBM team with Zyphra to accelerate open-source, enterprise superintelligence. Zyphra’s Maia, a multi-modal AI agent, will unify knowledge discovery, communication, and work into one platform. A large cluster of AMD Instinct™ MI300X GPUs on IBM Cloud will be used by Zyphra to train frontier multimodal foundation models. Learn more: https://lnkd.in/gJm8a6r2 #TogetherWeAdvance_Zyphra
382
9 Comments -
Ajay Joshi
CipherSonic AI • 3K followers
One of the most common misconceptions about Fully Homomorphic Encryption (FHE) is around the keys: what gets generated, what gets shared, and what never leaves the data owner. Rashmi Agrawal does a great job breaking down the FHE key lifecycle in a clear, accessible way. #FHE #Keys #DataPrivacy #DataEncryption #EncryptedComputing
8
-
Rick Schoonmaker
IBM • 3K followers
Building an AI-Ready Infrastructure: The Foundation for Scalable AI Workloads AI is transforming industries, but success depends on having the right infrastructure to support every phase of the AI lifecycle: ✅ Training – Building models from massive datasets requires extreme parallel compute and storage throughput. ✅ Fine-Tuning – Adapting models to business-specific data demands a balance of compute and I/O for rapid iterations. ✅ Inferencing – Delivering real-time insights in production calls for low latency and high reliability. An AI-ready infrastructure includes: 🔹 Accelerators for AI Math – CPUs, GPUs, NPUs, and custom chips optimized for performance and cost efficiency. 🔹 Fast Memory & Fabric – High bandwidth, low latency, and non-blocking design to move data at scale. 🔹 Smart Data Pipelines – Tiered storage (hot, warm, cold) with prefetching for seamless data availability. 🔹 Secure & Governed Operations (MLOps) – Ensuring trust, compliance, and cost-effective innovation. Watch the full video here: https://lnkd.in/ehYpzqEx Question to consider: How is your organization preparing its infrastructure for AI at scale?
18
-
Partha Ranganathan
Google • 25K followers
10/25: Tenth in the series on 25 Google papers for 25 years of WSC: Google-Wide Profiling The "Always On" Monitor: Most companies turn profiling off in production because they fear the overhead. At Google, we keep it on. This paper (Google-Wide Profiling) explains how to get continuous, fleet-wide visibility with a tiny tax: just 0.01% overhead. Why bother, you might ask? Because "lab" environments can be misleading. By profiling in prod, you find: Hardware differences you didn't expect; "Cold" code that is actually running hot; Library bottlenecks across every binary in the fleet. It’s the difference between guessing where the performance went and _knowing_. And for today's nano-banana poster for the paper, I tried a different prompt: [Make an infographic academic-style poster for the following paper. Use a visual aesthetic similar to the architectural blueprint roll -- Large blue sheets with white technical lines, measurements, and room labels.]
382
12 Comments -
Ronnie Moodley (MBA)
IBM • 2K followers
Big News for AI Innovators! Excited to announce the availability of Red Hat OpenShift AI 3.0 on IBM Power—a major milestone in accelerating enterprise AI with open innovation. This release brings: ✅ A unified MLOps platform for building, training, and deploying AI models ✅ New capabilities like Feature Store, Model Registry, and KServe integration ✅ Guardrails for safer generative AI outputs ✅ Support for vLLM for efficient inference By running OpenShift AI on IBM Power, organizations gain performance, scalability, and resilience for their most demanding AI workloads—all while maintaining flexibility and avoiding vendor lock-in. Ready to take your AI strategy to the next level? 👉 Read the full announcement blog: https://ibm.biz/Bdb8wu #IBMPower
42
1 Comment
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top contentOthers named Amin Vahdat
13 others named Amin Vahdat are on LinkedIn
See others named Amin Vahdat