Sign in to view Steve’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Steve’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
San Luis Obispo, California, United States
Sign in to view Steve’s full profile
Steve can introduce you to 10+ people at Google
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
3K followers
500+ connections
Sign in to view Steve’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Steve
Steve can introduce you to 10+ people at Google
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Steve
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Steve’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
3K followers
-
Steve McGhee shared thisHey that's us!Steve McGhee shared thisSeason 7 of the Prodcast is coming soon, and we've put together a set of videos to share some of what happens behind the scenes. These videos introduce the hosts, engineers, and producers for Prodcast. New full-length episodes with Jordie G. Paul Guglielmino Steve McGhee Matthew Siegler drop in two weeks! Fire up your podcast apps and keep your eyes on https://lnkd.in/e8Fg-gUi
-
Steve McGhee reposted thisSteve McGhee reposted thisA couple of weeks ago someone in a forum I frequent someone asked what the "OWASP Top 10 for Reliability" was. As far as I know, there isn't one, and I wanted to know why. In the CWE list, each CWE has a stable ID, and scanners emit those IDs, so tools and rankings speak the same language. Reliability doesn't have anything like that. In our incident corpus, 2,309 contributing-condition fields hold 2,207 distinct values, and 97.6% of them occur exactly once. Everyone writes each cause from scratch. There were previous attempts, from Oppenheimer et al. in 2003 and Gunawi et al. in 2016. Neither became a shared vocabulary. I've tried to understand why, and what a replacement has to do differently. Initially, I thought a ranking might be simple, but it turns out you need something to rank. As a result, we're publishing the first edition of the causal factor enumeration catalog: 38 contributing conditions, each with a stable CF identifier, in seven categories adapted from CAST (Leveson, MIT). Each entry states what the system did. It never names an absence, and promotes blameless analysis. CF-0026, for example: "A protective control engaged and its action became the outage." Each entry also carries quotes from the public incident reports it covers, with a link to each report. This is a first edition, and it's knowingly incomplete. If an entry is wrong, or one of your incidents doesn't fit any of the 38, tell me. Three of the entries exist because a reviewer gave feedback on the draft. More details about the methodology and inspiration: https://lnkd.in/gAj9qTtG The full catalog: https://lnkd.in/g_3CkaFJBuilding a Causal Factor Enumeration for Reliability | RevelaraBuilding a Causal Factor Enumeration for Reliability | Revelara
-
Steve McGhee shared thisThe SRE backpack travel lives on. (Dolemites) Thanks again Sarah Coty and Daniel Hobe :)
-
Steve McGhee reposted thisSteve McGhee reposted thisI've talked to hundreds of engineers about reliability, and the same realization surfaced almost every time: everyone was reactive. On-call, incident response, war rooms, all of it aimed at cleaning up faster after something broke. They were pouring real money and real hours into it. And almost none of them felt like they were getting ahead, because nobody knew how to build the other half: a program that gets in front of the failures instead of chasing them. Reacting fast was good enough when code moved at human speed, but that's not true anymore. AI coding agents ship changes faster than any of us can reason about, and the consequences to reliability is where that shows up first. Not as a dramatic outage. As a quiet pile of risk nobody has looked at, or understands yet. Here's the part I don't love admitting: I've been the one who shipped the risky code. I've written it. I've missed seeing the failure mode before it hit. I've caught it three incidents later, in an incident review, wishing someone had said something while there was still time to do something about it. That's not a character flaw. It's a missing feedback loop. So I built the loop. Today Revelara is open to everyone. Revelara reads your codebase, correlates what it finds against a knowledge base of thousands of real incidents, and surfaces the reliability risks right inside the tools you're already using, like Claude Code and Cursor. It gives your agents a reliability grounding they wouldn't have otherwise. You get a shared risk register, recommended controls, and evidence that tracks itself. Reliability stops being tribal knowledge and becomes something you can actually govern proactively. One thing matters to me more than the demo: this is not AI finding your risks and quietly fixing them behind you. It hands judgment back to the person at the keyboard, with the context to use it well. Reliability was never really about uptime. It's about understanding, and understanding has to live in people. There's a 21-day trial, no credit card. Watch the demo below, and go see the product at revelara.ai. Is your team getting ahead of reliability risk, or just trying to get better at cleaning up after it? If you're not sure, check out Revelara and find out.
-
Steve McGhee shared thisThis was a fun one. I’ve worked with Adam for ummm some large number of years that we don’t need to get into. He’s a great SRE and cool dude. Also, he represents here one of my favorite parts of SRE that doesn’t get talked about enough: “IRTs” or teams of folks whose job is just: to help out. Experts you can call up, or sometimes just drop in and say “can I help?” This is actually my new role too, so I guess I’m biased. This wraps up another season of the Prodcast and it makes me happy to have these out in public. I’ve joked that my job was, for a while, to “exfiltrate Google SRE lore” (don’t worry, there’s a whole approval process) and this is a good one. Enjoy! See you next season.Steve McGhee shared thisGoogle's Tech Incident Response Team (IRT) manages the complex production environment during times of intensive change. In this episode, IRT member Adam Kramer talks about the importance of psychological safety to ensure that engineers can communicate effectively. Adam joins hosts Steve McGhee and Matthew Siegler on this episode of the Prodcast, Google's podcast on site reliability engineering and production software.Incident response: psychological safety and effective communicationIncident response: psychological safety and effective communication
-
Steve McGhee reposted thisSteve McGhee reposted thisGoogle's Tech Incident Response Team (IRT) manages the complex production environment during times of intensive change. In this episode, IRT member Adam Kramer talks about the importance of psychological safety to ensure that engineers can communicate effectively. Adam joins hosts Steve McGhee and Matthew Siegler on this episode of the Prodcast, Google's podcast on site reliability engineering and production software.Incident response: psychological safety and effective communicationIncident response: psychological safety and effective communication
-
Steve McGhee reposted thisSteve McGhee reposted thisJohn Allspaw of Adaptive Capacity Labs talks with the Prodcast about the challenges of dynamic systems, the value of learning, and much more. With hosts Steve McGhee, Florian Rathgeber, and Matthew Siegler, in two parts, live from SREcon!
-
Steve McGhee reposted thisCourtney is amazing! She has a unique combination of skills that I have not encountered elsewhere. From consuming her work on The VOID, I’ve witnessed her deep understanding of software reliability, which was reinforced for me when I heard her SREcon23 talk “Far from the Shallows: The Value of Deeper Incident Analysis”. In addition, she possesses the organizational skills that enable the real work to actually happen, whether that’s organizing the iconic O’Reilly Velocity conference series (where I first met Courtney), or leading the Resilience in Software Foundation’s blogathon effort to shepherd a steady stream of content for the foundation’s blog.Steve McGhee reposted thisWhelp, it turns out solopreneurship is hard and I am not immune to that particular truth. I had really hoped that I’d be able to continue to build out The VOID and do original research on software incidents and complex systems under that umbrella, but at least for now, the universe seems to have other ideas. My career hasn’t followed a straight line so there’s not always a direct correlation between my skills and job titles. Most of it has been spent helping explain complexity: whether it's humans, computers, or the messy place where they intersect. As both an editor and a researcher, my sweet spot is taking deeply technical ideas and making them not just understandable but actionable. I'm currently looking for roles that generally map to technical editorial leadership, incident research, and technical program management. I’m open to contract or full-time roles, the latter especially if the fit is right. (I'm also still taking consulting work if you need someone to help you improve and fine-tune your incident response and/or analysis processes!) If you're working on hard problems at the intersection of humans and complex systems and need someone who can research, synthesize, and communicate that work, I'd love to hear about it. I’d especially appreciate a warm intro over a link to a job post (though the latter still helps, especially if the role strikes you as particularly Courtney-shaped).
-
Steve McGhee reposted thisSteve McGhee reposted this"We don't focus enough on expertise, learning, how we can share that with each other. Those are the things that make the incidents less painful or not happen at all." Courtney Nash of The VOID discusses the role of human expertise in managing complex systems, and how SREs continue to bring critical value even as technology and AI evolve.
-
Steve McGhee liked thisSteve McGhee liked thisSeason 7 of the Prodcast is coming soon, and we've put together a set of videos to share some of what happens behind the scenes. These videos introduce the hosts, engineers, and producers for Prodcast. New full-length episodes with Jordie G. Paul Guglielmino Steve McGhee Matthew Siegler drop in two weeks! Fire up your podcast apps and keep your eyes on https://lnkd.in/e8Fg-gUi
-
Steve McGhee liked thisSteve McGhee liked thisA couple of weeks ago someone in a forum I frequent someone asked what the "OWASP Top 10 for Reliability" was. As far as I know, there isn't one, and I wanted to know why. In the CWE list, each CWE has a stable ID, and scanners emit those IDs, so tools and rankings speak the same language. Reliability doesn't have anything like that. In our incident corpus, 2,309 contributing-condition fields hold 2,207 distinct values, and 97.6% of them occur exactly once. Everyone writes each cause from scratch. There were previous attempts, from Oppenheimer et al. in 2003 and Gunawi et al. in 2016. Neither became a shared vocabulary. I've tried to understand why, and what a replacement has to do differently. Initially, I thought a ranking might be simple, but it turns out you need something to rank. As a result, we're publishing the first edition of the causal factor enumeration catalog: 38 contributing conditions, each with a stable CF identifier, in seven categories adapted from CAST (Leveson, MIT). Each entry states what the system did. It never names an absence, and promotes blameless analysis. CF-0026, for example: "A protective control engaged and its action became the outage." Each entry also carries quotes from the public incident reports it covers, with a link to each report. This is a first edition, and it's knowingly incomplete. If an entry is wrong, or one of your incidents doesn't fit any of the 38, tell me. Three of the entries exist because a reviewer gave feedback on the draft. More details about the methodology and inspiration: https://lnkd.in/gAj9qTtG The full catalog: https://lnkd.in/g_3CkaFJBuilding a Causal Factor Enumeration for Reliability | RevelaraBuilding a Causal Factor Enumeration for Reliability | Revelara
-
Steve McGhee liked thisSteve McGhee liked thisGreat sessions at Evolve NYC today https://lnkd.in/eJ_TMFMv Always great to see Casey West, and also to see Nathen Harvey in action!
-
Steve McGhee liked thisSteve McGhee liked thisThe rumours are true, I live in New York City now.
-
Steve McGhee reacted on thisSteve McGhee reacted on thisWhile I was traveling last week, an essay that I've been ruminating on for quite a while finally fell out of my head and onto a page. It's not a marketing piece. It's not me trying to put on my best impersonation of incident analysts much better than me and doing it poorly. It's not even a report on what companies had outages last week. It's just an essay born of personal interest and experience. It's really long, as in -- absolutely not AI generated and hopefully worth the read -- but realistically I don't expect many will given the many words this piece uses to compare how AI is impacting the world to the first industrial revolution. If you do, drop a note and let me know what you think. https://lnkd.in/g_ns6Ygm Also, in case there was ever any question, double-hyphens are the best form of emdash. IYKYK.
Experience & Education
-
Google
**** *********** ********
-
********* ****
************** *********
-
********* **** *** ********
***********
-
** ***** *******
** ******** ******* undefined
-
-
** ***** *******
** ******** *******
-
View Steve’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Courses
-
Distributed Systems
-
-
Network Security
-
-
Scalable Internet Services
-
Languages
-
Spanish
-
Organizations
-
ACM
-
View Steve’s full profile
-
See who you know in common
-
Get introduced
-
Contact Steve directly
Other similar profiles
-
Luther Hill (CISSP)
Luther Hill (CISSP)
CLEMSON UNIVERSITY FOUNDATION
16K followersGreenville-Spartanburg-Anderson, South Carolina Area
Explore more posts
-
Marcelo S.
Amazon Web Services (AWS) • 3K followers
AWS Network Engineering VP Matt Rehder shares how we're addressing AI's networking demands in Data Center Knowledge. The challenge: ML servers require 2-3x more bandwidth than traditional systems. Our approach includes a redesigned control plane enabling sub-second failure recovery, hollow-core fiber deployment in 5-10 locations for geographic flexibility, and custom networking hardware across our entire infrastructure. Real impact: customers get more capacity, lower latency, and less jitter—without thinking about the network at all. Dive into the technical details →
29
-
Sree Chadalavada
Open Compute Project… • 6K followers
Hope this initiative drives convergence of AI scale-up networking standards. The following enables clear delineation, ownership, and interoperability to make scale-up standards successful. The scale-up domain in XPU-based systems can be viewed in two primary areas: 1) network functionality, and 2) XPU-endpoint functionality
3
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content