Sign in to view Miguel’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Arlington, Texas, United States
Sign in to view Miguel’s full profile
Miguel can introduce you to 10+ people at AT&T
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
953 followers
500+ connections
Sign in to view Miguel’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Miguel
Miguel can introduce you to 10+ people at AT&T
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Miguel
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Miguel’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Activity
953 followers
-
Miguel Armenta posted thisAlmost 2 years of watching LiteLLM AI Gateway evolve, and today was a big one. I started deploying and evaluating this product at AT&T back in January 2025 as an option for an enterprise AI Gateway. Since then I've had the chance to contribute code back to the project, and to watch our own usage grow from the open source tier into a full enterprise deployment. The pace of improvement has been genuinely impressive: features, speed, hardening, and optimization that keep compounding release over release. Today marked another leap. I rolled out their latest architectural shift, moving core routing from a monolith to dedicated microservices, purpose-built for throughput. I'll be sharing performance numbers soon alongside the LiteLLM team, once we've had time to validate them in production. For now, credit to a team building infrastructure that keeps earning its place at the center of enterprise AI stacks. What an experience to build alongside Krrish D., Ishaan Jaffer and Yassin Kortam. #AIInfrastructure #EnterpriseAI #LLMOps #OpenSource #LifeAtATT
-
Miguel Armenta shared thisHeading to Ray Summit 2026 in San Francisco next week. I am very excited to catch the first-ever co-located vLLM Conference. vLLM is core to the inference platform I run at AT&T. If you're into Ray, PyTorch, LMCache Lab, llm-d, vLLM, NVIDIA, AMD or KServe let's connect! #RaySummit2026 #vLLM #AIInfrastructure #Ray https://lnkd.in/ggC7Gsjr
-
Miguel Armenta reposted thisMiguel Armenta reposted thisI'm hiring a Site Reliability Engineer at LiteLLM (YC W23). You will be engineer #6 - based in San Francisco, working directly with me (In Person Only) DM me or apply here if interested https://lnkd.in/g83XbEQR About Us: LiteLLM is an open-source LLM gateway with 34K stars on GitHub (YC-backed). Millions of LLM requests flow through our AI Gateway every day. We're growing fast and need someone who can make our infrastructure as solid as our adoption. You'd own reliability for a system that's critical infrastructure for a lot of teams. OOMs, connection pool tuning, race conditions, cache consistency, making the proxy self-heal under load. Real systems problems. Stack: Python, FastAPI, Postgres, Redis, Kubernetes. Small team, moving fast. You'd work directly with me.Site Reliability Engineer at LiteLLM | Y CombinatorSite Reliability Engineer at LiteLLM | Y Combinator
-
Miguel Armenta shared thisI'm thrilled to share that I recently completed the Corporate Executive Development Program from the SMU Cox School of Business at Southern Methodist University, a premier leadership experience for high-potential leaders. As a software engineer specializing in AI and ML Platforms, this program has been transformative. A huge thank you to my leadership team, Gus Mustakas, Abhay Dabholkar and Mark Austin, for their unwavering support and investment in my development, it truly made this possible. This achievement is propelling my career toward senior leadership roles, equipping me with advanced strategic thinking, team-building expertise, and the vision to lead innovative AI initiatives. I'm excited to apply these skills to drive cutting-edge ML platforms, foster high-performing engineering teams, and contribute to greater organizational impact. #LeadershipDevelopment #AILEADERSHIP #MachineLearning #ProfessionalGrowth #SMUCox #SMU #ExecutiveEducation #LifeAtATT https://lnkd.in/gHt8Kzh2Corporate Executive Development Program was issued by Cox School of Business at Southern Methodist University to Miguel Armenta.Corporate Executive Development Program was issued by Cox School of Business at Southern Methodist University to Miguel Armenta.
-
Miguel Armenta reposted thisMiguel Armenta reposted this🙏 Feeling extra grateful for our wonderful community today. Whatever LGTM+ means to you today — Leftovers Go To Me; Let’s Give Thanks, Mate; Limitless Gravy and Turkey Mayhem; or simply Loki, Grafana, Tempo, and Mimir — we wish you a happy Thanksgiving!
-
Miguel Armenta shared thisHey all! This is what I consider to be "the stack" for serving models. Let me explain: - Kubernetes, of course, to run at scale anywhere, on-prem or on any cloud provider. - KServe Inference Service, a clean solution to simplify model deployment, with easy autoscaling implementation. - vLLM, in my opinion, the best serving engine for LLMs, embedding models, and more. - LiteLLM (YC W23) Proxy Server, a lightweight proxy with support for vLLM and other providers, all under the OpenAI spec for compatibility and packed with many other features. This open-source stack is a killer combo, with great flexibility to support startups and large companies with high-availability requirements. I'm still working on putting together some documentation and a GitHub repo to share. Hopefully, in the next week or so, I'll wrap it up and post it here. Cheers! #llm #genai #scale #kubernetes #kserve #knative #infrastructure #mlops
-
Miguel Armenta posted thisGetting back to the topic of Generative AI workloads at scale... ...poor GPU utilization can silently and effectively drain your AI project’s budget and performance. On Kubernetes, getting Nvidia GPUs to work at full capacity for model inferencing can be like solving a puzzle with missing pieces. Autoscaling using Prometheus Adapter approach? It’s powerful but clunky—custom metrics, endless tweaking, and still, GPUs sit idle during low demand. I'm putting together a deployment strategy to simplify scaling, maximize GPU usage and making workloads highly available. Can't wait to post more about it! In the meantime, let's connect! Are you currently deploying any LLMs at scale? What is your tech stack? What pain points are you facing? #AI #Kubernetes #Nvidia #DataScience #CNCF
-
Miguel Armenta reposted thisMiguel Armenta reposted thisThrilled to enable 👋 AUDIT LOGS on LiteLLM (YC W23) for tracking key deletions on the Admin UI - https://lnkd.in/dUHG-GkD (+4 more updates 👇) 💪 Gemini - urlcontext tool support ✅ Gemini - disable thinking support s/o Jian Sheng Low 🚀 Redis - support lpop for Azure Redis s/o Miguel Armenta 🧹 Google Secret Manager - Respect GSM Project ID if set - enables working within cloud run environments
-
Miguel Armenta posted this🚨 Are you wrestling with scaling Nvidia GPU workloads for AI model inferencing on Kubernetes? The struggle is real: unpredictable demand, underutilized GPUs, and complex setups that eat up time and resources. I used to rely on Prometheus Adapter and custom metrics for autoscaling, but it felt like herding cats—complicated and never quite perfect. What if there’s a simpler, more efficient way to achieve high availability AND max GPU utilization? 🔥 I believe I’ve cracked the code and can’t wait to share! What’s your biggest scaling challenge? Drop it in the comments! 👇 #kubernetes #nvidia #ai #gpu #ML #devops
-
Miguel Armenta reacted on thisMiguel Armenta reacted on thisThrilled to launch LiteLLM Academy - https://lnkd.in/g7Yzs-8r This is designed for admins to ramp up to the product, and understand how the product works. It covers controlling access, routing, model selection, and how budgets + team management works. All in 1 place. s/o to Miguel Armenta for his suggestion and Moe Khalil for his work on this
-
Miguel Armenta liked thisMiguel Armenta liked thisGermany’s F-35 era starts now. Germany’s first F-35 officially rolled out in Fort Worth, marking an important milestone for Germany, the Luftwaffe and Europe as a whole. The F-35 strengthens allied interoperability and gives Germany proven 5th Gen capability to help deter threats and protect our shared security. I’m proud of our partnership with the German government, Luftwaffe, industry and the F-35 enterprise who made this milestone come to life.
-
Miguel Armenta reacted on thisMiguel Armenta reacted on thisReasoning models on the Responses API send back their reasoning encrypted, and only the deployment that wrote it can read it. Tin Lo explains how Auto Router keeps the readable summary when it moves a conversation to another model tier, so the next turn runs on the new tier. https://lnkd.in/gm2Fx6t6
-
Miguel Armenta reacted on thisMiguel Armenta reacted on thisYou can now send LiteLLM AI Gateway request logs to Databricks through Zerobus Ingest. Token usage, cost, latency, and request metadata land in Unity Catalog Delta tables, ready to query alongside the business data already in your lakehouse. For platform teams, this means: Breaking down AI spend by model, team, or application. Comparing token usage and latency across models. Connecting AI usage to business reporting with Databricks SQL. Our end-to-end quickstart includes screenshots, Python/JavaScript/cURL examples, and a query to verify your first request. https://lnkd.in/gC8SQbPa
-
Miguel Armenta reacted on thisMiguel Armenta reacted on thisWe’ve always embraced the potential of emerging technology. From the transistor and cell phone to AI, our teams have helped shape game-changing technologies that change how people connect, work, and live. Today, we’re bringing tested, scalable AI tools to employees across our offices, retail stores, call centers, and field operations. Generative AI is helping our teams find information faster, develop more complete solutions, simplify their work, and keep customers better connected. The next era of innovation is here… and we’re putting it to work. #LifeAtATT
-
Miguel Armenta reacted on thisMiguel Armenta reacted on thisSurreal seeing provider and model integrations I worked on get shipped to LiteLLM for millions of developers! It was great working with the Meta team to get their models integrated into LiteLLM. Looking forward to making the process even smoother for future launches!
-
Miguel Armenta reacted on thisMiguel Armenta reacted on thisYour agent shouldn’t use the same model for every task. Today we’re introducing liteagents SDK: the familiar query() interface from the Claude Agent SDK, with a built-in auto-router that decides which model to use for each turn, across providers. Ask it to “fix the login bug, add tests, and open a PR.” liteagents can distribute the work across models: Claude Opus 4.8 plans the fix, GPT-5.4 mini writes the code and tests, and Claude Sonnet 4.6 drafts the PR. The auto-router picks the best-fit model for each turn, balancing quality, speed, and cost
-
Miguel Armenta reacted on thisCongrats SpaceXAI! Day-0 support for Grok 4.7 on LiteLLM! Get started here: https://lnkd.in/dyUfG_dTMiguel Armenta reacted on thisGrok 4.7 is here. It's a notable improvement over Grok 4.6 at the same price and speed. Grok 4.7 works longer on difficult tasks, checks its work more carefully, and comes with our strongest safeguards to date. https://x.ai/news/grok-4-7
Experience & Education
-
AT&T
********* *** **** * ** ********
-
******* ** *******
*** ***********
-
*** ***********
****** *******
-
*** ********** ** ***** ** ** ****
****** ** ******* ****** ******* *********** undefined
-
-
*** ********** ** ***** ** ** ****
******** ** ******* ****** ********** ***********
-
View Miguel’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Licenses & Certifications
View Miguel’s full profile
-
See who you know in common
-
Get introduced
-
Contact Miguel directly
Other similar profiles
-
John Walz
John Walz
Product-focused Software Engineer with a deep passion for GenAI and the transformative potential of intelligent systems.<br><br>I'm a self-starter who thrives in the ambiguity of early-stage ventures, where my experience across multiple industries and tech disciplines allows me to be a highly-productive generalist. I believe the best products emerge when technical versatility meets genuine user needs.<br><br>With experience scaling both engineering teams and products from inception, I'm passionate about continuous learning through hands-on application of cutting-edge research and emerging technologies—bridging the gap between what's possible and what's practical.
10K followersUnited States
Explore more posts
-
Roberto Hortal
Wall Street English • 6K followers
Context engineering shifts LLMs from oracle to analyst—frame the task, curate tokens, and apply reusable patterns (RAG, tool calls, memory) to keep complex systems robust. A must-read for product and Agile teams building trustworthy AI-enabled workflows. Explore Chris Loy's take: https://buff.ly/Ca2GrrO #Product
1
-
Nalini Polavarapu
The Hartford • 10K followers
Will AlphaEvolve take off in 2026? When Google DeepMind first announced #AlphaEvolve, I was surprised it didn’t capture the same "mainstream" momentum as #AlphaFold or #AlphaGo. While those milestones solved major bottlenecks or complexities that are more relatable, AlphaEvolve is tackling something fundamental: the autonomous evolution of the code and algorithms that runs our world and industries. I’ve been bullish on its potential for a while now [https://lnkd.in/gmCmw_MA], it became much more accessible recently: AlphaEvolve is officially in public preview on Google Cloud. We are witnessing a shift from AI that simply "assists" with code or design of algorithms to agentic systems that “autonomously evolve it”. If you’re looking at this technology for your own roadmap, consider these results: • Infrastructure: Recovered 0.7% of global compute resources. • Speed: Accelerated AI training kernels by 23%. • Discovery: Solved math problems stagnant since 1969. By bringing this to the Cloud, Google Cloud is finally closing the gap between "research breakthrough" and "enterprise reality." I suspect 2026 will be the year this "silent" revolution becomes the new standard for efficiency. Or, perhaps it’s just me enamored by the influence of the patterns of #life and the #humanbrain—the very systems that drove our greatest #AI breakthroughs—are now being mirrored in this new development. Either way, the question remains: Are enterprise systems actually ready to harness this evolution, or will it take another year or two to bridge that gap? Let’s see! #AlphaEvolve #DeepMind #GoogleCloud #AlgorithmOptimization #AI2026 #TechLeadership #BioInspiration Alexander Novikov Matej Balog Pushmeet Kohli Oliver Parker Vladimir Vuskovic Anant Nawalgaria Thomas Kurian
44
1 Comment -
Firoz Subair
Emirates • 3K followers
🚀 Built for Your Most Demanding AI Use Cases If you're scaling Retrieval-Augmented Generation (RAG), recommendation engines, or enterprise-grade semantic search, - performance isolation and predictable low-latency are no longer “nice to have” rather they’re foundational. That’s why Pinecone’s new Dedicated Read Nodes caught my eye. They are purpose-built for environments where every millisecond matters and workloads can’t afford noisy neighbors. 🔧 Why Dedicated Read Nodes? Because modern AI systems demand: • Predictable latency under heavy load • Linear scaling as vectors, traffic, and QPS multiply • Full isolation so one workload never impacts another • Cost predictability at enterprise scale 🔥 Perfect for high-demand scenarios such as: • Billion-vector semantic search with strict latency targets • High-QPS recommender systems (feeds, ads, marketplaces) • Mission-critical AI services with hard SLOs • Large enterprise / multitenant platforms that need strong workload isolation As vector databases become the backbone of enterprise AI, innovations like these are essential for reliability, safety, and scale. If you're architecting for high throughput and real-time intelligence, this is worth a look. 🔗 https://lnkd.in/dWTY7CXx
10
-
Amy Morrison
Elastic • 11K followers
How do you observe generative AI (GenAI) systems in production when traditional metrics fall short? As part of the I/built with Elastic series at Elastic{ON} Singapore, Adrian Cole, Principal Engineer at Tetrate, takes the stage to demonstrate how Envoy AI Gateway, powered by Elastic, turns opaque large language model (LLM) traffic into observable systems. Explore AI-focused observability patterns, OpenTelemetry integration, and a live demo connecting GenAI and traditional observability. Join the Observability track and register today. https://gag.gl/wTgP3X
4
-
Gevorg A. Galstyan
Life360 • 3K followers
An internal platform should describe the capability it provides, not only the technology it operates. “We run Kafka” and “we run Postgres” tell consumers about the engine. They do not define what a team can accomplish, which guarantees it may rely on, where responsibility ends, or how the implementation can evolve. A useful capability contract has four parts: 1. Outcome: what can the consumer reliably accomplish? 2. Guarantees: which behavior can the consumer depend on? 3. Boundary: what does the consumer own, and what does the platform own? 4. Evolution: what can change behind the contract without coordinated consumer rewrites? Hiding the engine is not the goal. Consumers still need to see health, limits, failure behavior, and responsibilities. A vague abstraction can be as costly as exposed machinery. The migration test is straightforward: If the engine changes while the promised capability stays the same, which consumer work is unavoidable, and which work exists only because implementation details crossed the boundary? Technology choices matter. Platform leverage comes from a contract that absorbs machinery changes without hiding responsibility or tradeoffs.
4
-
Tactical Edge
930 followers
Your GenAI stack doesn’t fail in testing. It fails quietly after launch. When latency creeps up, retrieval breaks, and hallucinations slip through, That’s not an accident. It’s bad architecture. We’ve audited more than 40 enterprise AI systems this year. The same seven fail points keep showing up again and again. From weak retrieval logic to missing audit trails, Every failure maps back to one missing layer: observability. If you’re not watching these signals, You’re not in production. You’re still in demo mode. Swipe through the breakdown. Then ask yourself: Which layer in your GenAI stack is most fragile?
5
-
Saptarshi Banerjee
OpenAI • 11K followers
🚀 Anthropic Claude 4.6 Sonnet is now live on Amazon Bedrock! The Vending-Bench Arena simulation (pictured below) highlights a significant advancement in strategic reasoning. Unlike previous iterations, Sonnet 4.6 demonstrates the ability to prioritize long-term capacity investment over immediate gains, resulting in a dramatic pivot toward profitability in the final stages. For our customers, this means access to a model that doesn’t just execute tasks, but strategically optimizes for complex, long-horizon business outcomes. You can begin building and scaling with Claude 4.6 Sonnet on Amazon Bedrock today: https://lnkd.in/gmE68-9N
19
3 Comments -
James Henderson
Texas Integrated Services • 5K followers
Alignment research is often siloed. For teams like Ollie Matthews - OpenAI Alignment, the daily friction isn't just technical complexity. It's the isolation of working on high-stakes problems without a structured peer network to validate assumptions or share mitigation strategies. When you are focused on mitigating AI-related risks and ensuring accountability in decision-making, having a single point of contact for diverse perspectives is rare. Most practitioners are left to navigate these challenges independently, leading to duplicated efforts and missed connections with other researchers facing similar hurdles. AI Coalition exists to bridge that gap. We are a network bringing together AI practitioners, researchers, and organizations to collaborate on responsible AI development. By joining, you gain access to a community dedicated to shared learning and responsible practice, moving beyond hype to focus on the actual work of building safer systems. If you are looking to expand your collaborative network and connect with others prioritizing accountability in AI, consider joining the coalition. https://ai-coalition.net #AIAlignment #ResponsibleAI #AICoalition
-
James Henderson
Texas Integrated Services • 5K followers
Most AI automation projects stall because the data feeding them is messy. Before you build complex workflows, spend one hour auditing your last 20 missed calls. If you can’t clearly see why they were missed, your automation will just automate the chaos. Mike Cindric - AI Automation for SMB, this is a common trap in the space. You’re likely seeing clients struggle not with the tech, but with the baseline data hygiene required to make AI actually useful. That’s where Lieutenant James Henderson comes in. We handle the messy front-end stuff—phone answering, lead capture, and scheduling—so the data is clean and structured from day one. It’s about $1,400/mo in savings with a 2-week setup, giving you a solid foundation to build on without the admin drag. Learn more: https://lnkd.in/euvUEfkF #AIAutomation #SMB #Operations
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content