Sign in to view Ben (Xiaojun)’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Ben (Xiaojun)’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
San Francisco Bay Area
Sign in to view Ben (Xiaojun)’s full profile
Ben (Xiaojun) can introduce you to 10+ people at Microsoft
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
35K followers
500+ connections
Sign in to view Ben (Xiaojun)’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Ben (Xiaojun)
Ben (Xiaojun) can introduce you to 10+ people at Microsoft
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Ben (Xiaojun)
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Ben (Xiaojun)’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Articles by Ben (Xiaojun)
-
Frontier Interview Practice — May 2, 2026
Frontier Interview Practice — May 2, 2026
Three questions I would actually expect to see in OpenAI / Anthropic loops this month, grounded in what those labs and…
4
-
Context Engineering In Enterprise ApplicationsJul 5, 2025
Context Engineering In Enterprise Applications
What is Context Engineering? Context engineering is the art and science of structuring everything an AI model needs to…
5
-
Review of MIT NANDA: The Internet of AI AgentsApr 13, 2025
Review of MIT NANDA: The Internet of AI Agents
Introduction The rapid advancement of Artificial Intelligence (AI), particularly Large Language Models (LLMs) and…
45
6 Comments -
Deep Dive on Model Context Protocol (MCP) for Enterprise Messaging IntegrationApr 3, 2025
Deep Dive on Model Context Protocol (MCP) for Enterprise Messaging Integration
1. Executive Summary: Understanding the Model Context Protocol (MCP) The Model Context Protocol (MCP) has emerged as a…
21
4 Comments -
Ace Tutor App Beta Test InvitationJun 7, 2018
Ace Tutor App Beta Test Invitation
If you currently live in San Francisco Bay Area and have a need for finding local private tutors for your children, I…
11
Activity
35K followers
-
Ben (Xiaojun) Li shared thisOpenAI recently reported daily inference usage in its research organization exceeding $600 at the median and $7,000 at the 90th percentile, valued at API prices. For most people working within tighter budgets, choosing the right model for the right task can be a struggle. For real work, what matters is not the cost per LLM token, but the overall cost of successfully completing a task: including retries, verification, and human correction time. It’s a bit counterintuitive, but more advanced models can be cheaper overall for complex agentic tasks. The Terminal-Bench 4.0 leaderboard (https://www.tbench.ai/) shows that Astra (GPT-6, running in #Codex at max effort) has the highest reported resolution rate, with substantially lower token usage and total cost than Fable 5.1 and Opus 5. This closely matches my recent experience working with Astra, Opus 5, and GPT-5.6 Sol. Terminal-Bench measures how effectively AI agents perform complex, multi-step tasks through a real command-line interface (CLI). The benchmark covers software engineering, machine learning, systems operations, security, media, and scientific research. If your tasks resemble those in Terminal-Bench, my recommendation is to start with Astra and/or Fable 5.1 (or Opus 5 if your organization doesn’t allow Fable). For basic automation tasks, lower-cost models can make sense, provided they reliably meet your quality requirements.
-
Ben (Xiaojun) Li shared this▎ Ten years ago, in Apple's chip department, we were designing A-series chips two generations ahead of what customers had in their newest iPhones. We could predict the future iPhone, but we could not simply make it arrive any faster. ▎ Fast forward to now. Researchers at OpenAI and Anthropic are building models ahead of anything the rest of us can access, and they are using those models to build the next ones. That lead does not stay fixed. It compounds. ▎ OpenAI published their competitive advantages yesterday. By mid-August, the median researcher in its research organization was spending more than $600 a day on inference at API prices, and the 90th-percentile user more than $7,000 a day. The research org now spends 3.1 agent-workdays of effort for every workday of human labor. Most companies cannot fund that, let alone match it, so the frontier gap is more likely to widen than to close in near future. ▎ OpenAI's chief scientist Jakub Pachocki published an essay yesterday on what he sees in the internal models and what it costs him sleep: #RSI (recursive self-improvement) and automated alignment research. He says "This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence." ▎ Over the next several years, OpenAI will prioritize work in service of three north stars: ▎ Navigating the next period of AI progress, by building an automated AI researcher, iterating with it on the alignment problem and finding ways for people to remain part of the self-improvement loop. ▎ Delivering the benefits of scientific progress and economic growth that very intelligent machines enable. ▎ Empowering everyone individually with a personal AGI. ▎ Read one way, these are about alignment, enablement and empowerment. Read another way, they are about competition, profit and distribution. ▎ Jakub hopes voluntary slowdowns become common until shared safety bars exist. I think achieving that globally will be difficult. With OpenAI targeting an automated AI researcher by March 2028, watching from the sidelines is not an option. Grokking the frontier models and the 3C harness tools (#Claude, #Codex, #Copilot) is how we build the judgment to direct work we increasingly won’t do ourselves. "finding ways for people to remain part of the self-improvement loop" is not something a lab does for us. It is a seat we choose take. https://lnkd.in/gUEzAsSf
-
Ben (Xiaojun) Li shared thisOne year ago, I claimed the single most important benchmark for AGI is a simple proof of the Fermat’s Last Theorem (FLT). For me, true AGI should be able to prove FLT in one page, using elementary number theory, or inventing a new mathematical idea to do it elegantly. At that time, Anthropic was struggling and Claude Code wasn't a big thing, and I excluded them in the AGI race in the AI Slop image in my post. Today, Anthropic made a major advancement in that AGI benchmark: Claude AI Formalizes Fermat’s Last Theorem in Just 11 Days. Anthropic’s #Claude completed a full formalization of FLT in Lean 4 in only 11 days. The project included more than 13 million lines of code and verified about 29,500 intermediate results across areas such as algebra and geometry. The formal proof follows the same broad path used by Andrew Wiles at Princeton University, relying on Frey curves and modularity. Each step was checked against the underlying mathematical axioms, with dozens of AI agents working together on different parts of the proof. The project could mark an important step for automated theorem proving. Large-scale AI systems may help mathematicians convert existing proofs into machine-checkable form, verify complex arguments, and expand formal mathematics libraries much faster. Although this is not reaching the AGI level I wished, it is still a significant step toward the final proof. I am happy to see scientists in Anthropic share the same passion on FLT and hope in another year, there will be the final breakthrough.
-
Ben (Xiaojun) Li shared thisFree half-day workshop in San Francisco by Microsoft and Anthropic experts on Claude and Agent in Microsoft Foundry. This workshop is designed to accelerate your ability to ideate, optimize, and deploy AI solutions in real-world scenarios. Agenda-at-a-glance: - Welcome and expectations for the day - Hands-on labs: Building an Agent & Making Your Agent Work over a Long Horizon - Preconfigured Azure environments - Azure-direct model access and configuration - Best practices for prompting, safety, optimization, and fine-tuning - Scenario-based exercises tailored to real use cases This event is a unique opportunity to gain direct experience with Claude and accelerate your path to production. Attendance is limited—claim your spot now. https://luma.com/p6m4n0g7
-
Ben (Xiaojun) Li shared thisThree years ago, one of the highest achieving engineer on my team completed about 1000 pull requests within a year; last month, Lauren Tan (@poteto on X) at Cursor/SpaceXAI did 1000 PRs within a month using coding agents. Coding agents are so powerful, but why most companies do not see 10x productivity increase? Boris Cherny at Anthropic mapped it out. He lays out five steps in AI adoption based on the number of agents you can orchestrate and manage: Step 0. Gated. Zero agents. Security review and procurement still own the decision, and whatever you build lives on your laptop. Step 1. Assisted. One agent. You pair with it, read every diff, and sit and watch while it works. Most engineers live at this stage now. Step 2. Parallel. Around ten agents, each in its own worktree. You stop writing code and start reviewing six streams of it. Step 3. Supervised autonomy. Around a hundred. Claude kicks off Claude. The maintenance backlog that used to wait for a free afternoon now runs in the background. Step 4. AI-native. A thousand or more. You steer by intent and check exceptions. A quarter-long migration becomes something you start Monday and check on Friday. Anthropic is on step 3. Boris is entering step 4. Getting from 1 to 2, and from 2 to 3, is on you. It will cost billions of tokens, and you will have to live inside Claude Code, Codex, or Copilot CLI every day. You need to build your own CLI environment with lots of MCPs, customized plugins and skills. You build your trust in your agents and do not babysit their work. Step 4 is a different stage. You need deep domain expertise, large job scope, unlimited token budget, and AI native infra at your company. One motivated engineer can't get there alone. I am in the mid of step 3, largely pulled back sometimes due to orchestrating agents through 10 repos for given projects. I am curious which step are you on, and what's the one thing blocking the next? AI Adoption roadmap - https://lnkd.in/gRB7qA_M.
-
Ben (Xiaojun) Li shared thisWith the explosion of open source LLM models, inference engineering becomes the fastest growing domain in AI industry. Many software engineers are learning AI engineering now, but the mistake to avoid is not stop at the LLM APIs. Understanding inference will help develop much better applications and user experiences. Philip Kiely wrote a book about it and made it free with both online interactive course and PDF/ePub versions. It is a good resource for SDEs who are new in AI domain and interested in learning inference engineering.
-
Ben (Xiaojun) Li shared thisAnthropic's ARR jumped 14x in the most recent quarter, and 70% of it came from API charges, and it's on a path to $500 billion ARR by 2027. This is unprecedented growth but not a total surprise. The software market has gone through three revenue models: Software-on-CDROM charged for development cost, Software-in-Cloud charged for hosting cost, and Software-by-Intelligence charges for compute cost. From now on, no token is free and every task costs a fortune. Welcome to the new era! Recently I noticed that my fresh sessions in Claude Code CLI and GitHub Copilot CLI, both set to a 1 million token context window, started compacting after just a couple of turns of prompt and response. That leads to huge API cost. Sometimes a single "hi" as the first prompt in a fresh session cost me about $2. So I looked into what my sessions were actually carrying, and found they constantly held close to 500K of MCP tool data in the context window (I have about 80 MCP servers). These tools should be lazy-loaded on demand, so I spent some time investigating. It turns out there's a CLI setting bug that stops MCP lazy-loading. Coincidentally, Anthropic published a blog post, "Maximizing the value of your Claude Code sessions," with a lot of great suggestions on how to manage and optimize the context window to reduce token waste. The most important ones are (1) minimize startup context, (2) set the right model, (3) keep the KV cache warm, (4) @-mention files, and (5) compact before you pause. The attached PDF has more detail on how to apply these in your own tasks. I still find it very odd that LLM vendors charge for input tokens, and everyone seems to just follow along. Input tokens are the user's effort to tell and steer the LLM on what to do and how to do it. It would be very funny if you spent time collecting evidence and affidavits for your suit, went to talk to your lawyer, and your lawyer charged you for the time you spent gathering it before you walked into his office. Regardless, this is the Software-by-Intelligence era, we're charged for compute, and it's a seller's market, so as customers we don't have much negotiating power. What we can do is get really good at understanding context and agent runtime, so we can minimize token waste.
-
Ben (Xiaojun) Li shared thisJeff Dean left Google to found DiscoveryLoop to pursue AI for Science. According to Pathfounders, Demis Hassabis originally planned to leave together with Jeff, but Google persuaded him to stay. AI for Science becomes possible due to the fast advancement of Recursive self-improvement (RSI) in Agentic LLM. Both Andrej Karpathy (joining Anthropic) and Thinking Machines Lab cofounder Lilian Weng (joining OpenAI) are working on RSI. It looks like RSI will become the next frontier of intelligence. Stanford University has a new class "CS329A Self-Improving AI Agents" taught by Prof. Azalia Mirhoseini and Aakanksha Chowdhery, Ph.D. The course covers the latest self-improvement techniques of LLM agents that can continuously improve themselves through interaction with themselves and the environment. It's a good starting journey for people who are interested in RSI.
-
Ben (Xiaojun) Li shared thisDatabricks published a big post on how they cut their AI token spend by close to 90% — model switching, context reduction, smart routing, etc.. Impressive engineering. But I think it's premature. Here's the hard truth: outside the frontier labs (OpenAI and Anthropic, etc.), maybe only 1% of employees at any company are generating real ROI from the tokens they burn. The rest are spending tokens on simple tasks, not real intelligence. If you're reading this post, you're probably in that 1%. There's a piece going around on X right now called "AI Adoption is a Myth" that puts numbers on it. In a typical enterprise rollout: 5-10% become power users, 20% use it badly, 70% never touch it. Roughly 10% of people burn 90% of the tokens. So we should stop booking employee token use as an expense. It's an investment — part of the training budget for employee re-skilling. Agentic tools are incredibly capable but also wide open, and most people need a long time before they really grasp them and see any return. Which brings me back to Databricks. If a 90% cost cut lands harder on their 1% of real adopters than on the 90% who barely use it, they're not saving money. They're throttling the only cohort of the company that's actually innovating. Honest question: how long do you think it takes before AI actually makes most employees faster?
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisCareer update: after about a year and a half at OpenAI, I’m leaving to build something of my own. I joined the Codex team about a year and a half ago, 2 weeks before we launched the very first incarnation of Codex. Since then, I’ve had the chance to work across a ridiculous range of problems: reliability, product, agents, code review, infrastructure, capacity, safety, and a few things I had absolutely no idea how to do when I started. A year and a half at OpenAI somehow felt like five years, or maybe more, anywhere else. There was just so much to learn, absorb, build, get wrong, and figure out. To capture what the experience actually felt like, I wrote a piece about it, both as a reflection for myself and for founders and builders trying to make sense of this strange moment we’re in, as well as anyone who’s ever wondered what working on Codex at OpenAI actually feels like. Link - https://lnkd.in/gU44Pf2H As for what’s next, I’m joining Aishwarya Naresh Reganti at LevelUp Labs, where we’ll work closely with companies to figure out where AI can meaningfully change how they operate, build the systems that make that possible, and help teams learn to work with them.
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisLast week at OpenAI's annual DevDay, we launched AgentKit to help you build, deploy, and optimize agents. It's been a privilege to work with the dream team to bring AgentKit to life, and I can't wait to see all the agents people build! Watch the full talk here: https://lnkd.in/etxWPKHp
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisChatGPT is now live in Microsoft Teams 🎉 - https://lnkd.in/gaGnuNE6 Start with @ChatGPT in a channel, a group chat, or a DM, and it picks up relevant context right there. Answer the question in the thread. Unblock the decision. Turn a messy discussion into a brief, a plan, or a deck your team can use, without leaving Teams. This is what we built the Teams platform for. Conversation context, a user's identity and permissions, and admin-managed access to the tools a company already runs. That is what lets ChatGPT work natively, and seamlessly in Teams. Proud of where this landed and grateful to the people across Teams and OpenAI who did the hard work to make it real.. Excited for the partnership, with lots more to come, including the newly announced Dots integration! #MicrosoftTeams #OpenAI
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisToday, The Atlantic published the first excerpt from my upcoming book, THE AGI CHRONICLES. It's about the long-running feud between Sam Altman and Dario Amodei, who used to work together at OpenAI before Amodei and six of his colleagues left in 2020 to start a rival AI lab, Anthropic. The full story of the OpenAI/Anthropic split has never been told before, and it's pretty juicy. (Kanye West is involved, as is a stuffed panda named "Beary Bonds," and a secret plan to sell AGI to China and Russia.) Getting this story took months of digging and dozens of interviews, but I think it's important to understand why the leaders of the AI race are so motivated to beat each other. The book comes out next week! I've never worked harder on anything, and I'm very proud of how it turned out. https://lnkd.in/g6fR5nERInside the Biggest Feud in Artificial IntelligenceInside the Biggest Feud in Artificial Intelligence
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisSonnet 5.5 fixing a bug with Claude Code. 30% faster and 30% less usage.
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisSo proud of my daughter, Aria, for being named a 2026 Caroline D. Bradley Scholar! She’s one of 26 students nationwide selected to receive a merit-based, need-blind scholarship providing full high school tuition for four years through the Institute for Educational Advancement. More importantly for us, IEA works individually with Scholars to identify an optimally matched school or create an individualized educational plan. That’s particularly meaningful because Aria has never fit neatly into one lane - a quality I frequently see in the founders, entrepreneurs, students, and changemakers I have the privilege of meeting and advising. It’s been a gratifying, but at times challenging, journey to parent and support her. She loves mathematics deeply, studying Calculus BC as an eighth grader at Proof School and pursuing mathematical research in combinatorics with renowned mathematicians like Ken Ono. She also competes in freeride skiing, climbs and skis mountains, has performed at Carnegie Hall and danced with the San Francisco Ballet. Aria seems happiest when she’s discovering connections between worlds that adults usually put into separate boxes. For highly curious and passionate kids - and, I think, for unconventional founders and leaders - the challenge isn’t always finding more acceleration or more achievement. Sometimes it’s constructing a life in which intellectual depth, physical challenge, creativity, friendship, impact, and joy don’t have to compete with one another. One of Aria’s mentors, Paul Zeitz, once wrote that while her interests may seem orthogonal, they actually reinforce one another in unique ways. I see something similar in some of the most fulfilled founders and executives I meet: an ability to weave seemingly disparate obsessions and talents into a life tapestry that makes sense to them - even when it looks a little chaotic from the outside. Aria has been fortunate to encounter extraordinary people who have helped make that possible: teachers who took her questions seriously, mathematicians like Paul and Ken who treated her as a fellow explorer, coaches like Joe Stevens and Reine Barkered who pushed her while keeping skiing fun, artists like Sasha De Sola and Misa Kuranaga who showed her another language for expression, brands like The North Face and Atomic who have supported her, and communities that gave her places to belong. I’m enormously grateful to all of them, and to IEA, Deborah Monroe, Mallory Aldrich, and Daniel Martinez for now joining that village. Congratulations, Aria. Keep exploring all the things that make you, you. ❤️ IEA’s announcement and the full Caroline D. Bradley Class of 2031: https://lnkd.in/g5ftfHmr
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisReady to take AI from toy demos to enterprise infrastructure? We’re hosting two deep-dive technical workshops live in Sunnyvale in two weeks—built for anyone pushing the boundaries of compute, scale, and robotics. Choose your track (or conquer both): ⚡ Track 1: Reinforcement Learning & Custom Agents on GKE 🗓️ Tuesday, Oct 6 | 9:00 AM – 1:00 PM Cut API costs and scale production agent swarms. We're breaking down self-hosted Gemma on vLLM, hardware fungibility across GPUs/TPUs using GKE + Dynamic Workload Scheduler, and live, hands-on Ray RL loops for policy updates and DPO. 👉 RSVP: https://lnkd.in/gtwbqqV6 🤖 Track 2: Navigating the Cloud Physical AI Stack 🗓️ Thursday, Oct 8 | 9:00 AM – 1:00 PM Bridge the robotics data gap. Master NVIDIA Isaac Sim, Cosmos generative synthetic data, and Trillium TPU architectures powering sub-10ms multimodal reasoning. Features live embodied AI runs on Unitree robots and hands-on policy training. 👉 RSVP: https://lnkd.in/gr3EEU4d 📍 Where: Google Cloud Campus | Sunnyvale, CA 🛠️ Format: Architecture deep dives, live labs, and direct access to Google Cloud engineers (plus lunch on us). Grab your spot now!https://lnkd.in/gtwbqqV6 🤖 Track 2: Navigating the Cloud Physical AI Stack 🗓️ Thursday, Oct 8 | 9:00 AM – 1:00 PM Bridge the robotics data gap. Master NVIDIA Isaac Sim, Cosmos generative synthetic data, and Trillium TPU architectures powering sub-10ms multimodal reasoning. Features live embodied AI runs on Unitree robots and hands-on policy training. 👉 RSVP: https://lnkd.in/gr3EEU4d 📍 Where: Google Cloud Campus | Sunnyvale, CA 🛠️ Format: Architecture deep dives, live labs, and direct access to Google Cloud engineers (plus lunch on us). Grab your spot now!
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisLast week, we joined Amherst College in celebrating the grand opening of the new Ford Student Center—a 144,000-square-foot hub for campus life built through a highly coordinated blend of adaptive reuse and mass timber construction. Rather than starting from scratch, the team preserved and repurposed the existing concrete structure of the former Merrill Science Center, then built two new floors above in mass timber. The result brings dining, student life, cultural, recreation, and gathering spaces together under one roof—from Betty’s Kitchen and the Alumni Pub to cultural resource centers, performance spaces, lounges, terraces, and more. Congratulations to our partners at Herzog & de Meuron and Sasaki—and we can’t wait to share more from the project soon. 🔗 In the meantime, read more here: https://lnkd.in/e3b78EwJ #ShawmutBuilt
-
Ben (Xiaojun) Li liked thisBen (Xiaojun) Li liked thisWe’ve expanded OpenAI Academy with new courses for developers, leaders, educators, and college students 🎓 Learners practice applying AI to real tasks, whether they’re improving a business workflow, building with Codex and the OpenAI API, developing an AI strategy, or planning a class. Complete a course and pass the assessment to earn an OpenAI Academy course badge. Check out our blog post to explore the new learning paths and how to bring them to your organization: https://lnkd.in/ggKTN7jj
Experience & Education
-
Microsoft
********* ******** *********** *******
-
******** ****
********** **** ****** ***********
-
*****
*** ********* *** ***** ********
-
********** ** ******** ******* ****
****** ** ********** * *** *********** undefined
-
********** ** ********
****** ** ******* * ** ***********
View Ben (Xiaojun)’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Licenses & Certifications
Patents
-
System and Method for Managing Online Group Chat
Issued US 10,616,289
See patentPatent Approved
Big Group Collaboration US Patent:
System and Method for Managing Online Group Chat
Patent No. 10,616,289 Issue Date: April 7, 2020
Link: https://patentswarm.com/patents/US20170237785A1 -
AUTOMATED NOTIFICATION OF CONTENT UPDATE PROVIDING LIVE REPRESENTATION OF CONTENT INLINE THROUGH HOST SERVICE ENDPOINT(S)
US Pending
Languages
-
English
Native or bilingual proficiency
-
Chinese
Native or bilingual proficiency
Recommendations received
View Ben (Xiaojun)’s full profile
-
See who you know in common
-
Get introduced
-
Contact Ben (Xiaojun) directly
Other similar profiles
Explore more posts
-
Andrew Lombard
TESORO VC • 7K followers
Big news for deep-tech builders 🚀 TESORO VC × Parallel Works: Unlocking Scalable HPC & GPU Power for Founders We’re excited to share that Tesoro VC has entered a strategic ecosystem partnership with Parallel Works, bringing scalable HPC and GPU compute directly into the hands of our AI + Semiconductor Accelerator startups. This is a major unlock for the kinds of companies we support. ⚡ Why this matters for Tesoro founders Many of our startups aren’t building lightweight software — they’re: • Designing next-gen semiconductor architectures • Running advanced chiplet and packaging simulations • Training compute-intensive AI models • Performing large-scale physics, materials, and systems simulations These workloads demand serious compute — and managing that infrastructure can easily distract from core product development. Through Parallel Works’ ACTIVATE platform, our startups now get: ✅ Access to HPC + GPU resources across hybrid cloud and on-prem ✅ A ready-to-use compute framework with best practices built in ✅ Cost controls and budgeting tools (huge for early-stage teams) ✅ Repeatable workflows that scale from prototype → production Net effect: founders spend more time innovating and less time wrestling with infrastructure. 🧩 Why this strengthens the Tesoro ecosystem This partnership doesn’t just help individual startups — it reinforces the entire Tesoro platform. For industry partners Startups can design, simulate, and validate technologies in environments that more closely resemble real enterprise and manufacturing conditions. For semiconductor & manufacturing stakeholders Founders can run higher-fidelity simulations earlier, reducing technical risk before fabrication and scale-up. For AI and national security–aligned innovation Secure, scalable, portable compute workflows are essential in regulated and mission-critical environments. For investors This improves technical de-risking. Companies can validate performance and scalability earlier — leading to stronger milestones and clearer paths to commercialization. At Tesoro, we’re building more than an accelerator — we’re building an execution layer for the AI + Semiconductor economy. Pairing founder support, industry access, and capital pathways with advanced compute infrastructure creates a powerful bridge from: Idea → Simulation → Validation → Scale Proud of this step forward and excited to see what our founders build with this new level of computational firepower. 📰 Read the full announcement here: https://lnkd.in/giXf9_vZ
46
2 Comments -
Abhishek Kshirsagar
PROMAFT Partners • 12K followers
𝐒𝐞𝐦𝐢𝐜𝐨𝐧 101 𝐄𝐏1: 𝐁𝐞𝐟𝐨𝐫𝐞 NVIDIA, 𝐓𝐡𝐞𝐫𝐞 𝐰𝐚𝐬 𝐚 𝐒𝐰𝐢𝐭𝐜𝐡! Everyone is talking about AI chips, GPUs, fabs, and trillion-dollar valuations. The conversation usually starts in the present, as if intelligence suddenly appeared because we finally built fast enough machines. That framing sounds exciting, but it skips the most important part of the story. This was never really about speed. It was about control. Early electricity was like a wild river. Powerful, useful, and dangerous. We learned how to generate it and move it across long distances, but we had no good way to control it precisely. We could switch it on and off at large scales, but we could not guide it gently, reliably, or repeatedly. Without that level of control, complexity simply could not scale. That changed with semiconductors. Think of materials as personalities. Metals are oversharers. Electricity flows through them freely, whether you want it to or not. Insulators are introverts. Nothing gets through, no matter how politely you ask. Semiconductors sit in between. They listen. They wait. And then they decide. A semiconductor matters not because it carries electricity, but because it can be persuaded. With very small changes, the same material can behave like a wire, a wall, or a switch. Once electricity can be told when to flow, when to stop, and when to wait for permission, logic becomes possible. And once logic exists, everything we now call computing follows. What looks like intelligence today is really obedience at scale. Billions of tiny yes or no decisions, repeated perfectly, at extreme speed. The magic is not that machines are smart. It is that electrons are disciplined. I just published a long-form piece breaking this down from first principles. What a semiconductor actually is. Why diodes and transistors mattered. How dumb switches quietly create logic. Why silicon won. And why manufacturing discipline matters more than most people think. Check out my Substack for the complete post! #Semiconductors #AI #Chips #Technology #DeepTech #Hardware #FirstPrinciples #Computing #Manufacturing #Innovation NVIDIA Switch Semicon India PROMAFT Partners https://lnkd.in/d7SXQBYw
21
4 Comments -
Ankur Bhuva
Leadership Accelerator Co. • 9K followers
Some of Nvidia's best engineers have stayed with Jensen Huang for 30+ years. That's an unusual kind of retention in tech. Nvidia went through near-failure in its early years. The product struggled, the market was uncertain, and the company had every reason to become more controlling. Huang chose a different approach: give talented people ownership instead of managing them through fear. That philosophy scaled with the company. → Engineers were trusted to solve problems independently. → Senior talent stayed long enough to build deep institutional knowledge. → The company grew without replacing its engineering culture with layers of control. 3 leadership lessons: 1️⃣ Give your best people room to operate. High performers rarely need someone watching every decision. They need the authority to make them. 2️⃣ Retention is often a byproduct of trust. People don't stay for decades because of one great incentive. They stay when they consistently feel trusted, challenged and valuable. 3️⃣ Scale shouldn't mean more control. A common leadership mistake is becoming more bureaucratic as the company grows. Huang's case shows the opposite: if autonomy helped attract great talent early, taking it away later can be exactly what makes that talent leave. The real competitive advantage isn't just hiring exceptional people. It's building a company where exceptional people still want to be there decades later.
339
30 Comments -
Kit Yu
33K followers
NVIDIA’s 2027 fiscal Q2 earnings call spotlighted the rollout of its nextNVIDIA’s 2027 fiscal Q2 earnings call spotlighted the rollout of its next‑generation Vera CPU, which demonstrated 1.8x faster performance on SPEC genetic benchmarks and delivered five times the bandwidth per watt compared to any other data center CPU. Management emphasized that Vera will be widely deployed across hyperscale cloud AI labs and system OEMs, further expanding NVIDIA’s total addressable market. Alongside this, the company reinforced its full-stack AI factory strategy, integrating GPUs, CPUs, NVLink, InfiniBand, Spectrum networking, and the KuTa ecosystem to deliver fungibility and durability across all AI workloads. This combination of performance leadership and platform economics positions NVIDIA to capture a larger share of the data center market, with sequential revenue growth of 18% and accelerating adoption of AI infrastructure globally.
1
-
Axel Egger
1K followers
Summary of Jensen Huang’s Keynote at GTC 2026, March 16, 2026 Jensen Huang’s keynote addressed various topics, ranging from general AI trends, AI performance and energy optimizations, Nvidia-specific innovations, to new business strategies. I will examine a couple of these topics in more detail in separate posts. To start with, here is a general overview: 1. The “Vera Rubin” Platform: The New Powerhouse Nvidia has officially introduced the Vera Rubin architecture as the successor to Blackwell. This is not just a chip, but an entire system designed for extreme efficiency: • Six new chips: The platform includes, among others, the Rubin GPU (with HBM4 memory) and the custom-built Vera CPU. • Performance leap: Compared to Blackwell, inference token costs are expected to drop by a factor of 10. • Confidential Computing: This is a specific hardware technology built into the new chips. It ensures that data is isolated in an encrypted environment. 2. Focus on Inference and “Agentic AI” A strategic turning point is the clear focus on inference, does mean the application of AI: • OpenClaw & NemoClaw: OpenClaw was introduced as an “operating system for agents” and is positioned as the open standard (protocol) for interoperability and identity. NemoClaw is presented as the highly optimized Nvidia engine for maximum performance and hardware security. • Enterprise security: With the “Nvidia Agent Toolkit” and the OpenShell runtime, enterprises should be able to run these agents securely and with full data control in their own data centers. 3. “Physical AI” and Robotics Jensen Huang emphasized that the next big wave will be physical AI: • AI factories for robots: He presented blueprints for “AI Data Factories,” where robots, vision agents, and autonomous vehicles are trained using real and generated data. Simulation before physical reality. • Humanoid robots: Bots are not only meant to perform simple movements but to understand and execute complex tasks in industrial environments. 4. Technological Milestones: DLSS 5 and Photonics • DLSS 5: For the graphics sector, DLSS 5 has been announced, which uses “Neural Rendering” to significantly enhance details such as faces and textures through AI, rather than simply upscaling them. • Optical Connections: To handle the massive amounts of data between chips, Nvidia is increasingly relying on light (photonics/CPO) instead of electricity, which is expected to drastically reduce energy consumption in data centers. Watch the full-length video here (2:20h): https://lnkd.in/e3cYXsN5 Conclusion: Jensen Huang’s keynote is a milestone. In our context, it also demonstrates that autonomous agents are becoming increasingly important in the areas of cybersecurity and GRC. As agents begin to operate independently within enterprise systems, a “Data Centric Zero Trust” architecture, as we advocate, is indispensable. More on this and the other elements of his keynote in upcoming posts.
8
-
Daniel L.
BIP town • 28K followers
😎♐️👂 As Fridman points out, Huang has previously said the timeline for AGI depends on what defines it. At the 2023 New York Times DealBook Summit, Huang defined AGI as software capable of passing tests that approximate normal human intelligence at a reasonably competitive level. He expected AI to clear that bar within five years. For his part, Fridman offered Huang a generous definition to work with: true AGI, in Fridman's framing, would look like an AI capable of starting, growing, and running a technology company worth more than a billion dollars. He asked whether that was achievable in the next five to 20 years, given the recent proliferation of agentic AI tools like OpenClaw. https://lnkd.in/ea3waU88
2
1 Comment -
Rubén Domínguez Ibar
The VC Corner • 338K followers
Ownership with NO off switch 🔌 At NVIDIA, responsibility isn’t episodic. For Jensen Huang, problems don’t get delegated and forgotten. They live rent-free in his head until reality resolves them. That choice creates a very specific leadership model: ▫️ Decisions stay exposed to consequences ▫️ Weak signals are carried forward, not ignored ▫️ Thinking doesn’t end when the meeting ends ▫️ Feedback loops stay tight, even at massive scale The cost is personal bandwidth. The payoff is clarity most leaders never reach. This is what long-term ownership actually looks like. Uncomfortable. Demanding. Compounding. I broke down why Jensen Huang chose to lead this way, and the price he paid to build NVIDIA to last. Full breakdown here 👇 https://lnkd.in/ePbQi9ig
18
3 Comments -
Banyan Ventures
2K followers
Mitesh Agrawal took the stage today with Jay Jackson at AI Infra Summit in Santa Clara to discuss Positron AI and "The Making of an AI-Era Semiconductor Company" Key takeaways if you missed it: 1. Heterogeneous compute for AI is here. 2. Speed to market and time to the next chip are key to winning. 3. Differentiated technology is not enough. You need a path to deployment at scale, supported by both your supply chain and the data center footprint/infrastructure. Darren Chien Thomas Sohmers Sheryl Savage
31
1 Comment -
MAHESH YADAV
allNeurons • 19K followers
The AI Chip Power Struggle: Who Controls the Future of Compute? Just gave the following onboarding readiness notes to someone getting into director level role in ASIC AI space. The AI chip market looks chaotic, but in reality, it’s highly concentrated. A few hyperscalers, merchant vendors, and specialized startups control is all we have here.. 1. In 2026, AI chips are no longer just chip with flops and interconnect speed; they are infrastructure platforms. Success depends on compute, networking, cooling, and debugging tool maturaity. Nvidia exemplifies this with systems like the DGX SuperPOD. 2. Despite widespread AI adoption, global compute demand is dominated by a handful of hyperscalers: Google, Amazon, Microsoft, Meta Platforms, OCP and ByteDance. These companies operate the largest AI clusters and are driving 80+% of demand… that is it... if they stop then party stops. 3. GPUs remain dominant not because of raw performance, but because of software ecosystem lock-in. For example, CUDA powers training, deployment, and distributed workloads, creating high switching costs in past - and even with transformer (where its little easy to catch up on new architecture/operators) the close integration of chip to workload where you have large context can be better processed with Nvidia inference context memory storage vs expensive HBM in TPUv7. 4. The strategic divide is clear: hyperscalers build custom silicon to optimize internal workloads (Microsoft Maia , Amazon Trainium), while merchant vendors like Nvidia and AMD sell broadly compatible platforms. And both will have business and growth although some(TPU) will try to enter merchant business - but- it will be very hard for them.... 5. Custom accelerators only pay off at hyperscale utilization, making them suitable mainly for hyperscalers or large AI labs such as Anthropic or AWS. This means no one can just build out on one chip or one large customer demand and this will be a heterogeneous market ( good for neo cloud). 6. Modern AI clusters that are under pressure to scale fast from workloads of agents (clawbots/ claude code) face bottlenecks beyond chips: HBM memory, networking, cooling and datacenter power. Suppliers like TSMC, SK Hynix, and Samsung Electronics are critical in short run. 7. Startups like Graphcore, Cerebras Systems, and Groq can innovate rapidly, but ecosystem lock-in and software integration remain GTM blockers. The clearest opportunity lies in inference, where power, latency, and throughput dominate but that also has not played out for them and will not play well in near future as well… In the end, AI chip market will test the market patience, AI appetite and adaptability- this space will remain in high attention - in a world where attention is all you need. Hope it helps..
29
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content