Sign in to view Peng’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Bellevue, Washington, United States
Sign in to view Peng’s full profile
Peng can introduce you to 10+ people at ByteDance
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
3K followers
500+ connections
Sign in to view Peng’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Peng
Peng can introduce you to 10+ people at ByteDance
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Peng
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Peng’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
About
- Expert in…
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
Articles by Peng
-
Make optimized decision with Deep Learning
Make optimized decision with Deep Learning
As humans, we have been practicing to make good decisions since we were young. As we passed the skill to machine, more…
11
-
Setup Spark Cluster in Hyper-V with standby masterJun 26, 2017
Setup Spark Cluster in Hyper-V with standby master
Background This is the high level design of spark cluster. Spark currently supports three type of cluster managers:…
4
-
Setup Zookeeper Cluster in Hyper-VJun 26, 2017
Setup Zookeeper Cluster in Hyper-V
Background The Zookeeper cluster architecture is like this below, and detail can be found here https://zookeeper.apache.
2
-
Setup Kafka Cluster in hyper-VJun 26, 2017
Setup Kafka Cluster in hyper-V
Background Kafka Cluster is different than zookeeper or spark. For each topic, you can have multiple partitions.
2
Activity
3K followers
-
Peng Zhang posted this字节跳动 AML Engine 编排调度组最近在 Seattle 或 SJC 招聘如下岗位。 主要从事 TikTok 搜索,广告,推荐模型 的 GPU infra 编排调度 (scheduling and orchestration),管理百万卡量级GPU集群,训练和推理都招。 希望有 k8s operators/controllers, 资源池管理,新GPU硬件上线,动态扩缩容,A/B test 流量切换,QPS压力测试,训练workflow,模型封装和deployment, 跨region容灾 等等。 要求有 ML platform,GPU scheduling, K8s 集群管理,Compute resource orchestration, multi-host GPU training/inferernce 等相关工作/实习经验。 校招/实习也看中有 open source开发经验的同学, 比如 Ray,Airflow,Kubeflow, Flyte, Kueue,Temporal, Skypilot, SGLang, etc。 Prefer 可以中英双语交流 (常与中国区会议), 对OCI/GCP 有经验的同学。 - Sr/Staff SDE 或者 工作经验特别match的SDE2,请点下面link申请。也欢迎直接联系我了解情况。 ---- 西雅图: https://lnkd.in/gSNbQpQj ---- SJC:https://lnkd.in/gc-HrbFj - SWE New Grad 直接点以下链接申请 ---- 西雅图 BS/MS:https://lnkd.in/g-vWps7J ---- SJC BS/MS:https://lnkd.in/gEgDW_A7 ---- 西雅图PhD:https://lnkd.in/gjfUqP-r ---- SJC PhD:https://lnkd.in/gTftMgAf - SWE intern 直接点以下链接申请 ---- BS/MS intern in SEA/SJC: https://lnkd.in/g3WfdpzF ---- PHD intern in SEA/SJC: https://lnkd.in/gKpHCRFg ---- 非常match但没有CPT的同学,可帮���联系国内实习机会,之后转美国全职岗位
-
Peng Zhang reposted thisPeng Zhang reposted thisMost infrastructure providers buy GPUs and rent them out. We build every layer – the power, the data centers, the compute. Today, Nscale has agreed to acquire Anyscale, the platform behind Ray, built for scaling AI workloads across thousands of GPUs. We built the infrastructure AI runs on. Anyscale built the software developers actually live in. Together, we go from capacity supplier to the full-stack home for AI development – power to production. Anyscale will keep operating under its own brand, serving its customers as it does today. And we're joining the PyTorch Foundation, standing behind Ray's open governance and the community that built it. What excites me most: Anyscalers joining us. That's the caliber of talent shaping what comes next. I get to work with more talented people!! Most companies optimize one layer. We're building all of them, on purpose. More here: https://lnkd.in/gUbfnpa3
-
Peng Zhang shared thisMichelangelo’s submission to PyTorch Conf has been accepted. We’re excited to have the team present the impactful work we’ve accomplished on top of PyTorch, Ray and DeepSpeed open sources. Topic: “PyTorch-Native Feature Transformation and Training Framework for Uber Eats Recommendation” PS: Will co-present this work as ex-Uber.
-
Peng Zhang shared thisWe are hiring Software Engineers in multiple levels in Uber AI platform ( Michelangelo) team in Seattle, WA or Sunnyvale, CA. Please apply directly if interested. https://lnkd.in/g2MsBfSM https://lnkd.in/gVJmd7QV https://lnkd.in/gQGpKw2J https://lnkd.in/geZ-NmXw
-
Peng Zhang shared thisGreat to see Hotels in Uber besides foods and drinks. I really like the direction! Hope some day in future we can use Ube to book tickets like parks, events, museums, games, concerts, etc. That will make it a one-stop shop for a good travel.Peng Zhang shared thisHotels on Uber has arrived! Today, we announced the launch of Hotels on Uber at our premiere consumer event GO–GET. Our goal is to be ever more helpful in your travel journey. Most of us already rely on Uber to get to the airport, get to the hotel, go around the city, order food and groceries. With Hotels, now you can complete that journey all in one app. So what’s in it for you? We are bringing the same experience - easy, clear, uncomplicated, that you are accustomed to while booking a ride - to booking a hotel. And there is huge value that I think you’ll love. 20% off on most hotels and Uber One members get an additional 10% back in Uber One Credits. On your next vacation, try Uber Hotels, it will cover your rides to the airport and those meals that you order on your trip. We’re rolling this out in the US over the next few days – let me know what you think. I am very proud of the team behind this, who worked extra hard to bring this value to all of you!
-
Peng Zhang shared thisExcited to share that I recently presented at Ray Day Seattle 2026 on scaling large-scale machine learning training 🚀 In my talk, “Scale 10B+ Model Training with Ray,” I shared how we’re evolving Uber’s AI platform to support the next generation of deep learning workloads — scaling from millions to 10B+ parameter models. Some key takeaways: - Building a distributed training stack with Ray, PyTorch, and DeepSpeed to efficiently scale across heterogeneous clusters - Evolving from traditional pipelines to Ray Data to improve GPU utilization and eliminate data loading bottlenecks - Leveraging model parallelism (ZeRO), mixed precision, and Flash Attention to push model size and training efficiency - Achieving up to 20% improvement in GPU utilization and ~50% reduction in training time in large-scale pipelines - It’s exciting to see how the Ray ecosystem is maturing — enabling a more unified approach where data processing and training pipelines converge into a single system. Big thanks to the Uber AI Platform (Michelangelo) team for driving this work forward, and to Anyscale for hosting this fantastic event Robert Nishihara Brent Bain. If you’re working on large-scale ML systems or distributed training, I’d love to connect and exchange ideas. Also enjoyed connecting with and learning from industry peers in Seattle — Chi Wang Mickey Liu Haocheng B. Qia Wang Shujun Bian https://lnkd.in/gHqSbcGw https://lnkd.in/gCKGd5SJ https://lnkd.in/g77ZF4xs #MachineLearning #DistributedSystems #Ray #DeepLearning #MLOps #AIInfrastructure
-
Peng Zhang reposted thisPeng Zhang reposted thisExplore how we optimized Petastorm to cut deep learning training time by 6× and eliminated hidden randomness at Uber. Read more: https://lnkd.in/g4sXsKMM #Uber #UberEngineering #MachineLearningAccelerating Deep Learning: How Uber Optimized Petastorm for High-Throughput and Reproducible GPU TrainingAccelerating Deep Learning: How Uber Optimized Petastorm for High-Throughput and Reproducible GPU Training
-
Peng Zhang reposted thisPeng Zhang reposted thisLearn how Uber Eats scaled its homefeed discovery engine by transitioning to transformer-based generative recommenders and a modernized ML infrastructure to deliver real-time, personalized recommendations to millions. Read More: https://lnkd.in/g_pJwt2C #Uber #UberEngineering #MachineLearningNext-Gen Restaurant Recommendation with Generative Modeling and Real-Time FeaturesNext-Gen Restaurant Recommendation with Generative Modeling and Real-Time Features
-
Peng Zhang reposted thisPeng Zhang reposted thisSome of the great work our AI infra and customer teams (led by Himaanshu Gupta , Peng Chen and many other folks). Congrats to everyone! https://lnkd.in/gmn5mJunNext-Gen Restaurant Recommendation with Generative Modeling and Real-Time FeaturesNext-Gen Restaurant Recommendation with Generative Modeling and Real-Time Features
-
Peng Zhang liked thisPeng Zhang liked this
-
Peng Zhang liked thisPeng Zhang liked thisI’m excited to be speaking at AI Agenda Live in San Francisco on September 23, joining leaders across the AI ecosystem to explore what’s next for the technology—and what’s actually happening on the ground. If you’d like to attend, register here: https://lnkd.in/dEj9P_9F
-
Peng Zhang liked thisPeng Zhang liked thisHarvey is hiring a Staff Product Manager, Infrastructure (San Francisco). This is the person who'll own the roadmap for the systems powering every interaction on our legal AI platform — multi-region deployments, observability, and reliability/scalability/security work. 6+ years shipping and scaling production platforms, comfort going deep with engineers on distributed-systems trade-offs, and the ability to turn ambiguity into clear plans. AI/ML or legal-tech background is a plus, not a must. Know someone great? Send them our way https://lnkd.in/gRXV_qiv
-
Peng Zhang liked thisPeng Zhang liked thisGreat piece by @JoannMuller at Axios, going behind the scenes of what we're building at Uber. Data is going to be the key unlock to accelerating safe, reliable autonomy everywhere - and we're just getting started. AV Labs data collection vehicles hit the road this month 🚀 https://lnkd.in/giyBg9gb
-
Peng Zhang liked thisPeng Zhang liked thisAgentic coding is moving fast, and GitHub Copilot Inline Suggestions continue to be used and loved by millions of developers. Excited to share how our team is making this experience even better with a specialized, low-latency model and post-training. Great work by Julia Gong, Ben Liggett and Ulugbek Abdullaev! Part 1 is out now, with more coming in Part 2. Stay tuned! We’re also hiring applied researchers to join us. Links are in the comments. #GitHubCopilot #VSCode
-
Peng Zhang liked thisPeng Zhang liked thisHi friends - We're continuing to hire for AI PM and Eng on the team, and wanted to do a plug for some additional roles below. If you know anyone that's a good fit, please send them over. Principle AI Engineers (L8 at most companies, US & India) - ping directly if interested Sr. Staff AI Engineers (L7 at most companies, US & India) - https://lnkd.in/g-cHHsBK L4 AI PM, US - https://lnkd.in/gBhneTDt L5 AI PM, Bangalore - https://lnkd.in/grJNcp6H I wrote a bit about the team/context in my earlier post (https://lnkd.in/p/gtMPfn-i), but TLDR - once-in-a-lifetime AI transformation, lots of really interesting tech innovation, systems-level thinking on steroids, and the team is great! One of the most fun I've had in AI because the results are so measurable and visceral -- strong foundations but with enough gaps that good people have the opportunity to directly drive outsized impact. Come join us! cc Benigne Shashaank Verena Anthony Raghav Surabhi Anuj Divya AditiSr. Staff Engineer (GenAI), San Francisco, United StatesSr. Staff Engineer (GenAI), San Francisco, United States
-
Peng Zhang liked thisPeng Zhang liked thisCall for Papers: AI Medicine; <https://lnkd.in/gQ7AQ7Fs>
-
Peng Zhang liked thisPeng Zhang liked this🚀 We’re hiring — come build the infrastructure powering the next generation of AI at ByteDance! Our teams in Singapore are growing, and we’re looking for talented engineers and builders across AI Infrastructure, ML Systems, Backend Engineering, SRE and ModelArk. If you’re excited about solving large-scale engineering problems and building systems that power real-world AI applications, we’d love to hear from you. 👀 🔥 Experienced Hires 🔹 Machine Learning Storage Infrastructure Engineer Build next-generation storage infrastructure for large-scale ML workloads. 👉 https://lnkd.in/gbScTF4P 🔹 Senior Backend Engineer — AML Engine Orchestration Work on large-scale orchestration systems powering ML/AI workloads. 👉 https://lnkd.in/dtEfJuUy 🔹 Site Reliability Engineer, Machine Learning Systems K8s × GPU × AI Infrastructure — build highly reliable systems at scale. 👉 https://lnkd.in/geA8VmnQ 🔹 Product Solution Architect — ModelArk (Singapore) Help customers and businesses unlock the potential of AI through ModelArk. 👉 https://lnkd.in/gpgjxbvY 🎓 Campus Hiring We’re also opening opportunities for 2027 campus talent across AI/ML Systems, Backend, AI Infrastructure, ML Engineering and more! 👉 Check out our campus opportunities: https://lnkd.in/gQF9z8sS Whether you’re an experienced engineer looking for your next challenge or a student ready to start your AI journey — come build the future with us in Singapore. 🇸🇬🤖 📩 Feel free to reach out to me / Xiaoci Li Yizhi Li if you’d like to learn more or explore which role might be the best fit for you. #ByteDance #Hiring #SingaporeJobs #AI #ArtificialIntelligence #MachineLearning #AIInfrastructure #MLSystems #BackendEngineering #SRE #Kubernetes #GPU #CampusHiring #TechJobsBackend Engineer - Machine Learning Storage Infra (Singapore) - Join ByteDanceBackend Engineer - Machine Learning Storage Infra (Singapore) - Join ByteDance
Experience & Education
-
ByteDance
*********** ******* * ***** ** ******
-
****
*********** ******* * **** ** ******** **************
-
*********
******** ******** * ********* ********
-
******* ********* ** **********
-
-
******* ********* ** **********
-
View Peng’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Licenses & Certifications
Honors & Awards
-
Best Advanced AI Projects Award from Microsoft AI School
Microsoft AI School
Successful finished the PUE optimization project using deep neural network. Win the 1st place among 500+ AI project submitted to Microsoft AI School.
Languages
-
English
Full professional proficiency
-
Chinese
Full professional proficiency
Recommendations received
3 people have recommended Peng
Join now to viewView Peng’s full profile
-
See who you know in common
-
Get introduced
-
Contact Peng directly
Other similar profiles
Explore more posts
-
SemiWiki.com
12K followers
SK Hynix presented a recent IEEE paper describing an architecture combining High-Bandwidth Memory (HBM) speed and High-Bandwidth Flash (HBF) capacity on a single interposer connecting both to a GPU to accelerate AI model and agent inference processing. #Semiconductors #Semiconductor #Semiconductormanufacturing https://lnkd.in/gVPYkyqq
3
-
Coherence
254 followers
When a Berkeley systems architect says “we need to operate in a coherent manner,” it’s not just about teamwork — it’s about compute. Today at the AI Investment Summit, the leaders shaping the next compute stack — from NVIDIA, SambaNova, Google TPU, and Berkeley EECS — all echoed the same theme: Coherence is the new optimization frontier. From KV-cache locality to orchestration across agentic layers, the future of compute will be measured not just in FLOPs per watt — but in coherence per system. One panel said it best: we’re not just designing chips — we’re redesigning coordination. Jiantao Jiao (NVIDIA / UC Berkeley) reminded us that multi-agent systems aren’t just in software — they need to operate coherently all the way down to the hardware layer. Orchestration is no longer abstract — it’s physical. Suvinay Subramanian (Google TPU) added that algorithm and hardware are finally intersecting — and that “it’s just like traditional software.” Compute itself is becoming dynamic, programmable, and alive. Raghu Prabhakar (SambaNova) grounded the conversation in the economics of energy and specialization: “We extract maximum efficiency by specializing models — like fine-tuning LLaMA — and by increasing chip utilization through architectural coherence.” Sophia Shao (UC Berkeley / ex-NVIDIA) closed with the line that stayed with me: “The next bet is on talents.” Because as compute grows more distributed, the real bottleneck isn’t bandwidth — it’s alignment. We talk about compute, cache, and concurrency — but the next layer is coherence. Hardware achieves it through synchronization. Teams achieve it through trust. AI will need both. 🧠⚡ #AI #InferenceFactory #ComputeStack #CoherenceProtocol #Berkeley #NVIDIA #SambaNova #GoogleTPU Sophia Shao Jiantao Jiao Suvinay Subramanian Raghu Prabhakar
1
-
Frontline One Capital
259 followers
🧠 Cohesity and NVIDIA: how unstructured data is becoming fuel for AI agents Our portfolio company unveiled the updated Gaia AI platform at NVIDIA GTC 2026. Cohesity is transforming data stored in secure repositories and backup systems into an important layer of safe AI infrastructure for enterprise use. 🔑 Key points: 🥇 Leader at the intersection of AI and cybersecurity Gaia enables secure AI search and AI agents to run directly on customer backups without duplicating data or moving it beyond the company’s security perimeter, helping minimize risks and reduce the attack surface for malicious actors. 💡 Cohesity’s growth trajectory reinforces our investment thesis The company represents a unique combination of data protection / data management and AI functionality within a single platform and technology stack. We are looking forward to a successful IPO this year! 🤝 Deep partnership with Nvidia The platform is tightly integrated with Nvidia AI Enterprise products, using NIM, Nemotron Reranking, NeMo Guardrails, and also testing AI-Q Blueprint, NeMo Agent Toolkit, and cuVS. Integration with Nvidia accelerates model deployment, improves response accuracy, and reduces the share of toxic or inaccurate outputs. 🔗 A gateway for AI agents Gaia serves as a secure infrastructure layer for enterprise data access for Microsoft Copilot, Glean, and Google Gemini. Cohesity unlocks years of company history and archived files stored in backups without compromising compliance or requiring complex access policies. 🔒 Sovereign secure AI Cohesity has become a gold standard for highly regulated sectors, which is why the largest corporations and banks choose its platform (70% of Global 500 companies). Gaia allows customers to securely embed AI into their workflows and perform computation where the data already resides — on customer-owned servers or in sovereign cloud environments. This architecture is extremely difficult for cybercriminals to compromise. 🔗 : https://lnkd.in/efAjTP8D #Portfolio #Cohesity #NVIDIA #AI #Gaia #Cybersecurity #AgenticAI
7
-
Joe Pierce
NVIDIA • 19K followers
☁️ Scale frontier mixture-of-experts models faster with Google Cloud and NVIDIA. Google Cloud utilized A4X machines powered by NVIDIA GB200 NVL72 systems and NVIDIA Dynamo to achieve: ✅ Over 6,000 total tokens/sec/GPU ✅ 10ms inter-token latency for long sequence lengths This validated reference architecture delivered distributed runtime KV cache management and kernel scheduling with Wide Expert Parallelism (WideEP) across the full 72-GPU compute domain. Explore the technical details and get started with deployment recipes →
1
-
iSTART-TEK
273 followers
【iSTART Pick|UDA: Choose the Optimal Test Algorithm Yourself】 UDA (User-Defined Algorithm) allows teams to define and quickly deploy memory test algorithms, achieving high coverage and fast debugging across different memory architectures. This video highlights: practical UDA operations, how to improve test efficiency and coverage, and recommendations for integrating UDA into existing test flows. 🎥 Watch the full video: https://lnkd.in/gPunUiQ9 #UDA #memorytest #algorithm iSTART-TEK
1
-
NAND Research
887 followers
How do you efficiently extend LLM context memory at scale? You tier & extend the KV cache. NVIDIA provides explicit support w/ upcoming Vera Rubin, announcing its new Inference Context Memory Storage (ICMS) platform. The latest NAND Research Note from chief analyst Steve McDowell takes a look: → Creates new "G3.5" storage tier between local node storage and shared network storage using BlueField-4 data processing units to manage flash-based context memory at pod level → NVIDIA promises 5x higher tokens-per-second and 5x improved power efficiency (compared to traditional storage) for long-context inference workloads → Integrates with NVIDIA Dynamo orchestration framework and Spectrum-X Ethernet fabric for RDMA-based KV cache access → Initial support from broad range of vendors: AIC, Cloudian Inc, DDN, Dell Technologies, Hewlett Packard Enterprise, Hitachi Vantara, IBM, Nutanix, Pure Storage, Supermicro, VAST Data and WEKA TL;DR: NVIDIA ICMS addresses real constraints for multi-turn agentic AI systems that maintain context across sessions, while also showing off the true power of NVIDIA’s BlueField-4 DPUs. Our full take at the link. #AI #LLM #InferenceOptimization #DataCenter #EnterpriseAI #KVCache #NVIDIA #StorageInfrastructure #ArtificialIntelligence https://lnkd.in/eJjBAh4F
3
-
Vicharak
20K followers
We build computing platforms, but we also like to understand what happens underneath them. On the Vicharak Blog, we explore AI, embedded systems, Linux, SBCs, FPGA, and reconfigurable computing through practical guides, experiments, projects, and the things we learn while building. https://lnkd.in/eHvVMY4H
14
-
Edge Impulse
51K followers
Great session walking through a powerful edge AI workflow: train in Amazon Web Services (AWS) #SageMaker, optimize with Qualcomm AI Hub, prepare and package in Edge Impulse, then deploy and monitor fleets using AWS IoT #Greengrass and IoT Core, all running on Qualcomm #Dragonwing processors. Clear guidance for teams building production edge systems. Way to go Ashvin Roharia and Krishna Sridhar!
43
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content