Sign in to view Stephanie’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Stephanie’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Seattle, Washington, United States
Sign in to view Stephanie’s full profile
Stephanie can introduce you to 10+ people at University of Washington
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
1K followers
500+ connections
Sign in to view Stephanie’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Stephanie
Stephanie can introduce you to 10+ people at University of Washington
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Stephanie
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Stephanie’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Activity
1K followers
-
Stephanie Wang posted thisEveryone knows that RDMA is fast. But how do you actually realize the speedups in your own application? As we were developing Ray Direct Transport (RDT), we found that there a lot of nuances in how Python applications use RDMA libraries like NVIDIA's NIXL, to the point where just 𝘂𝘀𝗶𝗻𝗴 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝗹𝗶𝗯𝗿𝗮𝗿𝘆 𝗰𝗮𝗹𝗹𝘀 𝗰𝗮𝗻 𝗺𝗮𝗸𝗲 𝗮 7.5𝘅 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗶𝗻 𝗽𝗲𝗿𝗳𝗼𝗿𝗺𝗮𝗻𝗰𝗲. Our goal with RDT was to make RDMA easy and seamless to use for Ray applications like weight syncing for RL for LLMs. To do that, we also had to make sure performance was easy to achieve. Check out our blog post that walks through how to do this yourself using RDT. These steps are the key to the speedups that RDT achieves for weight syncing in end frameworks (𝘮𝘰𝘳𝘦 𝘰𝘯 𝘵𝘩𝘦𝘴𝘦 𝘪𝘯𝘵𝘦𝘨𝘳𝘢𝘵𝘪𝘰𝘯𝘴 𝘴𝘰𝘰𝘯): • 3.5s for Qwen3-235B-A22B with SkyRL - 18x faster than NCCL • 2.3s for GLM-4.5-Air with Miles - 2.2x faster than NCCL and 1.14x faster than Mooncake Blog link: https://lnkd.in/gCQSf-gz Joint work with my awesome teammates at Anyscale: Joshua Lee Xinyu Zhang Sumanth R Hegde Aaron Hao (and former teammate Dhyey Shah). Thanks also for feedback from Mengjin Yan Richard Liaw.
-
Stephanie Wang shared thisWhat do coding agents look like from the perspective of an LLM inference system? Find out yourself with TraceLab :) https://lnkd.in/gBh-jE4AStephanie Wang shared thisCoding agents have become one of the hottest LLM workloads. But serving them looks nothing like serving a chatbot: 294× more input than output, hundreds of thousands of tool calls, long autonomous loops, and extremely long-tailed latency. Introducing TraceLab: a large real-world coding agent trace for serving systems, plus an open-source pipeline to collect, sanitize, analyze, and replay your own traces. The trace contains roughly 4,300 real-world coding-agent sessions from our daily use of Claude Code and Codex, with 55B input tokens, 186.9M output tokens, and 433K tool calls. It is designed to help study coding agents as a serving workload, not just as isolated benchmark tasks. A few findings stood out: 1. Coding agents are autonomous, multi-step workloads. Each user request triggers an average of 8.8 self-directed LLM-tool steps and 10.8 tool calls. As a result, 88% of LLM rounds respond to tool results rather than directly to humans. End-to-end latency is highly skewed: the median is around 38s, but the mean is about 4 min and the p99 is near 44 min. 2. Coding agents are long-context, short-output. Across the trace, models read 52.56B cached input tokens and prefill 2.34B new input tokens, but generate only 186.9M output tokens. A median round carries a 119K-token prefix, appends only 875 new tokens, and decodes just 214 tokens. 3. Tool calls dominate the agent loop. Across 433K tool calls, about 76% are shell or command executions: builds, tests, git commands, scripts, and similar operations. Tool latency is extremely long-tailed: calls under 1s make up around 61% of all calls but only 1% of total tool time, while the roughly 4% of calls lasting over a minute consume about 85% of total tool-call time. 4. Prefix caching works well overall, but misses are still expensive. In aggregate, 95.7% of input tokens hit the prefix cache. But tool-result continuations hit around 97.5%, while user-initiated steps hit only about 84.4%. Of 2.3B append tokens, only 19% are genuinely new context; the other 81% is repeated re-prefill work. Our release includes trace examples, the collector/sanitizer, the full anonymized trace with analysis scripts and a chatbot, and a replay client for serving engines such as vLLM and SGLang. Check out TraceLab, try it on your own coding-agent logs, and get in touch if you are interested in coding-agent workloads, LLM serving, and other related research! Links: Live demo: https://lnkd.in/gBh-jE4A Paper: https://lnkd.in/gUcP8jeT Blog: https://lnkd.in/g3AUYTzx Trace and code: https://lnkd.in/gJqyemHi Joint work by Kan Zhu Mathew Jacob, Chenxi Ma, Yi Pan, Stephanie Wang, Arvind Krishnamurthy, and Baris Kasikci!
-
Stephanie Wang shared thisIt's SyFI release time - last week distributed training, this week multimodal inference! The main idea: New multimodal model architectures shouldn't require new serving systems. Introducing our work, M* (M-Star): a universal serving system for multimodal models that separates what a model computes - a dataflow graph - from how it runs: placement, scheduling, batching, and transport. The latest multimodal models have broken a key assumption in modern LLM serving systems: that inference is a single autoregressive loop. Today's models have composite architectures. One model can wire together structurally different components: vision/audio encoders, a transformer backbone, diffusion/flow heads, and action predictors. Within the same multimodal model, a chat, an image edit, and a speech request could each walk a different path. Multimodal inference systems like vLLM-Omni and SGLang-Omni support these models by extending from LLMs to model pipelines. However, that abstraction fails to capture the diversity and complexity of modern multimodal graphs, leaving significant performance on the table. M*'s abstraction is the Walk Graph: 1. The model is a dataflow graph of component nodes (vision encoders, transformer backbones, etc) and tensor edges. 2. The model author also declares a series of Walks: Each Walk is a labeled subgraph for one phase of model execution (prefill_text, decode, image_gen, …). 2. A request is a series of Walks. The Walk Graph captures arbitrary composite architectures and physical placements, allows the runtime to execute only the components needed by a given request, and maintains state-of-the-art per-component performance. With this abstraction, we show that M* matches or beats specialized systems built for each model family: - Qwen3-Omni text-to-speech: up to 2.7× higher throughput than vLLM-Omni and 4.0× higher than SGLang-Omni, at up to 2.9× lower real-time factor (RTF). - BAGEL: ~20% lower text-to-image latency and up to 2.6× faster image editing compared to vLLM-Omni. - V-JEPA 2-AC rollouts for robotic planning: up to 12.5× faster as the horizon grows compared to a hand-built baseline Check out M* and get in touch if you have a multimodal use case, if you’re struggling with the diversity of multimodal models, or if you’re interested in disaggregated and autoscaling inference! Links: Project and blog: https://m-star.org/ Code: https://lnkd.in/gtnFyXv4 Arxiv: arxiv.org/abs/2606.12688 Work led by our wonderful students Atindra J., Naomi Sagan, and Keisuke Kamahori!M*: A Modular, Extensible, Serving System for Multimodal ModelsM*: A Modular, Extensible, Serving System for Multimodal Models
-
Stephanie Wang shared thisIf you are attending OSDI '26 in Seattle next month, please consider signing up for the mentorship program! We are looking for both mentors (faculty, postdocs, industry) and students (undergrad, PhD) to participate. The mentorship activity is designed to give student attendees a chance to get to know more people at the symposium, get career advice from senior members of the community, and gain feedback on their research. Sign up at the link below by June 26, 2026 to take part! https://lnkd.in/gZMfa25h
-
Stephanie Wang shared thisThanks for having us! Don't miss the July 14 event in Seattle :)
-
Stephanie Wang shared thisExcited and proud to announce the open-source release of Piper, work led by my first PhD student Megan Frisella (coadvised with Gilbert Bernstein)! The main idea: New distributed training strategies and optimizations should not require new distributed runtimes. Large training jobs increasingly combine multiple parallelism strategies such as pipeline, data, and expert parallelism with ZeRO-style sharding, creating placement and GPU scheduling choices that current frameworks cannot express cleanly. Today, ML researchers and practitioners choose between building one-off specialized systems that perform well but are hard to adapt, and general-purpose frameworks that are easier to use but expose limited control. Piper is a user-controllable distributed training system for PyTorch that separates model placement and GPU scheduling from model code and runtime implementation. With lightweight model annotations and a small scheduling language, Piper lets users express, visualize, profile, and run high-performance training schedules such as DualPipe-style pipeline- and expert-parallel overlap. On PP for MoE models, Piper beats TorchTitan by at least 1.2x and matches Megatron-LM. Piper also beats Megatron-LM using pure PyTorch (no manually optimized kernels) by 6% by using a DualPipe strategy. More results coming soon! Check out the blog for the code and more info, and feel free to reach out if you have feedback! https://lnkd.in/gw5m8-_WIntroducing Piper: A Programmable Distributed Training SystemIntroducing Piper: A Programmable Distributed Training System
-
Stephanie Wang shared thisVery exciting stuff from our lab, UW SyFI! :)Stephanie Wang shared thisCan AI agents build an entire LLM serving system from scratch? We've spent the last few months exploring this question, and I'm excited to share what we found. 🚀 Introducing VibeServe: a multi-agent framework that synthesizes a complete LLM serving runtime end-to-end. Instead of a single generic stack for every deployment, VibeServe generates a bespoke system tailored to your model, hardware, and workload. Why now? LLM serving systems (like many other computer systems) have been built with a one-size-fits-all mindset. They run mainstream workloads with impressive performance, but there's always a long tail -- new model architectures, niche workload patterns, and exotic hardware are emerging every day, where generic stacks fall short. Coding agents flip the cost calculus and make per-target specialization practical for the first time. Building this with agents isn't easy, though: long-horizon task planning, reward-hacking under performance pressure, and context decay across many rounds. VibeServe tackles all three with two nested loops plus a persistent state kept outside any agent's context: 🔁 Outer loop: a search policy over validated git checkpoints, driven by an issue backlog and long-term memory. ⚙️ Inner loop: three agents in independent contexts (Implementer, Accuracy Judge, and Performance Evaluator). 📚 Skills library: distilled knowledge about serving-system design and performance optimization. Two main findings across six case studies: 1️⃣ No tax on mainstream setups. On Llama-3.1-8B + H100, a heavily optimized standard config, VibeServe matches vLLM and SGLang on throughput, TTFT, and TPOT. 2️⃣ Substantial wins on the long tail. VibeServe-generated runtimes substantially outperform human-optimized stacks on niche-yet-realistic deployments: - Predicted-output-based decoding for code editing: 5.95× speedup - Hybrid SSM+attention model with prompt caching: 3.45× speedup - Streaming ASR with per-stream encoder cache: 1.69× speedup - Constrained JSON decoding on a MacBook: 2.6× speedup - Multimodal model with complex architecture: 6.27× speedup Each speedup comes from an optimization specific to that target. This is the kind of work a generic runtime can't justify, but a bespoke, agent-generated one can. To our knowledge, this is one of the first demonstrations that AI agents can autonomously generate real systems that beat human-optimized ones on realistic deployments. And we don't think it stops at LLM serving -- the same tension between generic abstraction and target-specific performance shows up across many computer systems, and a target-agnostic harness plus a skills library is a recipe that should travel. Joint work with Vic Li, Simon Peter, and Baris Kasikci at the UW SyFI Lab. Huge thanks to Modal for the compute credits! 📄 Paper: https://lnkd.in/gy2JZFj3 📝 Blog: https://lnkd.in/gtcK5uCq 💻 Code: https://lnkd.in/gABqqKBv
-
Stephanie Wang reposted thisStephanie Wang reposted this#MLSys2026 is offering Student Travel Grants to help students and early-career researchers attend the conference in Bellevue in May! The grant can partially cover registration, lodging, and airfare. Apply by May 8, 2026! https://lnkd.in/gQQF5cSP
-
Stephanie Wang shared thisCalling all students, postdocs, and recent grads working in ML systems! We’re organizing the Young Professionals Symposium (YPS) at #MLSys2026, on Monday May 18 in Bellevue, WA. We have a full day of content including invited speakers, industry talks, panels, and a poster session. Invited speakers include Lisa Li (OpenAI, incoming assistant professor at UW), Esha Choukse (Principal Researcher at Microsoft), Yuhan Liu (final-year PhD candidate at UChicago and core author of LMCache), and Roger Wang (core maintainer of vLLM and cofounder at Inferact). Register for this event using the main conference registration page (https://lnkd.in/gcs3HDnG). We also strongly encourage students to submit a 1-page abstract for the poster session, due April 26, details at https://lnkd.in/g8nJ_mqy. Hope to see you there!
-
Stephanie Wang liked thisAnother AgentSys for the books! Last night at our Seattle office: 1/ Zihao Ye (creator of FlashInfer) introduced a new "agent-compiler" paradigm for GPU kernel programming. 2/ Prof. Simon Peter (UW Seattle) shared his group's work on "self-defining systems" and one of the early successes in automating the SDLC for inference engineering. Again, we were joined by leading professors, researchers, engineers and CTOs from NVIDIA, Google, Amazon, Meta MSL, Microsoft Research, and a number of startups. Thank you to our wonderful speakers. Thank you to each one who joined and thank you to Prof. Stephanie Wang (UW SyFi Lab), Claris Winston (Amazon), Cleah Winston (UW WEIRD Lab) for co-organizing and Arker for hosting and sponsoring. Here at Arker, we work on the virtualization layer for a completely new kind of cloud architected to run the most challenging production systems4agents and agents4systems workloads across CPU/GPU metal. It's genuinely an honor to help bring together folks passionate about the open problems in agents and systems as we are. If you couldn't make it, not to worry! The next AgentSys will be on November 17. Register early as space fills fast: https://agentsys.cc and luma.com/agentsys.
-
Stephanie Wang liked thisStephanie Wang liked this[new blog post] Specula: Scaling formal specifications for autonomous model checking of system code https://lnkd.in/g8c5NGxe
-
Stephanie Wang liked thisIncredible privilege to work with Stephanie Wang and the rest of team on RDT! Excited for people to try out all the improvements we made :)
-
Stephanie Wang liked thisStephanie Wang liked thisWant to dive in to building and harnessing agentic flows for computer architecture, systems, and chip design research for free and win prizes? Google is sponsoring a CHIA hackathon in the lead up to the A^3: Agentic Approaches to Architecture Workshop at the 2026 International Symposium on Microarchitecture (MICRO) in Athens, Greece. Google will provide free cloud and AI model access to hackathon participants solving cool architecture/systems/VLSI problems by building agentic loops in CHIA (https://chialoops.ai). Hackathon winners will receive cash prizes and give talks on their work at the A^3 Workshop! All hackathon participants will be invited to present a poster on their creative agentic loops at the workshop as well. To get involved in the hackathon, we require a very short (1 page max) proposal outlining some details about your planned hackathon project. The hackathon page outlines several potential project ideas, or you can propose your own! Proposals are due August 25 (but are very simple!). Learn more and find out how to submit your proposal here: https://lnkd.in/ge4nddhB If you have questions, reach out to me or my co-organizers (Abraham Gonzalez, Akanksha J., Amir Y., Dimitrios Skarlatos, Qijing Jenny Huang) at a3-workshop-hackathon@eecs.berkeley.edu.
-
Stephanie Wang liked thisStephanie Wang liked thisHave a GB200/GB300 lying around and feel like Ray is missing that extra oomph? One culprit might be where Ray is placing your workers. Today, we’re introducing NVLink Domain-Aware Placement Groups: a new topology-aware scheduling primitive in Ray designed in collaboration with NVIDIA to maximize the effectiveness of Ray on rack-scale hardware and keep your GPUs brrring. Thanks to Mengjin Yan Aaron Li Edward Oakes and our design partners at NVIDIA Connor Pedersen Mathew Wicks for helping this get to prod! :)Maximizing the Power of NVIDIA GB300 NVL72: NVLink Domain-Aware Placement Groups in RayMaximizing the Power of NVIDIA GB300 NVL72: NVLink Domain-Aware Placement Groups in Ray
-
Stephanie Wang liked thisStephanie Wang liked thisWe've raised money. It's time to spend it on the smartest people we know. We're hiring founding engineers. https://lnkd.in/g59Hhn2w DM if you want to kick some ass.
Experience & Education
-
University of Washington
********* *********
-
********
******** ********
-
********** ** *********** ********
******** ******* **********
-
********** ** *********** ********
****** ** ********** * *** undefined undefined
-
-
************* ********* ** **********
****** ** *********** ****** ******** *******
-
View Stephanie’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Courses
-
Advanced Data Structures
6.851
-
Algebra I
18.701
-
Algebra II
18.702
-
Algorithms
6.046
-
Computation Structures
6.004
-
Computer Systems
6.033
-
Computer Systems Security
6.858
-
Distributed Systems
6.824
-
Intro to Algorithms
6.006
-
Linear Algebra
18.06
-
Math for Computer Science
6.042
-
Performance Engineering
6.172
-
Real Analysis
18.100C
-
Software Engineering
6.005
-
Theory of Computation
18.404
Projects
-
HODOR: HODOR On-Disk Orthogonal Range Trees
See projectSerializing k-dimensional trees on disk to optimize query performance on large datasets
6.851 (Advanced data structures) final project -
Bookdrop
- Present
See projectFlask web app that syncs eBooks from a Dropbox folder with a Kindle. Deployed at getbookdrop.com with over 7,000 users and 200,000 books synced.
-
ParselTongue
- Present
Chrome extension for client-side encryption of chat applications to protect user privacy
Honorable mention at NYC Facebook Hackathon, Summer 2013 (demoed with Facebook Chat)Other creatorsSee project -
Stitch-It
Django web app providing a visualization and editing interface for crochet text patterns, built on a social network for crocheters and designers
Designed algorithm for positioning and rendering stitches, implemented in JavaScript and HTML Integrated interface with a text pattern parser to produce responsive visualizationOther creatorsSee project
View Stephanie’s full profile
-
See who you know in common
-
Get introduced
-
Contact Stephanie directly
Other similar profiles
Explore more posts
-
Tirth Shah
California State University… • 2K followers
I built a multi-agent debate loop for my master's thesis and never once checked what it cost me. I was at Cerebras Supernova on Tuesday for the CS-4 launch, listening to claims of up to 30x faster inference. My first reaction was skeptical. Every hackathon I have been to had people trimming prompts to save tokens. Nobody was wishing their chatbot answered one second sooner. Then I thought about my own thesis and stopped arguing. It runs two agents holding opposing personas, arguing over five rounds with RAG grounding underneath. Ten sequential model calls before a student sees anything. Each round carries the whole transcript forward, so the later rounds cost more than the earlier ones. And for a while it broke in the worst possible way. One agent restated itself. The other restated itself in response. Five full rounds of that. No crash, no error. A transcript that looked exactly like a debate and contained one round of real content. The only reason it stopped was a hard cap at five rounds, which I picked because five felt like a reasonable classroom debate length. An arbitrary teaching decision turned out to be my cost control. Here is the part I keep coming back to. I never measured the tokens, because the model ran locally and was free. Local inference removed my cost signal. Fast inference removes your time signal. Slowness is currently doing free monitoring work for all of us. A stuck loop announces itself when each turn takes seconds. At 30x, the same bug fails silently with a bill attached. Four things I would build into any agent loop before chasing speed: 1. A hard turn cap per run. Mine saved me by accident. 2. A cost ceiling that kills the job outright. 3. Loop detection that flags when two consecutive turns produce substantially the same output. 4. Checkpointing, so a killed run does not restart from zero. Full writeup in the comments, including where I think the speed pitch does not hold at all. Has anyone had a loop degenerate quietly, produce plausible looking output, and only find out later? #AIAgents #LLMOps #AIInfrastructure #ArtificialIntelligence
31
1 Comment -
Jai Palan
Tidra • 1K followers
UW engineers, what if every coding session quietly helped build your next resume? I’ve been working on a project called Cadence, and it's built specifically for UW engineers. The idea is simple: you give your coding agent a short instruction, and it quietly keeps a log of the projects, features, and technical work you complete over your internship. Cadence can then turn that log into a resume, update an existing resume, or tailor it to a job posting. Very useful when we keep switching between school and work terms + makes it easier when applying to tons of jobs during cycles. It includes: - A LaTeX editor with a live PDF preview (Similar to Overleaf) - AI-assisted resume editing - Targeted updates that don’t rewrite your entire resume - Job-specific resume tailoring - Support for Claude Code, Cursor, Copilot, Codex, and other agents - Your choice of Gemini, OpenAI, Anthropic, or another compatible provider It’s still early, so I’m mainly looking for honest feedback and people willing to tell me what feels confusing or broken. Try it here for free! https://cadencecv.ca/
31
5 Comments -
Rain
Oxide Computer Company • 2K followers
My colleague David Crespo and I were recently on the Oxide and Friends podcast, where we talked about some of the ways we've been able to use LLMs like Claude Opus 4.5 to achieve a higher level of rigor and robustness than possible with humans alone. If you're a systems engineer who is despondent at how the dominant narrative is slop and shipping as much code as possible, quality be damned, then please give the episode a listen! LLMs have a lot to offer to detail-oriented infrastructure engineers. YouTube: https://lnkd.in/gAcrpQNb Transistor: https://lnkd.in/gzhsEFuQ
259
4 Comments -
Zikun Wang
Northeastern University • 680 followers
VLDB is happening next week in Boston, and I have two exciting works to share on #VectorSearch, in collaboration with Hunter McCoy and my advisor Prashant Pandey. In the main conference, we are presenting Jasper, a GPU-accelerated ANNS engine that supports streaming inserts and deletes without rebuilding, at 1.84× CAGRA's query throughput and 7.0× faster construction. https://lnkd.in/g-txzgdq At the VecDB workshop, we are presenting a poster on Directional Beam Search. We store a compact LSH summary on each edge, cutting the beam search I/O cost per explored node from 1+R probes to exactly 1, and giving up to 28× higher throughput on host memory with zero extra GPU memory. https://lnkd.in/gAHqjbQ8 Code for both projects is available on GitHub: https://lnkd.in/giDsbU9F If you're at VLDB and care about vector search, come say hi! #VLDB2026 #ANNS #GPU
42
3 Comments -
Vivian Di Bai
Google DeepMind • 1K followers
What happens to all the attempts an AI agent makes on the way to discovering a better solution? Dream-RSI keeps them as a world it can replay, asking what a different search would have found — no new experiments, no retraining. Glad to have been part of this work. Worth a read if you think about discovery loops and self-improvement in agents.
24
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content