Sign in to view Aaron’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Aaron’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Palo Alto, California, United States
Sign in to view Aaron’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
626 followers
500+ connections
Sign in to view Aaron’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Aaron
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Aaron
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Aaron’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View Aaron’s full profile
-
See who you know in common
-
Get introduced
-
Contact Aaron directly
Other similar profiles
-
Asher Feldman
Asher Feldman
Hands-on Platform and Reliability Engineering leader with 20+ years of professional experience. I have a real passion for building scalable systems and teams.
2K followersGreater Sacramento
Explore more posts
-
Nabeel Shah
Game Plan Tech • 1K followers
One detail in NVIDIA’s Nemotron 3 Super technical report really stood out to me. Not the 1M-token context. Not even the benchmark story. The KV cache numbers. At 262k context, the BF16 KV cache estimates in my breakdown come out to: • Nemotron 3 Super: 2.0 GiB • Qwen3-235B-A22B: 47.0 GiB • Llama 3.3 70B: 80.0 GiB That’s the part I think matters most for real long-context systems. Because this is where long-context stops being a model-card feature and starts becoming a systems problem. When KV cache explodes, inference gets expensive fast: less batching, more VRAM pressure, more memory traffic, and much harder economics for production deployments. What’s interesting about Nemotron is that NVIDIA seems to be attacking that problem directly. The architecture combines: • Mamba-2 layers for sequence propagation • a small number of grouped-query attention anchors • LatentMoE for sparse conditional capacity • MTP for faster generation So the story here is not just “another big model with 1M context.” It’s a model designed around the bottlenecks that usually make long-context inference painful to serve. I wrote a deeper technical breakdown here: https://lnkd.in/eAssu6EQ
5
-
Mohammed Mansoor Ahmed
SoruslyAI • 2K followers
400,000 GPUs. Trillions of tokens per day. That's SGLang's current footprint in production. SGLang is a high-performance serving framework for LLMs and multimodal models, built for low-latency, high-throughput inference from a single GPU to massive distributed clusters. If you're running LLM inference at any scale, you've probably hit the ceiling on vLLM or rolled your own patchy solution. SGLang's designed to replace that. It's already the backbone for xAI, LinkedIn, Cursor, and a dozen cloud providers. - RadixAttention for prefix caching cuts redundant compute on repeated prefixes - Prefill-decode disaggregation, speculative decoding, and zero-overhead CPU scheduling baked in - Runs on NVIDIA, AMD, Intel CPUs, Google TPUs, and Ascend NPUs - OpenAI-compatible API, so you don't have to rewrite your client code It's also the rollout backend for several frontier model training pipelines, not just serving. How are you handling LLM inference in your stack right now, and what's the bottleneck you keep running into?
5
1 Comment -
Javier Ruiz
Synapse Optics • 1K followers
A small tool I maintain has had an update: ConstrainedOpt, a standalone application that drives Zemax OpticStudio through the ZOS-API to run bound-constrained least-squares optimization. The reason it exists is that the limits are bounds on the variables themselves, not penalty operands in the merit function. If a thickness has to stay above 1 mm, it stays above 1 mm, and the merit function you wrote is the one being minimized — nothing added to it to police the geometry. Version 2 changes how a bound is enforced. A step that would leave the feasible region is now reflected back inside rather than clipped to the limit. Clipping parks a variable on its wall, where it stops responding to the optimizer while still occupying a column of the normal equations; reflection keeps it in the interior. It also adds Don Dilworth's PSD II and PSD III as alternatives to Marquardt damping, with the derivation and references in the README. MIT licensed, and useful mainly if you have designs where the constraints are the hard part: https://lnkd.in/gnGjPv4V
37
4 Comments -
Brian Budge
Meta • 6K followers
dispenso 1.6 -- high-performance C++ parallelism from Meta The thread pool was restructured around locality. Mandelbrot is 6x faster than in the 1.5 release, and the new parallel_for chunking strategy can beat TBB by 3-12% on SpMM. kAdaptive, the new dynamic load balancing strategy for parallel_for, is based on Callisto-RTS (Harris & Kaestle, USENIX ATC 2015). Against TBB on AMD CPU it holds that edge from 8 to 32 threads, and is within noise from 64 to 192. Elsewhere: bulk scheduling is 19-49% faster, futures trees 4-12%, graph scenes 2-11% (on a 166-thread EPYC), and Windows picked up roughly 6% geomean across the tuning set on a 48-thread Xeon. The scheduling wins come from locality; the bulk-scheduling one is allocation elimination -- small functors now live inline in a cache-line-sized object instead of being pool-allocated. You can now try dispenso on Compiler Explorer -- no install, no build. It runs the same uneven workload serially, with the default kStatic chunking, and with kAdaptive -- per-element cost grows 64x across the range, exactly where deciding chunk boundaries up front goes wrong. It's a shared sandbox, so ignore the timings; the example is also in the repo, so you can try it locally for real timings. New this release: CpuSet for CPU affinity and some NUMA/cache topology, including cache-aware thread group construction (full support on Linux, Windows and FreeBSD; topology-only on macOS). parallel_invoke for fork-join over heterogeneous tasks. when_any. ChaseLevDeque and MpmcRingBuffer. And DistributedRWLock, a sharded reader-writer lock for the low-write, high-concurrency-read case, drawing loosely on Dice et al. that can win big for cases where reads happen orders of magnitude more often than writes. FreeBSD support for CpuSet and for futex-like behavior came from Gilbert Morgan Jr. -- thank you! fast_math is still experimental, still behind DISPENSO_BUILD_FAST_MATH. It gained pow, hypot, tanh, erf, sincos and better accuracy from Sollya-generated minimax polynomials. Same warning as last time: the API may churn -- it will get the same cross-minor-version stability guarantees as the rest of dispenso in a later release. dispenso 1.6.2 is on vcpkg, Conan, Homebrew, MacPorts and conda-forge. Repo and Compiler Explorer links are in the comments. Last release I asked whether C++14 support mattered for fast_math; here's a broader version, because I'd rather ask before deciding than apologise after. dispenso targets C++14 today. I'm considering moving the default standard to C++20 -- a minor-release change that drops nothing. Further out, a 2.0 release would drop C++14 entirely. So, would a default bump to C++20 cause you trouble? If C++14 support went away in 2.0, how disruptive would that be? And if it did, is C++20 a reasonable floor now, or are enough people still stuck on C++17 that it should be the target? #opensource #opensourcesoftware #cplusplus
90
10 Comments -
Nima Shams
Qualcomm • 8K followers
A few months ago, our team and I unveiled Snapdragon 𝐒𝐓𝐀𝐑𝐓, a platform we've been building at Qualcomm over the past year to power the next generation of Personal AI devices. 𝐒𝐓𝐀𝐑𝐓 is an end-to-end solution designed to help brands bring AI products to market faster while maintaining their unique identity, user experience, and ecosystem. For the end customer, 𝐒𝐓𝐀𝐑𝐓 enables choice in product design and agentic capabilities in the rapidly growing space of Artificial Intelligence and Agentic flow. At its core, 𝐒𝐓𝐀𝐑𝐓 is built around the 3 S's: 𝐒𝐢𝐥𝐢𝐜𝐨𝐧 – Fully integrated, production-ready System-in-Packages (SiPs) that deliver state-of-the-art performance, power efficiency, and connectivity. 𝐒𝐨𝐟𝐭𝐰𝐚𝐫𝐞 – A complete, optimized stack spanning on-device AI, white-label Android and iOS applications, and the 𝐒𝐓𝐀𝐑𝐓 Cloud, enabling model orchestration, scalability, and freedom of choice. Start platform gives the user and brand agentic choice, unified memory and an optimized adaptable platform that grows with the Persona AI ecosystem. 𝐒𝐜𝐚𝐥𝐞 – A growing global ecosystem of partners and manufacturers who have integrated the 𝐒𝐓𝐀𝐑𝐓 hardware and software platform to create industry-leading white-label AI products. These companies include partners such as Applied Materials, PEGATRON, Thundercomm, 佐臻股份有限公司 Jorjin Technologies Inc., GoerTek Inc., Quanta, JBD, GHH and Avegant. Next week at Snapdragon Summit, our CEO Cristiano R. Amon, our GM Ziad Asghar, and our Software Lead Cameron Sylvia will unveil the next evolution of the 𝐒𝐓𝐀𝐑𝐓 platform, including agentic AI orchestration, end-to-end multimodal experiences, and perhaps a few product surprises along the way. We're approaching launch, and I couldn't be more excited about what comes next for Personal AI. Tune in to the livestream to see what's ahead: https://lnkd.in/g7cFeGyE And to learn more about 𝐒𝐓𝐀𝐑𝐓 please watch this: https://lnkd.in/gzpiBc5x #Qualcomm #SnapdragonSummit #PersonalAI #AgenticAI #GenerativeAI #Wearables #SmartDevices #AI #Innovation #augentedreality #virtualreality #AI
152
15 Comments -
Jay Patel
Silver Turtle Ventures • 717 followers
As OS developers, we live and breathe C. It's the lingua franca of systems programming. But in an era defined by security breaches and complex, concurrent hardware, we should ask: is "tradition" a good enough reason to ignore tools provably better for reliability? We've normalized buffer overflows, use-after-frees, and entire classes of race conditions as "just the cost of doing business" in C. We spend an enormous amount of time and money on static analyzers, fuzzers, and runtime mitigations (like ASLR/DEP) to bolt safety onto a language that was never designed for it. What if we started with a language that builds safety in from the ground up? I'm talking about Ada and its formally verifiable subset, SPARK. Why Ada/SPARK for OS Development? * No More "Undefined Behavior": Ada's strict type system and runtime checks (which can be disabled in production) catch errors at compile time that C compilers would happily ignore. Think integer overflows, out-of-bounds array access, and type-mismatch errors. * Memory Safety by Default: Forget malloc/free headaches and buffer overflows. Ada provides much safer mechanisms for memory and resource management. * Built-in Concurrency: Ada was designed for real-time, concurrent systems from day one. Its tasking and protected-object features are part of the language core, not a bolted-on library, making it far easier to write correct, robust multi-core code. * Formal Verification with SPARK: This is the game-changer. SPARK allows you to mathematically prove properties of your code. You can prove the absence of all runtime errors, prove that a function adheres to its dataflow, and even prove functional correctness against a specification. Imagine a kernel where you can prove a driver will never cause a buffer overflow or a scheduler will never deadlock. "But is anyone really using this?" Yes. While C dominates the mainstream, Ada/SPARK are the go-to for systems where failure is not an option. * NVIDIA uses SPARK to develop and verify security-critical firmware for its GPUs, protecting some of the most complex hardware on the market. * Airbus, Boeing, Lockheed Martin, and BAE Systems rely on Ada for safety-critical avionics (fly-by-wire, engine control, etc.). These are essentially real-time operating systems in their own right. * The Muen Separation Kernel is a prominent open-source kernel written in SPARK, designed for high-assurance virtualization. * Projects like the MaRTE OS (a real-time POSIX OS) and CubitOS demonstrate the feasibility of building entire, general-purpose-style operating systems in Ada. It's time to challenge the status quo. If you're building a new kernel, a secure enclave, a hypervisor, or critical firmware, C shouldn't be the only option on the table. What do you think? Have you used Ada/SPARK for systems programming? What's stopping our industry from adopting these safer languages more broadly? #OperatingSystems #Ada #SPARK #C #Cybersecurity #SoftwareEngineering #SafetyCritical
23
10 Comments
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content