Sign in to view Simon’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Simon’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Berkeley, California, United States
Sign in to view Simon’s full profile
Simon can introduce you to 10+ people at Inferact
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
5K followers
500+ connections
Sign in to view Simon’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Simon
Simon can introduce you to 10+ people at Inferact
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Simon
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Sign in to view Simon’s full profile
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Experience & Education
-
Inferact
*** *** *********
-
****
****
-
********** ** *********** ********
****** ** ********** * *** undefined undefined
-
-
********** ** *********** ********
********** ****** undefined
-
View Simon’s full experience
See their title, tenure and more.
Welcome back
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
New to LinkedIn? Join now
or
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View Simon’s full profile
-
See who you know in common
-
Get introduced
-
Contact Simon directly
Other similar profiles
Explore more posts
-
George Z. Lin
5K followers
UniFreiburg and MSR team have explored the challenges associated with state-tracking in neural sequence models, particularly in the context of code execution. It highlights the limitations of traditional sequence-to-sequence tasks, which often do not align with the next-token prediction paradigm commonly used in training LLMs. By converting permutation composition tasks into code-like representations through Python REPL traces, the study establishes a more realistic framework for state-tracking. The findings indicate that linear recurrent neural networks (RNNs), particularly models like DeltaNet with extended eigenvalue ranges, demonstrate strong performance in state-tracking tasks, even with sparse supervision. In contrast, Transformers tend to struggle in these scenarios, requiring dense supervision to maintain their performance, which becomes challenging when state reveals are sparse or adversarial. The research introduces a formal model known as the Probabilistic Finite-State Automaton with State Reveals (PFSA-SR), designed to capture the complexities of state transitions and partial observability in code execution. The study notes that while linear RNNs can perform well in deterministic environments, they encounter significant challenges in probabilistic settings, particularly due to issues such as norm decay that can result in the loss of state information. Additionally, linear RNNs may underperform compared to their nonlinear counterparts in contexts involving probabilistic transitions. The implications of this research are relevant across various domains, including programming, game-playing, and world modeling, where effective state-tracking is crucial. This research contributes to a deeper understanding of state-tracking capabilities across different neural architectures and suggests that future work could focus on enhancing nonlinear RNNs and investigating hybrid architectures that integrate linear recurrence with periodic nonlinear normalization. Such advancements could lead to improvements in probabilistic state-tracking benchmarks and their practical applications. Arxiv: https://lnkd.in/emaQySJF
2
-
Koustava Goswami
Adobe • 6K followers
🚀 Excited to share our new paper — “Decomposition-Enhanced Training for Post-Hoc Attributions in Language Models” (https://lnkd.in/guHftWQH) We revisit attribution through a reasoning lens — reframing it as a decomposition-based reasoning task rather than post-hoc evidence retrieval. Our method, DECOMPTUNE, introduces a new framework designed specifically for structured reasoning supervision. 🔹 Novelty in RL design: We build a decomposition-aware reward model that scores both logical step quality and evidence grounding consistency. The RL signal integrates multiple reasoning metrics — correctness of intermediate steps, coherence of decomposition structure, and source alignment fidelity — creating a continuous gradient toward faithful reasoning. 🔹 Training pipeline: 1️⃣ Stage 1: Supervised fine-tuning (SFT) on decomposition traces from a strong LLM teacher. 2️⃣ Stage 2: GRPO-based reinforcement learning that jointly optimises for structural reasoning quality and attribution grounding. 🔹 Outcome: Our Qwen 7B and 14B models—trained entirely with this SFT + RL pipeline on a small-scale dataset—demonstrate GPT-level attribution reasoning, outperforming the GPT and base models by over +23 points and even surpassing larger open and closed-source baselines in faithfulness and interpretability. This work shows that scaling isn’t the only path to reasoning: with the right RL-driven training signal, small models can learn structured, interpretable reasoning that rivals massive systems. Great work led by Sriram Balasubramanian during the summer internship, with Samyadeep Basu and other co-authors. Excited to hear from the community on the work. #Reasoning #ExplainableAI #LLMs #NLP #DeepLearning
58
1 Comment -
Maarten Van Segbroeck, Ph.D.
NVIDIA • 4K followers
Our technical report on NVIDIA's NeMo Data Designer is now on arXiv. Data Designer is our open-source framework (Apache 2.0) for generating high-quality synthetic datasets. It is declarative by design: specify what you want, we handle the how, such as execution order, batching, parallelization at scale. To date Data Designer has processed 25T+ tokens (19T context, 6T+ generated) for teams building and fine-tuning models. This report details the architecture behind it: an extensible, multimodal generation engine spanning text, code, structured data, images, and embeddings, with a plugin system built for teams who need to customize and extend it to their own use cases. Congrats to everyone who contributed: Johnny Greco Nabin Mulepati Andre M. Eric Tramel Kirit Thadaka Michael Knepper Dhruv Nathawani Dane Corneil (PhD) Yev Meyer, Ph.D. Alex Watson and many others. Read the paper: https://lnkd.in/gbT_cPXs
143
3 Comments -
Hung-Yueh Chiang
NVIDIA • 1K followers
Throughout my Ph.D., I have developed numerous profiling tools to understand and diagnose latency bottlenecks in deep learning models. Most existing commercial profilers are either designed for very specific purposes and are difficult to customize, or they offer complex features with steep learning curves. As a result, building profiling tools often becomes one of the most painful parts of research. Surprisingly, almost every researcher ends up reinventing this wheel to showcase latency improvements in their own work. Encouraged by my advisor, Prof. Diana Marculescu, we decided to consolidate these tools and open-source them to benefit future researchers in this area. We compiled and released them as ELANA: A Simple Energy & Latency Analyzer for LLMs: 📄 Technical report: https://lnkd.in/g8HmscPP 💻 Github repo: https://lnkd.in/gDUA2_Ny ELANA is a lightweight, academic-friendly profiler built on Hugging Face APIs for analyzing LLM latency and energy efficiency. We hope it will serve as a foundation for reproducible benchmarking, enable fair comparisons across models and systems, and accelerate research on resource-efficient LLMs. Huge thanks to Bokun Wang and Prof. Diana Marculescu for their invaluable support on both the tool and the technical report.
50
1 Comment -
PyTorch
330K followers
We’re excited to welcome Ray to the PyTorch Foundation 👋 Ray is an open source distributed computing framework for AI workloads, including data processing, model training and inference at scale. By contributing Ray to the PyTorch Foundation, Anyscale reinforces its commitment to open governance and long-term sustainability for Ray and open source AI. “The PyTorch Foundation is committed to fostering an open, interoperable, and production-ready AI ecosystem. By bringing Ray under the PyTorch Foundation umbrella, alongside projects like vLLM and DeepSpeed, we are uniting the critical components needed to build next-generation AI systems. Ray’s inclusion strengthens our collective mission to support developers with the tools to efficiently train, serve, and deploy AI models at scale.” - Matt White GM of AI at The Linux Foundation Foundation and Executive Director of the PyTorch Foundation “With Ray, our goal is to make distributed computing as straightforward as writing Python code. Joining the PyTorch Foundation helps us stay true to that mission, ensuring Ray continues to be an open, community-driven backbone for developers and their organizations.” - Robert Nishihara, co-founder of Anyscale ➡️ Read the announcement: https://lnkd.in/e3e4mjds #Ray #PyTorchCon #PyTorch #AI #OpenSource #AICompute #DistributedComputing #AIWorkloads #DataProcessing #ModelTraining #Inference
669
5 Comments -
Gayathri G
elsai • 4K followers
#HunyuanOCR sets a new standard in end-to-end OCR. Built on Hunyuan’s native multimodal architecture, this lightweight 1B-parameter VLM is delivering state-of-the-art results across multiple industry benchmarks. What stands out is its ability to handle complex, multilingual document parsing with ease while staying efficient enough for real-world deployment. It performs strongly across a range of practical scenarios — from text spotting and open-field information extraction to video subtitle extraction and on-the-fly photo translation. A compact model with serious capability, and a solid step forward for production-grade OCR. Link : https://lnkd.in/gYwHvvAj
27
-
Nishantha Ruwan
IWROBOTX Software Inc. • 2K followers
The paper introduces VisPlay, a novel reinforcement-learning framework designed to allow vision-language models (VLMs) to improve themselves using only unlabeled image data. Existing RL approaches for VLMs typically require costly human annotations or carefully crafted heuristics to provide verifiable rewards, which limits scalability. VisPlay circumvents this by decomposing a base VLM into two interacting roles: an Image-Conditioned Questioner, which generates challenging yet answerable questions conditioned on images, and a Multimodal Reasoner, which attempts to answer those questions. These two agents are trained jointly via a technique called Group Relative Policy Optimization (GRPO), which uses intrinsic rewards—difficulty and diversity of questions, quality of answers—to drive co-evolution without external labels. In experiments the authors apply VisPlay to different VLM architectures (e.g., Qwen2.5-VL and MiMo-VL) and evaluate across multiple benchmarks covering general visual understanding, compositional reasoning, visual math, and hallucination detection. They report consistent performance gains in reasoning accuracy and generalization, along with reductions in hallucination compared to baseline training. The results suggest that VisPlay offers a scalable path toward autonomous multimodal intelligence, enabling VLMs to evolve their reasoning capabilities beyond the limits of human-curated supervision. The paper also acknowledges limitations: the experiments are confined to certain model families and there remains a challenge in certifying the correctness of self-generated training data. https://lnkd.in/gTQ2-UkG
-
Alan Kochukalam George
Myovine • 618 followers
⚙️ One library, real DSP speed. There’s a gap in the developer ecosystem. Teams doing DSP/ML often migrate to Python, because that’s where most tooling lives. But Python wasn’t built for high-throughput I/O — under load, it spawns new interpreters/processes, creating duplicated memory, slow cold starts, and more containers. Compute stays fast, but orchestration cost explodes. Meanwhile, Node.js/TS scale I/O well, but were never meant for serious compute. Pure-JS DSP hits performance walls. Even worse, serializing data between Redis ↔ DSP/ML adds latency and forces batching — slowing real-time pipelines. So devs are stuck: Python → strong math, weak concurrency Node → strong networking, weak compute dspx closes that gap — TypeScript DX + native C++/SIMD (AVX2/SSE3/SSE2/NEON) under the hood. 🧩 How dspx works FIR / FFT / Conv1D run in optimized C++/SIMD Memory-safe circular buffers → O(1) throughput TS handles Kafka / Redis / WebSockets Redis persistence < 0.5 ms Batched logging avoids I/O stalls At tiny batch sizes, N-API overhead can make naive JS look competitive. At realistic scale, batched native pipelines win. ⚡ Benchmarks Dell OptiPlex 3000 Micro (i5-12600T · AVX2 · Node 22) FFT: 2.6× faster than fft.js, ~9× faster than tfjs-node FIR (51-tap): ~4× faster than fili / naive JS Conv1D (128-kernel): ~3.5× faster than tfjs-node Moving Avg (O(1)): 559× more throughput/sec than naive JS Redis save/load: sub-ms Logging: <3% overhead 📊 Full tables + charts in carousel 💬 Observations SIMD FMA + contiguous memory → big FIR/Conv wins O(1) moving-average design → massive throughput gain Sub-ms Redis + low I/O overhead → real-time persistence in pure Node 🧠 Open Source dspx is Apache 2.0, free for commercial/academic use. Looking for contributors interested in: ARM NEON / Apple M / Graviton tuning Audio / sensor / biomedical DSP Visualization + benchmark tooling 📦 npm i dspx 🔗 https://lnkd.in/e-tWxAgu 🧠 Real-time DSP for Node, TypeScript & Redis 💼 Note I’m exploring opportunities in full-stack, real-time systems, performance engineering, and DSP. If your team works in this space, I’d love to connect. My next post will cover why sub-ms Redis latency isn’t just technical — it’s economic. Lower serialization + compute overhead → lower infra + energy cost. 🔖 Tags & Mentions #NodeJS #TypeScript #DSP #PerformanceEngineering #EdgeComputing #OpenSource #SIMD #Redis #Cplusplus #RealTime NodeJS Developer TensorFlow Google Microsoft JavaScript Developer Amazon Web Services (AWS) Vercel
5
1 Comment -
Russ Salakhutdinov
Sooth Labs • 10K followers
Check out the new Machine Learning Department at CMU blogpost: How to Explore to Scale RL Training of LLMs on Hard Problems? https://lnkd.in/egihtUvu Current on-policy RL methods fail to learn from hard problems as they rarely generate a single correct rollout, producing no reward signal and no learning. Including easy problems can also be harmful, as models tend to overfit to them and fail to improve on harder tasks. Distilling human-written solutions is not only costly, but also provides difficult targets for fine-tuning. This blogpost discusses various approaches and introduces a framework that uses existing human or model solutions as privileged guidance to unlock learning on hard problems. The key idea is simple: Prepend a minimal solution prefix to difficult prompts, enabling on-policy RL to obtain reward and learn behaviors that generalize back to the original, unconditioned tasks. This expands the set of solvable problems and results in significant gains on challenging reasoning benchmarks. With YUXIAO QU, Amrith Setlur, Virginia Smith, and Aviral Kumar. Paper/Code is coming soon.
172
1 Comment
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content