GSP: grammar-conditioned sparse LM-head projection - replication and SGLang integration feedback wanted #41638
StealthEyeLLC
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I am sharing an open research artifact that may be relevant to SGLang's structured-generation path.
GSP (Grammar-Conditioned Sparse Projection) consumes a packed grammar legality mask before/during the target-model LM-head projection instead of computing the full vocabulary projection and masking afterward.
The current artifact includes:
On the reference RTX 5060 Laptop GPU, 30 paired AB/BA trials gave median end-to-end ratios of 1.0756x (enum JSON), 1.0642x (tool-call JSON), and 1.0842x (nested-array JSON). The free-string control is not claimed as a reliable mean speedup.
This is not claiming first invention of batch-1 grammar-aware LM-head filtering; the contribution boundary is the packed-mask batched execution, heterogeneous routing, and controlled target-model evaluation.
I would especially value:
Release, verified preprint, source, raw evidence:
https://github.com/StealthEyeLLC/grammar-sparse-projection/releases/tag/v0.1.1
All reactions