Skip to content
View rayleizhu's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report rayleizhu

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. Tencent-Hunyuan/Simple-Attention-Sparsification Tencent-Hunyuan/Simple-Attention-Sparsification Public

    Code for our resarch paper "SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking"

    Python 63 1

  2. sglang sglang Public

    A fork of official SGLang that supports block-sparse attention. The backend and two algorithms (SeerAttention-R and MoBA) are implemented.

    Python 10

  3. BiFormer BiFormer Public

    [CVPR 2023] Official code release of our paper "BiFormer: Vision Transformer with Bi-Level Routing Attention"

    Python 586 42

  4. vllm-ra vllm-ra Public

    [ACL 2024] RelayAttention for Efficient Large Language Model Serving with Long System Prompts

    Python 40 6

  5. GLMix GLMix Public

    [NeurIPS 2024] official code release for our paper "Revisiting the Integration of Convolution and Attention for Vision Backbone".

    Python 43 3

  6. docker-cuda-codeserver docker-cuda-codeserver Public

    Forked from works-on-my-machine/pytorch-code-server

    Docker code server with CUDA development environment.

    Shell 3