Skip to content
#

model-offloading

Here are 8 public repositories matching this topic...

Language: All
Filter by language

Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.

  • Updated Oct 1, 2026
  • Swift

Quality-first local MiniMax H3 video-series generation and loopback API for dual RTX 4090 workstations, with native audio, references, P8/P9 continuity, and preserved artifacts.

  • Updated Sep 4, 2026
  • Python

Run BF16 LLMs larger than your GPU's memory without quantization: lossless weight compression, layer streaming, and speculative decoding. Code and frozen evidence for H6.5, schedule-aware exact weight placement.

  • Updated Sep 30, 2026
  • Python

Add this topic to your repo

To associate your repository with the model-offloading topic, visit your repo's landing page and select "manage topics."

Learn more