You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Tiny engine, immense models — run large MoE LLMs (gpt-oss, Mixtral, Qwen3-MoE) on ordinary machines by streaming experts from disk. OpenAI-compatible server with tool calling + hybrid cloud relay; CPU, Apple Silicon (MLX) & CUDA.
(consider giving us a star!) Lightweight, feature‑packed DeepSeek app built with PyWebView, no Electron bloat. Includes all features of the most‑starred builds plus custom UI polish, dynamic greetings, animations, auto‑updater, enhanced Markdown rendering, and a injection js with lots of features :] <3
Pingstream is a lightweight, Streamlit-based API testing tool using curl under the hood. Designed for low-RAM systems, it supports GET, POST, PUT, DELETE with JSON, headers, file uploads, and Postman/OpenAPI import.
Surgical reasoning on consumer silicon. Hybrid SSM + causal memory architecture with entropy-gated System 1/2 dispatch, O(1) inference memory, and continual learning — designed for 16 GB VRAM.
Run an agentic coding LLM on the GPU you already own. No cloud, no invoice — just llama.cpp on Vulkan, a few glue scripts, and a 7–8B model that calls tools. This repo turns the whole thing into a LAN OpenAI-compatible endpoint that opencode (or any OpenAI client) can drive.
Eidetic — a single-process, single-database memory server for AI agents. SQLite-native: vectors + FTS5 + knowledge graph + raw logs. Five-layer memory (L0 raw → L1–L4), three-tier recall, auto-healing. One-command deploy, local-first, works with any MCP-compatible agent (OpenClaw, Claude Code, etc.).
A minimal Python shell optimized for low memory usage (~21 MB RAM). Features tab completion, command history, I/O redirections, and built-in commands compiled to a native binary with Nuitka.