Today we're releasing IQuest-Q1 and opening the model weights. 320B total. 15B active. Built for code, software engineering, and complex agentic tasks. Technical report/HuggingFace/GitHub — available now.
RL training run misbehaves. The reward curve looks wrong.
The model reads the curves, pulls the logs, forms a hypothesis, finds the bug — a stray whitespace breaking the generation chain — and patches it.
Mean reward: 0.704 → 0.769.
IQuest-Q1 is trained in three stages:
Pretraining → mid-training → post-training.
Post-training focuses on software engineering, long-horizon agentic tasks and general reasoning, combining SFT + reinforcement learning.
Numerically solved the equations of motion (100k steps). Rendered the three orbits as glowing silk threads in Three.js.
Klimt-style. Gold, ochre, dark brown. Click anybody to inspect mass and velocity.
Prompt: a multiplayer 3D FPS set in a children's bedroom.
One generation: Full game. Boss raid, in-game economy, HUD, lobby, building-block waiting room, paper-airplane spectator view after death.
No scaffolding. No iteration.😀