Skip to content

Latest commit

 

History

7,238 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

⚠�� Nightly release for early testing. Expect rough edges. Stable version coming out soon — please open an issue if you hit anything.

Future AGI — make AI agents reliable

AI Agents hallucinate. Fix it faster.

The open-source platform for shipping self-improving AI agents. Evaluations, tracing, simulations, guardrails, gateway, optimization. Everything runs on one platform and one feedback loop, from first prototype to live deployment.

Apache 2.0 License PyPI npm Discord

Try Cloud (Free) · Self-Host · Docs · Blog · Discord · Discussions



Why Future AGI?

Most AI agents fail in production, and teams end up stitching together evals, observability, and guardrails that never close the loop. Future AGI collapses all of it into one platform and one feedback loop. Simulate edge cases before launch, evaluate what happens in production, protect users in real time, and turn every trace into signal for the next version. The result: agents that don't just get monitored, they self-improve.

All-in-one

No more stitching Langfuse + Braintrust + Helicone + Guardrails AI + a custom simulator. One platform covers the lifecycle: simulate → evaluate → protect → monitor → optimize, with data flowing back as a loop.

Open & self-hostable

Apache 2.0 core. Every evaluator, every prompt, every trace is inspectable — no black-box scoring. Self-host for data sovereignty or use our managed Cloud. Drop in your own stack at any layer via OTel / OpenAI-compatible HTTP.

Built for production

Go-based gateway with ~9.9 ns weighted routing, ~29 k req/s on t3.xlarge, P99 ≤ 21 ms with guardrails on. OpenTelemetry-native traces. 50+ framework instrumentors. Every claim reproducible via the committed benchmark harness.


🚀 Quickstart

Run the whole platform on your own machine in three steps. Rather not run anything? Try Cloud free.

You need Docker Desktop or Docker Engine with Compose v2.24 or newer, and for the default Standalone setup 2 vCPUs and 4 GB of memory given to Docker.

1. Install

git clone https://github.com/future-agi/future-agi.git
cd future-agi
./bin/install          # Windows (PowerShell): .\bin\install.ps1

The installer checks your machine, writes this install's secrets to .env, downloads the images (about 800 MB) and waits until everything answers. The first boot sets up the database and takes a few minutes; http://localhost:3000 shows its progress meanwhile.

2. Sign in and copy your keys

Open http://localhost:3000, create your account (or sign in with the one the installer made), then copy the API key and secret key from Keys in the sidebar.

3. Send your first trace

pip install fi-instrumentation-otel
export FI_API_KEY="<your API key>" FI_SECRET_KEY="<your secret key>" FI_BASE_URL="http://localhost:4318"
from fi_instrumentation import register
from fi_instrumentation.fi_types import ProjectType

tracer_provider = register(project_name="my-first-project", project_type=ProjectType.OBSERVE)
with tracer_provider.get_tracer("quickstart").start_as_current_span("hello-future-agi") as span:
    span.set_attribute("input.value", "Hello, Future AGI")
tracer_provider.force_flush()

Open Tracing in the sidebar: my-first-project holds your first span. http://localhost:4318 is this install's trace collector; it listens on this machine only.

Standalone (default) Distributed (at scale)
Install ./bin/install ./bin/install --distributed
Runs one app container, Postgres, ClickHouse one container per service, PeerDB, Kafka
Docker resources 2 vCPUs, 4 GB 4+ vCPUs, 12–16 GB
Kubernetes Helm chart: about 4 CPUs and 8 GiB free for an evaluation

Choose before you add data: there is no supported way to move a Standalone install's data to Distributed or Helm later (Switching).

  • Configure: every variable in .env is described in the configuration reference. Nothing is required for a local install; add LLM provider keys, a public URL or email when you need them.
  • Manage: docker compose logs -f app (Distributed: backend); stop with docker compose down or ./bin/uninstall (both keep your data); upgrade with git pull && ./bin/install; remove everything, data included, with ./bin/uninstall --purge.
  • Develop: on a branch other than main, add --from-source to build the images from your checkout; ./bin/dev runs it with hot reload (Local development).
  • More: the self-hosting guide, INSTALLATION.md (every option and troubleshooting) and deploy/README.md (production).

Instrument your first agent

Swap the hand-made span for an instrumentor to trace a real app, here OpenAI (pip install traceai-openai). Against your own install, keep the FI_* variables from step 3 set.

Python

from fi_instrumentation import register
from traceai_openai import OpenAIInstrumentor

register(project_name="my-agent")
OpenAIInstrumentor().instrument()

# Your existing OpenAI code is now traced.
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": query}],
)

TypeScript

import { register } from "@traceai/fi-core";
import { OpenAIInstrumentation } from "@traceai/openai";

register({ projectName: "my-agent" });
new OpenAIInstrumentation().instrument();

// Your existing OpenAI code is now traced.
const response = await openai.chat.completions.create({
  model: "gpt-4o",
  messages: [{ role: "user", content: query }],
});

Full docs → · Cookbooks → · API reference →


Core features

Six pillars. Each one replaces a tool you probably have.

🧪 Simulate

Thousands of multi-turn conversations against realistic personas, adversarial inputs, and edge cases. Text and voice (LiveKit, VAPI, Retell, Pipecat).

Docs →

📊 Evaluate

50+ metrics under one evaluate() call: groundedness, hallucination, tool-use correctness, PII, tone, custom rubrics. LLM-as-judge + heuristic + ML.

Docs →

🛡️ Protect

18 built-in scanners (PII, jailbreak, injection, …) + 15 vendor adapters (Lakera, Presidio, Llama Guard, …). Inline in gateway or standalone SDK.

Docs →

👁️ Monitor

OpenTelemetry-native tracing across 50+ frameworks (LangChain, LlamaIndex, CrewAI, DSPy…). Span graphs, latency, token cost, live dashboards. Zero-config.

Docs →

🎛️ Agent Command Center

OpenAI-compatible gateway. 100+ providers, 15 routing strategies, semantic caching, virtual keys, MCP, A2A. ~29k req/s, P99 ≤ 21ms with guardrails on.

Docs → · Benchmarks →

🔁 Optimize

Six prompt-optimization algorithms (GEPA, PromptWizard, ProTeGi, Bayesian, Meta-Prompt, Random). Production traces feed back as training data.

Docs →


Deployment options

Target Status Notes
Docker Compose: Standalone ✅ ./bin/install: one app container next to Postgres and ClickHouse, for a laptop or a single VM
Docker Compose: Distributed ✅ ./bin/install --distributed: one container per service, for scale on one host
Production Compose overlay ✅ ./deploy/setup.sh on the Distributed setup: --skip-up writes deploy/.env.production (generating only the secrets you do not supply; you give every image version), then, once the databases are initialized, --confirm-initialized pulls the images and starts the stack (deploy/README.md)
Kubernetes / Helm ✅ Distributed on Kubernetes: helm install futureagi oci://ghcr.io/future-agi/charts/futureagi --version X.Y.Z, one signed chart for the open-source and Enterprise editions (chart README)
AWS / GCP / Azure ✅ Docker Compose on a VM, or the Helm chart on a Kubernetes 1.27+ cluster
AWS Marketplace ⏳ Coming soon
Air-gapped / on-prem ✅ Mirror the images, set FUTURE_AGI_TELEMETRY_DISABLED=true and block outbound traffic (Telemetry); on Helm, set global.airgap=true and mirror the images the release lists (chart README); contact sales for support

Every image, tag and size: Container images. Every setting: Configuration reference.


Architecture

Every arrow is an open, documented interface: OpenTelemetry OTLP for traces, OpenAI-compatible HTTP for the gateway, Postgres / ClickHouse SQL for storage. Drop in your own stack at any layer.

Runtime: Python 3.11+ (Django 5.1 + Channels) · Go 1.25+ (gateway), 1.24+ (trace collector) · React 18 + Vite · Node 22.18+. Data: PostgreSQL (metadata) · ClickHouse (spans + time-series) · Redis (state, live updates) · Temporal (jobs).

Component breakdown (per-package)
Layer Component Code
Edge traceAI — OpenTelemetry instrumentation future-agi/traceAI
Edge Agent Command Center — OpenAI-compatible proxy agentcc-gateway/
Platform tracer — OTLP ingest, span graph futureagi/tracer/
Platform agentic_eval — 50+ metrics, LLM-as-judge futureagi/agentic_eval/
Platform simulate — persona-driven scenario generation futureagi/simulate/
Platform model_hub — LLM routing, embeddings, datasets futureagi/model_hub/
Platform accounts · usage · integrations — auth, orgs, metering, connectors futureagi/accounts/
Data PostgreSQL · ClickHouse · Redis · Temporal —

SDKs & integrations

Future AGI is an open-source ecosystem — each SDK is independently usable, independently packaged, Apache/MIT-licensed.

Client libraries

Repo Install Languages Purpose
traceAI pip install fi-instrumentation-otel
npm i @traceai/fi-core
Python · TS · Java · C# Zero-config OTel tracing for 50+ AI frameworks
ai-evaluation pip install ai-evaluation
npm i @future-agi/ai-evaluation
Python · TS 50+ evaluation metrics + guardrail scanners
futureagi pip install futureagi Python Platform SDK — datasets, prompts, KB, experiments
agent-opt pip install agent-opt Python 6 prompt-optimization algorithms (GEPA, PromptWizard, …)
simulate-sdk pip install agent-simulate Python Voice-agent simulation via LiveKit + Silero VAD
agentcc pip install agentcc
npm i @agentcc/client
Python · TS (+ LangChain · LlamaIndex · React · Vercel) Gateway client SDKs

Integrations

LLM providers (100+) OpenAI · Anthropic · Google Gemini · Vertex AI · AWS Bedrock · Azure OpenAI · Mistral · Groq · Cohere · Together · Perplexity · OpenRouter · Fireworks · xAI · Replicate · HuggingFace · + self-hosted Ollama · vLLM · LM Studio · TGI · Llamafile
Agent frameworks LangChain · LangGraph · LlamaIndex · CrewAI · AutoGen · Phidata · PydanticAI · Claude SDK · LiteLLM · Haystack · DSPy · Instructor · Smol-agents
Voice platforms VAPI · Retell · LiveKit · Pipecat
Vector DBs Pinecone · Weaviate · Chroma · Milvus · Qdrant · pgvector
Tools & infra Vercel AI SDK · n8n · MongoDB · MCP · A2A · Guardrails AI · Langfuse · HuggingFace Smol-agents

Full integrations catalog →


How Future AGI compares

Future AGI Langfuse Phoenix Braintrust Helicone
Open source✅
Apache 2.0
✅
MIT
✅
Elastic v2
❌✅
Apache 2.0
Self-host✅✅✅❌✅
LLM tracing (OpenTelemetry)✅✅✅✅⚠️
via OpenLLMetry
Evaluation suites✅
50+ metrics
✅✅✅⚠️
Limited
Agent simulation✅❌❌❌❌
Voice agent eval✅❌⚠️
Cookbook
❌❌
LLM gateway built in✅
100+ providers
❌❌✅✅
Guardrails built in✅
18 + 15 adapters
❌❌❌❌
Prompt optimization✅
6 algorithms
❌❌❌❌
Prompt management✅✅✅✅✅
Datasets & experiments✅✅✅✅✅
No-code eval builder✅⚠️⚠️⚠️⚠️

Based on publicly-documented features as of April 2026. Corrections welcome — open a PR.


Built for every kind of agent

  • Customer Support: Ship support AI that customers actually trust
  • Voice Agents: Test, evaluate, and improve voice AI end-to-end
  • Internal Tools: AI copilots your whole org can rely on
  • RAG & Search: Every answer grounded, every citation verified
  • Autonomous Agents: Multi-step agents you can actually trust in production
  • Computer-Use Agents (CUA): Agents that click with confidence
  • Coding Agents: AI that writes code you can actually ship

Roadmap

Vote on the public roadmap → · GitHub Discussions · Releases · Changelog

Recently shipped In progress Coming up Exploring
  • Prompt optimization engine
  • Taxonomy-based Feed Clustering
  • Agent Runs in Dataset Experiments
  • Simulate from Production Calls
  • LiveKit Configuration via UI
  • System Metric Filtering for Voice
  • Agent Playground
  • Dashboards
  • Access platform via MCP
  • Annotation Queues
  • Command Center
  • Open source Future AGI stack
  • Eval Explanation Output Size Control
  • Agent Changelog & Diff View
  • Smart Queue Assignment
  • Essential Node Library for Agent Builder
  • Full Execution Tracing for Agents
  • Multi-modal Support for Agents
  • Agent Changelog & Diff View
  • Smart Queue Assignment
  • Import agents to Agent Playground
  • Simulating CUA agents
  • Simulating Coding agents
  • Scheduled Simulations

🤝 Contributing

We love contributions — bug fixes, new evaluators, framework integrations, docs, examples, anything.

  1. Browse good first issue
  2. Read the Contributing Guide
  3. Say hi on Discord or Discussions
  4. Sign the CLA on your first PR (automatic bot)

🌍 Community & support

💬 Discord Real-time help from the team and community
🗨️ GitHub Discussions Ideas, questions, roadmap input
🐦 Twitter / X Release announcements
📝 Blog Engineering & research posts
📺 YouTube Walkthroughs & demos
📊 Status Cloud uptime + incident history
📧 support@futureagi.com Cloud account / billing
🔐 security@futureagi.com Private vulnerability disclosure (24h ack on weekdays — see SECURITY.md)

Telemetry

Self-hosted Future AGI sends deployment telemetry, on by default, so we can count installs and size release testing: one registration with the email addresses and email domains of the install's owner, admin and staff/superuser accounts, then usage counts on a schedule. No trace data, no prompts, no completions, no datasets, no API keys, ever.

To opt out, install with ./bin/install --no-telemetry, or set FUTURE_AGI_TELEMETRY_DISABLED=true in .env (deploy/.env.production for the production overlay, config.telemetry=false for Helm) and run docker compose up -d. Opting out still sends one registration, without email addresses; block api.futureagi.com to send nothing. Everything else that could leave your install (HubSpot, Slack, Mixpanel, PostHog, reCAPTCHA, Sentry, Mailgun) is off until you set its key.

Telemetry and outbound connections has the exact payloads, what the opt-out still sends, every setting, and every outbound connection, with what an install that allows no outbound traffic must also set.


⭐ Star history

Star history

📄 License

Future AGI is licensed under the Apache License 2.0. See LICENSE and NOTICE.

You own your evaluation logic and your data. Inspect every evaluator, every prompt, every trace — no black-box scoring, no vendor lock-in.


Built with ❤️ by the Future AGI team and contributors worldwide.

If Future AGI helps you ship better AI, a ⭐ helps more teams find us.

🌐 futureagi.com · 📖 docs.futureagi.com · ☁️ app.futureagi.com · 📊 status.futureagi.com

About

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2.1k stars

Watchers

9 watching

Forks

Releases

Packages

Contributors

Languages