Jevons paradox states that making a resource more efficient to use leads to an overall increase in the total consumption of that resource. This was an observation made in the late 18th century. When steam engines became more fuel-efficient, coal consumption did not drop. It skyrocketed. Cheaper power made steam engines practical for thousands of new industries that could never afford them before.
The exact same observation now governs AI coding agents on the cloud.
When an AI agent is unguided, every cloud interaction is expensive. The agent dumps hundreds of thousands of tokens into prompt context, crashes on interactive CLI prompts, guesses deprecated SDK methods, and burns tokens in error recovery loops. Faced with this friction, developers keep their prompts small and defensive. They ask the agent for isolated helper functions and write the cloud plumbing themselves.
Now there are agent plugins. With plugins, cost and friction actually plummet because they bundle skill discovery, grounded documentation and surgical CLI output formatting. Because of this, developers respond immediately. They do not use the agent less. They give the agent larger, end-to-end cloud orchestration tasks: building, wiring, and deploying multi-product pipelines across Cloud Run, Firestore, and BigQuery in a single turn.
The Google Cloud Developer Plugin, built on the open Agent Plugins specification, provides the operational foundation for this transition across Antigravity, Claude Code, and Codex.
This article explores the architecture of the Google Cloud Developer Plugin and how it eliminates context bottlenecks. It then walks through an end-to-end scenario, guiding an AI coding agent to build and deploy a multi-service pipeline across Cloud Run, Firestore and BigQuery.
Author: Balaji Subramaniam, DevRel Engineer, Google Cloud
The context bottleneck in cloud development
Building cloud backends requires deep context. But dumping all that context into an LLM prompt upfront ruins performance. In practice, unguided agents hit three bottlenecks:
Prompt bloat
Standard agent harnesses use progressive disclosure by preloading skill metadata (names and descriptions) on Turn 0 and deferring full instruction bodies until activation. But across large skill catalogs like the Google Cloud skills, even preloading metadata headers pollutes system prompt context with thousands of tokens on every turn.
Observation bloat
Unfiltered terminal commands flood the context window. Running a basic command like gcloud run services list or gcloud compute instances list without format projections returns massive JSON or YAML payloads packed with cluster configs, internal metadata, and verbose status blocks. An agent ingesting raw CLI dumps fills its context buffer in just a few turns.
Hallucination loop
Foundation models rely on frozen training data. When an agent tries to write deployment scripts or SDK code, it may invent obsolete flags or import deprecated library methods. The CLI rejects the command, the agent parses the error, hallucinates another bad flag, and gets trapped in a multi-turn failure spiral.
How the Google Cloud Developer Plugin eliminates the context bottleneck
The google-cloud-developer plugin organizes cloud capabilities into four disciplined layers:
1. Multi-tier progressive disclosure (finding-google-skills)
The google-cloud-developer plugin externalizes the broader catalog behind a compact discovery router (~1,680 tokens). Context is revealed in three disciplined tiers:
- Tier 1 (discovery router & foundational bundle): The agent holds only foundational cloud guardrails and the discovery router. Zero remote catalog metadata is preloaded in the prompt.
- Tier 2 (on-demand catalog discovery): When a request requires specialized services, the router searches the central Google Agent Skills repository and pulls only the target SKILL.md instructions into active context.
- Tier 3 (execution grounding via MCP): Granular SDK parameter syntax and schemas are retrieved on demand via the Developer Knowledge MCP server.
2. Grounded documentation via MCP (retrieving-developer-knowledge)
The plugin connects the agent to the Developer Knowledge MCP server (developerknowledge.googleapis.com). Through tools like answer_query and search_documents, the agent queries official, verified documentation in real time. It looks up exact constructor arguments, client methods, and configuration keys before writing a single line of code.
3. Execution guardrails & data reduction (gcloud skill)
The plugin enforces strict CLI execution rules:
- Leaf Help Validation: The agent runs gcloud help <command> before executing new commands to verify valid parameters.
- Non-Interactive Flags: Every command must include --quiet (or -q) so processes never freeze waiting for an interactive TTY confirmation.
- Strict Output Projection: Commands must specify --format="json(...)" and --limit=5. This compresses megabytes of raw CLI output into concise JSON payloads containing only the exact fields required.
4. Credential safety (google-cloud-recipe-auth)
The plugin steers the agent away from generating long-lived service account keys or hardcoding secrets in code. It guides the workflow toward Application Default Credentials (ADC) and short-lived IAM role impersonation.
End-to-end walkthrough: Real-time event ingestion pipeline for payload classification
The Jevons Paradox becomes visible when you watch developers prompt an agent equipped with the plugin. Instead of asking for a snippet of Firestore code, the developer asks for a complete multi-product architecture.
Install the Google Cloud Developer plugin
Follow these instructions to install the Google Cloud Developer plugin in the coding agent ecosystem of your choice.
The developer prompt
Submit the following prompt to your coding agent (Antigravity CLI, Claude Code, or Codex):
Build and deploy a lightweight, high-throughput event classification and routing microservice on Cloud Run.
The service exposes a POST /events endpoint that receives incoming JSON event payloads, performs fast in-memory classification to assign a priority tier, persists live document state in Cloud Firestore, and streams analytics records into BigQuery.
Incoming Event Payload Schema:
{
"event_id": "evt-84920-x",
"event_type": "payment_authorization",
"risk_score": 0.85,
"user_tier": "enterprise",
"metadata": {
"source_region": "us-east1",
"amount_usd": 1450.00
}
}
Requirements:
1. Cloud Run Service (FastAPI):
- Expose GET /healthz and POST /events.
- Parse and validate incoming event payloads.
- Classification Logic: Evaluate the payload and classify priority:
- Assign "HIGH" if risk_score > 0.7 or user_tier is "critical".
- Assign "STANDARD" for all other events.
2. Cloud Firestore (Real-Time State):
- Write/update the classified event in the 'active_events' collection keyed by event_id.
- Fields: event_id, priority (HIGH/STANDARD), risk_score, status ("PROCESSED"), and updated_at (UTC ISO timestamp).
3. BigQuery (Streaming Analytics):
- Stream analytical metrics to 'analytics_dataset.events'.
- Row schema: event_id, priority, risk_score, payload_size, timestamp.
4. Security & Deployment:
- Provide non-interactive gcloud deployment commands following Google Cloud safety guardrails:
- Explicit --project and --region.
- Non-interactive flag (--quiet).
- Clean format projection (--format="json(status.url)").
- Least-privilege IAM service account (--no-allow-unauthenticated).
5. Validation via Plugin:
- Use the Google Cloud Developer Plugin to verify SDK method signatures via Developer Knowledge MCP and validate CLI syntax with leaf help before generating commands.Step-by-step agent execution
Here is how the plugin coordinates the agent's actions:
Step 1: Dynamic discovery
The agent identifies the requirement for a multi-service event pipeline. Using its discovery router and foundational cloud rules, it checks whether specialized cloud solution patterns exist in the skills repository.
Step 2: Live grounding over MCP
The agent calls answer_query on the Developer Knowledge MCP server:
- Query: "google-cloud-bigquery python insert_rows_json best practices and error handling"
- Result: The MCP server returns the exact signature for bigquery.Client().insert_rows_json(table, rows). No deprecated methods.
Step 3: Syntax and guardrail validation
Before proposing deployment commands, the agent runs leaf validation:
gcloud help run deployIt confirms flag availability and constructs a guarded command using explicit parameters, non-interactive mode, and projection formatting:
gcloud run deploy event-routing-service \
--source . \
--region us-central1 \
--project my-project-id \
--no-allow-unauthenticated \
--service-account [email protected] \
--quiet \
--format="json(status.url)"Step 4: Clean code generation
The agent produces a self-contained Python implementation connecting FastAPI, Firestore, and BigQuery.
Build your next cloud architecture
The Google Cloud Developer Plugin and Data Agent Kit are available now in the open Google Agent Skills repository.
Try running the event ingestion walkthrough in your own environment with Antigravity CLI, Claude Code, or Codex. Explore the skills catalog, test custom service combinations, and let us know what end-to-end architectures and pipelines you are building with the Google Cloud Developer Plugin.



