Article cover image

The Jevons Paradox of AI coding agents: How reducing developer friction scales AI adoption

Google Cloud Tech
Google Cloud Tech@GoogleCloudTech

Jevons paradox states that making a resource more efficient to use leads to an overall increase in the total consumption of that resource. This was an observation made in the late 18th century. When steam engines became more fuel-efficient, coal consumption did not drop. It skyrocketed. Cheaper power made steam engines practical for thousands of new industries that could never afford them before.

The exact same observation now governs AI coding agents on the cloud.

When an AI agent is unguided, every cloud interaction is expensive. The agent dumps hundreds of thousands of tokens into prompt context, crashes on interactive CLI prompts, guesses deprecated SDK methods, and burns tokens in error recovery loops. Faced with this friction, developers keep their prompts small and defensive. They ask the agent for isolated helper functions and write the cloud plumbing themselves.

Now there are agent plugins. With plugins, cost and friction actually plummet because they bundle skill discovery, grounded documentation and surgical CLI output formatting. Because of this, developers respond immediately. They do not use the agent less. They give the agent larger, end-to-end cloud orchestration tasks: building, wiring, and deploying multi-product pipelines across Cloud Run, Firestore, and BigQuery in a single turn.

The Google Cloud Developer Plugin, built on the open Agent Plugins specification, provides the operational foundation for this transition across Antigravity, Claude Code, and Codex.

This article explores the architecture of the Google Cloud Developer Plugin and how it eliminates context bottlenecks. It then walks through an end-to-end scenario, guiding an AI coding agent to build and deploy a multi-service pipeline across Cloud Run, Firestore and BigQuery.

Author: Balaji Subramaniam, DevRel Engineer, Google Cloud


The context bottleneck in cloud development

Building cloud backends requires deep context. But dumping all that context into an LLM prompt upfront ruins performance. In practice, unguided agents hit three bottlenecks:

Prompt bloat

Standard agent harnesses use progressive disclosure by preloading skill metadata (names and descriptions) on Turn 0 and deferring full instruction bodies until activation. But across large skill catalogs like the Google Cloud skills, even preloading metadata headers pollutes system prompt context with thousands of tokens on every turn.

Observation bloat

Unfiltered terminal commands flood the context window. Running a basic command like gcloud run services list or gcloud compute instances list without format projections returns massive JSON or YAML payloads packed with cluster configs, internal metadata, and verbose status blocks. An agent ingesting raw CLI dumps fills its context buffer in just a few turns.

Hallucination loop

Foundation models rely on frozen training data. When an agent tries to write deployment scripts or SDK code, it may invent obsolete flags or import deprecated library methods. The CLI rejects the command, the agent parses the error, hallucinates another bad flag, and gets trapped in a multi-turn failure spiral.


How the Google Cloud Developer Plugin eliminates the context bottleneck

The google-cloud-developer plugin organizes cloud capabilities into four disciplined layers:

1. Multi-tier progressive disclosure (finding-google-skills)

The google-cloud-developer plugin externalizes the broader catalog behind a compact discovery router (~1,680 tokens). Context is revealed in three disciplined tiers:

  • Tier 1 (discovery router & foundational bundle): The agent holds only foundational cloud guardrails and the discovery router. Zero remote catalog metadata is preloaded in the prompt.
  • Tier 2 (on-demand catalog discovery): When a request requires specialized services, the router searches the central Google Agent Skills repository and pulls only the target SKILL.md instructions into active context.
  • Tier 3 (execution grounding via MCP): Granular SDK parameter syntax and schemas are retrieved on demand via the Developer Knowledge MCP server.

2. Grounded documentation via MCP (retrieving-developer-knowledge)

The plugin connects the agent to the Developer Knowledge MCP server (developerknowledge.googleapis.com). Through tools like answer_query and search_documents, the agent queries official, verified documentation in real time. It looks up exact constructor arguments, client methods, and configuration keys before writing a single line of code.

3. Execution guardrails & data reduction (gcloud skill)

The plugin enforces strict CLI execution rules:

  • Leaf Help Validation: The agent runs gcloud help <command> before executing new commands to verify valid parameters.
  • Non-Interactive Flags: Every command must include --quiet (or -q) so processes never freeze waiting for an interactive TTY confirmation.
  • Strict Output Projection: Commands must specify --format="json(...)" and --limit=5. This compresses megabytes of raw CLI output into concise JSON payloads containing only the exact fields required.

4. Credential safety (google-cloud-recipe-auth)

The plugin steers the agent away from generating long-lived service account keys or hardcoding secrets in code. It guides the workflow toward Application Default Credentials (ADC) and short-lived IAM role impersonation.


End-to-end walkthrough: Real-time event ingestion pipeline for payload classification

The Jevons Paradox becomes visible when you watch developers prompt an agent equipped with the plugin. Instead of asking for a snippet of Firestore code, the developer asks for a complete multi-product architecture.

Install the Google Cloud Developer plugin

Follow these instructions to install the Google Cloud Developer plugin in the coding agent ecosystem of your choice.

The developer prompt

Submit the following prompt to your coding agent (Antigravity CLI, Claude Code, or Codex):

text
Build and deploy a lightweight, high-throughput event classification and routing microservice on Cloud Run.

The service exposes a POST /events endpoint that receives incoming JSON event payloads, performs fast in-memory classification to assign a priority tier, persists live document state in Cloud Firestore, and streams analytics records into BigQuery.

Incoming Event Payload Schema:
{
  "event_id": "evt-84920-x",
  "event_type": "payment_authorization",
  "risk_score": 0.85,
  "user_tier": "enterprise",
  "metadata": {
    "source_region": "us-east1",
    "amount_usd": 1450.00
  }
}

Requirements:
1. Cloud Run Service (FastAPI):
   - Expose GET /healthz and POST /events.
   - Parse and validate incoming event payloads.
   - Classification Logic: Evaluate the payload and classify priority:
     - Assign "HIGH" if risk_score > 0.7 or user_tier is "critical".
     - Assign "STANDARD" for all other events.
2. Cloud Firestore (Real-Time State):
   - Write/update the classified event in the 'active_events' collection keyed by event_id.
   - Fields: event_id, priority (HIGH/STANDARD), risk_score, status ("PROCESSED"), and updated_at (UTC ISO timestamp).
3. BigQuery (Streaming Analytics):
   - Stream analytical metrics to 'analytics_dataset.events'.
   - Row schema: event_id, priority, risk_score, payload_size, timestamp.
4. Security & Deployment:
   - Provide non-interactive gcloud deployment commands following Google Cloud safety guardrails:
     - Explicit --project and --region.
     - Non-interactive flag (--quiet).
     - Clean format projection (--format="json(status.url)").
     - Least-privilege IAM service account (--no-allow-unauthenticated).
5. Validation via Plugin:
   - Use the Google Cloud Developer Plugin to verify SDK method signatures via Developer Knowledge MCP and validate CLI syntax with leaf help before generating commands.

Step-by-step agent execution

Here is how the plugin coordinates the agent's actions:

Step 1: Dynamic discovery

The agent identifies the requirement for a multi-service event pipeline. Using its discovery router and foundational cloud rules, it checks whether specialized cloud solution patterns exist in the skills repository.

Step 2: Live grounding over MCP

The agent calls answer_query on the Developer Knowledge MCP server:

  • Query: "google-cloud-bigquery python insert_rows_json best practices and error handling"
  • Result: The MCP server returns the exact signature for bigquery.Client().insert_rows_json(table, rows). No deprecated methods.

Step 3: Syntax and guardrail validation

Before proposing deployment commands, the agent runs leaf validation:

shell
gcloud help run deploy

It confirms flag availability and constructs a guarded command using explicit parameters, non-interactive mode, and projection formatting:

shell
gcloud run deploy event-routing-service \
  --source . \
  --region us-central1 \
  --project my-project-id \
  --no-allow-unauthenticated \
  --service-account [email protected] \
  --quiet \
  --format="json(status.url)"

Step 4: Clean code generation

The agent produces a self-contained Python implementation connecting FastAPI, Firestore, and BigQuery.


Build your next cloud architecture

The Google Cloud Developer Plugin and Data Agent Kit are available now in the open Google Agent Skills repository.

Try running the event ingestion walkthrough in your own environment with Antigravity CLI, Claude Code, or Codex. Explore the skills catalog, test custom service combinations, and let us know what end-to-end architectures and pipelines you are building with the Google Cloud Developer Plugin.