We recently announced the general availability of two new models from the latest generation of Anthropic’s Claude model family on Vertex AI: Claude Opus 4 and Claude Sonnet 4.
Claude Opus 4 is Anthropic’s most powerful model to date. Claude Opus 4 excels at coding, with sustained performance on complex, long-running tasks and agent workflows. Top use cases include advanced coding work, autonomous AI agents, agentic search and research, tasks that require complex problem solving, and long-running tasks that require precise context management.
Claude Sonnet 4 is Anthropic’s mid-size model that balances performance with cost. It surpasses its predecessor, Claude Sonnet 3.7, across coding and reasoning while responding more precisely to steering. Use cases include coding tasks such as code reviews and bug fixes, AI assistants, efficient research, and large-scale content generation and analysis.
These powerful Claude models are live on Google Cloud’s Vertex AI as a fully managed service. That means you can start building your applications without worrying about infrastructure provisioning or management. Vertex AI doesn’t just give you access to Claude, it wraps these models in an entire platform, providing you the tools you need to serve and scale your applications.
In this blog, we’ll guide you through building with the new Claude 4 models on Vertex AI. We’ll begin by demonstrating how to quickly get Claude Code running, then delve into using Claude effectively on Vertex AI, leveraging its platform capabilities like token counting, prompt caching (for cost and performance optimization), and batch predictions for scaling your applications. Finally, we’ll cover how to build agents by integrating Claude via MCP with the Google Cloud Agent Development Kit (ADK).
Let’s get started!
Use Claude Code as your code assistant
Using Claude Code through Vertex AI is only a 4-step process to get your local code assistant up and running. First, you need to get the Claude Code command-line interface (CLI) onto your machine. It’s a simple, single command using npm.
npm install -g @anthropic-ai/claude-code
Next, you’ll tell the CLI to route its requests through Vertex AI. This is done by creating a simple JSON configuration file (~/.claude/settings.json) that points to your Google Cloud project, the correct region and the model you want to use.
mkdir -p ~/.claude && cat < ~/.claude/settings.json
{
"env": {
"CLAUDE_CODE_USE_VERTEX": 1,
"ANTHROPIC_VERTEX_PROJECT_ID": "your-project-id", # Replace with your Project ID
"CLOUD_ML_REGION": "us-east5", # Or any other supported region for your chosen model
"ANTHROPIC_MODEL": "claude-sonnet-4@20250514", # Or your desired model
"DISABLE_PROMPT_CACHING": 1
}
}
EOF
With the CLI configured, you need to ensure your Google Cloud project is ready to receive requests. This involves two quick actions in your project:
- Enable the Vertex AI API.
- Enable the Claude 4 model you want to use from the Vertex AI Model Garden.
Finally, grant your local machine permission to talk to your Google Cloud project. The gcloud CLI helps make this easy and secure by setting up Application Default Credentials (ADC). You won’t have to handle any secret keys manually.
gcloud auth application-default login
And that’s it! With those four steps complete, you can run the claude command and start using your local code assistant through your own Google Cloud project as shown below.

Get started with Anthropic SDK for Vertex AI
Once you enable Claude on Vertex AI, you can make your application interact with the enabled model using the Anthropic SDK or direct cURL commands to the Vertex AI endpoint. Here’s a quick example of how you might set up the Anthropic SDK for Vertex AI in Python:
# Ensure you've installed the SDK: pip3 install -U 'anthropic[vertex]'
# And authenticated: gcloud auth application-default login
from anthropic import AnthropicVertex
PROJECT_ID = "your-project-id" # Replace with your Project ID
REGION = "us-east5" # Or any other supported region for your chosen model
client = AnthropicVertex(project_id=PROJECT_ID, region=REGION)
message = client.messages.create(
model="claude-sonnet-4@20250514", # Or your desired model
max_tokens=1024,
messages=[
{
"role": "user",
"content": "Hello, Claude! Tell me a fun fact about Large Language Models.",
}
],
)
print(message.model_dump_json(indent=2))
Remember to use the correct model names, which include a version suffix starting with an @ symbol to ensure consistent behavior (e.g., claude-opus-4@20250514 or claude-sonnet-4@20250514). Also for a streaming response, you can use client.messages.stream(...).
Optimize cost and performance with Claude on Vertex AI
When prototyping your application, Vertex AI provides tools to help you manage costs and optimize the performance of Claude models.
Before sending a potentially large or complex prompt to a Claude model, it’s wise to understand its token count, especially with reasoning models. By determining the number of tokens in a message beforehand, you can better predict costs, stay within token limits, and optimize your prompts for efficiency.
Vertex AI offers a dedicated count-tokens endpoint for this purpose. And the best part is there’s no cost for using this endpoint itself. To count your tokens, send a rawPredict request to the count-tokens endpoint, providing the model ID and the messages you want to count tokens for as shown below.
from anthropic import AnthropicVertex
PROJECT_ID = "your-project-id" # Replace with your Project ID
REGION = "us-east5" # Or any other supported region for your chosen model
CLAUDE_PRICING = {
"claude-opus-4@20250514": {
"input": 0.000015, # $15 per million input tokens
"output": 0.000075, # $75 per million output tokens
},
"claude-sonnet-4@20250219": {
"input": 0.000003, # $3 per million input tokens
"output": 0.000015, # $15 per million output tokens
},
}
CLAUDE_MODEL = "claude-sonnet-4@20250514" # Or your desired model
# Initialize the Anthropic client with your project ID and region
client = AnthropicVertex(project_id=PROJECT_ID, region=REGION)
input_tokens = client.messages.count_tokens(
model=CLAUDE_MODEL, messages=[{"role": "user", "content": "Hello, world"}]
).input_tokens
print(f"Input tokens: {input_tokens}")
# Estimate input cost
cost_input_tokens = input_tokens * CLAUDE_PRICING[CLAUDE_MODEL]["input"]
print(f"Cost for input tokens: ${cost_input_tokens:.6f}")
# Sent a request to the Claude model
message = client.messages.create(
model=CLAUDE_MODEL,
max_tokens=1024,
messages=[
{
"role": "user",
"content": "Hello, Claude! Tell me a fun fact about Large Language Models.",
}
],
)
# Estimate output cost
output_tokens = message.usage.output_tokens
print(f"Output tokens: {output_tokens}")
cost_output_tokens = output_tokens * CLAUDE_PRICING[CLAUDE_MODEL]["output"]
print(f"Cost for output tokens: ${cost_output_tokens:.6f}")
# Estimate total cost
total_cost = cost_input_tokens + cost_output_tokens
print(f"Total cost: ${total_cost:.6f}")
The response will give you the input_tokens count.
Faster and cheaper generation using prompt caching with Claude
Prompt caching is another powerful feature for Claude on Vertex AI, designed to reduce latency and costs for repeated content across multiple requests. When prompt caching is enabled, Vertex AI automatically caches parts of your input. Subsequent queries containing identical text, images, and the cache_control parameter will then leverage these cached results, avoiding redundant computation and network overhead. This process occurs automatically, as shown below.
from anthropic import AnthropicVertex
PROJECT_ID = "your-project-id" # Replace with your Project ID
REGION = "us-east5" # Or any other supported region for your chosen model
MODEL = "claude-sonnet-4@20250514" # Replace with your model of choice
# Initialize the Anthropic client with your project ID and region
client = AnthropicVertex(project_id=PROJECT_ID, region=REGION)
# Read markdown file you can create
with open("claude_docs.md", "r") as file:
claude_docs_content = file.read()
# Sent a request to the Claude model to use the cache feature
response = client.messages.create(
model=MODEL,
max_tokens=1024,
system=[
{
"type": "text",
"text": "You are a documentation assistant. Use the provided documentation to answer questions.",
},
{
"type": "text",
"text": claude_docs_content,
"cache_control": {"type": "ephemeral"},
},
],
messages=[{"role": "user", "content": "How do I use the cache feature?"}],
)
print(response.usage.model_dump_json())
print(response.model_dump_json(indent=2))
# Call the model again with the same inputs up to the cache checkpoint
response = client.messages.create(
model=MODEL,
max_tokens=1024,
system=[
{
"type": "text",
"text": "You are a documentation assistant. Use the provided documentation to answer questions.",
},
{
"type": "text",
"text": claude_docs_content,
"cache_control": {"type": "ephemeral"},
},
],
messages=[
{
"role": "user",
"content": "Can you prepare a script to use the cache feature?",
}
],
)
print(response.usage.model_dump_json())
print(response.model_dump_json(indent=2))
With prompt caching, you can start seeing cost reductions from the second use of a prompt. This benefit quickly grows, potentially reaching up to 90% savings for frequently used prompts. Keep in mind, the cache lasts for five minutes and refreshes each time it’s accessed.
Scaling your Claude applications with batch predictions
For systems that require tasks like automatically generating unit tests across a codebase or analyzing extensive logs from AI agent interactions for optimization—where throughput is prioritized over immediate responses—batch predictions with LLMs is the ideal solution.
With Vertex AI, instead of sending one prompt per request (online predictions), you can batch many prompts into a single request, which is highly efficient for bulk processing. To use Claude models with batch predictions, you need to prepare your input dataset following the Anthropic Claude API Schema JSON format and store it either as a BigQuery table or a JSONL file in Cloud Storage. Then, you can use the Vertex AI SDK for Python to initiate batch prediction jobs. Here’s a an example using a BigQuery source:
import time
from google.cloud import storage
from google import genai
from google.genai.types import CreateBatchJobConfig, JobState, HttpOptions
import pandas as pd
PROJECT_ID = "your-project-id" # Replace with your Project ID
REGION = "us-east5" # Or any other supported region for your chosen model
BUCKET_NAME = "your-bucket" # Replace with your GCS bucket name
BUCKET_URI = f"gs://{BUCKET_NAME}"
OUTPUT_FILE = "anthropic_test_data.jsonl" # Name of the output file
OUTPUT_FILE_URI = f"{BUCKET_URI}/{OUTPUT_FILE}"
# Define samples and job status
AGENT_TRACES_SCENARIOS = [
{
"custom_id": "eval-job-1-simple-query",
"request": {
"system": "You are an AI quality evaluation assistant. Your task is to evaluate the provided conversation for helpfulness, accuracy, and adherence to safety guidelines. Analyze the assistant's response to the user's query. Provide your evaluation in a JSON object with two keys: 'score' (a value from 1 to 5, where 5 is best) and 'justification' (a brief explanation for your score).",
"messages": [
{
"role": "user",
"content": "What is the primary function of the mitochondria?",
},
{
"role": "assistant",
"content": "The primary function of mitochondria is to generate most of the cell's supply of adenosine triphosphate (ATP), which is used as a source of chemical energy.",
},
],
"max_tokens": 300,
"anthropic_version": "vertex-2023-10-16",
},
},
...
]
COMPLETED_STATES = {
JobState.JOB_STATE_SUCCEEDED,
JobState.JOB_STATE_FAILED,
JobState.JOB_STATE_CANCELLED,
JobState.JOB_STATE_PAUSED,
}
# Save the varied scenarios to a JSONL file
df = pd.DataFrame(AGENT_TRACES_SCENARIOS)
df.to_json(OUTPUT_FILE_URI, orient="records", lines=True)
# Initialize the GenAI client
client = genai.Client(
http_options=HttpOptions(api_version="v1"),
vertexai=True,
project=PROJECT_ID,
location=REGION,
)
# Create a batch job to process the scenarios using Claude 3.5 Haiku model
job = client.batches.create(
model="publishers/anthropic/models/claude-sonnet-4@20250514", # Replace with your model of choice
src=OUTPUT_FILE_URI,
config=CreateBatchJobConfig(dest=BUCKET_URI),
)
while job.state not in COMPLETED_STATES:
time.sleep(30)
job = client.batches.get(name=job.name)
print(f"Job state: {job.state}")
if job.state == JobState.JOB_STATE_SUCCEEDED:
print(f"Batch job completed successfully. Output file: {OUTPUT_FILE_URI}")
df = pd.read_json(OUTPUT_FILE_URI, lines=True)
print(df.head())
A similar approach applies to Cloud Storage (GCS) sources, changing the src to a gs:// path. After submitting the job, you can monitor the status using its name. Once completed, the output (predictions) will be available in your specified BigQuery table or Cloud Storage location.
Building intelligent agents: Claude meets the Agent Development Kit (ADK)
Now for the most exciting part: building sophisticated agents! Google’s Agent Development Kit (ADK) is designed for flexibility, allowing you to integrate various large language models (LLMs), including Anthropic’s Claude on Vertex AI, into your agents.
A powerful way to extend your ADK agent’s capabilities is by leveraging the Model Context Protocol (MCP). MCP is an open standard that standardizes how LLMs communicate with external applications, data sources, and tools. Think of it as a universal connector. An ADK agent can act as an MCP client, consuming tools exposed by external MCP servers. This allows your agent to interact with a wide array of systems, from local file systems to remote APIs.
The ADK provides the MCPToolset class, which acts as a bridge to MCP servers. It handles connecting to an MCP server, discovering its available tools, adapting them for the ADK LlmAgent, and proxying calls.
Let’s imagine you want to build an agent powered by Claude on Vertex AI that can use multiple MCP tools for generating different media simultaneously. For example, an agent that can generate images using Imagen models. Here’s how you can integrate Claude on Vertex AI with ADK and MCP.
import os
import asyncio
from contextlib import AsyncExitStack
from google.adk.models.anthropic_llm import Claude
from google.adk.models.registry import LLMRegistry
from google.adk.agents import LlmAgent
from google.adk.tools.mcp_tool.mcp_toolset import MCPToolset, StdioServerParameters
from dotenv import load_dotenv
# Load env variables and register the Claude model
LLMRegistry.register(Claude)
load_dotenv()
# Define your Agent with MCPToolset
imagen_toolset = MCPToolset(
connection_params=StdioServerParameters(
command="mcp-imagen-go",
env={
"PROJECT_ID": os.getenv("GOOGLE_CLOUD_PROJECT", "your-project-id"),
"GENMEDIA_BUCKET": os.getenv("GENMEDIA_BUCKET", "your-bucket-name"),
"MCP_SERVER_REQUEST_TIMEOUT": os.getenv("MCP_SERVER_REQUEST_TIMEOUT", "500")
},
),
)
root_agent = LlmAgent(
name="claude_agent",
instruction=(
"You are a helpful assistant that can interact with the Imagen tool to generate images based on user requests. "
"You will receive requests in the form of text, and you should respond with the appropriate image generation command."
),
model=os.getenv("MODEL_NAME","claude-sonnet-4@20250514"),
tools=[imagen_toolset],
)
Before you start, ensure that your Vertex AI environment is configured with the right credentials and necessary environment variables, you’ve installed the anthropic[vertex] and google-ads-adk Python libraries (along with Node.js/npx and model-context-protocol if using community or custom MCP servers). You can find more details in the documentation and in this github repo.
To use Claude models directly via Vertex AI with ADK, register the Claude model you intend to use. This allows the ADK’s to handle that model with LlmAgent (LLMRegistry.register(Claude)). Next, you can use MCPTools to orchestrate the integration. This class makes sure that external tools on an MCP server are available to your agent, handling the connection, discovery, and proxying automatically.
Finally, to extend your chosen Claude model on Vertex AI with diverse capabilities from all connected MCP sources, you can now combine all retrieved tools and pass them to a new LlmAgent instance. This instance is specifically configured for this purpose.
In this example below, you can see how an ADK agent powered by Claude 4 was able to generate a beautiful landscape using MCP Servers for Google Cloud Gen media APIs.
What’s next
We encourage you to try the new Claude models on Vertex AI, and start building new agentic applications!

