Skip to content

Model Configuration#

The AG2 framework provides an explicit, predictable, and type-safe way to configure Large Language Models (LLMs) for your agents. The configuration API is designed to provide a consistent developer experience across different model providers while maintaining strong typing support.

Supported Providers#

AG2 supports multiple LLM providers through dedicated configuration classes. Each provider requires its respective optional dependencies to be installed.

Provider Configuration Class Installation Command
OpenAI Responses OpenAIResponsesConfig pip install "ag2[openai]"
OpenAI OpenAIConfig pip install "ag2[openai]"
Anthropic AnthropicConfig pip install "ag2[anthropic]"
Gemini GeminiConfig pip install "ag2[gemini]"
Gemini on Vertex AI VertexAIConfig pip install "ag2[gemini]"
Amazon Bedrock BedrockConfig pip install "ag2[bedrock]"
Ollama OllamaConfig pip install "ag2[ollama]"
DashScope DashScopeConfig pip install "ag2[dashscope]"
xAI XAIConfig pip install "ag2[xai]"
Z.AI ZAIConfig pip install "ag2[zai]"
Mistral MistralConfig pip install "ag2[mistral]"
TypeSafe (Jev) TypeSafeConfig pip install "ag2[typesafe]"

(Note: OpenAIConfig is also available for OpenAI-compatible endpoints).


How to Configure a Model#

Basic Configuration#

To configure a model, import the specific provider's configuration class and initialize it with your desired parameters. The most common parameters are model, api_key, and base_url.

1
2
3
4
5
6
7
8
from ag2.config import OpenAIResponsesConfig

# Configure an OpenAI Responses API model
config = OpenAIResponsesConfig(
    model="gpt-4.1-nano",
    api_key="sk-...",
    streaming=True
)
1
2
3
4
5
6
7
8
9
from ag2.config import OpenAIConfig

# Configure an OpenAI model
config = OpenAIConfig(
    model="gpt-4o-mini",
    api_key="sk-...",
    temperature=0.2,
    streaming=True
)
1
2
3
4
5
6
7
8
from ag2.config import AnthropicConfig

# Configure an Anthropic model
config = AnthropicConfig(
    model="claude-haiku-4-5-20251001",
    api_key="sk-ant-...",
    streaming=True
)

ag2[anthropic] requires anthropic>=1.6.0,<2. See Anthropic Configuration for the httpx2 HTTP client and the sampling parameters (temperature, top_p, top_k).

1
2
3
4
5
6
7
8
from ag2.config import GeminiConfig

# Configure a Gemini model
config = GeminiConfig(
    model="gemini-3-flash-preview",
    api_key="...",
    streaming=True
)
1
2
3
4
5
6
7
8
from ag2.config import BedrockConfig

# Configure an Amazon Bedrock model (Converse API)
config = BedrockConfig(
    model="global.anthropic.claude-sonnet-5",
    region_name="us-east-1",
    streaming=True
)

Credentials follow the standard AWS resolution chain: explicit aws_access_key_id / aws_secret_access_key, a named profile_name, environment variables, shared config files, or instance roles. model accepts a Bedrock model id or an inference-profile ARN. See Amazon Bedrock authentication for the API-key (bearer token) alternative.

1
2
3
4
5
6
7
from ag2.config import OllamaConfig

# Configure an Ollama model
config = OllamaConfig(
    model="qwen3.5:latest",
    streaming=True
)
1
2
3
4
5
6
7
8
from ag2.config import DashScopeConfig

# Configure a DashScope model
config = DashScopeConfig(
    model="qwen-plus",
    api_key="...",
    streaming=True
)
1
2
3
4
5
6
7
8
from ag2.config import XAIConfig

# Configure an xAI Grok model
config = XAIConfig(
    model="grok-4",
    api_key="xai-...",
    streaming=True
)
1
2
3
4
5
6
7
8
from ag2.config import ZAIConfig

# Configure a Z.AI GLM model
config = ZAIConfig(
    model="glm-4.6",
    api_key="...",
    streaming=True
)
1
2
3
4
5
6
7
8
from ag2.config import MistralConfig

# Configure a Mistral model
config = MistralConfig(
    model="mistral-large-latest",
    api_key="...",
    streaming=True
)
1
2
3
4
5
6
7
from ag2.config import TypeSafeConfig

# Configure TypeSafe's Jev decision model (non-streaming)
config = TypeSafeConfig(
    model="jev-latest",
    api_key="...",
)

Jev answers a typed question instead of generating text, so the agent must set a decision response_schema. See TypeSafe (Jev) Configuration.

Tip

AG2 is designed to be async and streaming-first, so for the best user experience it is recommended to enable streaming on the model provider configurations for models that support it. As shown above, Streaming has been set to True in each config that supports it; TypeSafe's Jev answers in a single response and has no streaming option.

Using Environment Variables#

For security and convenience, you don't need to hardcode your API keys. If api_key is not explicitly provided, the configuration will automatically attempt to load it from your environment variables.

The system looks for provider-specific keys (e.g., OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, XAI_API_KEY, ZAI_API_KEY, TYPESAFE_API_KEY).

1
2
3
4
from ag2.config import OpenAIConfig

# Automatically falls back to OPENAI_API_KEY from the environment
config = OpenAIConfig(model="gpt-5")

Anthropic Configuration#

AnthropicConfig connects to the Claude API through the official anthropic SDK. Install the optional dependency with pip install "ag2[anthropic]", which requires anthropic>=1.6.0,<2.

If api_key is not passed explicitly, it is resolved from the ANTHROPIC_API_KEY environment variable.

1
2
3
4
from ag2.config import AnthropicConfig

# Auth resolved from ANTHROPIC_API_KEY when omitted
config = AnthropicConfig(model="claude-haiku-4-5-20251001", streaming=True)

Custom HTTP Client#

anthropic>=1 is built on httpx2, so http_client takes an httpx2.AsyncClient:

1
2
3
4
5
6
7
8
import httpx2
from ag2.config import AnthropicConfig

config = AnthropicConfig(
    model="claude-haiku-4-5-20251001",
    api_key="sk-ant-...",
    http_client=httpx2.AsyncClient(proxy="http://proxy.example.com:8080"),
)

An httpx client is rejected

The anthropic SDK keeps no compatibility path for a legacy httpx.AsyncClient — it raises TypeError as soon as the client is built, where the OpenAI SDK would still accept one. Rebuild yours against httpx2; the two packages share an API, so this is usually just the import line.

Sampling Parameters#

temperature, top_p and top_k are ordinary fields on AnthropicConfig. The 1.x Messages API dropped them from its method signature, so AG2 sends them in the request body, where the API still reads them:

1
2
3
4
5
6
7
from ag2.config import AnthropicConfig

config = AnthropicConfig(
    model="claude-haiku-4-5-20251001",
    api_key="sk-ant-...",
    temperature=0.2,
)

Recent Claude models reject sampling parameters

Claude Fable 5, Opus 5, Opus 4.8, Opus 4.7 and Sonnet 5 removed temperature, top_p and top_k from the Messages API — a request carrying any of them is answered with HTTP 400. Leave all three unset (the default) on those models. Claude Opus 4.6, Sonnet 4.6, Haiku 4.5 and earlier still accept them.

extra_body precedence

Pass extra_body to forward provider-specific keys that have no dedicated field. A key you write there wins over the same key AG2 derived from a field — an extra_body={"temperature": 0.9} overrides temperature=0.2 — while unrelated keys are merged through untouched.

Amazon Bedrock Authentication#

BedrockConfig authenticates in either of two ways, both resolved by the underlying AWS SDK:

1. AWS credentials (SigV4) — explicit keys, a profile_name, or the standard environment variables (AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN). Recommended for production; credentials refresh automatically through aiobotocore.

2. Bedrock API keys (bearer token) — set the Amazon Bedrock API key as an environment variable and botocore uses bearer-token auth for Bedrock calls automatically (no other credentials needed):

export AWS_BEARER_TOKEN_BEDROCK=<your-bedrock-api-key>
export AWS_DEFAULT_REGION=us-east-1
1
2
3
4
from ag2.config import BedrockConfig

# Auth from AWS_BEARER_TOKEN_BEDROCK, region from AWS_DEFAULT_REGION
config = BedrockConfig(model="global.anthropic.claude-sonnet-5")

A region is always required — pass region_name= or set AWS_DEFAULT_REGION. Notes on API keys:

  • Short-term keys expire with the console session that minted them (max 12 hours) and are region-bound — generate the key in the same region you call.
  • Long-term keys are backed by an auto-created IAM user; AWS recommends them for exploration only.
  • API keys work only for Bedrock / Bedrock Runtime actions. If both a bearer token and AWS credentials are present, the bearer token wins for Bedrock calls.

Custom Session#

BedrockConfig is built on aiobotocore, so session takes an aiobotocore.session.AioSession — a boto3.Session is not accepted. Sharing one session across configs resolves credentials and loads the Bedrock service model once:

from aiobotocore.session import AioSession
from ag2.config import BedrockConfig

session = AioSession()

planner = BedrockConfig(
    model="global.anthropic.claude-sonnet-5",
    region_name="us-east-1",
    session=session,
)
writer = BedrockConfig(
    model="global.anthropic.claude-sonnet-5",
    region_name="us-east-1",
    streaming=True,
    session=session,
)

Google Vertex AI (Gemini)#

For Gemini on Vertex AI (Google Cloud), use the dedicated VertexAIConfig class. GeminiConfig covers the public Developer API (api_key); VertexAIConfig covers the Vertex path (GCP project, location, and Google-issued credentials).

Authentication accepts any of the following:

from ag2.config import VertexAIConfig

config = VertexAIConfig(
    model="gemini-3-flash-preview",
    project="my-gcp-project",
    location="us-central1",
    credentials="/path/to/service-account-key.json",
    # Path to a service-account JSON key file downloaded from
    # GCP Console -> IAM & Admin -> Service Accounts -> Keys.
)

The service account needs the Vertex AI User (roles/aiplatform.user) IAM role on the project.

from ag2.config import VertexAIConfig

# Run `gcloud auth application-default login` first, or ensure
# GOOGLE_APPLICATION_CREDENTIALS points to a key file. With nothing
# passed to `credentials`, google-genai resolves ADC automatically.
config = VertexAIConfig(
    model="gemini-3-flash-preview",
    project="my-gcp-project",
    location="us-central1",
)
import google.auth
from ag2.config import VertexAIConfig

creds, _ = google.auth.default(
    scopes=["https://www.googleapis.com/auth/cloud-platform"],
)

config = VertexAIConfig(
    model="gemini-3-flash-preview",
    project="my-gcp-project",
    location="us-central1",
    credentials=creds,
)

Use this path for impersonated credentials, workload identity, or any other google.auth.credentials.Credentials source.

Environment variables#

Instead of passing parameters explicitly, the underlying google-genai SDK resolves any field left unset from the following environment variables:

Environment variable Used by Equivalent parameter Notes
GOOGLE_API_KEY GeminiConfig api_key Takes precedence over GEMINI_API_KEY if both are set.
GEMINI_API_KEY GeminiConfig api_key Developer API key.
GOOGLE_CLOUD_PROJECT VertexAIConfig project GCP project ID.
GOOGLE_CLOUD_LOCATION VertexAIConfig location GCP region (or global).
GOOGLE_APPLICATION_CREDENTIALS VertexAIConfig credentials Path to a service-account JSON key file, read via ADC.

With the three Vertex variables set in the environment, configuration collapses to just the model name:

export GOOGLE_CLOUD_PROJECT=my-gcp-project
export GOOGLE_CLOUD_LOCATION=us-central1
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json
1
2
3
4
from ag2.config import VertexAIConfig

# All Vertex auth parameters resolved from the environment.
config = VertexAIConfig(model="gemini-3-flash-preview")

Controlling Gemini Thinking#

Gemini 3 Pro models default to dynamic / unbounded thinking, which can cause individual calls to spend a large internal token budget before responding. Both GeminiConfig and VertexAIConfig accept thinking controls that map directly to Google's Thinking API.

Use thinking_level for Gemini 3 models, or thinking_budget for Gemini 2.5 models:

from ag2.config import GeminiConfig, VertexAIConfig

# Gemini 3 — bound thinking with a level
gemini3 = GeminiConfig(
    model="gemini-3-flash-preview",
    thinking_level="low",  # "low" | "medium" | "high"
)

# Gemini 2.5 — bound thinking with an explicit token budget
gemini25 = VertexAIConfig(
    model="gemini-2.5-pro",
    project="my-gcp-project",
    location="us-central1",
    thinking_budget=512,  # 0 disables thinking entirely
)

For full control (e.g. enabling include_thoughts), pass a google.genai.types.ThinkingConfig directly via thinking_config. When set, it takes precedence over the shorthand fields.

The number of thinking tokens consumed is reported on ModelResponse.usage.thinking_tokens and emitted as the gen_ai.usage.thinking_tokens OpenTelemetry attribute by TelemetryMiddleware.

Z.AI (GLM) Configuration#

ZAIConfig connects to Z.AI's GLM models through the official zai-sdk. Install the optional dependency with pip install "ag2[zai]".

If api_key and base_url are not passed explicitly, the SDK resolves them from the ZAI_API_KEY and ZAI_BASE_URL environment variables; base_url otherwise defaults to the international endpoint (https://api.z.ai/api/paas/v4).

1
2
3
4
from ag2.config import ZAIConfig

# Auth and endpoint resolved from ZAI_API_KEY / ZAI_BASE_URL when omitted
config = ZAIConfig(model="glm-4.6", streaming=True)

Controlling GLM Thinking#

GLM models accept reasoning controls that map directly to the Z.AI API. Set thinking to toggle the reasoning pass (True enables it, False disables it, unset uses the model default) and reasoning_effort to bound how much the model thinks:

1
2
3
4
5
6
7
8
from ag2.config import ZAIConfig

config = ZAIConfig(
    model="glm-4.6",
    thinking=True,
    reasoning_effort="high",
    streaming=True,
)

Reasoning content is streamed as ModelReasoning events, and reasoning tokens are reported on ModelResponse.usage.thinking_tokens.

extra_body precedence

Pass extra_body to forward provider-specific keys that have no dedicated field. Explicit top-level fields always win over the same key inside extra_body — e.g. a top-level thinking=True overrides a thinking entry inside extra_body, while unrelated extra_body keys are merged through untouched.

Mistral Configuration#

MistralConfig connects to Mistral's chat-completions API through the official mistralai SDK. Install the optional dependency with pip install "ag2[mistral]".

If api_key is not passed explicitly, the SDK resolves it from the MISTRAL_API_KEY environment variable.

1
2
3
4
from ag2.config import MistralConfig

# Auth resolved from MISTRAL_API_KEY when omitted
config = MistralConfig(model="mistral-large-latest", streaming=True)

Point server_url at a different deployment (for example a private or regional endpoint) when you are not calling https://api.mistral.ai.

Reasoning Models#

Magistral and other reasoning models return their thinking trace alongside the answer. AG2 splits the two: the trace is emitted as ModelReasoning events, and only the answer text lands on ModelResponse.message.

1
2
3
4
5
6
7
from ag2.config import MistralConfig

config = MistralConfig(
    model="magistral-medium-latest",
    reasoning_effort="high",
    streaming=True,
)

Image Generation#

ImageGenerationTool is executed by Mistral, not locally. The call and its result are surfaced as BuiltinToolCallEvent and BuiltinToolResultEvent, and the generated image arrives as a UrlInput:

from ag2 import Agent
from ag2.config import MistralConfig
from ag2.tools.builtin.image_generation import ImageGenerationTool

agent = Agent(
    "illustrator",
    config=MistralConfig(model="mistral-medium-latest"),
    tools=[ImageGenerationTool()],
)
reply = await agent.ask("Draw a red circle over a black square.")

The image URL is a short-lived signed link, so download it promptly if you need to keep it. Mistral's tool takes no options, so size, quality, and output_format are ignored. The model may reply with the image and no commentary, leaving reply.body empty — the image is on the tool result either way.

Other builtin tools

Apart from image generation, Mistral's server-side tools (web_search, code_interpreter, document_library) belong to its Agents API and are rejected by chat-completions. Passing those AG2 builtin tools raises UnsupportedToolError rather than failing at request time.

OpenTelemetry version ceiling

The mistralai SDK pins opentelemetry-semantic-conventions<0.61, which transitively caps opentelemetry-api at 1.39.1. A clean install of ag2[mistral,tracing] resolves to a consistent OpenTelemetry 1.39.x stack and traces normally.

Adding ag2[mistral] to an environment that already has a newer OpenTelemetry downgrades opentelemetry-api on its own, leaving it mismatched against the newer SDK. Reinstall the pair together, or hold the newer versions explicitly — the SDK itself works fine with current OpenTelemetry, the pin is just over-tight:

pip install "mistralai>=2.8.0" "opentelemetry-api>=1.43" "opentelemetry-sdk>=1.43" "opentelemetry-semantic-conventions>=0.64b0"

TypeSafe (Jev) Configuration#

TypeSafeConfig connects to TypeSafe AI's Jev decision model through the official typesafe-sdk. Install the optional dependency with pip install "ag2[typesafe]".

Jev is not a text generator. It answers one typed question about the conversation so far — yes or no, one label out of several, or a level on a rubric — and returns the answer together with its confidence and the probability of every option. That makes it a cheap, predictable front for a generative agent: route a ticket, gate a reply, grade a draft, then hand the text work to another provider.

If api_key and base_url are not passed explicitly, the SDK resolves them from the TYPESAFE_API_KEY and TYPESAFE_BASE_URL environment variables. model defaults to jev-latest.

from enum import Enum

from ag2 import Agent
from ag2.config import TypeSafeConfig

class Department(Enum):
    """Which team should handle this ticket?"""

    BILLING = "billing"
    """Payments, invoicing, refunds."""
    TECHNICAL = "technical"
    """Bugs, outages, integrations."""

router = Agent(
    "router",
    prompt="You triage customer support tickets.",
    config=TypeSafeConfig(),  # auth resolved from TYPESAFE_API_KEY
    response_schema=Department,
)

reply = await router.ask("My webhook integration has been down for 3 days.")
department = await reply.content()        # Department.TECHNICAL
reply.response.metadata["confidence"]     # e.g. 0.93
reply.response.metadata["probabilities"]  # {"billing": 0.07, "technical": 0.93}

The response schema is the question#

There is no free-text path, so an agent on TypeSafeConfig must set response_schema, and the schema picks which of Jev's three primitives is asked:

response_schema Jev asks await reply.content() returns
bool a yes/no question True once the yes-probability reaches boolean_threshold (default 0.5)
a number schema bounded to 0..1, e.g. ResponseSchema.from_schema({"type": "number", "minimum": 0, "maximum": 1}, name="churn", description="How likely is churn?") a yes/no question the yes-probability as JSON text, unparsed (from_schema schemas are not validated); reply.response.metadata["noul"] has it as a float
an Enum of strings a choice between its members the selected member
an IntEnum numbered 0..n-1 with 2–10 levels a score on that rubric the member nearest Jev's expected score

Anything else — str, int, a dataclass, a Pydantic model, a PromptedSchema — raises UnsupportedResponseSchemaError before any request is sent.

The agent prompt frames the question; the schema's own description asks it. For an Enum that is the class docstring, and for everything else an explicit ResponseSchema(..., description=...). The two are joined, so the router above asks "You triage customer support tickets. Which team should handle this ticket?".

A yes/no question has to be asked

response_schema=bool with no prompt, no description and no criteria raises ValueError locally. The API rejects a bare yes/no question, and failing before the request keeps the error actionable.

Describing the options#

The string literal under each Enum / IntEnum member describes that option to Jev. Python discards those strings at runtime, so AG2 reads them back from the class source.

1
2
3
4
5
6
7
8
9
from enum import IntEnum

class Severity(IntEnum):
    LOW = 0
    """Cosmetic; nothing is blocked."""
    MEDIUM = 1
    """Degraded but usable."""
    HIGH = 2
    """An outage or data loss."""

A score rubric requires a description for every level; a choice works without them. criteria on the config overrides or supplies descriptions without touching the type — keyed by choice label, by score level as a string, or by "true" / "false" for a yes/no question:

1
2
3
4
5
6
7
8
9
from ag2.config import TypeSafeConfig

config = TypeSafeConfig(
    boolean_threshold=0.8,
    criteria={
        "true": "The customer is asking for money back.",
        "false": "Anything else, including questions about pricing.",
    },
)

Member docstrings need the source file

Docstrings are read with inspect.getsource, so an Enum defined in a REPL, a notebook cell or a frozen app has none. Pass criteria there.

Confidence and probabilities#

The SDK answer, minus its type, is kept on reply.response.metadata. A yes/no answer is just noul, the yes-probability. A choice carries choice, confidence and probabilities per label. A score carries score (the expected value on the rubric, before snapping), confidence, the rubric legend and probabilities per level.

What Jev does not do#

  • Tools. Passing any tool raises UnsupportedToolError; keep tools on the generative agent Jev routes to.
  • Streaming. Requests are non-streaming; there is no streaming field, and the answer arrives as one ModelMessage.
  • Files and media. create_files_client() raises NotImplementedError, and only text and structured DataInput parts are sent — anything else raises UnsupportedInputError.
  • Free text. reply.body is the answer as JSON, not prose.

retry takes a typesafe_sdk.RetryPolicy and http_client an httpx2.AsyncClient; both are forwarded to the SDK untouched. Token counts land on ModelResponse.usage as prompt_tokens / completion_tokens when the API reports them.

Structured data extraction, one of the decision shapes in TypeSafe's docs, is the one this provider cannot express: an object schema is rejected, since Jev answers one decision per question.

Self-Hosted and OpenAI-Compatible Models (vLLM, LM Studio, etc.)#

If you are using a self-hosted model or an API that is compatible with the OpenAI format (such as vLLM, LM Studio, FastChat, or Together AI), you can use the OpenAIConfig class and specify a custom base_url.

1
2
3
4
5
6
7
8
9
from ag2.config import OpenAIConfig

# Configure a vLLM or other OpenAI-compatible endpoint
config = OpenAIConfig(
    model="qwen-3",
    base_url="http://localhost:8000/v1",
    # Some endpoints don't require an API key, but the client expects a non-empty string
    api_key="NotRequired",
)

Tip

If you are running a self-hosted server via HTTPS without a valid SSL certificate (e.g., a local self-signed certificate), you can disable SSL checks by passing a custom httpx2.AsyncClient with verify=False to the configuration:

1
2
3
4
5
6
7
8
9
import httpx2
from ag2.config import OpenAIConfig

config = OpenAIConfig(
    model="qwen-3",
    base_url="https://localhost:8000/v1",
    api_key="NotRequired",
    http_client=httpx2.AsyncClient(verify=False)
)

http_client is an httpx2 client

From openai>=3 the OpenAI SDK uses httpx2, which ag2[openai] installs for you. OpenAIConfig and OpenAIResponsesConfig annotate http_client accordingly, and ag2 forwards whatever you pass to the SDK untouched.

A legacy httpx.AsyncClient still works at runtime: the SDK keeps a compatibility path for it, though it fails static type checking and the SDK's migration guide calls that path a migration aid that may be discontinued. Build the client with httpx2.AsyncClient to be done with it.

Each provider takes the client its own SDK takes, so the annotation differs by provider — AnthropicConfig.http_client, for instance, follows whichever package the pinned anthropic release is built on.

TLS certificates come from the operating system

httpx2 verifies certificates against the operating system trust store, not a bundled one. Nothing extra is needed on a normal machine or a standard base image, but a minimal container without system CA certificates will fail to connect. Install the distribution's CA bundle (e.g. ca-certificates), or point the client at one explicitly:

import httpx2
from ag2.config import OpenAIConfig

config = OpenAIConfig(model="gpt-4o", http_client=httpx2.AsyncClient(verify="/path/to/ca-bundle.crt"))

certifi remains installed — ag2's own core depends on httpx — but the OpenAI client no longer reads it. Its presence is not evidence that TLS is configured there.

Extra Body Parameters#

Some OpenAI API-compatible providers require additional, provider-specific parameters in the request body. Use the extra_body parameter on OpenAIConfig to pass these through directly to the API call.

This is useful for enabling features like extended thinking on self-hosted or third-party models:

1
2
3
4
5
6
7
8
from ag2.config import OpenAIConfig

# NVIDIA NIM
nemotron = OpenAIConfig(
    model="nvidia/nemotron-3-super-120b-a12b",
    base_url="https://integrate.api.nvidia.com/v1",
    extra_body={"chat_template_kwargs": {"thinking": True}},
)

Reusing and Overriding Configurations#

Model configurations are immutable. If you need to reuse a configuration for multiple agents with slight variations (e.g., changing the model version or adjusting the temperature), use the .copy() method. This creates a new updated instance without mutating the original configuration.

from ag2.agent import Agent
from ag2.config import OpenAIConfig

base_config = OpenAIConfig(model="gpt-5")

agent1 = Agent(
    "Assistant",
    # Create a new configuration with updated temperature
    config=base_config.copy(temperature=0.2),
)

agent2 = Agent(
    "AnotherAssistant",
    # Create a new configuration with updated model and temperature
    config=base_config.copy(model="gpt-5-mini", temperature=0.8),
)

Delaying Model Configuration#

In many use cases, you may want to separate the logic of defining your agent (tools, system messages, instructions) from configuring the specific model it uses. This allows you to construct an agent once and dynamically provide the model configuration later during execution.

You can accomplish this by passing the configuration to the .ask() method when interacting with the agent. This is especially useful for applications like web servers where the user might bring their own API key or choose a different model on the fly.

from ag2.agent import Agent
from ag2.config import OpenAIConfig

# Define an agent without an initial model config,
# or with a default one you plan to override later
agent = Agent(
    "Assistant",
    prompt="You are a helpful assistant.",
    # other tools and settings...
)

# Ask the agent, passing the explicit model configuration
response = await agent.ask(
    "Hello!",
    config=OpenAIConfig(
        model="gpt-5",
        api_key="sk-user-specific-key"
    )
)

Warning

Providing a configuration or client directly to the ask() method completely overrides the original model configuration assigned to the agent for that specific turn.

from ag2.agent import Agent
from ag2.config import OpenAIConfig

agent = Agent(
    "Assistant",
    config=OpenAIConfig(model="gpt-5"),
)

response = await agent.ask(
    "Hello!",
    # overrides the original model configuration
    config=OpenAIConfig(model="gpt-5-mini")
)