Skip to main content
Version: Latest

AI Clients

Connect to one or more LLM providers. Each provider registers an IAIClientProvider that creates typed AI clients.

Quick Start​

Register the providers you use:

builder.Services
.AddCoreAIServices()
.AddCoreAIOrchestration()
.AddCoreAIOpenAI() // OpenAI (api.openai.com)
.AddCoreAIAzureOpenAI() // Azure OpenAI Service
.AddCoreAIOllama() // Ollama (local models)
.AddCoreAIAzureAIInference(); // Azure AI Inference / GitHub Models

You only need to register the providers you actually use.

Architecture​

Each provider follows the same pattern:

  1. Registers an IAIClientProvider — Creates chat clients, embedding generators, image generators, etc.
  2. Registers an IAICompletionClient — Handles completion requests for that provider
  3. Registers a connection source — Provides connection metadata (API keys, endpoints)
IAIClientFactory
│
├── OpenAIClientProvider
│ └── Creates OpenAI.ChatClient
│
├── AzureOpenAIClientProvider
│ └── Creates AzureOpenAI.ChatClient
│
├── OllamaAIClientProvider
│ └── Creates Ollama ChatClient
│
└── AzureAIInferenceClientProvider
└── Creates Azure.AI.Inference ChatClient

Client Connection​

Each provider needs at least one connection that stores credentials:

public class AIProviderConnectionEntry
{
public string Name { get; set; } // Unique connection name
public string ClientName { get; set; } // "OpenAI", "Azure", "Ollama", etc.
public string GetApiKey(); // API key
public Uri GetEndpoint(); // Endpoint URL (optional for OpenAI)
}

Connections are typically stored in a configuration file or database and loaded at startup. See the MVC Example for a complete setup.

Adding a Custom Provider​

Implement these interfaces:

  1. IAIClientProvider — Creates client instances
  2. IAICompletionClient — Handles completions
public sealed class MyProviderClientProvider : IAIClientProvider
{
public bool CanHandle(string clientName)
{
return string.Equals(clientName, "MyProvider", StringComparison.OrdinalIgnoreCase);
}

public ValueTask<IChatClient> GetChatClientAsync(
AIProviderConnectionEntry connection, string deploymentName)
{
// Create and return your chat client
return ValueTask.FromResult<IChatClient>(new MyProviderChatClient());
}

// Implement other client creation methods...
}

public sealed class MyProviderCompletionClient : IAICompletionClient
{
public async Task<ChatResponse> CompleteAsync(
AICompletionContext context,
CancellationToken cancellationToken = default)
{
// Send completion request to your provider
}
}

Register:

builder.Services.AddScoped<IAIClientProvider, MyProviderClientProvider>();
builder.Services.AddCoreAICompletionClient<MyProviderCompletionClient>("MyProvider");
builder.Services.AddCoreAIConnectionSource("MyProvider", configure => { /* ... */ });

Available Clients​

ClientExtensionClientNameDocumentation
OpenAIAddCoreAIOpenAI()"OpenAI"OpenAI
Azure OpenAIAddCoreAIAzureOpenAI()"Azure"Azure OpenAI
OllamaAddCoreAIOllama()"Ollama"Ollama
Azure AI InferenceAddCoreAIAzureAIInference()"AzureAIInference"Azure AI Inference

Real-time (speech-to-speech) clients​

A real-time client exchanges audio (and optionally text) with the provider over a persistent, bidirectional session, enabling low-latency speech-to-speech conversations without separate speech-to-text and text-to-speech steps. The client is the Microsoft.Extensions.AI IRealtimeClient / IRealtimeClientSession type, created from a resolved deployment:

#pragma warning disable MEAI001 // The Microsoft.Extensions.AI realtime API is evaluation-only.
IRealtimeClient client = await clientFactory.CreateRealtimeClientAsync(deployment);

await using var session = await client.CreateSessionAsync(new RealtimeSessionOptions
{
Instructions = "You are a friendly voice assistant.",
Voice = "alloy",
InputAudioFormat = new RealtimeAudioFormat("audio/pcm", 24000),
OutputAudioFormat = new RealtimeAudioFormat("audio/pcm", 24000),
OutputModalities = ["audio"],
VoiceActivityDetection = new VoiceActivityDetectionOptions { Enabled = true },
});
  • IAIClientFactory.CreateRealtimeClientAsync(AIDeployment) resolves the provider and connection, then returns the realtime client. Under the hood each provider implements IAIClientProvider.GetRealtimeClientAsync(connection, deploymentName); providers without a realtime API throw NotSupportedException.
  • OpenAI and Azure OpenAI support realtime (both use the OpenAI gpt-realtime models). Ollama and Azure AI Inference do not.
  • Declare the realtime model capability on the deployment; it then fills the realtime deployment slot. A chat profile or interaction becomes a voice conversation simply by selecting that deployment as its chat deployment; there is no separate realtime mode or field.

Realtime voices​

Realtime voices are a fixed, provider-defined set (the realtime API has no enumeration endpoint). Resolve them the same way as speech voices — through a resolver that delegates to the provider:

SpeechVoice[] voices = await realtimeVoiceResolver.GetVoicesAsync(deployment);
  • IRealtimeVoiceResolver.GetVoicesAsync(AIDeployment) (registered by default) mirrors ISpeechVoiceResolver and returns SpeechVoice[]. It delegates to IAIClientProvider.GetRealtimeVoicesAsync(...).
  • OpenAI and Azure OpenAI return the gpt-realtime voice set — alloy, ash, ballad, cedar, coral, echo, marin, sage, shimmer, verse — declared in OpenAIRealtimeVoices. Each carries a best-effort (unofficial) gender to help group a voice selector.
note

Realtime sessions apply the profile's system prompt, tools, retrieval-augmented knowledge, and orchestration. Use Chat Interactions, AI Profile chat sessions, or the admin chat widget to exercise realtime sessions end to end.

Provider Comparison​

CapabilityOpenAIAzure OpenAIOllamaAzure AI Inference
Chat completions✅✅✅✅
Streaming✅✅✅✅
Function calling✅✅⚠️ Model-dependent⚠️ Model-dependent
Embeddings✅✅✅✅
Image generation✅ (DALL·E)✅ (DALL·E)❌❌
Speech-to-text✅ (Whisper)✅ (via Azure Speech)❌❌
Text-to-speech✅✅ (via Azure Speech)❌❌
Realtime (speech-to-speech)✅✅❌❌
Vision (image input)✅✅⚠️ Model-dependent⚠️ Model-dependent
Managed identity❌✅N/A✅
Data residency❌✅ (per region)✅ (local)✅ (per region)
Cost tierPay-per-tokenPay-per-tokenFree (self-hosted)Pay-per-token

When to Choose Which Provider​

ScenarioRecommended ProviderWhy
Prototyping / getting startedOpenAISimplest setup — just an API key
Enterprise productionAzure OpenAIData residency, SLAs, managed identity, VNET support
Local developmentOllamaNo API costs, fast iteration, offline capable
Privacy-sensitive workloadsOllamaData never leaves your infrastructure
Multi-model explorationAzure AI InferenceAccess GPT, Llama, Mistral, Cohere through a single endpoint
GitHub-integrated workflowsAzure AI InferenceUse your GitHub token to access models via GitHub Models
Image generationOpenAI or Azure OpenAIOnly providers with DALL·E support
Speech capabilitiesOpenAI or Azure OpenAIOnly providers with Whisper/TTS support
tip

You can register multiple providers simultaneously and assign different profiles to different providers. For example, use Ollama for development and Azure OpenAI for production by switching connection names per environment.

Custom Provider Walkthrough​

To add a provider for a service not covered by the built-in providers, implement three components:

Step 1: Implement IAIClientProvider​

The client provider creates typed AI clients (chat, embedding, image) from a connection entry:

public sealed class MyProviderClientProvider : IAIClientProvider
{
public bool CanHandle(string clientName)
{
return string.Equals(clientName, "MyProvider", StringComparison.OrdinalIgnoreCase);
}

public ValueTask<IChatClient> GetChatClientAsync(
AIProviderConnectionEntry connection,
string deploymentName)
{
var apiKey = connection.GetApiKey();
var endpoint = connection.GetEndpoint()
?? new Uri("https://api.myprovider.com");

// Use Microsoft.Extensions.AI abstractions
return ValueTask.FromResult<IChatClient>(
new MyProviderChatClient(endpoint, apiKey, deploymentName));
}

public IEmbeddingGenerator<string, Embedding<float>> CreateEmbeddingGenerator(
AIProviderConnectionEntry connection,
string deploymentName)
{
var apiKey = connection.GetApiKey();
var endpoint = connection.GetEndpoint()
?? new Uri("https://api.myprovider.com");

return new MyProviderEmbeddingGenerator(endpoint, apiKey, deploymentName);
}

// Return null for capabilities the provider does not support
public object CreateImageGenerator(
AIProviderConnectionEntry connection,
string deploymentName)
=> null;
}

Step 2: Implement IAICompletionClient​

The completion client handles the request/response cycle:

public sealed class MyProviderCompletionClient(
IAIClientFactory clientFactory,
ILogger<MyProviderCompletionClient> logger) : IAICompletionClient
{
public async Task<ChatResponse> CompleteAsync(
AICompletionContext context,
CancellationToken cancellationToken = default)
{
var chatClient = clientFactory.GetChatClient(context);

if (chatClient is null)
{
logger.LogWarning("No chat client available for connection '{Name}'.",
context.ConnectionName);
return ChatResponse.Empty;
}

var options = new ChatOptions
{
Temperature = context.Profile.Temperature,
MaxOutputTokens = context.Profile.MaxOutputTokens,
};

// Delegate to the Microsoft.Extensions.AI IChatClient
return await chatClient.GetResponseAsync(
context.Messages,
options,
cancellationToken);
}
}

Step 3: Register Connection Source and Services​

public static class MyProviderServiceExtensions
{
public static AIServiceBuilder AddMyProvider(this AIServiceBuilder builder)
{
var services = builder.Services;

// Register the client provider
services.AddScoped<IAIClientProvider, MyProviderClientProvider>();

// Register the completion client for this client name
services.AddCoreAICompletionClient<MyProviderCompletionClient>("MyProvider");

// Register the connection source (how credentials are loaded)
services.AddCoreAIConnectionSource("MyProvider", options =>
{
// Connections can be loaded from configuration, database, etc.
options.Connections.Add(new AIProviderConnectionEntry
{
Name = "my-connection",
ClientName = "MyProvider",
});
});

return builder;
}
}

Use it:

builder.Services
.AddCoreAIServices()
.AddCoreAIOrchestration()
.AddMyProvider();

Fallback Strategies​

The framework does not include automatic provider failover, but you can implement fallback logic at the application level:

Connection-Level Fallback​

Register multiple connections for different providers and switch on failure:

public sealed class FallbackCompletionService(
IEnumerable<IAICompletionClient> clients,
ILogger<FallbackCompletionService> logger)
{
private readonly string[] _providerOrder = ["Azure", "OpenAI", "Ollama"];

public async Task<ChatResponse> CompleteWithFallbackAsync(
AICompletionContext context,
CancellationToken cancellationToken)
{
foreach (var providerName in _providerOrder)
{
var client = clients.FirstOrDefault(
c => c.GetType().Name.Contains(providerName));

if (client is null)
{
continue;
}

try
{
return await client.CompleteAsync(context, cancellationToken);
}
catch (Exception ex)
{
logger.LogWarning(ex,
"Provider '{Provider}' failed, trying next.", providerName);
}
}

throw new InvalidOperationException("All AI providers failed.");
}
}

Profile-Level Fallback​

Assign a primary and fallback connection at the profile level:

{
"Profiles": {
"my-chat": {
"ConnectionName": "azure-primary",
"FallbackConnectionName": "openai-backup"
}
}
}
warning

When implementing fallback logic, be mindful of token format differences between providers. A conversation started with one provider's tokenizer may behave differently when sent to another provider mid-stream.