AI Clients
Connect to one or more LLM providers. Each provider registers an
IAIClientProviderthat creates typed AI clients.
Quick Start
Register the providers you use:
builder.Services
.AddCoreAIServices()
.AddCoreAIOrchestration()
.AddCoreAIOpenAI() // OpenAI (api.openai.com)
.AddCoreAIAzureOpenAI() // Azure OpenAI Service
.AddCoreAIOllama() // Ollama (local models)
.AddCoreAIAzureAIInference(); // Azure AI Inference / GitHub Models
You only need to register the providers you actually use.
Architecture
Each provider follows the same pattern:
- Registers an
IAIClientProvider— Creates chat clients, embedding generators, image generators, etc. - Registers an
IAICompletionClient— Handles completion requests for that provider - Registers a connection source — Provides connection metadata (API keys, endpoints)
IAIClientFactory
│
├── OpenAIClientProvider
│ └── Creates OpenAI.ChatClient
│
├── AzureOpenAIClientProvider
│ └── Creates AzureOpenAI.ChatClient
│
├── OllamaAIClientProvider
│ └── Creates Ollama ChatClient
│
└── AzureAIInferenceClientProvider
└── Creates Azure.AI.Inference ChatClient
Client Connection
Each provider needs at least one connection that stores credentials:
public class AIProviderConnectionEntry
{
public string Name { get; set; } // Unique connection name
public string ClientName { get; set; } // "OpenAI", "Azure", "Ollama", etc.
public string GetApiKey(); // API key
public Uri GetEndpoint(); // Endpoint URL (optional for OpenAI)
}
Connections are typically stored in a configuration file or database and loaded at startup. See the MVC Example for a complete setup.
Adding a Custom Provider
Implement these interfaces:
IAIClientProvider— Creates client instancesIAICompletionClient— Handles completions
public sealed class MyProviderClientProvider : IAIClientProvider
{
public bool CanHandle(string clientName)
{
return string.Equals(clientName, "MyProvider", StringComparison.OrdinalIgnoreCase);
}
public ValueTask<IChatClient> GetChatClientAsync(
AIProviderConnectionEntry connection, string deploymentName)
{
// Create and return your chat client
return ValueTask.FromResult<IChatClient>(new MyProviderChatClient());
}
// Implement other client creation methods...
}
public sealed class MyProviderCompletionClient : IAICompletionClient
{
public async Task<ChatResponse> CompleteAsync(
AICompletionContext context,
CancellationToken cancellationToken = default)
{
// Send completion request to your provider
}
}
Register:
builder.Services.AddScoped<IAIClientProvider, MyProviderClientProvider>();
builder.Services.AddCoreAICompletionClient<MyProviderCompletionClient>("MyProvider");
builder.Services.AddCoreAIConnectionSource("MyProvider", configure => { /* ... */ });
Available Clients
| Client | Extension | ClientName | Documentation |
|---|---|---|---|
| OpenAI | AddCoreAIOpenAI() | "OpenAI" | OpenAI |
| Azure OpenAI | AddCoreAIAzureOpenAI() | "Azure" | Azure OpenAI |
| Ollama | AddCoreAIOllama() | "Ollama" | Ollama |
| Azure AI Inference | AddCoreAIAzureAIInference() | "AzureAIInference" | Azure AI Inference |
Real-time (speech-to-speech) clients
A real-time client exchanges audio (and optionally text) with the provider over a persistent, bidirectional
session, enabling low-latency speech-to-speech conversations without separate speech-to-text and text-to-speech
steps. The client is the Microsoft.Extensions.AI IRealtimeClient / IRealtimeClientSession type, created from a
resolved deployment:
#pragma warning disable MEAI001 // The Microsoft.Extensions.AI realtime API is evaluation-only.
IRealtimeClient client = await clientFactory.CreateRealtimeClientAsync(deployment);
await using var session = await client.CreateSessionAsync(new RealtimeSessionOptions
{
Instructions = "You are a friendly voice assistant.",
Voice = "alloy",
InputAudioFormat = new RealtimeAudioFormat("audio/pcm", 24000),
OutputAudioFormat = new RealtimeAudioFormat("audio/pcm", 24000),
OutputModalities = ["audio"],
VoiceActivityDetection = new VoiceActivityDetectionOptions { Enabled = true },
});
IAIClientFactory.CreateRealtimeClientAsync(AIDeployment)resolves the provider and connection, then returns the realtime client. Under the hood each provider implementsIAIClientProvider.GetRealtimeClientAsync(connection, deploymentName); providers without a realtime API throwNotSupportedException.- OpenAI and Azure OpenAI support realtime (both use the OpenAI
gpt-realtimemodels). Ollama and Azure AI Inference do not. - Declare the
realtimemodel capability on the deployment; it then fills therealtimedeployment slot. A chat profile or interaction becomes a voice conversation simply by selecting that deployment as its chat deployment; there is no separate realtime mode or field.
Realtime voices
Realtime voices are a fixed, provider-defined set (the realtime API has no enumeration endpoint). Resolve them the same way as speech voices — through a resolver that delegates to the provider:
SpeechVoice[] voices = await realtimeVoiceResolver.GetVoicesAsync(deployment);
IRealtimeVoiceResolver.GetVoicesAsync(AIDeployment)(registered by default) mirrorsISpeechVoiceResolverand returnsSpeechVoice[]. It delegates toIAIClientProvider.GetRealtimeVoicesAsync(...).- OpenAI and Azure OpenAI return the
gpt-realtimevoice set —alloy,ash,ballad,cedar,coral,echo,marin,sage,shimmer,verse— declared inOpenAIRealtimeVoices. Each carries a best-effort (unofficial) gender to help group a voice selector.
Realtime sessions apply the profile's system prompt, tools, retrieval-augmented knowledge, and orchestration. Use Chat Interactions, AI Profile chat sessions, or the admin chat widget to exercise realtime sessions end to end.
Provider Comparison
| Capability | OpenAI | Azure OpenAI | Ollama | Azure AI Inference |
|---|---|---|---|---|
| Chat completions | ✅ | ✅ | ✅ | ✅ |
| Streaming | ✅ | ✅ | ✅ | ✅ |
| Function calling | ✅ | ✅ | ⚠️ Model-dependent | ⚠️ Model-dependent |
| Embeddings | ✅ | ✅ | ✅ | ✅ |
| Image generation | ✅ (DALL·E) | ✅ (DALL·E) | ❌ | ❌ |
| Speech-to-text | ✅ (Whisper) | ✅ (via Azure Speech) | ❌ | ❌ |
| Text-to-speech | ✅ | ✅ (via Azure Speech) | ❌ | ❌ |
| Realtime (speech-to-speech) | ✅ | ✅ | ❌ | ❌ |
| Vision (image input) | ✅ | ✅ | ⚠️ Model-dependent | ⚠️ Model-dependent |
| Managed identity | ❌ | ✅ | N/A | ✅ |
| Data residency | ❌ | ✅ (per region) | ✅ (local) | ✅ (per region) |
| Cost tier | Pay-per-token | Pay-per-token | Free (self-hosted) | Pay-per-token |
When to Choose Which Provider
| Scenario | Recommended Provider | Why |
|---|---|---|
| Prototyping / getting started | OpenAI | Simplest setup — just an API key |
| Enterprise production | Azure OpenAI | Data residency, SLAs, managed identity, VNET support |
| Local development | Ollama | No API costs, fast iteration, offline capable |
| Privacy-sensitive workloads | Ollama | Data never leaves your infrastructure |
| Multi-model exploration | Azure AI Inference | Access GPT, Llama, Mistral, Cohere through a single endpoint |
| GitHub-integrated workflows | Azure AI Inference | Use your GitHub token to access models via GitHub Models |
| Image generation | OpenAI or Azure OpenAI | Only providers with DALL·E support |
| Speech capabilities | OpenAI or Azure OpenAI | Only providers with Whisper/TTS support |
You can register multiple providers simultaneously and assign different profiles to different providers. For example, use Ollama for development and Azure OpenAI for production by switching connection names per environment.
Custom Provider Walkthrough
To add a provider for a service not covered by the built-in providers, implement three components:
Step 1: Implement IAIClientProvider
The client provider creates typed AI clients (chat, embedding, image) from a connection entry:
public sealed class MyProviderClientProvider : IAIClientProvider
{
public bool CanHandle(string clientName)
{
return string.Equals(clientName, "MyProvider", StringComparison.OrdinalIgnoreCase);
}
public ValueTask<IChatClient> GetChatClientAsync(
AIProviderConnectionEntry connection,
string deploymentName)
{
var apiKey = connection.GetApiKey();
var endpoint = connection.GetEndpoint()
?? new Uri("https://api.myprovider.com");
// Use Microsoft.Extensions.AI abstractions
return ValueTask.FromResult<IChatClient>(
new MyProviderChatClient(endpoint, apiKey, deploymentName));
}
public IEmbeddingGenerator<string, Embedding<float>> CreateEmbeddingGenerator(
AIProviderConnectionEntry connection,
string deploymentName)
{
var apiKey = connection.GetApiKey();
var endpoint = connection.GetEndpoint()
?? new Uri("https://api.myprovider.com");
return new MyProviderEmbeddingGenerator(endpoint, apiKey, deploymentName);
}
// Return null for capabilities the provider does not support
public object CreateImageGenerator(
AIProviderConnectionEntry connection,
string deploymentName)
=> null;
}
Step 2: Implement IAICompletionClient
The completion client handles the request/response cycle:
public sealed class MyProviderCompletionClient(
IAIClientFactory clientFactory,
ILogger<MyProviderCompletionClient> logger) : IAICompletionClient
{
public async Task<ChatResponse> CompleteAsync(
AICompletionContext context,
CancellationToken cancellationToken = default)
{
var chatClient = clientFactory.GetChatClient(context);
if (chatClient is null)
{
logger.LogWarning("No chat client available for connection '{Name}'.",
context.ConnectionName);
return ChatResponse.Empty;
}
var options = new ChatOptions
{
Temperature = context.Profile.Temperature,
MaxOutputTokens = context.Profile.MaxOutputTokens,
};
// Delegate to the Microsoft.Extensions.AI IChatClient
return await chatClient.GetResponseAsync(
context.Messages,
options,
cancellationToken);
}
}
Step 3: Register Connection Source and Services
public static class MyProviderServiceExtensions
{
public static AIServiceBuilder AddMyProvider(this AIServiceBuilder builder)
{
var services = builder.Services;
// Register the client provider
services.AddScoped<IAIClientProvider, MyProviderClientProvider>();
// Register the completion client for this client name
services.AddCoreAICompletionClient<MyProviderCompletionClient>("MyProvider");
// Register the connection source (how credentials are loaded)
services.AddCoreAIConnectionSource("MyProvider", options =>
{
// Connections can be loaded from configuration, database, etc.
options.Connections.Add(new AIProviderConnectionEntry
{
Name = "my-connection",
ClientName = "MyProvider",
});
});
return builder;
}
}
Use it:
builder.Services
.AddCoreAIServices()
.AddCoreAIOrchestration()
.AddMyProvider();
Fallback Strategies
The framework does not include automatic provider failover, but you can implement fallback logic at the application level:
Connection-Level Fallback
Register multiple connections for different providers and switch on failure:
public sealed class FallbackCompletionService(
IEnumerable<IAICompletionClient> clients,
ILogger<FallbackCompletionService> logger)
{
private readonly string[] _providerOrder = ["Azure", "OpenAI", "Ollama"];
public async Task<ChatResponse> CompleteWithFallbackAsync(
AICompletionContext context,
CancellationToken cancellationToken)
{
foreach (var providerName in _providerOrder)
{
var client = clients.FirstOrDefault(
c => c.GetType().Name.Contains(providerName));
if (client is null)
{
continue;
}
try
{
return await client.CompleteAsync(context, cancellationToken);
}
catch (Exception ex)
{
logger.LogWarning(ex,
"Provider '{Provider}' failed, trying next.", providerName);
}
}
throw new InvalidOperationException("All AI providers failed.");
}
}
Profile-Level Fallback
Assign a primary and fallback connection at the profile level:
{
"Profiles": {
"my-chat": {
"ConnectionName": "azure-primary",
"FallbackConnectionName": "openai-backup"
}
}
}
When implementing fallback logic, be mindful of token format differences between providers. A conversation started with one provider's tokenizer may behave differently when sent to another provider mid-stream.