deploying-scalable-agents

作成者: microsoft

動作するエージェントプロトタイプを、Microsoft Foundry上でのスケーラブルで観測可能な本番デプロイメントに移行します。デプロイメントパターン(クライアントホスト型、ホスト型エージェント、…)をカバーします。

npx skills add https://github.com/microsoft/ai-agents-for-beginners --skill deploying-scalable-agents

Deploying Scalable Agents with Microsoft Foundry

Companion skill for Lesson 16 – Deploying Scalable Agents. Use it to help a learner move an agent from prototype to a scalable, observable production deployment. Ground every recommendation in the lesson content and the runnable notebook; do not invent Foundry APIs.

Triggers

Activate this skill when a learner wants to:

  • Deploy an agent to Microsoft Foundry as a hosted agent and make it versioned/observable.
  • Choose between client-hosted, hosted-agent, and agent-workflow deployment patterns.
  • Add model routing, response caching, or bounded concurrency to control latency and cost.
  • Add an evaluation gate so a bad agent version cannot ship.
  • Add a human-in-the-loop approval step for high-risk actions.
  • Instrument an agent with OpenTelemetry tracing for production observability.
  • Smoke-test a deployed agent as a fast post-deploy gate.

Core mental model

A production agent is mostly the operational skeleton around the model (~80%), not the model itself. Map every recommendation to one of these concerns:

ConcernPrototype → Production
Hostingnotebook → versioned hosted service
Identityyour az login → managed identity + scoped RBAC
Statein-memory → externalised thread/memory store
Failuretraceback → retries, fallbacks, alerts
Cost"a few cents" → tracked, routed, cached, budgeted
Qualityeyeballing → automated evaluation gate
Trustyou approve → policy + human-in-the-loop

Deployment patterns (pick one, or combine)

  1. Client-hosted — the reasoning loop runs in your process. Max control; you own scaling/state.
  2. Hosted agent (Foundry Agent Service) — Foundry hosts the loop, stores threads, enforces RBAC/content safety, shows the agent in the portal. Less control, far less operational surface.
  3. Agent workflow — multiple agents/tools composed into a graph with branching, approval nodes, and durable checkpoints.

Lifecycle (the loop that ships an agent)

create → version → evaluate (gate) → deploy hosted → observe online → collect failures → repeat. Offline evaluation is a gate, not an afterthought — a version does not ship unless it clears the threshold. Online observability feeds real failures back into the offline test set.

Scaling and cost levers (in priority order)

  1. Right-size the model — use the smallest model that passes the evaluation gate.
  2. Route by complexity — small/fast model for simple requests, large model for real reasoning (DIY classifier or Foundry Model Router).
  3. Cache — serve near-duplicate requests without a model call.
  4. Stateless design + bounded concurrency — externalise state; retry with backoff.

Key patterns to reproduce

Point the learner at these from the notebook 16-python-agent-framework.ipynb:

  • Request handler: cache → route by complexity → trace span → run → cache.
  • Evaluation gate: score an offline test set; return pass_rate >= threshold and only deploy if true.
  • Human approval: @tool(approval_mode="always_require") for actions like large refunds.
  • Tracing: wrap each request in tracer.start_as_current_span(...) and set attributes like routed.model, customer.id.

Smoke-testing a deployed agent

After deploy, verify the endpoint actually answers (a green deploy can still be silent). Use the AI Smoke Test action via .github/workflows/smoke-test.yml with the catalog in tests/. The runner POSTs each prompt to POST {project_endpoint}/agents/{agent_name}/endpoint/protocols/openai/responses and asserts on the reply text. The identity needs the Azure AI User role at Foundry project scope; the token audience must be https://ai.azure.com/.

Layer the gates: smoke test (reachable/responding, every deploy) → offline evaluation (good enough to ship, before promotion) → online evaluation (how is it doing in the wild, continuous).

Enterprise controls

  • RBAC: give each hosted agent a managed identity with least privilege.
  • MCP in production: treat every MCP server as an untrusted boundary — pin the version, scope its identity, validate outputs, rate-limit, never expose secrets.

Guardrails for the assistant

  • Prefer the canonical FoundryChatClient(...) + provider.as_agent(...) pattern used across the course.
  • Do not promise live-Azure results you have not verified; recommend the smoke-test workflow to confirm a deployment.
  • Keep evaluation and cost advice tied together: evaluation sets the quality floor, routing/caching keep cost near that floor.

microsoftのその他のスキル

oss-growth
microsoft
OSS成長ハッカーのペルソナ
agent-framework-azure-ai-py
microsoft
Microsoft Agent Framework Python SDK(agent-framework-azure-ai)を使用してAzure AI Foundryエージェントを構築します。AzureAIAgentsProviderを使用した永続的なエージェントの作成、ホスト型ツール(コードインタープリター、ファイル検索、ウェブ検索)の使用、MCPサーバーの統合、会話スレッドの管理、ストリーミング応答の実装時に使用します。関数ツール、構造化出力、マルチツールエージェントをカバーします。
development
airunway-aks-setup
microsoft
AKS上でAI Runwayをセットアップ — ベアクラスターからモデル実行まで。クラスター検証、コントローラーインストール、GPU評価、プロバイダー設定、初回デプロイをカバー。対象: 「AI Runwayのセットアップ」「AKSクラスターのオンボード」「AI Runwayのインストール」「airunway setup」「AKSへのモデルデプロイ」「AKSでのGPU推論」「AKSでのKAITOセットアップ」「AKSでのLLM実行」「AKSでのvLLM」「AKSでのモデルサービング設定」「AI Runwayコントローラー」。
devops
appinsights-instrumentation
microsoft
Azure Application Insightsを使用したWebアプリのインストルメンテーションに関するガイダンス。テレメトリパターン、SDKセットアップ、構成リファレンスを提供します。対象: アプリのインストルメンテーション方法、App Insights SDK、テレメトリパターン、App Insightsとは何か、Application Insightsガイダンス、インストルメンテーション例、APMベストプラクティス。
devops
applicationinsights-web-ts
microsoft
Application Insights JavaScript SDK(@microsoft/applicationinsights-web)を使用してブラウザ/Webアプリを計測します。Real User Monitoring(RUM)— ページビュー、クリック、AJAX/fetch依存関係、例外、カスタムイベント、およびバックエンドのOpenTelemetryトレースに関連付けられたブラウザ側のGenAIエージェントトレースに使用します。SDKローダースクリプトとnpmセットアップ、フレームワーク拡張機能(React、React Native、Angular)、Click Analytics、テレメトリ初期化子、およびブラウザから生成されるエージェント/ツール/モデルスパンのOTel GenAIセマンティック規約をカバーします。
devops
azure-ai-anomalydetector-java
microsoft
Azure AI Anomaly Detector SDK for Javaを使用して異常検出アプリケーションを構築します。単変量/多変量異常検出、時系列分析、またはAIを活用したモニタリングを実装する際に使用します。
development
azure-ai-language-conversations-py
microsoft
azure-ai-language-conversations Python SDKを使用して会話言語理解(CLU)を実装します。ConversationAnalysisClientを使用して会話の意図とエンティティを分析する場合、NLP機能を構築する場合、またはアプリケーションに言語理解を統合する場合に使用します。
development
azure-ai-ml-py
microsoft
Azure Machine Learning SDK v2 for Python。MLワークスペース、ジョブ、モデル、データセット、コンピュート、パイプラインに使用します。 トリガー: 「azure-ai-ml」、「MLClient」、「ワークスペース」、「モデルレジストリ」、「トレーニングジョブ」、「データセット」。
development