deploying-scalable-agents
एक कार्यशील एजेंट प्रोटोटाइप को Microsoft Foundry पर स्केलेबल, अवलोकन योग्य प्रोडक्शन डिप्लॉयमेंट में ले जाएं। डिप्लॉयमेंट पैटर्न (क्लाइंट-होस्टेड, होस्टेड एजेंट,…) को शामिल करता है।
npx skills add https://github.com/microsoft/ai-agents-for-beginners --skill deploying-scalable-agentsDeploying Scalable Agents with Microsoft Foundry
Companion skill for Lesson 16 – Deploying Scalable Agents. Use it to help a learner move an agent from prototype to a scalable, observable production deployment. Ground every recommendation in the lesson content and the runnable notebook; do not invent Foundry APIs.
Triggers
Activate this skill when a learner wants to:
- Deploy an agent to Microsoft Foundry as a hosted agent and make it versioned/observable.
- Choose between client-hosted, hosted-agent, and agent-workflow deployment patterns.
- Add model routing, response caching, or bounded concurrency to control latency and cost.
- Add an evaluation gate so a bad agent version cannot ship.
- Add a human-in-the-loop approval step for high-risk actions.
- Instrument an agent with OpenTelemetry tracing for production observability.
- Smoke-test a deployed agent as a fast post-deploy gate.
Core mental model
A production agent is mostly the operational skeleton around the model (~80%), not the model itself. Map every recommendation to one of these concerns:
| Concern | Prototype → Production |
|---|---|
| Hosting | notebook → versioned hosted service |
| Identity | your az login → managed identity + scoped RBAC |
| State | in-memory → externalised thread/memory store |
| Failure | traceback → retries, fallbacks, alerts |
| Cost | "a few cents" → tracked, routed, cached, budgeted |
| Quality | eyeballing → automated evaluation gate |
| Trust | you approve → policy + human-in-the-loop |
Deployment patterns (pick one, or combine)
- Client-hosted — the reasoning loop runs in your process. Max control; you own scaling/state.
- Hosted agent (Foundry Agent Service) — Foundry hosts the loop, stores threads, enforces RBAC/content safety, shows the agent in the portal. Less control, far less operational surface.
- Agent workflow — multiple agents/tools composed into a graph with branching, approval nodes, and durable checkpoints.
Lifecycle (the loop that ships an agent)
create → version → evaluate (gate) → deploy hosted → observe online → collect failures → repeat.
Offline evaluation is a gate, not an afterthought — a version does not ship
unless it clears the threshold. Online observability feeds real failures back
into the offline test set.
Scaling and cost levers (in priority order)
- Right-size the model — use the smallest model that passes the evaluation gate.
- Route by complexity — small/fast model for simple requests, large model for real reasoning (DIY classifier or Foundry Model Router).
- Cache — serve near-duplicate requests without a model call.
- Stateless design + bounded concurrency — externalise state; retry with backoff.
Key patterns to reproduce
Point the learner at these from the notebook
16-python-agent-framework.ipynb:
- Request handler: cache → route by complexity → trace span → run → cache.
- Evaluation gate: score an offline test set; return
pass_rate >= thresholdand only deploy if true. - Human approval:
@tool(approval_mode="always_require")for actions like large refunds. - Tracing: wrap each request in
tracer.start_as_current_span(...)and set attributes likerouted.model,customer.id.
Smoke-testing a deployed agent
After deploy, verify the endpoint actually answers (a green deploy can still be
silent). Use the AI Smoke Test
action via .github/workflows/smoke-test.yml
with the catalog in tests/. The runner POSTs each
prompt to POST {project_endpoint}/agents/{agent_name}/endpoint/protocols/openai/responses
and asserts on the reply text. The identity needs the Azure AI User role at
Foundry project scope; the token audience must be https://ai.azure.com/.
Layer the gates: smoke test (reachable/responding, every deploy) → offline evaluation (good enough to ship, before promotion) → online evaluation (how is it doing in the wild, continuous).
Enterprise controls
- RBAC: give each hosted agent a managed identity with least privilege.
- MCP in production: treat every MCP server as an untrusted boundary — pin the version, scope its identity, validate outputs, rate-limit, never expose secrets.
Guardrails for the assistant
- Prefer the canonical
FoundryChatClient(...)+provider.as_agent(...)pattern used across the course. - Do not promise live-Azure results you have not verified; recommend the smoke-test workflow to confirm a deployment.
- Keep evaluation and cost advice tied together: evaluation sets the quality floor, routing/caching keep cost near that floor.