capacity

Découvre la capacité disponible des modèles Azure OpenAI dans les régions et les projets. Analyse les limites de quota, compare la disponibilité et recommande un déploiement optimal…

npx skills add https://github.com/microsoft/skills --skill capacity

Capacity Discovery

Finds available Azure OpenAI model capacity across all accessible regions and projects. Recommends the best deployment location based on capacity requirements.

Quick Reference

PropertyDescription
PurposeFind where you can deploy a model with sufficient capacity
ScopeAll regions and projects the user has access to
OutputRanked table of regions/projects with available capacity
ActionRead-only analysis — does NOT deploy. Hands off to preset or customize
AuthenticationAzure CLI (az login)

When to Use This Skill

  • ✅ User asks "where can I deploy gpt-4o?"
  • ✅ User specifies a capacity target: "find a region with 10K TPM for gpt-4o"
  • ✅ User wants to compare availability: "which regions have gpt-4o available?"
  • ✅ User got a quota error and needs to find an alternative location
  • ✅ User asks "best region and project for deploying model X"

After discovery → hand off to preset or customize for actual deployment.

Scripts

Pre-built scripts handle the complex REST API calls and data processing. Use these instead of constructing commands manually.

ScriptPurposeUsage
scripts/discover_and_rank.ps1Full discovery: capacity + projects + rankingPrimary script for capacity discovery
scripts/discover_and_rank.shSame as above (bash)Primary script for capacity discovery
scripts/query_capacity.ps1Raw capacity query (no project matching)Quick capacity check or version listing
scripts/query_capacity.shSame as above (bash)Quick capacity check or version listing

Workflow

Phase 1: Validate Prerequisites

az account show --query "{Subscription:name, SubscriptionId:id}" --output table

Phase 2: Identify Model and Version

Extract model name from user prompt. If version is unknown, query available versions:

.\scripts\query_capacity.ps1 -ModelName <model-name>
./scripts/query_capacity.sh <model-name>

This lists available versions. Use the latest version unless user specifies otherwise.

Phase 3: Run Discovery

Run the full discovery script with model name, version, and minimum capacity target:

.\scripts\discover_and_rank.ps1 -ModelName <model-name> -ModelVersion <version> -MinCapacity <target>
./scripts/discover_and_rank.sh <model-name> <version> <min-capacity>

💡 The script automatically queries capacity across ALL regions, cross-references with the user's existing projects, and outputs a ranked table sorted by: meets target → project count → available capacity.

Phase 3.5: Validate Subscription Quota

After discovery identifies candidate regions, validate that the user's subscription actually has available quota in each region. Model capacity (from Phase 3) shows what the platform can support, but subscription quota limits what this specific user can deploy.

# For each candidate region from discovery results:
$usageData = az cognitiveservices usage list --location <region> --subscription $SUBSCRIPTION_ID -o json 2>$null | ConvertFrom-Json

# Check quota for each SKU the model supports
# Quota names follow pattern: OpenAI.<SKU>.<model-name>
$usageEntry = $usageData | Where-Object { $_.name.value -eq "OpenAI.<SKU>.<model-name>" }

if ($usageEntry) {
  $quotaAvailable = $usageEntry.limit - $usageEntry.currentValue
} else {
  $quotaAvailable = 0  # No quota allocated
}
# For each candidate region from discovery results:
usage_json=$(az cognitiveservices usage list --location <region> --subscription "$SUBSCRIPTION_ID" -o json 2>/dev/null)

# Extract quota for specific SKU+model
quota_available=$(echo "$usage_json" | jq -r --arg name "OpenAI.<SKU>.<model-name>" \
  '.[] | select(.name.value == $name) | .limit - .currentValue')

Annotate discovery results:

Add a "Quota Available" column to the ranked output from Phase 3:

RegionAvailable CapacityMeets TargetProjectsQuota Available
eastus2120K TPM3✅ 80K
westus390K TPM1❌ 0 (at limit)
swedencentral100K TPM0✅ 100K

Regions/SKUs where quotaAvailable = 0 should be marked with ❌ in the results. If no region has available quota, hand off to the quota skill for increase requests and troubleshooting.

Phase 4: Present Results and Hand Off

After the script outputs the ranked table (now annotated with quota info), present it to the user and ask:

  1. 🚀 Quick deploy to top recommendation with defaults → route to preset
  2. ⚙️ Custom deploy with version/SKU/capacity/RAI selection → route to customize
  3. 📊 Check another model or capacity target → re-run Phase 2
  4. ❌ Cancel

Phase 5: Confirm Project Before Deploying

Before handing off to preset or customize, always confirm the target project with the user. See the Project Selection rules in the parent router.

If the discovery table shows a sample project for the chosen region, suggest it as the default. Otherwise, query projects in that region and let the user pick.

Error Handling

ErrorCauseResolution
"No capacity found"Model not available or all at quotaHand off to quota skill for increase requests and troubleshooting
Script auth erroraz login expiredRe-run az login
Empty version listModel not in region catalogTry a different region: ./scripts/query_capacity.sh <model> "" eastus
"No projects found"No AI Services resourcesGuide to project/create skill or Azure Portal

Related Skills

  • preset — Quick deployment after capacity discovery
  • customize — Custom deployment after capacity discovery
  • quota — For quota viewing, increase requests, and troubleshooting quota errors, defer to this skill instead of duplicating guidance

Plus de skills de microsoft

oss-growth
microsoft
Persona de growth hacker OSS
agent-framework-azure-ai-py
microsoft
Créez des agents Azure AI Foundry à l’aide du SDK Python Microsoft Agent Framework (agent-framework-azure-ai). À utiliser lors de la création d’agents persistants avec AzureAIAgentsProvider, de l’utilisation d’outils hébergés (interpréteur de code, recherche de fichiers, recherche web), de l’intégration de serveurs MCP, de la gestion de fils de conversation ou de l’implémentation de réponses en streaming. Couvre les outils de fonction, les sorties structurées et les agents multi-outils.
development
airunway-aks-setup
microsoft
Configurez AI Runway sur AKS — du cluster nu au modèle en cours d'exécution. Couvre la vérification du cluster, l'installation du contrôleur, l'évaluation GPU, la configuration du fournisseur et le premier déploiement. QUAND : « configurer AI Runway », « intégrer un cluster AKS », « installer AI Runway », « configuration airunway », « déployer un modèle sur AKS », « inférence GPU sur AKS », « configuration KAITO sur AKS », « exécuter LLM sur AKS », « vLLM sur AKS », « configurer le service de modèles sur AKS », « contrôleur AI Runway ».
devops
appinsights-instrumentation
microsoft
Guidance for instrumenting webapps with Azure Application Insights. Provides telemetry patterns, SDK setup, and configuration references. WHEN: how to instrument app, App Insights SDK, telemetry patterns, what is App Insights, Application Insights guidance, instrumentation examples, APM best practices.
devops
applicationinsights-web-ts
microsoft
Instrumentez les applications navigateur/web avec le SDK JavaScript Application Insights (@microsoft/applicationinsights-web). Utilisez-le pour la surveillance des utilisateurs réels (RUM) — vues de page, clics, dépendances AJAX/fetch, exceptions, événements personnalisés et traces d’agents GenAI côté navigateur corrélées aux traces OpenTelemetry backend. Couvre le script de chargement du SDK et la configuration npm, les extensions de framework (React, React Native, Angular), Click Analytics, les initialiseurs de télémétrie et les conventions sémantiques OTel GenAI pour les spans d’agents/outils/modèles émises depuis le navigateur.
devops
azure-ai-anomalydetector-java
microsoft
Créez des applications de détection d'anomalies avec le SDK Azure AI Anomaly Detector pour Java. Utilisez-le lors de l'implémentation de la détection d'anomalies univariées/multivariées, de l'analyse de séries temporelles ou de la surveillance basée sur l'IA.
development
azure-ai-language-conversations-py
microsoft
Implémentez la compréhension du langage conversationnel (CLU) à l’aide du SDK Python azure-ai-language-conversations. Utilisez-le lorsque vous travaillez avec ConversationAnalysisClient pour analyser l’intention et les entités d’une conversation, créer des fonctionnalités de NLP ou intégrer la compréhension du langage dans des applications.
development
azure-ai-ml-py
microsoft
SDK v2 d’Azure Machine Learning pour Python. Utiliser pour les espaces de travail ML, les tâches, les modèles, les jeux de données, le calcul et les pipelines. Déclencheurs : « azure-ai-ml », « MLClient », « espace de travail », « registre de modèles », « tâches d’entraînement », « jeux de données ».
development