finetuning

作者: microsoft

在 Microsoft Foundry 上使用 SFT(監督式)、DPO(偏好)或 RFT(搭配評分者的強化)微調模型。涵蓋資料集準備、訓練作業…

npx skills add https://github.com/microsoft/skills --skill finetuning

Fine-Tuning on Microsoft Foundry

Fine-tune models using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset prep, training, deployment, and evaluation.

When to Use

Use this sub-skill when the user asks about:

  • Fine-tuning a model (SFT, DPO, or RFT)
  • Preparing, validating, or formatting training data
  • Submitting, monitoring, or diagnosing training jobs
  • Calibrating graders or pass thresholds for RFT
  • Deploying or evaluating a fine-tuned model
  • Choosing between training types (SFT vs DPO vs RFT)
  • Distillation, synthetic data generation, or dataset quality scoring
  • Large file uploads for training data
  • Cleaning up fine-tuning resources (files, deployments)

Do NOT use for: General model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).

Workflows

StageGuide
Quick startworkflows/quickstart.md
Full pipelineworkflows/full-pipeline.md
Create dataworkflows/dataset-creation.md
Iterateworkflows/iterative-training.md
Diagnoseworkflows/diagnose-poor-results.md

References

TopicFile
SFT vs DPO vs RFTreferences/training-types.md
Hyperparametersreferences/hyperparameters.md
Data formatsreferences/dataset-formats.md
Grader design (RFT)references/grader-design.md
Reward hackingreferences/reward-hacking.md
Agentic RFT (tools)references/agentic-rft.md
Deploymentreferences/deployment.md
Training curvesreferences/training-curves.md
Evaluationreferences/evaluation.md
Vision fine-tuningreferences/vision-fine-tuning.md
Large file uploadsreferences/large-file-uploads.md
Platform gotchasreferences/platform-gotchas.md

Scripts

ScriptPurpose
scripts/submit_training.pySubmit SFT/DPO/RFT jobs
scripts/monitor_training.pyPoll job until completion
scripts/calibrate_grader.pyFind optimal RFT pass_threshold
scripts/check_training.pyAnalyze curves, list checkpoints
scripts/deploy_model.pyDeploy via ARM REST API
scripts/evaluate_model.pyLLM judge evaluation
scripts/convert_dataset.pyConvert between SFT/DPO/RFT formats
scripts/generate_distillation_data.pyGenerate synthetic training data
scripts/score_dataset.pyQuality scoring on training data
scripts/cleanup.pyDelete old files and deployments
scripts/validate/Data validators (SFT, DPO, RFT) + stats

Rules

  1. Always baseline first — evaluate the base model before fine-tuning
  2. Validate data before submitting — run scripts/validate/validate_sft.py
  3. Calibrate RFT graders — target 25-50% failure rate on the base model
  4. Evaluate checkpoints — don't blindly deploy the final one
  5. Measure token cost alongside accuracy when comparing models

Quick Reference

TaskCommand
Validate SFT datapython scripts/validate/validate_sft.py data.jsonl
Submit SFT jobpython scripts/submit_training.py --model gpt-4.1-mini --training-file train.jsonl --validation-file val.jsonl --type sft
Monitor jobpython scripts/monitor_training.py --job-id ftjob-xxx
Analyze curvespython scripts/check_training.py --job-id ftjob-xxx
Deploy modelpython scripts/deploy_model.py --model-id ft:gpt-4.1-mini:... --name my-eval
Evaluate modelpython scripts/evaluate_model.py --deployment-name my-eval --test-file test.jsonl

Error Handling

ErrorCauseFix
"API version not supported"Older openai SDK on /v1/ endpointUpgrade to openai>=1.0
"does not support fine-tuning with Standard TrainingType"OSS model needs globalStandardUse --use-rest flag or script auto-falls back
Job stuck in post-training evalUnder-provisioned tool endpoint (RFT)Scale to S2+, enable Always On
"DeploymentNotReady" after ARM succeedsARM/data-plane race conditionDelete and recreate deployment, wait 5 min
Content safety block at deploymentPII-dense training dataRemove problematic document types

來自 microsoft 的更多技能

oss-growth
microsoft
開源增長駭客角色
agent-framework-azure-ai-py
microsoft
使用Microsoft Agent Framework Python SDK(agent-framework-azure-ai)构建Azure AI Foundry代理。适用于使用AzureAIAgentsProvider创建持久化代理、使用托管工具(代码解释器、文件搜索、网络搜索)、集成MCP服务器、管理对话线程或实现流式响应。涵盖函数工具、结构化输出和多工具代理。
development
airunway-aks-setup
microsoft
在AKS上設定AI Runway——從裸叢集到執行模型。涵蓋叢集驗證、控制器安裝、GPU評估、供應商設定及首次部署。時機:「設定AI Runway」、「上線AKS叢集」、「安裝AI Runway」、「airunway設定」、「部署模型至AKS」、「在AKS上進行GPU推論」、「在AKS上設定KAITO」、「在AKS上執行LLM」、「在AKS上使用vLLM」、「在AKS上設定模型服務」、「AI Runway控制器」。
devops
appinsights-instrumentation
microsoft
使用Azure Application Insights檢測Web應用程式的指南。提供遙測模式、SDK設定與組態參考。適用時機:如何檢測應用程式、App Insights SDK、遙測模式、什麼是App Insights、Application Insights指南、檢測範例、APM最佳實踐。
devops
applicationinsights-web-ts
microsoft
使用Application Insights JavaScript SDK(@microsoft/applicationinsights-web)為瀏覽器/Web應用程式進行檢測。適用於真實使用者監控(RUM)——頁面檢視、點擊、AJAX/fetch依賴、例外、自訂事件,以及與後端OpenTelemetry追蹤關聯的瀏覽器端GenAI代理追蹤。涵蓋SDK載入器指令碼與npm設定、框架擴充(React、React Native、Angular)、點擊分析、遙測初始化器,以及從瀏覽器發出的代理/工具/模型span的OTel GenAI語意慣例。
devops
azure-ai-anomalydetector-java
microsoft
使用適用於 Java 的 Azure AI 異常偵測器 SDK 建置異常偵測應用程式。在實作單變量/多變量異常偵測、時間序列分析或 AI 驅動監控時使用。
development
azure-ai-language-conversations-py
microsoft
使用 azure-ai-language-conversations Python SDK 實作對話語言理解(CLU)。當使用 ConversationAnalysisClient 分析對話意圖與實體、建置 NLP 功能,或將語言理解整合至應用程式時使用。
development
azure-ai-ml-py
microsoft
Azure Machine Learning SDK v2 for Python。用於機器學習工作區、作業、模型、資料集、計算資源與管線。 觸發詞:「azure-ai-ml」、「MLClient」、「workspace」、「model registry」、「training jobs」、「datasets」。
development