finetuning

作者: microsoft

在 Microsoft Foundry 上使用 SFT(监督)、DPO(偏好)或 RFT(带评估器的强化学习)微调模型。涵盖数据集准备、训练作业…

npx skills add https://github.com/microsoft/skills --skill finetuning

Fine-Tuning on Microsoft Foundry

Fine-tune models using SFT (supervised), DPO (preference), or RFT (reinforcement with graders). Covers dataset prep, training, deployment, and evaluation.

When to Use

Use this sub-skill when the user asks about:

  • Fine-tuning a model (SFT, DPO, or RFT)
  • Preparing, validating, or formatting training data
  • Submitting, monitoring, or diagnosing training jobs
  • Calibrating graders or pass thresholds for RFT
  • Deploying or evaluating a fine-tuned model
  • Choosing between training types (SFT vs DPO vs RFT)
  • Distillation, synthetic data generation, or dataset quality scoring
  • Large file uploads for training data
  • Cleaning up fine-tuning resources (files, deployments)

Do NOT use for: General model deployment without fine-tuning (use deploy-model), agent creation (use agents), prompt optimization without training (use prompt-optimizer).

Workflows

StageGuide
Quick startworkflows/quickstart.md
Full pipelineworkflows/full-pipeline.md
Create dataworkflows/dataset-creation.md
Iterateworkflows/iterative-training.md
Diagnoseworkflows/diagnose-poor-results.md

References

TopicFile
SFT vs DPO vs RFTreferences/training-types.md
Hyperparametersreferences/hyperparameters.md
Data formatsreferences/dataset-formats.md
Grader design (RFT)references/grader-design.md
Reward hackingreferences/reward-hacking.md
Agentic RFT (tools)references/agentic-rft.md
Deploymentreferences/deployment.md
Training curvesreferences/training-curves.md
Evaluationreferences/evaluation.md
Vision fine-tuningreferences/vision-fine-tuning.md
Large file uploadsreferences/large-file-uploads.md
Platform gotchasreferences/platform-gotchas.md

Scripts

ScriptPurpose
scripts/submit_training.pySubmit SFT/DPO/RFT jobs
scripts/monitor_training.pyPoll job until completion
scripts/calibrate_grader.pyFind optimal RFT pass_threshold
scripts/check_training.pyAnalyze curves, list checkpoints
scripts/deploy_model.pyDeploy via ARM REST API
scripts/evaluate_model.pyLLM judge evaluation
scripts/convert_dataset.pyConvert between SFT/DPO/RFT formats
scripts/generate_distillation_data.pyGenerate synthetic training data
scripts/score_dataset.pyQuality scoring on training data
scripts/cleanup.pyDelete old files and deployments
scripts/validate/Data validators (SFT, DPO, RFT) + stats

Rules

  1. Always baseline first — evaluate the base model before fine-tuning
  2. Validate data before submitting — run scripts/validate/validate_sft.py
  3. Calibrate RFT graders — target 25-50% failure rate on the base model
  4. Evaluate checkpoints — don't blindly deploy the final one
  5. Measure token cost alongside accuracy when comparing models

Quick Reference

TaskCommand
Validate SFT datapython scripts/validate/validate_sft.py data.jsonl
Submit SFT jobpython scripts/submit_training.py --model gpt-4.1-mini --training-file train.jsonl --validation-file val.jsonl --type sft
Monitor jobpython scripts/monitor_training.py --job-id ftjob-xxx
Analyze curvespython scripts/check_training.py --job-id ftjob-xxx
Deploy modelpython scripts/deploy_model.py --model-id ft:gpt-4.1-mini:... --name my-eval
Evaluate modelpython scripts/evaluate_model.py --deployment-name my-eval --test-file test.jsonl

Error Handling

ErrorCauseFix
"API version not supported"Older openai SDK on /v1/ endpointUpgrade to openai>=1.0
"does not support fine-tuning with Standard TrainingType"OSS model needs globalStandardUse --use-rest flag or script auto-falls back
Job stuck in post-training evalUnder-provisioned tool endpoint (RFT)Scale to S2+, enable Always On
"DeploymentNotReady" after ARM succeedsARM/data-plane race conditionDelete and recreate deployment, wait 5 min
Content safety block at deploymentPII-dense training dataRemove problematic document types

来自 microsoft 的更多技能

oss-growth
microsoft
OSS增长黑客角色
agent-framework-azure-ai-py
microsoft
使用Microsoft Agent Framework Python SDK(agent-framework-azure-ai)构建Azure AI Foundry代理。在创建使用AzureAIAgentsProvider的持久化代理、使用托管工具(代码解释器、文件搜索、网络搜索)、集成MCP服务器、管理对话线程或实现流式响应时使用。涵盖函数工具、结构化输出和多工具代理。
development
airunway-aks-setup
microsoft
Set up AI Runway on AKS — from bare cluster to running model. Covers cluster verification, controller install, GPU assessment, provider setup, and first deployment. WHEN: "setup AI Runway", "onboard AKS cluster", "install AI Runway", "airunway setup", "deploy model to AKS", "GPU inference on AKS", "KAITO setup on AKS", "run LLM on AKS", "vLLM on AKS", "set up model serving on AKS", "AI Runway controller".
devops
appinsights-instrumentation
microsoft
使用Azure Application Insights对Web应用进行插桩的指南。提供遥测模式、SDK设置和配置参考。适用场景:如何对应用进行插桩、App Insights SDK、遥测模式、什么是App Insights、Application Insights指南、插桩示例、APM最佳实践。
devops
applicationinsights-web-ts
microsoft
使用Application Insights JavaScript SDK(@microsoft/applicationinsights-web)为浏览器/Web应用添加检测。用于真实用户监控(RUM)——页面视图、点击、AJAX/fetch依赖项、异常、自定义事件,以及与后端OpenTelemetry追踪关联的浏览器端GenAI代理追踪。涵盖SDK加载器脚本和npm设置、框架扩展(React、React Native、Angular)、点击分析、遥测初始化器,以及从浏览器发出的代理/工具/模型跨度所遵循的OTel GenAI语义约定。
devops
azure-ai-anomalydetector-java
microsoft
使用适用于 Java 的 Azure AI 异常检测器 SDK 构建异常检测应用程序。在实现单变量/多变量异常检测、时间序列分析或 AI 驱动的监控时使用。
development
azure-ai-language-conversations-py
microsoft
使用azure-ai-language-conversations Python SDK实现对话语言理解(CLU)。当使用ConversationAnalysisClient分析对话意图和实体、构建NLP功能或将语言理解集成到应用程序中时使用。
development
azure-ai-ml-py
microsoft
Azure Machine Learning SDK v2 for Python。用于机器学习工作区、作业、模型、数据集、计算资源和管道。 触发词:“azure-ai-ml”、“MLClient”、“工作区”、“模型注册表”、“训练作业”、“数据集”。
development