eval-engineering

作成者: langchain-ai

Iteratively inspect an agent repository and optional user-provided traces, interview the user, and create, run, and audit Harbor evals one at a time. Use for…

npx skills add https://github.com/langchain-ai/langchain-skills --skill eval-engineering

Eval Engineering

Work with the user to define, build, run, and audit Harbor tasks.

map harness + environment -> propose directions -> user chooses
-> draft specs -> user approves -> build + run + audit -> repeat

Use the latest Harbor release. Put task source under evals/. Build sequentially while a later task depends on an unproven Harness, Environment, or Verifier. Build independent tasks in parallel when the user requests it.

Boundaries

  • Task: instruction.md plus an Environment and Verifier.
  • Harness: the complete agent Harbor runs: model, prompts, loop, repository-defined tools, middleware/hooks, memory/session behavior, and Harbor adapter. Harbor calls this the Agent.
  • Environment: the container/world around the Harness: OS, files, backing data, services, identity, permissions, network, clock, and mutable state.
  • Verifier: the test script that independently scores final artifacts or resulting Environment state; it uses trajectory only when final state cannot provide the required evidence.

Repository-defined tool code belongs to the Harness. The data or service behind it belongs to the Environment. Example: a docs agent's search_docs definition and result parsing stay in the Harness; the frozen search index and its error behavior live in the Environment. If production supplies a tool server dynamically, keep that server in the Environment and preserve how the Harness discovers and calls it.

Score the requested outcome

  • For stateful work, score independently observed final Environment state first. Example: a booking exists for the requested room and no conflicting booking exists.
  • Keep ATIF as diagnostic evidence by default. Use trajectory or session evidence only when final state cannot establish the requirement, such as proving later user turns used the same session.
  • Do not require a tool name, subagent, retry count, exact number of updates, or exact wording unless that is the user-facing requirement.
  • Before building, state what the agent can see, the required user-visible outcome, prohibited effects, and materially equivalent outcomes that must pass. Do not score a hidden evaluator preference.

References

Read each reference when its decision appears:

  • Trace sourcing: select and analyze traces only when the user supplies a source.
  • Harness: identify the actual agent Harbor will run and preserve its behavior.
  • Task design: turn one selected capability into a judgeable request.
  • Environment building: choose live, frozen, or simulated backing data and services.
  • Multi-turn simulation: run scripted or LLM-generated user turns through one Harness session.
  • Verifier design: define independent evidence, scoring, and calibration.
  • Harbor: create, run, and inspect the Harbor task.

1. Map the Harness and production Environment

Start at the public agent entrypoint and follow reachable code.

Harness: entrypoint; prompts; models; loop; routing; retries; hooks; memory;
         repository-defined tools, inputs, outputs, and effects
Environment: files; records; indexes; services behind tools; identity;
             permissions; network; time; mutable state
Purpose: intended users, jobs, and useful outcomes
Evidence: tests, fixtures, issues, existing evals, and documented failures

Do not start services, install packages, or use credentials during mapping. Explain the map in the conversation and ask only what code cannot answer, such as “Which user job matters most?” or “What failure must this eval catch?”

If the user provides traces, read Trace sourcing. Use trace evidence only when it changes an eval direction, dependency behavior, realistic request, or failure case. Never treat the recorded answer as truth.

2. Propose eval directions

Offer two or three capabilities grounded in the map and any supplied traces:

Name: choose the correct account lookup
Example request: “What plan is account A on?”
Tests: looks up A, uses the returned plan, and does not invent account details
Needs: known account records behind the existing read-only lookup

Recommend one and explain why. The user chooses before implementation.

3. Draft and approve the specs

Read the Harness, task, Environment, and Verifier references. After the user chooses a direction, write:

evals/<task-id>/
├── harness.md
├── environment.md
└── task.md

These are control-plane review files beside the runnable task. Never copy or mount them into the Harness workspace or task image. task.md is the review spec; Harbor's instruction.md is the Harness-visible request created from the approved spec.

  • harness.md: entrypoint, preserved behavior, adapter, sessions, credentials, recorded evidence, and reconstruction differences.
  • environment.md: live/frozen/simulated dependencies, backend contracts, generated or copied data, schemas and relationships, storage, effects, reset, and fidelity limits.
  • task.md: capability, request, initial conditions, pass condition, Verifier evidence, and accepted alternatives.

For each dependency, recommend live, frozen, or simulated use. Read-only, low-cost services backed by hard-to-reproduce data are strong live candidates. Stable copied data is a strong frozen candidate. Writes, unstable services, and resettable state are strong simulation candidates. State required credential names for live use.

Print the full contents of all three specs in the terminal, keeping them concise. Show their paths and your recommendation, then ask the user to approve or revise them. Mark each spec approved only after explicit user approval. Do not build the Harbor task until all three are approved. If user feedback or implementation changes the request, Harness, Environment, or Verifier boundary, update the affected spec, show the change, and obtain approval again.

For multiple user turns, prefer fixed follow-ups when they do not depend on Harness responses. Use an LLM user only when replies must react, correct, reject, or stop; read the multi-turn reference and include simulator credentials in the proposal.

4. Build one Harbor task

evals/<task-id>/
├── task.toml
├── instruction.md
├── task.md
├── harness.md
├── environment.md
├── environment/
└── tests/

Use the approved Harness unchanged when possible. Add an adapter only when Harbor needs one to invoke it. Do not expose hidden truth, simulator instructions, verifier criteria, or judge credentials to the Harness.

Prefer programmatic checks for final state, artifacts, tests, and independently recomputed facts. Use an LLM judge only for meaning code cannot reasonably decide. Run deterministic checks before the judge; give the judge only the final artifact and independent evidence for that unresolved semantic question. Emit one primary reward.

5. Run and audit

Calibrate the Verifier with realistic cases from supplied traces, prior eval runs, or production-like task variants: a valid paraphrase, a plausible wrong result, and any known boundary case. Run them through the same Verifier command Harbor uses. Run the Harness through Harbor, then inspect:

  • Harness-recorded messages, model/tool calls, results, retries, and errors;
  • Environment-observed service results, initial/final state, and reset;
  • Verifier evidence, decision, reason, reward, and errors;
  • resolved Harness and Environment configuration.

For every zero reward, classify the evidence as a fair agent failure, Verifier defect, Environment defect or leak, or infrastructure error. Fix and rerun non-agent failures before treating them as evaluation results. If the Environment leaked the answer, a wrong answer passed, or a valid result failed, the eval is not complete.

For an LLM user, inspect representative correct, wrong, clarification, and stop paths. Revise its contract or model when its replies are implausible. Simulator termination is not success; the Verifier alone assigns reward.

6. Review and repeat

Explain the task path and exact Harbor command, request, Harness, Environment, run behavior, Verifier decision, and main limitation. Completion requires a real Harbor run, evidence that the Verifier measured the intended capability, and user approval. If continuing, reuse the available evidence and propose a distinct capability.

Invariants

  • One capability per Harbor task.
  • No production writes; reset mutable state between trials.
  • Keep hidden truth and simulator/judge credentials unavailable to the Harness.
  • Treat build, credential, reset, timeout, judge, and Verifier failures as infrastructure errors, not failed agent work.

langchain-aiのその他のスキル

langgraph-docs
langchain-ai
LangGraphのドキュメントにアクセスして、ステートフルなエージェントやマルチエージェントワークフローを構築できます。公式のLangGraph Pythonドキュメントを取得し、ステートマシン、グラフベースのエージェント設計、ヒューマンインザループパターンをカバーします。クエリの種類に応じて関連ドキュメントを優先します:ハウツー質問には実装ガイド、理論にはコンセプトページ、エンドツーエンドの例にはチュートリアル、技術詳細にはAPIリファレンスを提供します。自動的に2~4個の最も関連性の高いドキュメントURLを選択し、その内容を取得して回答します。
official
langgraph-human-in-the-loop
langchain-ai
グラフの実行を一時停止し、人間によるレビュー、承認、検証を経て、その入力を反映させて再開します。これには3つのコンポーネントが必要です:チェックポインター(InMemorySaverまたはPostgresSaver)、設定内のスレッドID、そしてJSONシリアライズ可能なインタラプトペイロードです。interrupt(value)は一時停止してデータを表示し、Command(resume=value)は再開してその値を一時停止したノードに返します。interrupt()より前のすべてのコードは再開時に再実行されるため、副作用は冪等である必要があります(insertではなくupsertを使用)。承認ワークフローをサポートします。
official
web-research
langchain-ai
このスキルはウェブリサーチに関連するリクエストに使用します。包括的なウェブリサーチを実施するための構造化されたアプローチを提供します。
official
langchain-oss-primer
langchain-ai
LangChain、Deep Agents、またはLangGraphエージェント構築プロジェクトでは、必ずここから始めてください。他のスキルを選択したり、何かを記述する前に必要な出発点です。
official
skill-creator
langchain-ai
エージェントの機能を専門知識、ワークフロー、ツール連携で拡張する効果的なスキルを作成するためのガイド。このスキルは、ユーザーが…のときに使用します。
official
social-media
langchain-ai
プラットフォーム固有のソーシャルメディア投稿を、調査に基づいたコンテンツと生成された補完画像とともに下書きします。LinkedInの投稿(1,300文字、プロフェッショナルなトーン)とTwitter/Xのスレッド(1ツイートあたり280文字、1/🧵形式)に対応。執筆前にサブエージェントに調査を委任し、その後調査結果を読み、正確性と関連性を確認します。generate_social_imageツールを使用して、小さな画面向けに最適化された大胆でコントラストの高い構図の、目を引くソーシャル画像を自動生成します。
official
deep-agents-memory
langchain-ai
Deep Agents向けのプラグイン可能なメモリおよびファイルバックエンド。エフェメラル、永続、ハイブリッドルーティングオプションを備えています。4種類のバックエンドタイプ:StateBackend(スレッドスコープ、エフェメラル)、StoreBackend(セッションをまたいだ永続)、FilesystemBackend(ローカル開発用の実際のディスクアクセス)、CompositeBackend(異なるパスを異なるバックエンドにルーティング)。FilesystemMiddlewareは6つのファイル操作ツールを提供:ls、read_file、write_file、edit_file、glob、grep。CompositeBackendは最長プレフィックス一致を使用してルーティングします...
official
deep-agents-orchestration
langchain-ai
サブエージェントを調整し、複数ステップのタスクを計画し、機密操作には人間の承認を必要とします。タスクツールを介して専門サブエージェントに作業を委任します。カスタムサブエージェントは独立したツールセットとシステムプロンプトをサポートし、デフォルトの「汎用」サブエージェントはメインエージェント設定を継承します。write_todosを使用して複雑なワークフローを計画・追跡し、保留中、進行中、完了済みの状態でタスクを整理します。呼び出し間での永続性のためにthread_idが必要です。実装...
official