eval-engineering

작성자: langchain-ai

Iteratively inspect an agent repository and optional user-provided traces, interview the user, and create, run, and audit Harbor evals one at a time. Use for…

npx skills add https://github.com/langchain-ai/langchain-skills --skill eval-engineering

Eval Engineering

Work with the user to define, build, run, and audit Harbor tasks.

map harness + environment -> propose directions -> user chooses
-> draft specs -> user approves -> build + run + audit -> repeat

Use the latest Harbor release. Put task source under evals/. Build sequentially while a later task depends on an unproven Harness, Environment, or Verifier. Build independent tasks in parallel when the user requests it.

Boundaries

  • Task: instruction.md plus an Environment and Verifier.
  • Harness: the complete agent Harbor runs: model, prompts, loop, repository-defined tools, middleware/hooks, memory/session behavior, and Harbor adapter. Harbor calls this the Agent.
  • Environment: the container/world around the Harness: OS, files, backing data, services, identity, permissions, network, clock, and mutable state.
  • Verifier: the test script that independently scores final artifacts or resulting Environment state; it uses trajectory only when final state cannot provide the required evidence.

Repository-defined tool code belongs to the Harness. The data or service behind it belongs to the Environment. Example: a docs agent's search_docs definition and result parsing stay in the Harness; the frozen search index and its error behavior live in the Environment. If production supplies a tool server dynamically, keep that server in the Environment and preserve how the Harness discovers and calls it.

Score the requested outcome

  • For stateful work, score independently observed final Environment state first. Example: a booking exists for the requested room and no conflicting booking exists.
  • Keep ATIF as diagnostic evidence by default. Use trajectory or session evidence only when final state cannot establish the requirement, such as proving later user turns used the same session.
  • Do not require a tool name, subagent, retry count, exact number of updates, or exact wording unless that is the user-facing requirement.
  • Before building, state what the agent can see, the required user-visible outcome, prohibited effects, and materially equivalent outcomes that must pass. Do not score a hidden evaluator preference.

References

Read each reference when its decision appears:

  • Trace sourcing: select and analyze traces only when the user supplies a source.
  • Harness: identify the actual agent Harbor will run and preserve its behavior.
  • Task design: turn one selected capability into a judgeable request.
  • Environment building: choose live, frozen, or simulated backing data and services.
  • Multi-turn simulation: run scripted or LLM-generated user turns through one Harness session.
  • Verifier design: define independent evidence, scoring, and calibration.
  • Harbor: create, run, and inspect the Harbor task.

1. Map the Harness and production Environment

Start at the public agent entrypoint and follow reachable code.

Harness: entrypoint; prompts; models; loop; routing; retries; hooks; memory;
         repository-defined tools, inputs, outputs, and effects
Environment: files; records; indexes; services behind tools; identity;
             permissions; network; time; mutable state
Purpose: intended users, jobs, and useful outcomes
Evidence: tests, fixtures, issues, existing evals, and documented failures

Do not start services, install packages, or use credentials during mapping. Explain the map in the conversation and ask only what code cannot answer, such as “Which user job matters most?” or “What failure must this eval catch?”

If the user provides traces, read Trace sourcing. Use trace evidence only when it changes an eval direction, dependency behavior, realistic request, or failure case. Never treat the recorded answer as truth.

2. Propose eval directions

Offer two or three capabilities grounded in the map and any supplied traces:

Name: choose the correct account lookup
Example request: “What plan is account A on?”
Tests: looks up A, uses the returned plan, and does not invent account details
Needs: known account records behind the existing read-only lookup

Recommend one and explain why. The user chooses before implementation.

3. Draft and approve the specs

Read the Harness, task, Environment, and Verifier references. After the user chooses a direction, write:

evals/<task-id>/
├── harness.md
├── environment.md
└── task.md

These are control-plane review files beside the runnable task. Never copy or mount them into the Harness workspace or task image. task.md is the review spec; Harbor's instruction.md is the Harness-visible request created from the approved spec.

  • harness.md: entrypoint, preserved behavior, adapter, sessions, credentials, recorded evidence, and reconstruction differences.
  • environment.md: live/frozen/simulated dependencies, backend contracts, generated or copied data, schemas and relationships, storage, effects, reset, and fidelity limits.
  • task.md: capability, request, initial conditions, pass condition, Verifier evidence, and accepted alternatives.

For each dependency, recommend live, frozen, or simulated use. Read-only, low-cost services backed by hard-to-reproduce data are strong live candidates. Stable copied data is a strong frozen candidate. Writes, unstable services, and resettable state are strong simulation candidates. State required credential names for live use.

Print the full contents of all three specs in the terminal, keeping them concise. Show their paths and your recommendation, then ask the user to approve or revise them. Mark each spec approved only after explicit user approval. Do not build the Harbor task until all three are approved. If user feedback or implementation changes the request, Harness, Environment, or Verifier boundary, update the affected spec, show the change, and obtain approval again.

For multiple user turns, prefer fixed follow-ups when they do not depend on Harness responses. Use an LLM user only when replies must react, correct, reject, or stop; read the multi-turn reference and include simulator credentials in the proposal.

4. Build one Harbor task

evals/<task-id>/
├── task.toml
├── instruction.md
├── task.md
├── harness.md
├── environment.md
├── environment/
└── tests/

Use the approved Harness unchanged when possible. Add an adapter only when Harbor needs one to invoke it. Do not expose hidden truth, simulator instructions, verifier criteria, or judge credentials to the Harness.

Prefer programmatic checks for final state, artifacts, tests, and independently recomputed facts. Use an LLM judge only for meaning code cannot reasonably decide. Run deterministic checks before the judge; give the judge only the final artifact and independent evidence for that unresolved semantic question. Emit one primary reward.

5. Run and audit

Calibrate the Verifier with realistic cases from supplied traces, prior eval runs, or production-like task variants: a valid paraphrase, a plausible wrong result, and any known boundary case. Run them through the same Verifier command Harbor uses. Run the Harness through Harbor, then inspect:

  • Harness-recorded messages, model/tool calls, results, retries, and errors;
  • Environment-observed service results, initial/final state, and reset;
  • Verifier evidence, decision, reason, reward, and errors;
  • resolved Harness and Environment configuration.

For every zero reward, classify the evidence as a fair agent failure, Verifier defect, Environment defect or leak, or infrastructure error. Fix and rerun non-agent failures before treating them as evaluation results. If the Environment leaked the answer, a wrong answer passed, or a valid result failed, the eval is not complete.

For an LLM user, inspect representative correct, wrong, clarification, and stop paths. Revise its contract or model when its replies are implausible. Simulator termination is not success; the Verifier alone assigns reward.

6. Review and repeat

Explain the task path and exact Harbor command, request, Harness, Environment, run behavior, Verifier decision, and main limitation. Completion requires a real Harbor run, evidence that the Verifier measured the intended capability, and user approval. If continuing, reuse the available evidence and propose a distinct capability.

Invariants

  • One capability per Harbor task.
  • No production writes; reset mutable state between trials.
  • Keep hidden truth and simulator/judge credentials unavailable to the Harness.
  • Treat build, credential, reset, timeout, judge, and Verifier failures as infrastructure errors, not failed agent work.

langchain-ai의 다른 스킬

langgraph-docs
langchain-ai
LangGraph 문서에 접근하여 상태 기반 에이전트 및 멀티 에이전트 워크플로우를 구축합니다. 공식 LangGraph Python 문서를 가져오며, 상태 머신, 그래프 기반 에이전트 설계, 인간 개입 패턴을 다룹니다. 쿼리 유형에 따라 관련 문서를 우선시합니다: 방법 질문에는 구현 가이드, 이론에는 개념 페이지, 종단 간 예제에는 튜토리얼, 기술 세부 사항에는 API 참조를 제공합니다. 자동으로 가장 관련성 높은 2~4개의 문서 URL을 선택하고 해당 콘텐츠를 검색하여 답변합니다...
official
langgraph-human-in-the-loop
langchain-ai
그래프 실행을 일시 중지하여 사람의 검토, 승인 또는 검증을 받은 후, 입력을 받아 다시 실행합니다. 세 가지 구성 요소가 필요합니다: 체크포인터(InMemorySaver 또는 PostgresSaver), config의 스레드 ID, JSON 직렬화 가능한 인터럽트 페이로드. interrupt(value)는 데이터를 일시 중지하고 표시하며, Command(resume=value)는 다시 시작하여 일시 중지된 노드에 해당 값을 반환합니다. interrupt() 이전의 모든 코드는 다시 시작 시 재실행되므로, 부작용은 멱등성을 가져야 합니다(insert 대신 upsert 사용). 승인 워크플로우를 지원합니다,...
official
web-research
langchain-ai
웹 리서치와 관련된 요청에 이 스킬을 사용하세요. 포괄적인 웹 리서치를 수행하기 위한 체계적인 접근 방식을 제공합니다.
official
langchain-oss-primer
langchain-ai
LangChain, Deep Agents 또는 LangGraph 에이전트 구축 프로젝트를 시작할 때는 항상 여기서 시작하세요. 다른 스킬을 선택하거나 코드를 작성하기 전에 반드시 거쳐야 하는 시작점입니다.
official
skill-creator
langchain-ai
에이전트의 기능을 확장하기 위한 효과적인 스킬을 만드는 가이드로, 특화된 지식, 워크플로우 또는 도구 통합을 포함합니다. 사용자가...
official
social-media
langchain-ai
플랫폼별 소셜 미디어 게시물을 초안 작성하며, 연구 기반 콘텐츠와 함께 생성된 보조 이미지를 제공합니다. 링크드인 게시물(1,300자, 전문적인 어조)과 트위터/X 스레드(트윗당 280자, 1/🧵 형식)를 지원합니다. 작성 전에 하위 에이전트에 연구를 위임한 후, 결과를 읽어 정확성과 관련성을 확인해야 합니다. generate_social_image 도구를 사용하여 자동으로 눈에 띄는 소셜 이미지를 생성하며, 작은 화면에 최적화된 대담하고 대비가 높은 구성을 사용합니다.
official
deep-agents-memory
langchain-ai
Deep Agents를 위한 플러그형 메모리 및 파일 백엔드로, 임시, 영구 및 하이브리드 라우팅 옵션을 제공합니다. 네 가지 백엔드 유형: StateBackend(스레드 범위, 임시), StoreBackend(세션 간 영구), FilesystemBackend(로컬 개발을 위한 실제 디스크 액세스), CompositeBackend(다른 경로를 다른 백엔드로 라우팅). FilesystemMiddleware는 ls, read_file, write_file, edit_file, glob, grep의 여섯 가지 파일 작업 도구를 제공합니다. CompositeBackend는 최장 접두사 일치를 사용하여 라우팅합니다...
official
deep-agents-orchestration
langchain-ai
서브 에이전트를 조율하고, 다단계 작업을 계획하며, 민감한 작업에 대해 인간의 승인을 요구합니다. task 도구를 통해 전문화된 서브 에이전트에 작업을 위임합니다. 맞춤형 서브 에이전트는 격리된 도구 세트와 시스템 프롬프트를 지원하며, 기본 "범용" 서브 에이전트는 메인 에이전트 구성을 상속받습니다. write_todos를 사용하여 복잡한 워크플로우를 계획 및 추적하고, 보류 중, 진행 중, 완료 상태로 작업을 구성합니다. 호출 간 지속성을 위해 thread_id가 필요합니다. 구현...
official