langsmith-online-eval-engineering

Iteratively inspect traces, interview the user, and create LangSmith online evaluators one at a time. Use specifically for creating online evaluators for use…

npx skills add https://github.com/langchain-ai/langchain-skills --skill langsmith-online-eval-engineering

Online Eval Engineering

Build online evaluators iteratively:

inspect traces and interview user -> propose directions -> user chooses
-> build evaluator -> test, attach, verify -> review and repeat

Read references/langsmith-api.md before creating or modifying evaluators.

1. Inspect traces

Ask the user for their LangSmith project name. Fetch recent root-level traces and print their structure. Read references/trace-inspection.md. Find:

  • run name and type;
  • available input and output field names;
  • the shape and content of the data (truncated samples);
  • which fields carry the data an evaluator would need.

Summarize the trace structure in the conversation:

Project: name
Run type: chain | llm | tool | ...
Input fields: field names and what they contain
Output fields: field names and what they contain
Sample: one representative input/output pair (truncated)

Keep the user involved: explain the trace structure and what it implies, then ask only for information the traces cannot establish. For example: "What does this application do?", "What quality concern matters most?", or "What failure should never happen?"

Ask whether the user wants a naming prefix for evaluators in this session (e.g., myapp-, v2-, dogfood-). If they provide one, apply it to all evaluator names, prompt hub handles, and run rule display names. If they decline, use plain descriptive names.

Do not propose evaluators until the trace structure is understood and the user has described their concerns.

2. Discuss and choose an eval direction

Read references/evaluator-design.md. Propose two or three evaluation criteria grounded in the trace data. Apply the naming prefix from step 1 if the user provided one. For each, give:

Name: descriptive evaluator name (with prefix if set)
Type: LLM-as-judge or code
Measures: what quality dimension this evaluates
Scoring: bool, float (0-1), or int; what pass/fail means
Fields needed: which trace fields are used and how
Rationale: why this type and approach

Example:

Name: response-relevance
Type: LLM-as-judge
Measures: whether the response addresses the user's question
Scoring: bool; True = relevant, False = off-topic or non-responsive
Fields needed: input (user question), output (assistant response)
Rationale: relevance is semantic and requires reading comprehension; not decidable by code

Recommend one and ask the user which to build. Do not implement until the user chooses.

3. Build one evaluator

Read references/langsmith-api.md. Build the selected evaluator. Show the full configuration to the user and get approval before executing any API calls.

LLM-as-judge path. Define a ResponseSchema with reasoning first, then the score field. Write prompt messages with a clear rubric that assesses the result, not whether it matches a reference answer. Set variable_mapping using field names discovered in step 1. Present the schema, prompt, variable mapping, and evaluator name for approval. On approval, push the prompt and create the evaluator. Report the evaluator ID.

Code evaluator path. Write a perform_eval(run, example=None) function. It must be self-contained (only builtins and standard library), access run as a dict (run.get("outputs")), and return {"key": ..., "score": ..., "comment": ...}. Present the function code and evaluator name for approval. On approval, create the evaluator. Report the evaluator ID.

4. Test, attach, and verify

Before attaching, ask the user what sampling rate they want (1.0 = every trace, 0.5 = half, 0.1 = 10%, or custom). Do not default silently. If the user is unsure, recommend 1.0 for initial testing.

Ask the user whether they want to test the evaluator against a few existing traces before attaching. Run rules only fire on new traces, so historical testing is the only way to verify before new traffic arrives.

For code evaluators, execute perform_eval directly against fetched root-level traces, passing a dict with inputs, outputs, and attachments keys. This catches runtime errors (wrong field names, dict-vs-object access, missing data) before production. For LLM evaluators, verify the configuration: confirm variable_mapping keys match prompt placeholders, confirm mapped trace fields exist, and check that the mapped data is meaningful.

If testing reveals errors, fix and recreate before attaching. If the user declines testing, proceed to attach.

Create a run rule to connect the evaluator to the tracing project. Apply the user's naming prefix to the display_name. Confirm the evaluator appears in the evaluator list with the correct project attachment. Inspect:

  • evaluator attachment and run rule status;
  • recent trace feedback and scores (from historical testing or new traces);
  • whether scores match expectations for the traces inspected;
  • edge case handling (empty output, errored runs, unexpected structure).

Fix and reattach when the evaluator crashes, scores incorrectly, or fails on edge cases. Before approval, confirm the evaluator scored the intended quality dimension, not an infrastructure or data-shape failure.

5. Review with the user

Explain the evaluator name and ID, quality dimension and scoring approach, trace fields used, sampling rate, and any limitation. Ask the user to approve, revise, drop, or choose the next direction. If continuing, reuse the trace findings, then propose a distinct quality dimension.

Invariants

  • One quality dimension per evaluator.
  • No guessing field names; always inspect traces before implementing.
  • Show configuration and get user approval before making API calls.
  • Code evaluators must be self-contained: only builtins and standard library.
  • Code evaluators receive run as a plain dict; use run.get("inputs") and run.get("outputs"), not attribute access. The example parameter must default to None.
  • Treat API failures, auth errors, and run rule failures as infrastructure errors, not evaluator bugs.

Thêm skills từ langchain-ai

langgraph-docs
langchain-ai
Truy cập tài liệu LangGraph để xây dựng tác nhân có trạng thái và quy trình làm việc đa tác nhân. Lấy tài liệu Python chính thức của LangGraph bao gồm máy trạng thái, thiết kế tác nhân dựa trên đồ thị và các mẫu có sự can thiệp của con người. Ưu tiên tài liệu phù hợp theo loại truy vấn: hướng dẫn triển khai cho câu hỏi cách làm, trang khái niệm cho lý thuyết, hướng dẫn cho ví dụ từ đầu đến cuối và tham chiếu API cho chi tiết kỹ thuật. Tự động chọn 2–4 URL tài liệu phù hợp nhất và truy xuất nội dung của chúng để trả lời...
official
langgraph-human-in-the-loop
langchain-ai
Tạm dừng thực thi đồ thị để con người xem xét, phê duyệt hoặc xác thực, sau đó tiếp tục với đầu vào của họ. Yêu cầu ba thành phần: một bộ kiểm tra điểm dừng (InMemorySaver hoặc PostgresSaver), một ID luồng trong cấu hình và tải trọng ngắt có thể tuần tự hóa JSON. interrupt(value) tạm dừng và hiển thị dữ liệu; Command(resume=value) tiếp tục và trả về giá trị đó cho nút đã tạm dừng. Tất cả mã trước interrupt() sẽ thực thi lại khi tiếp tục, vì vậy các tác dụng phụ phải có tính chất đơn giản (sử dụng upsert, không phải insert). Hỗ trợ quy trình phê duyệt,...
official
web-research
langchain-ai
Sử dụng kỹ năng này cho các yêu cầu liên quan đến nghiên cứu web; nó cung cấp một cách tiếp cận có cấu trúc để thực hiện nghiên cứu web toàn diện
official
langchain-oss-primer
langchain-ai
LUÔN BẮT ĐẦU TỪ ĐÂY cho bất kỳ dự án xây dựng agent LangChain, Deep Agents hoặc LangGraph nào. Điểm khởi đầu bắt buộc trước khi chọn các kỹ năng khác hoặc viết bất kỳ…
official
skill-creator
langchain-ai
Hướng dẫn tạo kỹ năng hiệu quả để mở rộng khả năng của tác nhân với kiến thức chuyên môn, quy trình làm việc hoặc tích hợp công cụ. Sử dụng kỹ năng này khi người dùng…
official
social-media
langchain-ai
Soạn thảo bài đăng mạng xã hội theo từng nền tảng với nội dung dựa trên nghiên cứu và hình ảnh đồng hành được tạo tự động. Hỗ trợ bài đăng LinkedIn (1.300 ký tự với giọng văn chuyên nghiệp) và chuỗi Twitter/X (280 ký tự mỗi tweet theo định dạng 1/🧵). Yêu cầu ủy quyền nghiên cứu cho một trợ lý phụ trước khi viết, sau đó đọc lại kết quả để đảm bảo độ chính xác và phù hợp. Tự động tạo hình ảnh mạng xã hội bắt mắt bằng công cụ generate_social_image với bố cục đậm, tương phản cao, tối ưu cho kích thước nhỏ...
official
deep-agents-memory
langchain-ai
Các backend bộ nhớ và tệp có thể cắm cho Deep Agents với các tùy chọn định tuyến tạm thời, bền vững và kết hợp. Bốn loại backend: StateBackend (theo luồng, tạm thời), StoreBackend (bền vững xuyên phiên), FilesystemBackend (truy cập đĩa thực cho phát triển cục bộ) và CompositeBackend (định tuyến các đường dẫn khác nhau đến các backend khác nhau). FilesystemMiddleware cung cấp sáu công cụ thao tác tệp: ls, read_file, write_file, edit_file, glob, grep. CompositeBackend sử dụng so khớp tiền tố dài nhất để định tuyến...
official
deep-agents-orchestration
langchain-ai
Điều phối các tác nhân phụ, lập kế hoạch tác vụ đa bước và yêu cầu phê duyệt của con người cho các thao tác nhạy cảm. Ủy quyền công việc cho các tác nhân phụ chuyên biệt thông qua công cụ tác vụ; các tác nhân phụ tùy chỉnh hỗ trợ bộ công cụ và lời nhắc hệ thống riêng biệt, trong khi tác nhân phụ "đa năng" mặc định kế thừa cấu hình của tác nhân chính. Lập kế hoạch và theo dõi các quy trình phức tạp với write_todos, sắp xếp tác vụ qua các trạng thái đang chờ, đang tiến hành và đã hoàn thành; yêu c
official