langsmith-code-eval

Tạo các bộ đánh giá dựa trên mã cho các tác nhân được theo dõi bởi LangSmith. Sử dụng khi xây dựng logic đánh giá tùy chỉnh, kiểm tra các mẫu sử dụng công cụ hoặc chấm điểm đầu ra của tác nhân…

npx skills add https://github.com/langchain-ai/lca-skills --skill langsmith-code-eval

LangSmith Code Evaluator Creation

Creates evaluators for LangSmith experiments through structured inspection and implementation.

Prerequisites

  • langsmith Python package installed
  • LANGSMITH_API_KEY environment variable set (check project's .env file)

Workflow

Copy this checklist and track progress:

Evaluator Creation Progress:
- [ ] Step 1: Gather info from user
- [ ] Step 2: Inspect trace and dataset structure
- [ ] Step 3: Read agent code
- [ ] Step 4: Write evaluator
- [ ] Step 5: Write experiment runner
- [ ] Step 6: Run and iterate

Step 1: Gather Info from User

IMPORTANT: Do NOT search or explore the codebase. Ask the user all of these questions upfront using AskUserQuestion before doing anything else.

Ask the user the following in a single AskUserQuestion call:

  1. Python command: How do you run Python in this project? (e.g., python, python3, uv run python, poetry run python)
  2. Agent file path: What is the path to your agent file?
  3. LangSmith project name: What is your LangSmith project name (where traces are logged)?
  4. LangSmith dataset name: What is the name of the dataset to evaluate against?
  5. Evaluation goal: What behavior should pass vs fail? Common types:
    • Tool usage: Did the agent call the correct tool?
    • Output correctness: Does output match expected format/content?
    • Policy compliance: Did it follow specific rules?
    • Classification: Did it categorize correctly?

Step 2: Inspect Trace and Dataset Structure

Using the info from Step 1, run the inspection scripts located in this skill's directory:

{python_cmd} {skill_dir}/scripts/inspect_trace.py PROJECT_NAME [RUN_ID]
{python_cmd} {skill_dir}/scripts/inspect_dataset.py DATASET_NAME

Replace {python_cmd} with the command from Step 1, and {skill_dir} with this skill's directory path.

Verify the trace matches the agent:

  • Does the trace type match? (e.g., OpenAI trace for OpenAI agent)
  • Does it contain the data needed for evaluation?
  • If mismatched, clarify before proceeding.

From the dataset inspection, note:

  • Input schema (what gets passed to the agent)
  • Output schema (reference/expected outputs)
  • Metadata fields (e.g., expected_tool, difficulty, labels)

The dataset metadata often contains ground truth for evaluation (e.g., which tool should be called, expected classification).

Step 3: Read Agent Code

Read the agent file provided in Step 1 to identify:

  • Entry point function (look for @traceable decorator)
  • Available tools
  • Output format (what the function returns)

Step 4: Write the Evaluator

Create evaluator functions based on trace and dataset structure. See EVALUATOR_REFERENCE.md for function signatures and return formats.

Step 5: Write Experiment Runner

Create a script that:

  1. Imports the agent's entry function
  2. Wraps it as a target function
  3. Runs evaluate() or aevaluate() against the dataset

See EVALUATOR_REFERENCE.md for evaluate() usage.

Step 6: Run and Iterate

Execute the experiment, review results in LangSmith, refine evaluators as needed.

Thêm skills từ langchain-ai

deepagents-thread-inspector
langchain-ai
Kiểm tra và giải thích các cuộc hội thoại trong kho lưu trữ phiên SQLite cục bộ của Deep Agents Code. Sử dụng như phương án dự phòng khi công cụ theo dõi LangSmith không khả dụng, cho…
deepagents-python-quickstart
langchain-ai
Tạo khung một Deep Agent cục bộ tối thiểu bằng Python bằng cách làm theo hướng dẫn khởi động nhanh chính thức, sử dụng tìm kiếm web gốc của nhà cung cấp thay vì Tavily. Sử dụng khi người dùng muốn…
deepagents-typescript-quickstart
langchain-ai
Tạo khung một Deep Agent tối thiểu cục bộ bằng TypeScript theo hướng dẫn khởi động nhanh chính thức, sử dụng tìm kiếm web gốc của nhà cung cấp thay vì Tavily. Sử dụng khi người dùng…
eval-engineering
langchain-ai
Lặp lại kiểm tra kho lưu trữ agent và các dấu vết do người dùng cung cấp tùy chọn, phỏng vấn người dùng, và tạo, chạy, cũng như kiểm toán các đánh giá Harbor từng cái một. Sử dụng cho…
LangChain RAG Pipeline
langchain-ai
GỌI KỸ NĂNG NÀY khi xây dựng BẤT KỲ hệ thống tạo sinh tăng cường truy xuất (RAG) nào. Bao gồm bộ tải tài liệu, RecursiveCharacterTextSplitter, embeddings (OpenAI),…
LangChain Structured Output & HITL
langchain-ai
langchain-structured-output-&-hitl — một kỹ năng có thể cài đặt cho các tác nhân AI, được xuất bản bởi langchain-ai/langchain-skills.
LangSmith Datasets
langchain-ai
GỌI KỸ NĂNG NÀY khi tạo bộ dữ liệu đánh giá từ trace HOẶC tải bộ dữ liệu lên LangSmith HOẶC truy vấn bộ dữ liệu. Bao gồm các loại bộ dữ liệu (final_response,…
langsmith-evaluator
langchain-ai
GỌI KỸ NĂNG NÀY khi xây dựng pipeline đánh giá cho LangSmith. Bao gồm ba thành phần cốt lõi: (1) Tạo Evaluator - LLM-as-Judge, mã tùy chỉnh; (2)…