langsmith-code-eval

作者: langchain-ai

為LangSmith追蹤的代理建立基於程式碼的評估器。適用於建構自訂評估邏輯、測試工具使用模式或評分代理輸出時使用…

npx skills add https://github.com/langchain-ai/lca-skills --skill langsmith-code-eval

LangSmith Code Evaluator Creation

Creates evaluators for LangSmith experiments through structured inspection and implementation.

Prerequisites

  • langsmith Python package installed
  • LANGSMITH_API_KEY environment variable set (check project's .env file)

Workflow

Copy this checklist and track progress:

Evaluator Creation Progress:
- [ ] Step 1: Gather info from user
- [ ] Step 2: Inspect trace and dataset structure
- [ ] Step 3: Read agent code
- [ ] Step 4: Write evaluator
- [ ] Step 5: Write experiment runner
- [ ] Step 6: Run and iterate

Step 1: Gather Info from User

IMPORTANT: Do NOT search or explore the codebase. Ask the user all of these questions upfront using AskUserQuestion before doing anything else.

Ask the user the following in a single AskUserQuestion call:

  1. Python command: How do you run Python in this project? (e.g., python, python3, uv run python, poetry run python)
  2. Agent file path: What is the path to your agent file?
  3. LangSmith project name: What is your LangSmith project name (where traces are logged)?
  4. LangSmith dataset name: What is the name of the dataset to evaluate against?
  5. Evaluation goal: What behavior should pass vs fail? Common types:
    • Tool usage: Did the agent call the correct tool?
    • Output correctness: Does output match expected format/content?
    • Policy compliance: Did it follow specific rules?
    • Classification: Did it categorize correctly?

Step 2: Inspect Trace and Dataset Structure

Using the info from Step 1, run the inspection scripts located in this skill's directory:

{python_cmd} {skill_dir}/scripts/inspect_trace.py PROJECT_NAME [RUN_ID]
{python_cmd} {skill_dir}/scripts/inspect_dataset.py DATASET_NAME

Replace {python_cmd} with the command from Step 1, and {skill_dir} with this skill's directory path.

Verify the trace matches the agent:

  • Does the trace type match? (e.g., OpenAI trace for OpenAI agent)
  • Does it contain the data needed for evaluation?
  • If mismatched, clarify before proceeding.

From the dataset inspection, note:

  • Input schema (what gets passed to the agent)
  • Output schema (reference/expected outputs)
  • Metadata fields (e.g., expected_tool, difficulty, labels)

The dataset metadata often contains ground truth for evaluation (e.g., which tool should be called, expected classification).

Step 3: Read Agent Code

Read the agent file provided in Step 1 to identify:

  • Entry point function (look for @traceable decorator)
  • Available tools
  • Output format (what the function returns)

Step 4: Write the Evaluator

Create evaluator functions based on trace and dataset structure. See EVALUATOR_REFERENCE.md for function signatures and return formats.

Step 5: Write Experiment Runner

Create a script that:

  1. Imports the agent's entry function
  2. Wraps it as a target function
  3. Runs evaluate() or aevaluate() against the dataset

See EVALUATOR_REFERENCE.md for evaluate() usage.

Step 6: Run and Iterate

Execute the experiment, review results in LangSmith, refine evaluators as needed.

來自 langchain-ai 的更多技能

langgraph-docs
langchain-ai
存取 LangGraph 文件以建構具狀態代理與多代理工作流程。擷取官方 LangGraph Python 文件,涵蓋狀態機、基於圖形的代理設計及人機協作模式。根據查詢類型優先提供相關文件:實作指南用於操作問題、概念頁面用於理論、教學用於端到端範例、API 參考用於技術細節。自動選取 2 至 4 個最相關的文件 URL 並擷取內容以回答...
official
langgraph-human-in-the-loop
langchain-ai
暫停圖形執行以進行人工審查、批准或驗證,然後根據其輸入繼續執行。需要三個組件:檢查點儲存器(InMemorySaver 或 PostgresSaver)、配置中的執行緒 ID,以及可序列化為 JSON 的中斷負載。interrupt(value) 會暫停並顯示資料;Command(resume=value) 會繼續執行,並將該值返回給暫停的節點。所有 interrupt() 之前的程式碼在恢復時會重新執行,因此副作用必須是冪等的(使用 upsert,而非 insert)。支援審批工作流程,...
official
web-research
langchain-ai
使用此技能處理與網路研究相關的請求;它提供了一種結構化方法來進行全面的網路研究
official
langchain-oss-primer
langchain-ai
務必從此處開始任何 LangChain、Deep Agents 或 Lang
official
skill-creator
langchain-ai
建立有效技能的指南,透過專業知識、工作流程或工具整合來擴展代理功能。當使用者…時,請使用此技能。
official
social-media
langchain-ai
根據研究內容撰寫特定平台的社群媒體貼文,並生成搭配圖片。支援LinkedIn貼文(1,300字元,專業語氣)與Twitter/X推文串(每則280字元,採用1/🧵格式)。寫作前需將研究任務委派給子代理,並閱讀其發現以確保準確性與相關性。使用generate_social_image工具自動生成吸睛的社群圖片,採用大膽高對比構圖,針對小螢幕進行優化。
official
deep-agents-memory
langchain-ai
為Deep Agents提供可插拔的記憶體與檔案後端,支援短暫、持久及混合路由選項。四種後端類型:StateBackend(執行緒範圍內短暫)、StoreBackend(跨工作階段持久)、FilesystemBackend(本地開發的真實磁碟存取)及CompositeBackend(將不同路徑路由至不同後端)。FilesystemMiddleware提供六種檔案操作工具:ls、read_file、write_file、edit_file、glob、grep。CompositeBackend使用最長前綴匹配進行路由...
official
deep-agents-orchestration
langchain-ai
協調子代理、規劃多步驟任務,並在敏感操作時要求人類批准。透過任務工具將工作委派給專業子代理;自訂子代理支援隔離的工具集與系統提示,而預設的「通用」子代理則繼承主代理配置。使用 write_todos 規劃與追蹤複雜工作流程,將任務組織為待處理、進行中與已完成狀態;需提供 thread_id 以在多次調用間保持持續性。實作...
official