langsmith-code-eval

โดย langchain-ai

สร้างตัวประเมินที่ใช้โค้ดสำหรับเอเจนต์ที่ถูกติดตามด้วย LangSmith ใช้เมื่อสร้างตรรกะการประเมินแบบกำหนดเอง ทดสอบรูปแบบการใช้งานเครื่องมือ หรือให้คะแนนผลลัพธ์ของเอเจนต์…

npx skills add https://github.com/langchain-ai/lca-skills --skill langsmith-code-eval

LangSmith Code Evaluator Creation

Creates evaluators for LangSmith experiments through structured inspection and implementation.

Prerequisites

  • langsmith Python package installed
  • LANGSMITH_API_KEY environment variable set (check project's .env file)

Workflow

Copy this checklist and track progress:

Evaluator Creation Progress:
- [ ] Step 1: Gather info from user
- [ ] Step 2: Inspect trace and dataset structure
- [ ] Step 3: Read agent code
- [ ] Step 4: Write evaluator
- [ ] Step 5: Write experiment runner
- [ ] Step 6: Run and iterate

Step 1: Gather Info from User

IMPORTANT: Do NOT search or explore the codebase. Ask the user all of these questions upfront using AskUserQuestion before doing anything else.

Ask the user the following in a single AskUserQuestion call:

  1. Python command: How do you run Python in this project? (e.g., python, python3, uv run python, poetry run python)
  2. Agent file path: What is the path to your agent file?
  3. LangSmith project name: What is your LangSmith project name (where traces are logged)?
  4. LangSmith dataset name: What is the name of the dataset to evaluate against?
  5. Evaluation goal: What behavior should pass vs fail? Common types:
    • Tool usage: Did the agent call the correct tool?
    • Output correctness: Does output match expected format/content?
    • Policy compliance: Did it follow specific rules?
    • Classification: Did it categorize correctly?

Step 2: Inspect Trace and Dataset Structure

Using the info from Step 1, run the inspection scripts located in this skill's directory:

{python_cmd} {skill_dir}/scripts/inspect_trace.py PROJECT_NAME [RUN_ID]
{python_cmd} {skill_dir}/scripts/inspect_dataset.py DATASET_NAME

Replace {python_cmd} with the command from Step 1, and {skill_dir} with this skill's directory path.

Verify the trace matches the agent:

  • Does the trace type match? (e.g., OpenAI trace for OpenAI agent)
  • Does it contain the data needed for evaluation?
  • If mismatched, clarify before proceeding.

From the dataset inspection, note:

  • Input schema (what gets passed to the agent)
  • Output schema (reference/expected outputs)
  • Metadata fields (e.g., expected_tool, difficulty, labels)

The dataset metadata often contains ground truth for evaluation (e.g., which tool should be called, expected classification).

Step 3: Read Agent Code

Read the agent file provided in Step 1 to identify:

  • Entry point function (look for @traceable decorator)
  • Available tools
  • Output format (what the function returns)

Step 4: Write the Evaluator

Create evaluator functions based on trace and dataset structure. See EVALUATOR_REFERENCE.md for function signatures and return formats.

Step 5: Write Experiment Runner

Create a script that:

  1. Imports the agent's entry function
  2. Wraps it as a target function
  3. Runs evaluate() or aevaluate() against the dataset

See EVALUATOR_REFERENCE.md for evaluate() usage.

Step 6: Run and Iterate

Execute the experiment, review results in LangSmith, refine evaluators as needed.

Skills เพิ่มเติมจาก langchain-ai

langgraph-docs
langchain-ai
เข้าถึงเอกสาร LangGraph เพื่อสร้างเอเจนต์ที่มีสถานะและเวิร์กโฟลว์แบบหลายเอเจนต์ ดึงข้อมูลเอกสาร Python อย่างเป็นทางการของ LangGraph ซึ่งครอบคลุมเครื่องจักรสถานะ การออกแบบเอเจนต์แบบกราฟ และรูปแบบมนุษย์ในวงจร จัดลำดับความสำคัญของเอกสารที่เกี่ยวข้องตามประเภทคำถาม: คู่มือการใช้งานสำหรับคำถามวิธีทำ หน้าคอนเซปต์สำหรับทฤษฎี บทช่วยสอนสำหรับตัวอย่างแบบครบวงจร และเอกสารอ้างอิง API สำหรับรายละเอียดทางเทคนิค เลือก URL เอกสารที่เกี่ยวข้องมากที่สุด 2–4 รายการโดยอัตโนมัติและดึงเนื้อหามาเพื่อตอบ...
official
langgraph-human-in-the-loop
langchain-ai
หยุดการทำงานของกราฟเพื่อให้มนุษย์ตรวจสอบ อนุมัติ หรือตรวจสอบความถูกต้อง จากนั้นดำเนินการต่อด้วยข้อมูลที่มนุษย์ป้อนเข้าไป ต้องมีสามองค์ประกอบ: ตัวตรวจสอบจุด (InMemorySaver หรือ PostgresSaver), ID เธรดใน config, และเพย์โหลดขัดจังหวะที่แปลงเป็น JSON ได้ interrupt(value) จะหยุดและแสดงข้อมูล; Command(resume=value) จะดำเนินการต่อและส่งคืนค่านั้นไปยังโหนดที่ถูกหยุด โค้ดทั้งหมดก่อน interrupt() จะถูกเรียกใช้ใหม่เมื่อดำเนินการต่อ ดังนั้นผลข้างเคียงต้องเป็น idempotent (ใช้ upsert ไม่ใช่ insert) รองรับเวิร์กโฟลว์การอนุมัติ...
official
web-research
langchain-ai
ใช้ทักษะนี้สำหรับคำขอที่เกี่ยวข้องกับการค้นคว้าทางเว็บ โดยมีแนวทางที่มีโครงสร้างเพื่อดำเนินการค้นคว้าทางเว็บอย่างครอบคลุม
official
langchain-oss-primer
langchain-ai
เริ่มต้นที่นี่เสมอสำหรับโปรเจกต์สร้างเอเจนต์ LangChain, Deep Agents หรือ LangGraph ใดๆ จุดเริ่มต้นที่จำเป็นก่อนเลือกสกิลอื่นหรือเขียนอะไรก็ตาม…
official
skill-creator
langchain-ai
คู่มือสำหรับสร้างสกิลที่มีประสิทธิภาพเพื่อขยายความสามารถของเอเจนต์ด้วยความรู้เฉพาะทาง เวิร์กโฟลว์ หรือการรวมเครื่องมือ ใช้สกิลนี้เมื่อผู้ใช้…
official
social-media
langchain-ai
ร่างโพสต์โซเชียลมีเดียเฉพาะแพลตฟอร์มพร้อมเนื้อหาที่มีงานวิจัยรองรับและภาพประกอบที่สร้างขึ้น รองรับโพสต์ LinkedIn (1,300 ตัวอักษรด้วยน้ำเสียงมืออาชีพ) และเธรด Twitter/X (280 ตัวอักษรต่อทวีตในรูปแบบ 1/🧵) ต้องมอบหมายการวิจัยให้กับซับเอเจนต์ก่อนเขียน จากนั้นอ่านผลลัพธ์เพื่อให้แน่ใจว่าถูกต้องและเกี่ยวข้อง สร้างภาพโซเชียลที่สะดุดตาโดยอัตโนมัติโดยใช้เครื่องมือ generate_social_image ด้วยองค์ประกอบที่โดดเด่นและคอนทราสต์สูงซึ่งปรับให้เหมาะสมสำหรับขนาดเล็ก...
official
deep-agents-memory
langchain-ai
ปลั๊กอินหน่วยความจำและแบ็กเอนด์ไฟล์สำหรับ Deep Agents พร้อมตัวเลือกการกำหนดเส้นทางแบบชั่วคราว ถาวร และแบบผสม แบ็กเอนด์สี่ประเภท: StateBackend (ขอบเขตเธรด, ชั่วคราว), StoreBackend (คงอยู่ข้ามเซสชัน), FilesystemBackend (เข้าถึงดิสก์จริงสำหรับการพัฒนาในเครื่อง) และ CompositeBackend (กำหนดเส้นทางพาธที่แตกต่างไปยังแบ็กเอนด์ที่แตกต่างกัน) FilesystemMiddleware มีเครื่องมือปฏิบัติการไฟล์หกอย่าง: ls, read_file, write_file, edit_file, glob, grep CompositeBackend ใช้การจับคู่คำนำหน้าที่ยาวที่สุดเพื่อกำหนดเส้นทาง...
official
deep-agents-orchestration
langchain-ai
จัดระเบียบเอเจนต์ย่อย วางแผนงานหลายขั้นตอน และต้องได้รับการอนุมัติจากมนุษย์สำหรับการดำเนินการที่ละเอียดอ่อน มอบหมายงานให้กับเอเจนต์ย่อยเฉพาะทางผ่านเครื่องมืองาน เอเจนต์ย่อยแบบกำหนดเองรองรับชุดเครื่องมือและพรอมต์ระบบที่แยกออกจากกัน ในขณะที่เอเจนต์ย่อย "วัตถุประสงค์ทั่วไป" เริ่มต้นจะสืบทอดการกำหนดค่าเอเจนต์หลัก วางแผนและติดตามเวิร์กโฟลว์ที่ซับซ้อนด้วย write_todos โดยจัดระเบียบงานในสถานะรอดำเนินการ กำลังดำเนินการ และเสร็จสมบูรณ์ ต้องใช้ thread_id เพื่อความต่อเนื่องในการเรียกใช้หลายครั้ง ดำเนินการ...
official