langsmith-code-eval

Membuat evaluator berbasis kode untuk agen yang dilacak LangSmith. Gunakan saat membangun logika evaluasi kustom, menguji pola penggunaan alat, atau menilai keluaran agen…

npx skills add https://github.com/langchain-ai/lca-skills --skill langsmith-code-eval

LangSmith Code Evaluator Creation

Creates evaluators for LangSmith experiments through structured inspection and implementation.

Prerequisites

  • langsmith Python package installed
  • LANGSMITH_API_KEY environment variable set (check project's .env file)

Workflow

Copy this checklist and track progress:

Evaluator Creation Progress:
- [ ] Step 1: Gather info from user
- [ ] Step 2: Inspect trace and dataset structure
- [ ] Step 3: Read agent code
- [ ] Step 4: Write evaluator
- [ ] Step 5: Write experiment runner
- [ ] Step 6: Run and iterate

Step 1: Gather Info from User

IMPORTANT: Do NOT search or explore the codebase. Ask the user all of these questions upfront using AskUserQuestion before doing anything else.

Ask the user the following in a single AskUserQuestion call:

  1. Python command: How do you run Python in this project? (e.g., python, python3, uv run python, poetry run python)
  2. Agent file path: What is the path to your agent file?
  3. LangSmith project name: What is your LangSmith project name (where traces are logged)?
  4. LangSmith dataset name: What is the name of the dataset to evaluate against?
  5. Evaluation goal: What behavior should pass vs fail? Common types:
    • Tool usage: Did the agent call the correct tool?
    • Output correctness: Does output match expected format/content?
    • Policy compliance: Did it follow specific rules?
    • Classification: Did it categorize correctly?

Step 2: Inspect Trace and Dataset Structure

Using the info from Step 1, run the inspection scripts located in this skill's directory:

{python_cmd} {skill_dir}/scripts/inspect_trace.py PROJECT_NAME [RUN_ID]
{python_cmd} {skill_dir}/scripts/inspect_dataset.py DATASET_NAME

Replace {python_cmd} with the command from Step 1, and {skill_dir} with this skill's directory path.

Verify the trace matches the agent:

  • Does the trace type match? (e.g., OpenAI trace for OpenAI agent)
  • Does it contain the data needed for evaluation?
  • If mismatched, clarify before proceeding.

From the dataset inspection, note:

  • Input schema (what gets passed to the agent)
  • Output schema (reference/expected outputs)
  • Metadata fields (e.g., expected_tool, difficulty, labels)

The dataset metadata often contains ground truth for evaluation (e.g., which tool should be called, expected classification).

Step 3: Read Agent Code

Read the agent file provided in Step 1 to identify:

  • Entry point function (look for @traceable decorator)
  • Available tools
  • Output format (what the function returns)

Step 4: Write the Evaluator

Create evaluator functions based on trace and dataset structure. See EVALUATOR_REFERENCE.md for function signatures and return formats.

Step 5: Write Experiment Runner

Create a script that:

  1. Imports the agent's entry function
  2. Wraps it as a target function
  3. Runs evaluate() or aevaluate() against the dataset

See EVALUATOR_REFERENCE.md for evaluate() usage.

Step 6: Run and Iterate

Execute the experiment, review results in LangSmith, refine evaluators as needed.

Lebih banyak skill dari langchain-ai

langgraph-docs
langchain-ai
Mengakses dokumentasi LangGraph untuk membangun agen stateful dan alur kerja multi-agen. Mengambil dokumentasi resmi LangGraph Python yang mencakup mesin state, desain agen berbasis grafik, dan pola human-in-the-loop. Memprioritaskan dokumentasi yang relevan berdasarkan jenis kueri: panduan implementasi untuk pertanyaan cara, halaman konsep untuk teori, tutorial untuk contoh ujung ke ujung, dan referensi API untuk detail teknis. Secara otomatis memilih 2–4 URL dokumentasi yang paling relevan dan mengambil kontennya untuk menjawab...
official
langgraph-human-in-the-loop
langchain-ai
Jeda eksekusi graf untuk peninjauan, persetujuan, atau validasi manusia, lalu lanjutkan dengan masukan mereka. Membutuhkan tiga komponen: checkpointer (InMemorySaver atau PostgresSaver), ID thread dalam konfigurasi, dan payload interupsi yang dapat diserialisasi JSON. interrupt(value) menjeda dan menampilkan data; Command(resume=value) melanjutkan dan mengembalikan nilai tersebut ke node yang dijeda. Semua kode sebelum interrupt() akan dieksekusi ulang saat melanjutkan, sehingga efek samping harus idempoten (gunakan upsert, bukan insert). Mendukung alur kerja persetujuan,...
official
web-research
langchain-ai
Gunakan keterampilan ini untuk permintaan yang terkait dengan riset web; ini menyediakan pendekatan terstruktur untuk melakukan riset web yang komprehensif.
official
langchain-oss-primer
langchain-ai
SELALU MULAI DI SINI untuk proyek pembuatan agen LangChain, Deep Agents, atau LangGraph apa pun. Titik awal yang diperlukan sebelum memilih keterampilan lain atau menulis apa pun…
official
skill-creator
langchain-ai
Panduan untuk membuat skill yang efektif guna memperluas kemampuan agen dengan pengetahuan khusus, alur kerja, atau integrasi alat. Gunakan skill ini ketika pengguna…
official
social-media
langchain-ai
Menyusun draf posting media sosial khusus platform dengan konten berbasis riset dan gambar pendamping yang dihasilkan. Mendukung posting LinkedIn (1.300 karakter dengan nada profesional) dan utas Twitter/X (280 karakter per tweet dengan format 1/🧵). Memerlukan delegasi riset ke subagen sebelum menulis, kemudian membaca temuan untuk memastikan akurasi dan relevansi. Menghasilkan gambar sosial yang menarik secara otomatis menggunakan alat generate_social_image dengan komposisi tebal dan kontras tinggi yang dioptimalkan untuk ukuran kecil...
official
deep-agents-memory
langchain-ai
Backend memori dan file yang dapat dipasang untuk Deep Agents dengan opsi perutean sementara, persisten, dan hibrida. Empat jenis backend: StateBackend (berlaku dalam thread, sementara), StoreBackend (persisten lintas sesi), FilesystemBackend (akses disk nyata untuk pengembangan lokal), dan CompositeBackend (merutekan jalur berbeda ke backend berbeda). FilesystemMiddleware menyediakan enam alat operasi file: ls, read_file, write_file, edit_file, glob, grep. CompositeBackend menggunakan pencocokan prefiks terpanjang untuk merutekan...
official
deep-agents-orchestration
langchain-ai
We need to translate the given English text into Indonesian. The text describes an agent skill for orchestrating subagents, planning tasks, requiring human approval, delegating work, etc. We must preserve product names, protocol names, URLs, numbers, technical terms. The name "deep-agents-orchestration" is not in the text, so we don't include it. We translate only the text inside <text>. No extra commentary, labels, etc. Let's translate step by step: "Orchestrate subagents, plan multi-step tasks, and require human approval for sensitive operations." -> "Orkestrasi subagen, rencanakan tugas multi-langkah, dan minta persetujuan manusia untuk operasi sensitif." "Delegate work to specialized subagents via the task tool; custom subagents support isolated tool sets and system prompts, while the default "general-purpose" subagent inherits main agent configuration" -> "Delegasikan pekerjaan ke subagen khusus melalui alat tugas; subagen kustom mendukung set alat dan prompt sistem yang terisolasi, s
official