langsmith-code-eval

Crea evaluadores basados en código para agentes rastreados con LangSmith. Úsalo al construir lógica de evaluación personalizada, probar patrones de uso de herramientas o puntuar salidas de agentes…

npx skills add https://github.com/langchain-ai/lca-skills --skill langsmith-code-eval

LangSmith Code Evaluator Creation

Creates evaluators for LangSmith experiments through structured inspection and implementation.

Prerequisites

  • langsmith Python package installed
  • LANGSMITH_API_KEY environment variable set (check project's .env file)

Workflow

Copy this checklist and track progress:

Evaluator Creation Progress:
- [ ] Step 1: Gather info from user
- [ ] Step 2: Inspect trace and dataset structure
- [ ] Step 3: Read agent code
- [ ] Step 4: Write evaluator
- [ ] Step 5: Write experiment runner
- [ ] Step 6: Run and iterate

Step 1: Gather Info from User

IMPORTANT: Do NOT search or explore the codebase. Ask the user all of these questions upfront using AskUserQuestion before doing anything else.

Ask the user the following in a single AskUserQuestion call:

  1. Python command: How do you run Python in this project? (e.g., python, python3, uv run python, poetry run python)
  2. Agent file path: What is the path to your agent file?
  3. LangSmith project name: What is your LangSmith project name (where traces are logged)?
  4. LangSmith dataset name: What is the name of the dataset to evaluate against?
  5. Evaluation goal: What behavior should pass vs fail? Common types:
    • Tool usage: Did the agent call the correct tool?
    • Output correctness: Does output match expected format/content?
    • Policy compliance: Did it follow specific rules?
    • Classification: Did it categorize correctly?

Step 2: Inspect Trace and Dataset Structure

Using the info from Step 1, run the inspection scripts located in this skill's directory:

{python_cmd} {skill_dir}/scripts/inspect_trace.py PROJECT_NAME [RUN_ID]
{python_cmd} {skill_dir}/scripts/inspect_dataset.py DATASET_NAME

Replace {python_cmd} with the command from Step 1, and {skill_dir} with this skill's directory path.

Verify the trace matches the agent:

  • Does the trace type match? (e.g., OpenAI trace for OpenAI agent)
  • Does it contain the data needed for evaluation?
  • If mismatched, clarify before proceeding.

From the dataset inspection, note:

  • Input schema (what gets passed to the agent)
  • Output schema (reference/expected outputs)
  • Metadata fields (e.g., expected_tool, difficulty, labels)

The dataset metadata often contains ground truth for evaluation (e.g., which tool should be called, expected classification).

Step 3: Read Agent Code

Read the agent file provided in Step 1 to identify:

  • Entry point function (look for @traceable decorator)
  • Available tools
  • Output format (what the function returns)

Step 4: Write the Evaluator

Create evaluator functions based on trace and dataset structure. See EVALUATOR_REFERENCE.md for function signatures and return formats.

Step 5: Write Experiment Runner

Create a script that:

  1. Imports the agent's entry function
  2. Wraps it as a target function
  3. Runs evaluate() or aevaluate() against the dataset

See EVALUATOR_REFERENCE.md for evaluate() usage.

Step 6: Run and Iterate

Execute the experiment, review results in LangSmith, refine evaluators as needed.

Más skills de langchain-ai

langgraph-docs
langchain-ai
Accede a la documentación de LangGraph para construir agentes con estado y flujos de trabajo multiagente. Obtiene la documentación oficial de LangGraph en Python que cubre máquinas de estado, diseño de agentes basado en grafos y patrones de intervención humana. Prioriza la documentación relevante según el tipo de consulta: guías de implementación para preguntas de procedimiento, páginas conceptuales para teoría, tutoriales para ejemplos completos y referencias de API para detalles técnicos. Selecciona automáticamente de 2 a 4 URL de documentación más relevantes y recupera su contenido para responder...
official
langgraph-human-in-the-loop
langchain-ai
Pausar la ejecución del grafo para revisión, aprobación o validación humana, luego reanudar con su entrada. Requiere tres componentes: un checkpointer (InMemorySaver o PostgresSaver), un ID de hilo en la configuración y cargas de interrupción serializables en JSON. interrupt(value) pausa y expone datos; Command(resume=value) reanuda y devuelve ese valor al nodo pausado. Todo el código antes de interrupt() se re-ejecuta al reanudar, por lo que los efectos secundarios deben ser idempotentes (usar upsert, no insert). Soporta flujos de aprobación,...
official
web-research
langchain-ai
Usa esta habilidad para solicitudes relacionadas con investigación web; proporciona un enfoque estructurado para realizar investigaciones web exhaustivas.
official
langchain-oss-primer
langchain-ai
SIEMPRE EMPIEZA AQUÍ para cualquier proyecto de construcción de agentes LangChain, Deep Agents o LangGraph. Punto de partida requerido antes de elegir otras habilidades o escribir cualquier…
official
skill-creator
langchain-ai
Guía para crear habilidades efectivas que extiendan las capacidades del agente con conocimientos especializados, flujos de trabajo o integraciones de herramientas. Usa esta habilidad cuando el usuario…
official
social-media
langchain-ai
Redacta publicaciones para redes sociales específicas de cada plataforma, con contenido respaldado por investigación e imágenes complementarias generadas. Admite publicaciones de LinkedIn (1.300 caracteres con tono profesional) e hilos de Twitter/X (280 caracteres por tuit con formato 1/🧵). Requiere delegar la investigación a un subagente antes de redactar, y luego leer los hallazgos para garantizar precisión y relevancia. Genera automáticamente imágenes llamativas para redes sociales usando la herramienta generate_social_image con composiciones audaces y de alto contraste optimizadas para tamaños pequeños...
official
deep-agents-memory
langchain-ai
Backends de memoria y archivos conectables para Deep Agents con opciones de enrutamiento efímero, persistente e híbrido. Cuatro tipos de backend: StateBackend (efímero, con ámbito de hilo), StoreBackend (persistente entre sesiones), FilesystemBackend (acceso real a disco para desarrollo local) y CompositeBackend (enruta diferentes rutas a diferentes backends). FilesystemMiddleware proporciona seis herramientas de operación de archivos: ls, read_file, write_file, edit_file, glob, grep. CompositeBackend utiliza coincidencia de prefijo más largo para enrutar...
official
deep-agents-orchestration
langchain-ai
Orquestrar subagentes, planificar tareas de múltiples pasos y requerir aprobación humana para operaciones sensibles. Delegar trabajo a subagentes especializados mediante la herramienta de tareas; los subagentes personalizados admiten conjuntos de herramientas aislados y mensajes del sistema, mientras que el subagente "de propósito general" predeterminado hereda la configuración del agente principal. Planificar y rastrear flujos de trabajo complejos con write_todos, organizando tareas en estados pendientes, en curso y completados; requiere un thread_id para la persistencia entre invocaciones. Implementar...
official