langsmith

Trace, evaluate, and deploy AI agents and LLM applications with LangSmith. Use when adding observability, running evaluations, engineering prompts, or…

npx skills add https://github.com/langchain-ai/docs --skill langsmith

LangSmith

LangSmith is a framework-agnostic platform for building, debugging, and deploying AI agents and LLM applications. Trace requests, evaluate outputs, test prompts, and manage deployments all in one place at smith.langchain.com.

When to use

Use LangSmith when you need to:

  • Trace and debug LLM calls, agent steps, retrieval, and tool use
  • Evaluate LLM outputs with automated or human-in-the-loop scoring
  • Engineer prompts with a visual playground and version control
  • Deploy agents to production with the LangGraph-based agent server
  • Monitor production systems with dashboards, alerts, and cost tracking

When NOT to use

  • To build agent logic or LLM pipelines, use LangChain, LangGraph, or Deep Agents instead
  • LangSmith is the platform layer that complements these frameworks

Quick setup

Set two environment variables to enable tracing from any supported framework:

export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY="your-api-key"  # from smith.langchain.com/settings

Install the SDK

# Python
pip install langsmith

# JavaScript/TypeScript
npm install langsmith

Verify tracing

from langsmith import traceable

@traceable
def my_function(query: str) -> str:
    # Your LLM logic here—all calls inside are traced automatically
    return "result"

Core capabilities

CapabilityDescription
ObservabilityTrace every step of your LLM app with automatic or manual instrumentation
EvaluationRun evaluations with code, LLM-as-judge, or composite evaluators
Prompt engineeringCreate, version, and test prompts in a visual playground
Agent deploymentDeploy LangGraph agents with streaming, human-in-the-loop, and durable execution
MonitoringDashboards, alerts, and cost tracking for production workloads

Key documentation

API reference

For SDK class and method details, use the LangChain API Reference site:

  • Browse: https://reference.langchain.com/python/langsmith
  • MCP server: https://reference.langchain.com/mcp

Related skills

  • langchain—Build agents with prebuilt architecture and model integrations
  • langgraph—Orchestrate stateful, durable agent workflows
  • deep-agents—Batteries-included agent harness with planning and subagents

More skills from langchain-ai

deepagents-thread-inspector
langchain-ai
Inspect and explain conversations in the local Deep Agents Code SQLite session store. Use as a fallback when LangSmith trace tooling is unavailable, for…
deepagents-python-quickstart
langchain-ai
Scaffold a minimal local Deep Agent in Python by following the official quickstart, using provider-native web search instead of Tavily. Use when the user wants…
deepagents-typescript-quickstart
langchain-ai
Scaffold a minimal local Deep Agent in TypeScript by following the official quickstart, using provider-native web search instead of Tavily. Use when the user…
eval-engineering
langchain-ai
Iteratively inspect an agent repository and optional user-provided traces, interview the user, and create, run, and audit Harbor evals one at a time. Use for…
LangChain RAG Pipeline
langchain-ai
INVOKE THIS SKILL when building ANY retrieval-augmented generation (RAG) system. Covers document loaders, RecursiveCharacterTextSplitter, embeddings (OpenAI),…
LangChain Structured Output & HITL
langchain-ai
langchain-structured-output-&-hitl — an installable skill for AI agents, published by langchain-ai/langchain-skills.
LangSmith Datasets
langchain-ai
INVOKE THIS SKILL when creating evaluation datasets from trace OR uploading datasets to LangSmith OR querying datasets. Covers dataset types (final_response,…
langsmith-evaluator
langchain-ai
INVOKE THIS SKILL when building evaluation pipelines for LangSmith. Covers three core components: (1) Creating Evaluators - LLM-as-Judge, custom code; (2)…