langsmith-dataset

GỌI KỸ NĂNG NÀY khi tạo bộ dữ liệu đánh giá, tải bộ dữ liệu lên LangSmith hoặc quản lý các bộ dữ liệu hiện có. Bao gồm các loại bộ dữ liệu (final_response,…)

npx skills add https://github.com/langchain-ai/skills-benchmarks --skill langsmith-dataset
Create, manage, and upload evaluation datasets to LangSmith for testing and validation. Environment Variables
LANGSMITH_API_KEY=lsv2_pt_your_api_key_here          # REQUIRED
LANGSMITH_PROJECT=your-project-name                   # Check this to know which project has traces
LANGSMITH_WORKSPACE_ID=your-workspace-id              # Optional: for org-scoped keys

Authentication is REQUIRED: either set the LANGSMITH_API_KEY environment variable, or pass the --api-key flag to CLI commands (preferred):

langsmith dataset list --api-key $LANGSMITH_API_KEY

IMPORTANT: Always check the environment variables or .env file for LANGSMITH_PROJECT before querying or interacting with LangSmith. This tells you which project contains the relevant traces and data. If the LangSmith project is not available, use your best judgement to identify the right one.

Python Dependencies

pip install langsmith

JavaScript Dependencies

npm install langsmith

CLI Tool

curl -sSL https://raw.githubusercontent.com/langchain-ai/langsmith-cli/main/scripts/install.sh | sh
Use the `langsmith` CLI to manage datasets and examples.

Dataset Commands

  • langsmith dataset list - List datasets in LangSmith
  • langsmith dataset get <name-or-id> - View dataset details
  • langsmith dataset create --name <name> - Create a new empty dataset
  • langsmith dataset delete <name-or-id> - Delete a dataset
  • langsmith dataset export <name-or-id> <output-file> - Export dataset to local JSON file
  • langsmith dataset upload <file> --name <name> - Upload a local JSON file as a dataset

Example Commands

  • langsmith example list --dataset <name> - List examples in a dataset
  • langsmith example create --dataset <name> --inputs <json> - Add an example to a dataset
  • langsmith example delete <example-id> - Delete an example

Experiment Commands

  • langsmith experiment list --dataset <name> - List experiments for a dataset
  • langsmith experiment get <name> - View experiment results

Common Flags

  • --limit N - Limit number of results
  • --yes - Skip confirmation prompts (use with caution)

IMPORTANT - Safety Prompts:

  • The CLI prompts for confirmation before destructive operations (delete, overwrite)
  • If you are running with user input: ALWAYS wait for user input; NEVER use --yes unless the user explicitly requests it
  • If you are running non-interactively: Use --yes to skip confirmation prompts

<dataset_types_overview> Common evaluation dataset types:

  • final_response - Full conversation with expected output. Tests complete agent behavior.
  • single_step - Single node inputs/outputs. Tests specific node behavior (e.g., one LLM call or tool).
  • trajectory - Tool call sequence. Tests execution path (ordered list of tool names).
  • rag - Question/chunks/answer/citations. Tests retrieval quality. </dataset_types_overview>

<creating_datasets>

Creating Datasets

Datasets are JSON files with an array of examples. Each example has inputs and outputs.

From Exported Traces (Programmatic)

Export traces first, then process them into dataset format using code:

# 1. Export traces to JSONL files
langsmith trace export ./traces --project my-project --limit 20 --full --api-key $LANGSMITH_API_KEY
```python import json from pathlib import Path from langsmith import Client

client = Client()

2. Process traces into dataset examples

examples = [] for jsonl_file in Path("./traces").glob("*.jsonl"): runs = [json.loads(line) for line in jsonl_file.read_text().strip().split("\n")] root = next((r for r in runs if r.get("parent_run_id") is None), None) if root and root.get("inputs") and root.get("outputs"): examples.append({ "trace_id": root.get("trace_id"), "inputs": root["inputs"], "outputs": root["outputs"] })

3. Save locally

with open("/tmp/dataset.json", "w") as f: json.dump(examples, f, indent=2)

</python>

<typescript>
```typescript
import { Client } from "langsmith";
import { readFileSync, writeFileSync, readdirSync } from "fs";
import { join } from "path";

const client = new Client();

// 2. Process traces into dataset examples
const examples: Array<{trace_id?: string, inputs: Record<string, any>, outputs: Record<string, any>}> = [];
const files = readdirSync("./traces").filter(f => f.endsWith(".jsonl"));

for (const file of files) {
  const lines = readFileSync(join("./traces", file), "utf-8").trim().split("\n");
  const runs = lines.map(line => JSON.parse(line));
  const root = runs.find(r => r.parent_run_id == null);
  if (root?.inputs && root?.outputs) {
    examples.push({ trace_id: root.trace_id, inputs: root.inputs, outputs: root.outputs });
  }
}

// 3. Save locally
writeFileSync("/tmp/dataset.json", JSON.stringify(examples, null, 2));

Upload to LangSmith

# Upload local JSON file as a dataset
langsmith dataset upload /tmp/dataset.json --name "My Evaluation Dataset" --api-key $LANGSMITH_API_KEY

Using the SDK Directly

```python from langsmith import Client

client = Client()

Create dataset and add examples in one step

dataset = client.create_dataset("My Dataset", description="Evaluation dataset")

client.create_examples( inputs=[{"query": "What is AI?"}, {"query": "Explain RAG"}], outputs=[{"answer": "AI is..."}, {"answer": "RAG is..."}], dataset_name="My Dataset", )

</python>

<typescript>
```typescript
import { Client } from "langsmith";

const client = new Client();

// Create dataset and add examples
const dataset = await client.createDataset("My Dataset", {
  description: "Evaluation dataset",
});

await client.createExamples({
  inputs: [{ query: "What is AI?" }, { query: "Explain RAG" }],
  outputs: [{ answer: "AI is..." }, { answer: "RAG is..." }],
  datasetName: "My Dataset",
});

<dataset_structures>

Dataset Structures by Type

Final Response

{"trace_id": "...", "inputs": {"query": "What are the top genres?"}, "outputs": {"response": "The top genres are..."}}

Single Step

{"trace_id": "...", "inputs": {"messages": [...]}, "outputs": {"content": "..."}, "metadata": {"node_name": "model"}}

Trajectory

{"trace_id": "...", "inputs": {"query": "..."}, "outputs": {"expected_trajectory": ["tool_a", "tool_b", "tool_c"]}}

RAG

{"trace_id": "...", "inputs": {"question": "How do I..."}, "outputs": {"answer": "...", "retrieved_chunks": ["..."], "cited_chunks": ["..."]}}

</dataset_structures>

<script_usage>

CLI Usage

# List all datasets
langsmith dataset list --api-key $LANGSMITH_API_KEY

# Get dataset details
langsmith dataset get "My Dataset" --api-key $LANGSMITH_API_KEY

# Create an empty dataset
langsmith dataset create --name "New Dataset" --description "For evaluation" --api-key $LANGSMITH_API_KEY

# Upload a local JSON file
langsmith dataset upload /tmp/dataset.json --name "My Dataset" --api-key $LANGSMITH_API_KEY

# Export a dataset to local file
langsmith dataset export "My Dataset" /tmp/exported.json --limit 100 --api-key $LANGSMITH_API_KEY

# Delete a dataset
langsmith dataset delete "My Dataset" --api-key $LANGSMITH_API_KEY

# List examples in a dataset
langsmith example list --dataset "My Dataset" --limit 10 --api-key $LANGSMITH_API_KEY

# Add an example
langsmith example create --dataset "My Dataset" \
  --inputs '{"query": "test"}' \
  --outputs '{"answer": "result"}' --api-key $LANGSMITH_API_KEY

# List experiments
langsmith experiment list --dataset "My Dataset" --api-key $LANGSMITH_API_KEY
langsmith experiment get "eval-v1" --api-key $LANGSMITH_API_KEY

</script_usage>

<example_workflow> Complete workflow from traces to uploaded LangSmith dataset:

# 1. Export traces from LangSmith
langsmith trace export ./traces --project my-project --limit 20 --full --api-key $LANGSMITH_API_KEY

# 2. Process traces into dataset format (using Python/JS code)
# See "Creating Datasets" section above

# 3. Upload to LangSmith
langsmith dataset upload /tmp/final_response.json --name "Skills: Final Response" --api-key $LANGSMITH_API_KEY
langsmith dataset upload /tmp/trajectory.json --name "Skills: Trajectory" --api-key $LANGSMITH_API_KEY

# 4. Verify upload
langsmith dataset list --api-key $LANGSMITH_API_KEY
langsmith dataset get "Skills: Final Response" --api-key $LANGSMITH_API_KEY
langsmith example list --dataset "Skills: Final Response" --limit 3 --api-key $LANGSMITH_API_KEY

# 5. Run experiments
langsmith experiment list --dataset "Skills: Final Response" --api-key $LANGSMITH_API_KEY

</example_workflow>

**Dataset upload fails:** - Verify LANGSMITH_API_KEY is set - Check JSON file is valid: each element needs `inputs` (and optionally `outputs`) - Dataset name must be unique, or delete existing first with `langsmith dataset delete`

Empty dataset after upload:

  • Verify JSON file contains an array of objects with inputs key
  • Check file isn't empty: langsmith example list --dataset "Name"

Export has no data:

  • Ensure traces were exported with --full flag to include inputs/outputs
  • Verify traces have both inputs and outputs populated

Example count mismatch:

  • Use langsmith dataset get "Name" to check remote count
  • Compare with local file to verify upload completeness

Thêm skills từ langchain-ai

langgraph-docs
langchain-ai
Truy cập tài liệu LangGraph để xây dựng tác nhân có trạng thái và quy trình làm việc đa tác nhân. Lấy tài liệu Python chính thức của LangGraph bao gồm máy trạng thái, thiết kế tác nhân dựa trên đồ thị và các mẫu có sự can thiệp của con người. Ưu tiên tài liệu phù hợp theo loại truy vấn: hướng dẫn triển khai cho câu hỏi cách làm, trang khái niệm cho lý thuyết, hướng dẫn cho ví dụ từ đầu đến cuối và tham chiếu API cho chi tiết kỹ thuật. Tự động chọn 2–4 URL tài liệu phù hợp nhất và truy xuất nội dung của chúng để trả lời...
official
langgraph-human-in-the-loop
langchain-ai
Tạm dừng thực thi đồ thị để con người xem xét, phê duyệt hoặc xác thực, sau đó tiếp tục với đầu vào của họ. Yêu cầu ba thành phần: một bộ kiểm tra điểm dừng (InMemorySaver hoặc PostgresSaver), một ID luồng trong cấu hình và tải trọng ngắt có thể tuần tự hóa JSON. interrupt(value) tạm dừng và hiển thị dữ liệu; Command(resume=value) tiếp tục và trả về giá trị đó cho nút đã tạm dừng. Tất cả mã trước interrupt() sẽ thực thi lại khi tiếp tục, vì vậy các tác dụng phụ phải có tính chất đơn giản (sử dụng upsert, không phải insert). Hỗ trợ quy trình phê duyệt,...
official
web-research
langchain-ai
Sử dụng kỹ năng này cho các yêu cầu liên quan đến nghiên cứu web; nó cung cấp một cách tiếp cận có cấu trúc để thực hiện nghiên cứu web toàn diện
official
langchain-oss-primer
langchain-ai
LUÔN BẮT ĐẦU TỪ ĐÂY cho bất kỳ dự án xây dựng agent LangChain, Deep Agents hoặc LangGraph nào. Điểm khởi đầu bắt buộc trước khi chọn các kỹ năng khác hoặc viết bất kỳ…
official
skill-creator
langchain-ai
Hướng dẫn tạo kỹ năng hiệu quả để mở rộng khả năng của tác nhân với kiến thức chuyên môn, quy trình làm việc hoặc tích hợp công cụ. Sử dụng kỹ năng này khi người dùng…
official
social-media
langchain-ai
Soạn thảo bài đăng mạng xã hội theo từng nền tảng với nội dung dựa trên nghiên cứu và hình ảnh đồng hành được tạo tự động. Hỗ trợ bài đăng LinkedIn (1.300 ký tự với giọng văn chuyên nghiệp) và chuỗi Twitter/X (280 ký tự mỗi tweet theo định dạng 1/🧵). Yêu cầu ủy quyền nghiên cứu cho một trợ lý phụ trước khi viết, sau đó đọc lại kết quả để đảm bảo độ chính xác và phù hợp. Tự động tạo hình ảnh mạng xã hội bắt mắt bằng công cụ generate_social_image với bố cục đậm, tương phản cao, tối ưu cho kích thước nhỏ...
official
deep-agents-memory
langchain-ai
Các backend bộ nhớ và tệp có thể cắm cho Deep Agents với các tùy chọn định tuyến tạm thời, bền vững và kết hợp. Bốn loại backend: StateBackend (theo luồng, tạm thời), StoreBackend (bền vững xuyên phiên), FilesystemBackend (truy cập đĩa thực cho phát triển cục bộ) và CompositeBackend (định tuyến các đường dẫn khác nhau đến các backend khác nhau). FilesystemMiddleware cung cấp sáu công cụ thao tác tệp: ls, read_file, write_file, edit_file, glob, grep. CompositeBackend sử dụng so khớp tiền tố dài nhất để định tuyến...
official
deep-agents-orchestration
langchain-ai
Điều phối các tác nhân phụ, lập kế hoạch tác vụ đa bước và yêu cầu phê duyệt của con người cho các thao tác nhạy cảm. Ủy quyền công việc cho các tác nhân phụ chuyên biệt thông qua công cụ tác vụ; các tác nhân phụ tùy chỉnh hỗ trợ bộ công cụ và lời nhắc hệ thống riêng biệt, trong khi tác nhân phụ "đa năng" mặc định kế thừa cấu hình của tác nhân chính. Lập kế hoạch và theo dõi các quy trình phức tạp với write_todos, sắp xếp tác vụ qua các trạng thái đang chờ, đang tiến hành và đã hoàn thành; yêu c
official