SkillTotal
Security scanner for MCP servers, agent skills and npm/PyPI packages that reports tool poisoning and exfiltration paths with file:line evidence.
Documentation
SkillTotal
Scan packages before Claude Code installs them. SkillTotal is a static security scanner for what coding agents install: npm and PyPI packages, MCP servers, agent skills and git repositories. It runs on your machine, never executes the code it reads and never calls an LLM. Every finding cites the file and line it came from.
Quick start
pipx install skilltotal
Then, inside Claude Code:
/plugin marketplace add pezhik/skilltotal
/plugin install skilltotal@skilltotal
/reload-plugins
From then on, when the agent runs npx, npm install, pip install, uvx or claude mcp add,
the plugin first scans the package the command names. If the package has malicious indicators, the
command is denied and the agent is told what was found and where:
SkillTotal blocked this install: it found malicious indicators. npm:some-pkg@1.0.3:
Decode-and-execute (obfuscated execution) at `scripts/setup.js:4`. Run
`skilltotal scan npm:some-pkg@1.0.3` to see every finding with its file:line evidence.
Without the plugin, scan anything from a terminal:
skilltotal scan npm:some-package # or pypi:name, a git URL, a local folder or an archive
skilltotal guard npm:some-mcp-server # exit code 2 if it should not be installed
Try it online (no install, no account): www.skilltotal.ai runs the same engine.
SkillTotal analyzes only the component itself, not the user, company, environment, deployment or runtime context around it. Every score and finding comes from the files inside the component.
Core principle: every confirmed finding carries evidence (file, line range, code snippet). Anything that cannot be tied to evidence goes to
needs_reviewinstead offindingsand never affects the score.
Why SkillTotal
- Checks packages before your agent installs them. The Claude Code plugin reads each
npx,npm install,pip installorclaude mcp addcommand the agent is about to run. It blocks packages with malicious indicators, and asks you first about high- or critical-risk packages and anything it could not check. Set it up. - Runs on your machine. It needs no account or API token and uploads nothing. It uses the network only to fetch the package or repository you asked it to scan.
- Safe to point at untrusted components. The engine reads them and never runs them. (Dynamic analysis in an isolated sandbox is planned as a separate paid service.)
- Zero runtime dependencies. It uses only the Python standard library, so it is easy to audit, vendor or run air-gapped.
- Deterministic. Detection is regex and AST matching with no LLM, so the same input always yields the same findings and score.
- Evidence-anchored, with few false positives. Every finding points at an exact file:line.
- Standards-aligned. Every component gets a behavioral trait fingerprint mapped to the Cloud Security Alliance (CSA) agentic threat model, MAESTRO threat-model layers and MITRE ATLAS tactics. Three of the traits record how the component authenticates its tool calls (an embedded static credential, delegated OAuth/OIDC or a least-privilege scoped identity), which shows the blast radius of a compromise, not only whether a secret exists.
- Free and open source (Apache-2.0). The full static report is free, forever.
Measured, not asserted
Detection claims are cheap, so the numbers behind them are published with the data and the code that produced them.
- The MCP registry, scanned: 15,341 of the 17,535 distinct components in the official registry (87.5%), scanned with rulesets 56 to 59 and finished with engine 0.49.0 on 2026-09-19. Most of the rest were no longer reachable or exceeded a time or size bound, and every exclusion is counted. Of the components scanned, 88.2% expose tools to an agent, 64.8% can reach the network and 21.1% can execute shell commands, yet 99.3% of them score low risk, because a capability on its own scores zero. Raw JSON · the harness.
- Detection efficacy: recall and precision on a labeled corpus we built and publish (64 malicious and 46 benign samples), regenerated every release and enforced by CI as a floor. A perfect score on our own corpus guards against regressions; it is not a claim about every attack in the wild.
Both reproduce: same input, same engine, same output. Nothing is executed and no LLM is involved.
Install
Requires Python 3.10+. Zero runtime dependencies. git is required only for scanning
remote URLs.
For the CLI, pipx is recommended. It installs into an isolated environment
and also works on Debian/Ubuntu and Homebrew Python, where PEP 668 blocks a bare pip install:
pipx install skilltotal
To install into a virtual environment, or to use it as a library:
pip install skilltotal
You can also run it with npx. The npm package only starts this engine, so it needs uv, pipx, or
Python 3.11+ with skilltotal installed:
npx -y skilltotal scan https://github.com/owner/repo
From source (development):
pip install -e ".[dev]"
Claude Code plugin
Coding agents install packages in the middle of a task, and the permission prompt, if there is
one, shows only the command. The plugin checks each package on your machine, with the same engine
as the CLI, before the install command runs. Its hook recognizes npx, bunx,
pnpm dlx, npm/pnpm/yarn/bun add or install, pip, uv, uvx, pipx,
claude mcp add and claude mcp add-json. It reads a command the way the shell would, so it also finds
installs behind sudo, env, bash -c '...', cmd /c and $(...), in groups like (npm i y)
and in chains like cd x && npm i y. On Windows, Claude Code runs most commands through its
PowerShell tool, and the hook also follows & { ... }, iex '...',
Start-Process npm -ArgumentList ... and powershell -EncodedCommand.
What happens to the install:
- Malicious indicators: the command is denied, and the agent is told the finding and its file:line.
- High or critical risk: Claude Code asks you to approve the install and shows the top finding.
- Anything SkillTotal could not check: Claude Code asks you. This covers a package from a
custom registry or index (
--registry,--index-url,--extra-index-url), a direct archive URL, a git ref the scan cannot reproduce exactly, a check that failed or did not finish within the 20 seconds all checks of one command share, the sixth and later packages in one command, a name a letter or two from a popular package, and an install script SkillTotal could not read. SkillTotal never denies an install because of its own failure. - Clean: the package installs as usual, and the agent gets a one-line note with the version that was checked and its score.
Packages from GitHub (github:owner/repo#v1.2.0, git+https://github.com/...@ref) are scanned
at the ref that will be installed. A verdict is reused for the same artifact (an exact npm release
or PyPI file for a week, a git repository for an hour), so repeated npx tsc or npx prettier
calls don't trigger a rescan. When the engine version changes, everything is scanned again. For
large packages, set SKILLTOTAL_HOOK_BUDGET (in seconds, up to 100) to wait longer than 20
seconds. In an unattended run (claude -p), Claude Code turns a question nobody can answer into a
denial, so a package the hook could not check is not installed there; give large packages more
time with the same variable. The plugin also adds a /skilltotal:scan <target> command and
registers the MCP server described below.
The plugin calls the CLI, so install the CLI where Claude Code can find it: skilltotal --version
should work in the terminal you start Claude Code from. If the CLI is missing or fails to start,
install commands run unchecked and nothing is blocked. Claude Code then shows a warning on each
one, so you can tell a broken setup from a clean check.
The hook loads along with the plugins, so run /reload-plugins (or restart Claude Code) after
/plugin install. To see it work without touching a real package, ask the agent to run
npm install --registry https://registry.example.invalid left-pad. Claude Code should stop and ask
you, with SkillTotal's reason. Nothing is installed unless you approve.
What the plugin does not check
- Dependencies. It scans the package a command names, not the packages it pulls in.
- Installs that name no package.
npm install,npm ci,pip install -r requirements.txtanduv syncinstall from a lockfile or a requirements file, which the hook does not read. - Commands that build the package name at run time, and scripts the agent downloads and runs. The hook sees only what the command spells out.
- Minified bundles. They are listed in the report but not analyzed. The scripts npm runs at install are the exception: they are analyzed even when minified.
- Platform-specific wheels. For a PyPI release that ships them, SkillTotal scans the source distribution; the wheel pip picks for your machine can differ from it. A release whose only wheel is pure Python is scanned as that wheel.
- Behavior in Go, Rust, Java, Ruby and PHP code. Secrets, sensitive paths and hidden Unicode are still checked there.
It is a static guardrail, not a sandbox. For code you don't trust, run the agent in a container.
Usage
# Human-readable report
skilltotal scan ./path/to/component
# Scan a remote repository (shallow git clone)
skilltotal scan https://github.com/owner/repo
# Scan a project archive or a single file (e.g. an AI-generated project downloaded as a ZIP)
skilltotal scan ./my-project.zip
skilltotal scan ./app.tar.gz
skilltotal scan ./suspicious.py
# Scan a package from a registry (latest, or a pinned version)
skilltotal scan npm:left-pad
skilltotal scan npm:left-pad@1.3.0
skilltotal scan pypi:requests
skilltotal scan pypi:requests==2.31.0
# JSON to stdout
skilltotal scan ./component --json
# SARIF 2.1.0 (GitHub Code Scanning / IDE)
skilltotal scan ./component --sarif --output report.sarif
# Write the report to a file (SARIF if --sarif, else JSON)
skilltotal scan ./component --output report.json
# CI gate: exit code 2 by severity level or by risk score
skilltotal scan ./component --fail-on-high # alias for --fail-on high
skilltotal scan ./component --fail-on medium
skilltotal scan ./component --fail-on-score 50
# Skip paths (repeatable; combined with the config file's `exclude`)
skilltotal scan ./component --exclude "vendor/*" --exclude "*.min.js"
# Opt-in provenance for npm:/pypi: sources (registry metadata -> needs_review, never scored)
skilltotal scan npm:some-lib --provenance
# Baseline: snapshot current findings, then suppress them on later scans
skilltotal scan ./component --write-baseline .skilltotal-baseline.json
skilltotal scan ./component --baseline .skilltotal-baseline.json --fail-on-high
# Diff two versions of a component: what changed between them?
# Each side is any scannable source (path/archive/git/npm:/pypi:) or a saved --json report.
skilltotal diff npm:some-lib@1.2.3 npm:some-lib@1.2.4
skilltotal diff ./old-checkout ./new-checkout --json
skilltotal diff old-report.json new-report.json
# CI gate: fail (exit 2) if the new version INTRODUCES a high/critical finding
skilltotal diff npm:some-lib@1.2.3 npm:some-lib@1.2.4 --fail-on-new high
# Pre-install guard: allow/block decision (exit 2 on block) you can chain before installing
skilltotal guard npm:some-mcp-server && claude mcp add some-mcp-server -- npx some-mcp-server
skilltotal guard --installed # check every AI component already on this machine
skilltotal guard npm:x --block-on malicious # block only on malicious indicators
# Inventory: discover AI components already installed on this machine and scan them
# (reads agent configs for Claude Desktop/Code, Cursor, Windsurf, VS Code, Gemini, and
# local skills; derives an npm:/pypi:/local source per MCP server and runs the engine)
skilltotal inventory
skilltotal inventory --json
skilltotal inventory --no-scan # list only, do not scan
skilltotal inventory --project . # also include this project's agent configs
skilltotal inventory --sbom # AI-BOM: CycloneDX 1.6 JSON of your agent stack,
# scan verdicts attached as component properties
# List every detection rule
skilltotal rules list
skilltotal rules list --json
Baseline suppresses findings by a stable fingerprint of
(rule id, file, code snippet) — independent of line numbers, so it survives edits.
Suppressed findings are removed before scoring and do not affect the risk score.
Diff reports new / resolved / changed findings, evidence-level additions and removals
(matched by the same line-independent fingerprint as the baseline, so pure line shifts are
not noise), capability changes, and the risk-score delta. --fail-on-new LEVEL gates only
on risk the new version introduces — existing accepted findings never trip it, so it fits
upgrade reviews ("is 1.2.4 riskier than the 1.2.3 we already vetted?") without a baseline
file.
Guard is the install-time answer to "should I trust this component right now?".
Malicious indicators always block; scored risk at/above --block-on blocks;
capabilities alone never block — a legitimate MCP server with shell/network access
passes, so the guard stays quiet enough to leave enabled everywhere (unlike a raw
--fail-on high gate, which would trip on most of the ecosystem's honest capability
findings).
Provenance (--provenance, opt-in) adds registry-metadata signals for npm: /
pypi: sources: recently published, deprecated / yanked, no recent releases, no
repository link. Metadata is context about a component, not component content — so these
signals go to needs_review and never affect the score or verdict, and the default
scan stays component-only.
Project config (optional) — commit a .skilltotal.toml instead of repeating flags
(CLI flags override it):
fail_on = "high" # low | medium | high | critical
fail_on_score = 50 # or gate on the 0-100 risk score
exclude = ["vendor/*", "*.min.js"]
ignore = ["ST-NET-PY"] # rule ids to drop
baseline = ".skilltotal-baseline.json"
# Per-rule policy: reviewable gate decisions that live in the repo, not in a dashboard.
[policy]
"ST-SHELL-PIPE-EXEC" = "block" # gate trips (exit 2) whenever this rule fires,
# even with no fail_on configured
"ST-DYN-PY" = "warn" # explicit accept-but-show: reported, still counts toward
# the risk score, but exempt from the fail_on severity gate
"ST-SENS-WORD" = "ignore" # suppressed entirely (same effect as `ignore`)
To suppress a single finding, put a # skilltotal:ignore (or # skilltotal:ignore[ST-ID])
comment on its line. Inline markers count only when you scan a local directory (your own code,
as in CI). In a package or repository that SkillTotal downloads they are ignored, so a package
cannot silence its own findings.
python -m skilltotal ... works identically to the skilltotal console script.
Exit codes
| Code | Meaning |
|---|---|
| 0 | Success |
| 1 | Usage or collection error (for example an unknown flag value, a missing path or a failed clone) |
| 2 | A configured gate tripped (--fail-on/--fail-on-high severity, --fail-on-score or diff --fail-on-new), or guard blocked the component |
Gate semantics:
--fail-on/--fail-on-hightrip on the severity of any single finding, not the aggregaterisk_score. A component can reportrisk_level: low(score 0) and still fail the gate if it has a high-severity finding, including a capability such as shell or network access, which is reported but never scored as malicious. To gate on the score instead, use--fail-on-score; to accept known findings, use a baseline, an inline# skilltotal:ignore[ST-ID], or a per-rule[policy]action (block/warn/ignore).
CI / GitHub Action
Run SkillTotal in CI and surface findings in your repository's Security → Code scanning tab.
# .github/workflows/skilltotal.yml
name: SkillTotal
on: [push, pull_request]
permissions:
contents: read
security-events: write # required to upload SARIF to Code Scanning
pull-requests: write # required only for comment-on-pr (optional)
jobs:
scan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: pezhik/skilltotal@v0.62.0
with:
source: . # a path, a git URL, or an npm:/pypi:<name> spec
fail-on: high # fail the build on a high/critical finding (or 'none')
comment-on-pr: 'true' # post a sticky summary comment on pull requests (optional)
The action runs the engine that ships with the tag you pin, so nothing is downloaded from PyPI. It
scans source, uploads SARIF (so findings appear inline on pull requests and in Code Scanning) and
fails the job on a high/critical finding unless fail-on: none. With comment-on-pr: 'true', it
posts a single summary comment on the pull request (risk level, score, findings, capabilities) and
updates it in place on later runs. The comment is off by default and needs pull-requests: write.
Pin the action to a released tag (see Releases),
and set the version: input only to run a different engine release from PyPI. Without the action,
the same scan is skilltotal scan . --sarif --output skilltotal.sarif --fail-on-high.
Use as a pre-commit hook
Run SkillTotal on every commit via pre-commit:
# .pre-commit-config.yaml
repos:
- repo: https://github.com/pezhik/skilltotal
rev: v0.62.0
hooks:
- id: skilltotal
args: [".", "--fail-on-high"] # scan the repo; block the commit on a high/critical finding
Then pre-commit install. The hook installs the CLI in its own environment and scans the repo
on commit; tune the scan with the same flags as the CLI (e.g. --exclude, --fail-on).
Use as an MCP server
Let your agent check a component before installing it. skilltotal mcp runs the engine
as a stdio MCP server (stdlib-only, still zero dependencies) — register it in Claude
Code/Desktop, Cursor, Windsurf, or any MCP client:
{ "mcpServers": { "skilltotal": { "command": "skilltotal", "args": ["mcp"] } } }
If you have not installed it locally, use "command": "npx", "args": ["-y", "skilltotal", "mcp"]
instead. This needs uv, pipx, or Python 3.11+ with skilltotal.
Tools exposed: scan_component (full report for a path / git URL / npm: / pypi:
source), diff_components (upgrade review: what changed between two versions), and
list_rules. Scans run locally with the same never-execute static engine — the component's
code is not uploaded anywhere.
Add a status badge
Scan a component on skilltotal.ai and each report offers an "Add this badge" snippet — a small SVG that always reflects the component's latest scan and links back to the full report. Drop it in your README so visitors see the risk at a glance:
[](https://www.skilltotal.ai)
Copy the exact, ready-to-paste markdown from the report page — it fills in the badge URL for you.
Methodology
SkillTotal performs static security analysis of AI components — MCP servers, agent skills/plugins, npm and PyPI packages, and AI-generated projects/repositories. The engine combines capability analysis, dangerous-pattern detection, privilege analysis, supply-chain (install-time) analysis, prompt-surface analysis, and data-flow correlation (e.g. secret access combined with network egress). Findings are mapped to risk categories and contribute to a 0–100 risk score; capabilities are reported but never inflate the score — capability ≠ risk. Nothing is executed and no LLM is called, so results are deterministic and reproducible.
What it detects
| Category | Examples |
|---|---|
| Shell execution | subprocess.*, os.system, child_process.exec |
| Filesystem access | open, read_text/write_text, fs.readFile/writeFile |
| Sensitive paths | ~/.ssh, ~/.aws, .env, id_rsa, credentials, secrets |
| Network egress | requests, urllib, aiohttp, fetch, axios |
| Install-time execution | npm preinstall/postinstall/prepare, setup.py hooks |
| Dynamic code execution | eval, exec, compile, new Function, vm.runInNewContext |
| Obfuscation | decode-and-execute chains, base64 blobs, hex escaping, minification |
| MCP risks | manifests, dangerous tools (shell/fs/network/credential), server commands |
| Prompt surface | "ignore previous instructions", "reveal system prompt", exfiltration phrasing |
Coverage by component type
Legend: ✅ analyzed by default for this component type · ⚠️ the engine detects this, but that surface is uncommon for this type — so it is flagged only when the component actually contains it (e.g. prompt-injection text inside an npm/PyPI package) · ❌ not applicable to this type · 🚧 planned (SkillTotal Cloud).
Columns are the component types SkillTotal scans. AI project = a scanned repository or folder — an agent skill/plugin, an AI-generated codebase, or a set of prompts/configs — that is not a published npm/PyPI package.
| Category | MCP | npm | PyPI | AI project |
|---|---|---|---|---|
| Prompt injection / instruction override | ✅ | ⚠️ | ⚠️ | ✅ |
| Tool poisoning (MCP tool metadata) | ✅ | ❌ | ❌ | ⚠️ |
| Dangerous capabilities (shell / fs / network) | ✅ | ✅ | ✅ | ⚠️ |
| Data exfiltration (secret access + egress) | ✅ | ✅ | ✅ | ⚠️ |
| Secret theft / sensitive-path access | ✅ | ✅ | ✅ | ⚠️ |
| Dynamic code execution | ✅ | ✅ | ✅ | ⚠️ |
| Obfuscation (decode-and-execute) | ✅ | ✅ | ✅ | ✅ |
| Hidden-Unicode smuggling | ✅ | ✅ | ✅ | ✅ |
| Embedded secrets (hardcoded keys/tokens) | ✅ | ✅ | ✅ | ✅ |
| Install-time / supply-chain hooks | ⚠️ | ✅ | ✅ | ❌ |
| Overprivileged / auto-approved tools | ✅ | ❌ | ❌ | ⚠️ |
| Runtime behavior analysis | 🚧 | 🚧 | 🚧 | 🚧 |
| Sandbox analysis | 🚧 | 🚧 | 🚧 | 🚧 |
Typical findings
- An MCP tool can execute arbitrary shell commands
- A package downloads and runs code from an external URL
- Access to credential locations (
~/.aws,~/.ssh,.env) detected - Dynamic code execution (
eval/exec) detected - Prompt-injection / instruction-override phrasing in a tool description or skill
- Sensitive-data access combined with outbound network egress
- Hardcoded API keys or tokens
- An MCP server with auto-approved or overprivileged tools
- Untrusted input (environment,
sys.argv, a request/response body) flowing intoexecor a shell — a proven injection path, not just a dangerous API in isolation - An agent skill does more than its declared
allowed-toolsallow (undeclared capability / least-privilege violation)
Out of scope
SkillTotal statically analyzes a single component's own files. It does not execute code, observe runtime behavior, or assess your environment, deployment, or infrastructure. It is not a substitute for:
- a penetration test
- an application-security (app-sec) review
- an architecture / design review
- a cloud-security or infrastructure assessment
- a Kubernetes / container runtime audit
- a business-logic review
- a manual code review
Runtime behavior and sandbox analysis are planned for SkillTotal Cloud (paid).
Output
A normalized report containing the component identity, a risk score (0–100) and risk level (low / medium / high / critical), detected capabilities (each evidence-backed), a behavioral trait fingerprint (with a CSA / MAESTRO / MITRE ATLAS crosswalk), findings, needs_review, and metadata. See docs/report-schema.md and docs/scoring.md.
Every finding also carries its OWASP Agentic Skills Top 10 category ids (owasp), emitted in
both the JSON report and SARIF (native taxonomies/relationships);
docs/owasp-agentic-skills-mapping.md explains the coverage
(AST01–AST05) and the honest gaps. For MCP servers,
docs/mcp-owasp-mapping.md maps SkillTotal's checks to the OWASP MCP
Security Cheat Sheet (and names the runtime controls a static engine can't cover).
The report's traits array is a behavioral fingerprint — a higher-level projection over the
findings (e.g. execution_authority, embedded_credential, untrusted_perception, and the
emergent exfil_correlation combination) — each mapped to the Cloud Security Alliance
trait-based model, a MAESTRO threat-model layer, and a MITRE ATLAS tactic where there is an honest
fit. It is descriptive and never affects the score; see
docs/trait-crosswalk.md.
Architecture
The package under skilltotal/ (except cli.py) is a pure, side-effect-free library so the
same engine powers the web app at www.skilltotal.ai. See
docs/architecture.md.
Development
pip install -e ".[dev]"
pytest
Accuracy notes
- Python is analyzed via an AST (resolves import aliases, tells
open(p,'w')from a read, ignores API names that only appear in strings/comments). Node.js/config use regex. - Languages: shell execution, network access, file access and dynamic code are detected in
Python and JavaScript/TypeScript. Go, Rust, Java, Ruby and PHP files still get the secret,
sensitive-path and hidden-Unicode checks, but their behavior is not analyzed yet. If a component
ships code in those languages, the report adds a
needs_reviewnote, and a low verdict reads "Partially analyzed" instead of "No significant risks found". - Test code (
__tests__/,*.test.*,tests/,conftest.py, …) is demoted toneeds_review— it is not executed by consumers, so it does not affect the score. - Ambiguous signals (bare
secrets/credentialswords, lone base64 blobs, "before answering" phrasing, minified files) go toneeds_review, never tofindings. - Hidden Unicode (ASCII-smuggling tag characters, Trojan-Source bidi overrides,
zero-width chars) is detected and decoded — a real evasion used to smuggle instructions
past human review. See
tests/manual_eval/for calibration against real-world attacks. - Shell execution covers
subprocess/os.system,asyncio.create_subprocess_*, Nodechild_process, and common process-spawning libraries (Pythonsh/plumbum/pexpect/invoke/fabric; Nodezx/execa/cross-spawn/shelljs/tinyexec/node-pty). - MCP dangerous tools are classified by name/description both in JSON manifests and when
defined in code (
server.tool("run_command", …),@mcp.tooloverdef read_file). - Limitations: detection is at the call/import level. Capability via an unrecognized higher-level library (e.g. a git library that writes files internally, a browser library) may not be flagged as a raw filesystem/shell call. Capabilities indicate presence, not proven misuse.
Open source vs SkillTotal Cloud
SkillTotal is open core. This engine (the analysis, all detection rules and the CLI) is open source and complete on its own: you can run it locally or in CI, for free, with zero runtime dependencies. It tells you what a component does, with evidence.
Paid features are planned for SkillTotal Cloud (the website) and will explain why it matters: LLM interpretation and prioritization of findings, dynamic sandbox execution, scan history and monitoring. They will run as server-side services on top of this engine, and their code is not part of this repository. See docs/open-core.md.
License
Apache-2.0. See also NOTICE.