See what your npm dependencies actually do when they run — and gate your build on it.
Quick start · Two engines · What it catches · Measured results · CLI · Docs
A package that quietly reads ~/.npmrc and POSTs it to an unknown host — caught at runtime, scored, and failed in CI:
$ npx bheeshma --enforce -- node app.js
══════════════════════════════════════════════════════════════════════
BHEESHMA Runtime Dependency Behavior Report
══════════════════════════════════════════════════════════════════════
📦 chalk-helper@2.4.1 Trust Score: 29/100 🔴 CRITICAL
Observed behaviors
🔐 ENV ACCESS 📖 FS READ ~/.npmrc 🌍 HTTPS REQUEST 🧭 DNS QUERY
Pattern analysis — behavioral correlations (2 threats)
📤 Data exfiltration CRITICAL read .npmrc + outbound HTTPS
🔑 Credential theft HIGH CREDENTIAL_FILE_READ — .npmrc
──────────────────────────────────────────────────────────────────────
✗ POLICY VIOLATION: chalk-helper@2.4.1 is CRITICAL → exit 1Static and registry-graph scanners (Socket, Snyk, Dependabot, npm audit) reason about a package before it runs. BHEESHMA adds the missing, complementary view — what your dependencies actually do when they execute — and gates CI on it.
flowchart LR
A["📦 your deps run"] --> B["🪝 hooks / syscalls"]
B --> C["📡 signals<br/>env · fs · net · dns · exec"]
C --> D["🎯 attribute<br/>→ package"]
D --> E["🔗 correlate<br/>patterns"]
E --> F["💯 trust score<br/>0–100"]
F --> G{"🚦 fail-level<br/>gate"}
G -->|"≥ high"| H["❌ fail build<br/>+ SARIF"]
G -->|ok| I["✅ pass"]
Warning
BHEESHMA's default in-process engine runs with the same privileges as the code it watches — it is telemetry and a detection aid, not a containment boundary, and a motivated attacker can evade or disable it. The optional out-of-process engine provides a real boundary. It is defense-in-depth, meant to run alongside static tooling. See docs/THREAT_MODEL.md for exactly what each engine does and does not catch.
flowchart TB
subgraph IN["🟢 In-process engine (default — runs anywhere)"]
direction LR
H1["monkey-patched<br/>Node APIs"] --> S1["signals +<br/>stack attribution"]
end
subgraph OUT["🔵 Out-of-process engine (bheeshma-sandbox · Linux + strace)"]
direction LR
K1["strace / ptrace<br/>(kernel)"] --> S2["syscalls +<br/>process-lineage attribution"]
end
S1 --> SC["⚖️ scoring + correlated patterns"]
S2 --> SC
SC --> OUT2["📤 CLI · JSON · HTML · SARIF + exit code"]
| 🟢 In-process (default) | 🔵 Out-of-process (bheeshma-sandbox) |
|
|---|---|---|
| How | Wraps Node APIs, attributes by stack | Kernel syscall tracing (strace/ptrace) |
| Per-package attribution of JS behavior | ✅ precise | ◑ per-process (great for installs) |
Sees native subprocess egress (curl…) |
❌ | ✅ |
| Resists the watched code disabling it | ❌ | ✅ |
| Can prevent, not just detect | ❌ | ✅ --block-network |
| Runs anywhere / zero deps | ✅ | Linux + strace |
Tip
Use the in-process engine for fast, precise CI signal everywhere; add the out-of-process engine where native egress, evasion-resistance, or prevention matter — e.g. monitoring npm install. Details in docs/ARCHITECTURE.md.
GitHub Actions — annotates every PR via Code Scanning:
# .github/workflows/ci.yml
- uses: bb1nfosec/bheeshma/.github/actions/bheeshma@v3.0.0
with:
command: 'npm test'
fail-level: 'high' # default; see docs/ENTERPRISE.md for tuningLocally:
npx bheeshma -- node app.js # monitor any command
npx bheeshma install # monitor `npm install` (postinstall scripts)
npx bheeshma --format sarif -o results.sarif -- npm test
# out-of-process engine (Linux + strace): observe + optionally PREVENT egress
npx -p bheeshma bheeshma-sandbox --enforce -- npm ci
npx -p bheeshma bheeshma-sandbox --block-network -- npm ci| Behavior | Signal | Detects |
|---|---|---|
| 🔐 Env access | ENV_ACCESS |
Credential / API-key theft |
| 📖 File read | FS_READ |
Recon, credential-file access |
| 📝 File write | FS_WRITE |
Persistence, backdoor install |
| 🌐 TCP connect | NET_CONNECT |
Reverse shells, C2 |
| 🌍 HTTP / HTTPS | HTTP(S)_REQUEST |
Data exfiltration — http/https modules (incl. .get) and global fetch |
| 🧭 DNS query | DNS_QUERY |
DNS tunneling / encoded-subdomain exfil |
| ⚡ Shell exec | SHELL_EXEC |
Arbitrary code execution |
| 🌀 Obfuscation | OBFUSCATION_DETECTED |
Hidden payloads (eval/Function/hex) |
🔗 Correlated patterns — combinations that cap the trust score into a risk band
- Data exfiltration — credential read + outbound connection →
CRITICAL/HIGH - DNS tunneling — high-entropy / known-exfil query names →
HIGH - Backdoor — reverse-shell command or suspicious port →
CRITICAL/HIGH - Persistence — writes to shell rc / cron /
~/.ssh/authorized_keys/ systemd →HIGH - Credential theft — secret-env read + egress; context-aware
.envreads (dotenv = expected) →HIGH/LOW - Crypto mining — wallet / pool indicators →
CRITICAL - Typosquat — name 1 edit from a popular package →
MEDIUM - Obfuscation + network — obfuscated source making outbound calls →
HIGH
Trust score & gate — each package gets a deterministic score [0–100] → <30 🔴 CRITICAL · <60 🟠 HIGH · <80 🟡 MEDIUM · else 🟢 LOW. The CI gate fails at or above fail-level, which defaults to high.
Evidence, not adjectives. Reproduce:
npm run benchmark·node benchmark/fp-real.js·npm run perf. Full detail + caveats in benchmark/FINDINGS.md.
| Gate | 🎯 Detection (modeled attacks) | 🟢 False positives (71 real packages) |
|---|---|---|
critical |
29% | 0% |
high (default) |
💯 100% | 0% |
⚙️ Overhead ≈ 6.7× on a worst-case pure-hook microbenchmark (real workloads are far lower — most time is non-hooked work).
Note
Detection corpora are synthetic/limited — validate against your own dependency tree before relying on the gate. This measures real-world npm malware classes, not nation-state evasion.
bheeshma -- <command> # monitor (in-process)
bheeshma-ci -- <command> # CI-tuned: SARIF + exit codes
bheeshma install [ci] # monitor npm install / npm ci
bheeshma-sandbox -- <command> # out-of-process (strace); --block-network to prevent egress
bheeshma --format <cli|json|html|sarif> -o report.ext -- <command>
bheeshma --enforce --fail-level high -- npm test # exit 1 at/above fail-level (default high)📦 Programmatic API
const bheeshma = require('bheeshma');
bheeshma.init();
require('./your-app');
console.log(bheeshma.generateReport('json'));
const result = bheeshma.enforcePolicy({ failLevel: 'high' });
if (!result.passed) process.exit(1);Types ship in the package (src/index.d.ts), validated against the runtime by a drift-guard test.
🤖 GitHub Action inputs (SARIF v2.1.0 → Code Scanning annotations)
| Input | Default | Description |
|---|---|---|
command |
(required) | Command to run under monitoring |
fail-level |
high |
Min risk to fail the build (critical/high/medium/low) |
config |
'' |
Path to .bheeshmarc.json |
sarif-output |
bheeshma-results.sarif |
SARIF output path |
skip-low |
true |
Skip LOW signals in SARIF (less noise) |
upload-sarif |
true |
Upload to Code Scanning |
🛠️ Configuration (.bheeshmarc.json, entirely optional)
{
"thresholds": { "critical": 30, "high": 60, "medium": 80 },
"packageThresholds": { "axios": 40 },
"whitelist": ["express@*", "@types/*"],
"blacklist": ["known-malicious-package"],
"patterns": { "enabled": true },
"performance": { "maxSignals": 10000, "deduplicateSignals": true }
}- 🚫 No telemetry, local-only — all analysis runs on your machine (the only outbound is an optional alert webhook you configure).
- 🏷️ Metadata only — records hosts, ports, paths, env-var names; never secret values, file contents, or request bodies.
- 📭 Zero dependencies and no install scripts — a security tool shouldn't run the very thing it warns about.
- 🛟 Fail-safe — hook errors never break your application.
| 🧬 Architecture | the two-engine design and trade-offs |
| what it does and does not catch — read before relying on it | |
| 🏢 Enterprise guide | deployment, gating, assurance, evaluation checklist |
| 📊 Benchmark findings | measured detection / false-positive results |
| 📋 Roadmap · Release checklist | what's next, how releases are cut |
A live threat-intel dashboard tracking known npm supply-chain attacks and how bheeshma's signals map to them:
Auto-updates daily from the OSV API + a curated threat set (scripts/fetch-threats.js).
npm test # 83 tests: in-process harness + CLI integration (offline, deterministic)
npm run benchmark # efficacy benchmark (detection / FP by gate)
npm run perf # monitoring-overhead microbenchmark
node benchmark/fp-real.js # false-positive sweep on real packages (needs network)CI runs the unit, CLI-integration, and benchmark suites across Node 14 / 18 / 20 / 22.
PRs welcome — see CONTRIBUTING.md. Especially valuable: real-world attack replay scripts (demos/), false-positive reports with repros, new hook/syscall coverage, and out-of-process engine work (eBPF, seccomp/Landlock).
BHEESHMA — trust, but verify. At runtime. 🛡️