Docs Test Architecture

Test Architecture

The test tiers, what each one asserts, how to regenerate its fixtures, and the self-enforcing gate tests.

On this page

Over two hundred integration-test binaries under tests/ — every tests/*.rs file is one, because nothing turns Cargo’s auto-discovery off — plus unit tests inside src/, plus doctests. The exact total is not repeated here on purpose: the homepage stat card carries it and a pre-commit gate holds it, so any number written on this page would be wrong within a week. The opening sentence read “fifty” until 2026-09-06, when a sweep of these pages counted 238.

What follows is the map: what each tier asserts, how to regenerate its artifacts, and — most importantly — the roster of gate tests that fail you for reasons that have nothing to do with the code you were writing.

Tiers

TierFilesAsserts
CLI surfacecli_test, cli_options_test, cli_help_test, cli_defaults_test, cli_flag_behavior_testParsing, the default value of every parameter, --help grouping across ~160 flags, and behavior for flags that were once untested.
CLI goldencli_goldenstrycmd goldens under tests/cli/ — exact stdout/stderr for a command line.
Integrationintegration_test, capture_test, rtp_integration_test, pipeline_test, parse_path_test, bootstrap_test, app_servers_test, config_testCapture-to-output pipeline, and the library facades WS2 extracted from main.rs — bootstrap planning is a pure Cli + Config → RunPlan function precisely so these tests can reach it. config_test covers the step before planning: config discovery (-f, SIPNAB_CONFIG, --no-config, unknown-key warnings, the missing-file error) driven through the real binary’s --dump-config.
TUItui_snapshot_test, tui_state_test, tui_e2e_testRendered buffers via ratatui’s TestBackend + insta; state-machine transitions; and end-to-end drives of the real binary inside tmux. See TUI testing.
Serversapi_test, api_token_test, mcp_stdio_test, mcp_http_test, mcp_token_test, mcp_token_rotation_test, metrics_test, hep_testREST, MCP (stdio and HTTP), signed-token auth and rotation, Prometheus scrapes, HEP ingestion — all end to end against a spawned process.
Securitysecurity_test, privilege_drop_test, resource_bounds_test, crash_testAudit regressions, the never-continue-as-root guarantee, attacker-keyed map caps, and crash-report handling with the real binary.
Property & fuzzproperty_test, smoke_fuzz_test, fuzz_corpus_replayproptest invariants; the always-on stable-toolchain fuzz floor; and — despite its name — a deterministic sweep over an adversarial seed set defined in the file itself, not a replay of fuzz/corpus/ (it opens no files).
Contractjson_schema_test, summary_consistency_test, output_behavior_test, error_types_test, api_guidelines_test, wasm_exports_testThe shapes other people’s code depends on: JSON schemas, cross-surface summary agreement, machine-readable output flags, typed errors, #[non_exhaustive] on growth-prone enums, WASM export list.
Docs enforcementdocs_drift_test, dev_docs_drift_test, link_integrity_test, doc_example_coverage_test, config_examples_test, site_journey_test, mockup_alignment_testSee the gate roster below.
Governanceflag_coverage_test, keybinding_drift_test, feature_gate_test, config_wiring_testRules about how the project grows, not about a specific behavior.
Metasupport_selftestThe shared test helpers have their own tests.

tests/support/

Files in a subdirectory of tests/ are not compiled as their own test binaries, so an explicit path attribute pulls shared helpers in:

#[path = "support/run.rs"]
mod run;

That idiom is why the same module appears in several test files with no mod.rs chain — each test binary compiles its own copy.

HelperProvides
mod.rsShared normalization (normalize()) used to compare output across platforms; self-tested by support_selftest.
run.rsThe canonical binary-spawn helper for CLI/output/config integration tests — one place that knows how to find and invoke the built binary.
server.rsREST API spawn harness: start the server on an ephemeral port, wait for readiness, tear down.
mcp.rsThe same for HTTP MCP, including the JSON-RPC framing.
schema.rsJSON-Schema validation against tests/schemas/.
tui_fixtures.rsSIP fixture builders shared by the TUI snapshot and state tests.
fuzz.rsA deterministic xorshift PRNG shared by the stable-toolchain fuzzers, so a failure reproduces from its seed.

Fixtures and corpora

Every command below ran against the current tree and left it unchanged — that is the point of documenting them: if one produces a diff, something has drifted.

ArtifactRegenerate with
tests/fixtures/ — synthetic pcapscargo run --features native --bin gen_fixture
tests/snapshots/ — TUI bufferscargo insta test --features tui --accept (needs cargo install cargo-insta)
tests/cli/ — trycmd goldensTRYCMD=overwrite cargo test --features full --test cli_goldens
tests/schemas/ — JSON SchemasHand-maintained; json_schema_test validates output against them.
tests/pcap-samples/ — capture fixturesThirteen are synthetic: python3 tests/gen-pcap-samples.py writes ten and python3 tests/gen-link-type-samples.py writes the three link-layer framings (DLT_LOOP, PPPoE inside Linux cooked capture v1 and v2); each script in check mode reports any of its own that have drifted from it. Nothing regenerates the rest; they stay as checked in. New ones arrive through harness/scripts/capture.sh then promote.sh, which record the run in PROVENANCE.md beside them — every_committed_capture_fixture_says_where_it_came_from refuses a fixture with no entry, and the 36 that predate that rule sit on an enumerated list which only shrinks.
tests/install-sh/ — installer casesHand-maintained; exercised by the install-sh CI job.
fuzz/corpus/ — fuzz seedsGrown by cargo fuzz run (nightly). Note nothing on the stable toolchain reads this directory: fuzz_corpus_replay drives its own in-file seed set. To make a reproducer run in every cargo test, add it to smoke_fuzz_test as well.

Accepting a snapshot or overwriting a golden is a decision, not a fix. Read the diff first: these files are the record of what the tool promised its users.

The gate-test roster

These fail on changes you thought had nothing to do with them. Each one exists because the thing it guards silently rotted at least once.

GateTrips when
docs_drift_testA --flag named on a published page does not exist in the CLI. That corpus derives from git ls-files rather than a hand-kept list, and covers every tracked docs/ page except design/, research/ and superpowers/, plus website/content/, README.md, SECURITY.md and CONTRIBUTING.md — these pages included, so a phantom flag here fails too. A flag belonging to another tool goes in FOREIGN_FLAGS, scoped to the label of each page naming it. Also: a version marker in the docs or man page disagrees with Cargo.toml; the README feature table misses a Cargo feature; a [theme] slot in ThemeConfig has no documentation, or the slot count quoted in either config reference is wrong. Also the benchmark reproducibility contract: the bench/ harness the benchmarks page tells readers to run must exist and be executable, bench/carrier.py must still produce the corpus composition the page quotes (checked at 1/100 scale), and both doc trees must name the same measured artifact and date.
dev_docs_drift_testA page under docs/internals/ links to a path that no longer exists, names a fn that no longer exists, uses an absolute GitHub URL instead of a relative one, is not registered in build-wiki.py, or breaks a mermaid convention. It also builds the wiki and fails if any relative link survives into the output — the wiki is flat and has no repo tree, so such a link publishes dead. That check runs the generator rather than reading it: the version that only asserted CODE_LINK_RE appeared in the script passed while ](../bench/) was shipping broken.
link_integrity_testAny relative link or heading anchor in either doc tree does not resolve; Zola content uses a plain relative .md link that would render as a dead URL. On the wiki-source side the scan is the top-level docs/*.md plus docs/internals/docs/design/, docs/research/ and docs/superpowers/ are planning material outside the published journey and are deliberately not walked, though a link into them from a scanned page still has to resolve. A third scope covers the root community files (README.md, SUPPORT.md, MAINTAINERS.md, CONTRIBUTING.md, SECURITY.md, CODE_OF_CONDUCT.md), which neither doc tree walks — these are what GitHub renders in the sidebar, so a rename that broke their cross-references used to go unnoticed.
doc_example_coverage_testA user-facing CLI flag appears in fewer than two documented examples. A ratchet — the exemption list may only shrink.
flag_coverage_testA new long flag ships with no test referencing it. Also a ratchet: adding a test for a grandfathered flag fails until you remove it from the baseline list.
keybinding_drift_testA controller handles a key that the F1 help never mentions.
config_wiring_testA config key exists but is never read, or a CLI flag has no config fallback where its peers do.
feature_gate_testA flag whose subsystem is not compiled in fails late or silently instead of fast and clearly.
api_guidelines_testA growth-prone public enum loses #[non_exhaustive], or a shared store stops being Debug.
summary_consistency_testTwo serializing surfaces disagree about a dialog or stream field.
wasm_exports_testThe WASM binding loses an export the browser analyzer calls.
site_journey_test / mockup_alignment_testA website journey breaks, or a terminal mockup on the site drifts from real output. Also the homepage’s advertised numbers: the binary-size ceiling, the glibc floor, the throughput tiles (which must quote a figure that appears on the benchmarks page they link to), and the two automated-test counts, which must agree with each other and — via quality.yml — with the measured total. It holds the standards cards to a canonical table naming the code symbol behind every item, the field the program emits for every quality metric, and a citation in docs/ and the built site mirror, one card per standard. It also holds the homepage’s four MCP examples to the files gen-mcp-examples.sh generated, and holds each of those to the claim the surrounding copy makes about it. That covers the second half of live binary -> website/data/mcp-examples/*.json -> index.html; demos/gen-mcp-examples.sh --check covers the first and needs a built binary, so CI cannot run it.
config_examples_testA config sample in the docs no longer parses.
doc_commands_run_testA command shown in docs/ is one sipnab refuses. It extracts every sipnab … line from every shell block and RUNS 323 of the 352 against a capture that ships with the repository, failing on a usage error rather than on a non-zero exit – plenty of documented commands exit non-zero for honest reasons. A device or root example runs with a device name that cannot exist substituted for the one the page names, so its flags still parse in full; a server example runs under a shared wall-clock bound, because clap refuses in milliseconds and a process still alive after it necessarily parsed its arguments. Only two kinds never run, and it names both: a shell program rather than one invocation, and a command that would exec something of its own or attach to the kernel. A command matching neither fails the test. It prints its own coverage, because the difference between “323 ran” and “3 ran, 349 skipped” is the difference between a gate and a decoration.
homepage_claim_truth_testA claim the homepage states in PROSE stops being true: the MCP tool counts and their read-only split, which released builds carry the bpf feature, what the installer actually puts on a host, a capability the static musl build does not have, the filter language’s field and operator counts, a TUI key the page names that nothing binds, and a CLI flag the page names that does not exist, plus the unsafe count and the safety claim resting on it — the page said “memory-safe by construction” beside 88 unsafe blocks and no forbid(unsafe_code). site_journey_test gates the tiles; this gates the sentences, because the two disagreed on one page for months.
support_selftestThe shared normalization helper changes behavior under the tests that depend on it.

The rule for all of them: the gate is not the problem. If flag_coverage_test fails, the flag needs a test. If dev_docs_drift_test fails, a page now lies. Adding an exemption is the last resort, and every exemption list in this repo works as a ratchet so it cannot quietly grow.

The development loop

Logging. SIPNAB_LOG is a tracing EnvFilter, so it takes levels and per-module directives (SIPNAB_LOG=debug, SIPNAB_LOG=sipnab::rtp=trace). In TUI mode sipnab suppresses logging unless SIPNAB_LOG says otherwise — the alternative is log lines painting over the interface. -q lowers the default in CLI mode. Configured by init_logging().

Benches. cargo bench --profile profiling. Plain cargo bench cannot build: the cdylib crate type needed for the WASM build forces panic = "abort" into the lib unit while bench units build with unwind, so shared dependencies get built twice with incompatible panic strategies and fail to unify. The profiling profile is release codegen with panic = "unwind" and debug symbols kept, which is also what callgrind wants. Baselines live in benches/BASELINES.md.

nextest. .config/nextest.toml defines three profiles: default (no retries, 30s slow timeout), ci (no retries, no fail-fast, immediate-final output) and e2e (2 retries, 60s timeout) for the timing-sensitive tmux tests. Worth knowing: no workflow currently invokes nextest — CI runs plain cargo test. The config is there for local use and for the day the e2e shim goes away.

The docker lab. harness/ is a docker-compose stack — OpenSIPS, rtpengine, SIPp — for generating real traffic. make up in that directory builds and starts it, and make down tears it down.

WASM. One pre-commit gate covers the browser analyzer: it checks that website/static/wasm/sipnab.js still exports every function the site calls.

A second gate once demanded a freshly built bundle alongside any staged src/wasm.rs. It no longer exists, nor does the binary it guarded: the published analyzer went eleven releases stale while that gate stayed green, because src/wasm.rs had no commits in the window — the interface held still while the implementation behind it moved. The Pages workflow builds the bundle at deploy time now, so the build produces what ships rather than the tree carrying it, and it cannot go stale.

Feature matrices. A green cargo test --features full is not proof: CI also builds reduced feature sets, and code behind #[cfg(not(feature = ...))] is invisible to the full build. Before pushing, cargo clippy --workspace --all-features --all-targets -- -D warnings and the fuzz workspace check are what the pre-push hook runs for exactly this reason.