Engineering notes
Engineering notes
Walkthroughs, new features, and problems worth writing down: how to do a thing with sipnab, what a release added, and what broke along the way.
Each entry carries a label that says which kind it is:
- How-to: do one job, end to end, with the commands that actually ran.
- Feature: what a release added and what each part is for.
- Post-mortem: a real problem with real numbers. A regression that shipped, a protocol assumption that turned out to be wrong, or a gate that passed when it should not have. These sit in a collapsed list at the end.
For what changed in each release, read the changelog. For reference material, read the documentation.
How-tos
-
A finding is a floor, not a verdict
The same problem query answered 0 and then 127 on one capture, thirty seconds apart. Four questions to ask of any finding before you act on it, and the fields that answer each one.
-
Hand over evidence somebody else can check
You found the answer. Now a carrier that does not trust you has to confirm it. Four artifacts, each answering a different doubt, and the one thing sipnab deliberately refuses to claim.
-
One call, two nodes, and neither saw the whole thing
The SBC says it sent the call, the PBX says it never arrived, and a B2BUA changed the Call-ID in between. How to join the two legs, and how to tell an identifier match from a guess that scored well.
-
Read SIP over TLS on a box you cannot restart
No private key, no maintenance window, and the proxy has to keep serving calls. Three methods survive those constraints, they cost different things, and one command tells you which of them your host can actually run.
-
Rule out the relay before you blame the NAT
One-way audio has four ordinary causes and they look identical in a complaint ticket. The order to eliminate them costs four commands, and the first one rules out the answer everybody reaches for first.
-
Share the capture, not the customer
A vendor wants the packets and the file holds everybody's traffic. What sipnab redacts, what it deliberately does not, and the one narrowing that actually reaches the pcap writer.
-
The call answered and the audio never started
A 200 OK, an ACK, and then nothing. Half the work is proving your capture point could have seen the media, because sipnab deliberately refuses to call a signaling-only tap a silent call.
-
What to switch on when an agent drives sipnab
Fifty-seven tools register by default and none of them changes anything. Six flags open the doors that do, ranked by what an agent could do with each, plus the one flag to leave on whatever else you decide.
-
Keep the audio your store refuses
A vCon container carries its audio inline, and sipnab caps that at 5 MiB by default. Here is where the number came from, when to raise it, and how to say never inline media at all.
Features
-
A total that cannot say what it missed is a floor
Kernel drops, interface drops, unusable timestamps, frames no decoder read, and SIP the port gate set aside — five separate channels, kept separate because their remedies disagree. Plus the run record that ties an artifact back to the command that made it.
-
An observer's vCon, and its third verdict
sipnab writes a conversation container from a tap: nothing signed, no vouched-for name, audio inline only when the run kept it, and a block stating what the capture missed. Plus why the container validator has three verdicts rather than two.
-
Fifty-seven tools, and the two that stay off
The MCP surface an agent sees: what the full and core profiles register, why every answer carries how much of the capture it rests on, and why the two opt-in flags are opt-in — one of them makes sipnab transmit.
-
One filter language, every door
The same expression narrows --filter, decides which calls export as vCons, scopes a CI expectation and pages an MCP tool. Why reusing one language beats growing a flag per policy, and what an unmeasured value matches.
-
Reading SIP over TLS without a certificate
The wire never yields encrypted SIP on its own. sipnab takes session keys from wherever you can get them — a key log, a pipe, an eCapture probe on a daemon you cannot restart — and states plainly what none of that recovers.
-
Where an endpoint came from, and what that path is worth
Three tools answer what no other surface does: how sipnab learned a media endpoint, whether the path that carried it authenticated anybody, and whether anyone asked the relay at all.
-
What 0.5.128 added to the vCon exporter
Seven fields the format defines and sipnab was not emitting, transfer objects for observed REFERs, a configurable media ceiling, tombstones for withheld dialogs, and RFC 9457 problem-detail errors. What each one is for.
Post-mortems 18 what broke, and what the gate did not catch
-
The parser worked, its tests passed, and the feature was dead
SIPREC metadata had a parser, unit tests and a field on the dialog. It was never filled on a live capture, because the fault sat between the parser and the store where no unit test could see it.
-
Two tests that only failed on macOS, and the thing they had in common
Both encoded when a kernel reports a peer's disconnect as if it were the behavior under test. Both passed on Linux, and on the one platform the author could not run, both were wrong.
-
An instruction with no gate is a suggestion
A written rule said to delete every temp file in the same turn. The checkout held 21 stray logs, three abandoned worktrees and a 1.6 TB target directory. What the enforcement had to prove before it could delete anything.
-
Seventy words somebody thought of
A gate had enforced US English since 0.5.105 by checking a list. It caught the word written that morning and had never heard of 67 British forms already in the tree.
-
An assumption nobody timed
The commit hook linted a narrower scope than CI, and widening it stayed on the backlog because it would 'roughly double the wall clock'. Measured: 245 ms against 517 ms — and the middle option nobody chose would have taken 40 seconds.
-
Three tests failed on macOS and the cleaner was innocent
A GNU-only flag in a test fixture silently did nothing on macOS. The code under test behaved correctly, three assertions about it failed anyway, and every message pointed at the wrong file.
-
A popup that cut off the only key that closes it
A 60-column constant against a 66-column hint dropped the words Esc cancel off the right edge, silently. Writing the general gate found three more dialogs with the same defect and two crashes, one of them at 66x12.
-
The process that died writing its own log
CI ran out of disk, and the thing that could not write was the logger. There was no readable job log to diagnose it from, and the cause was four copies of one rule that had drifted apart.
-
A tool list that advertised what the build could not run
One feature combination shipped an MCP server listing two tools, with full schemas, that refused every call. The report the refusal told operators to consult had never named the feature it was about.
-
Three surfaces, three different calls
The homepage still, the animation it swaps to, and the sample /analyze offers were three separate captures. Every asset was valid, every gate was green, and no diff put the three facts side by side.
-
Two ways a scan reports nothing
A regex sweep declared the tree clean and missed fourteen spellings. The gate that replaced it reported twenty passing tests while its own file was invisible to it. Neither failure looks any different from a pass.
-
Three quarters of a test binary was debug info
A CI linker ran out of disk after two test files landed. Measuring before changing anything found 315 MB of DWARF in a 422 MB binary, and a profile setting the release build had used for years.
-
The test measured the container, not the field it named
A macOS job ran red for four commits over a 1024-byte limit. The cap held at 327 bytes on both platforms, and the string the assertion measured was the JSON around the field.
-
A release is not done when the binaries are
0.5.128 published twenty-three assets, every workflow went green, and the website went on offering 0.5.127 to everyone who visited. A gate permitted the gap, and it was working exactly as designed.
-
Four scripts, one blind spot, and a fence that was never closed
A markdown fence is three or more backticks, and only a run at least as long closes it. Four scripts in this repository did not know that, and one of them was a gate quietly checking less than it claimed.
-
sipnab recorded every successful call as a failure
sipnab typed its vCon dialog objects "incomplete", a value the draft reserves for calls that never reached conversation. The reasoning behind it was local, careful, and wrong in a way no test could see.
-
A 27% regression hid behind a gate built to catch it
sipnab shipped a quarter of its offline throughput away for ten releases. The benchmark gate ran on every one of them and passed. Here is why, and what the fix says about ratchets.
-
The INVITE that arrived before its keys
A TLS capture that decrypted everything except the one message that mattered, and reported a NAT problem instead. Three bugs, none of them the race everyone assumed.