NADIR / Technology / 03 of 04

From raw disagreement to a tier you can act on.

The path from a stream of sensor frames to a single word an operator can dispatch against — with every step chosen so the number it produces still means something under audit.

Signal

The residual.

A residual is the difference between what one sensor claims and what the others imply. NADIR normalises those differences into a single cross-sensor sigma per vehicle per window, then plots it against days since the repair order closed.

The shape matters more than any single value. For seventy days this vehicle sits near 0.7 sigma with ordinary scatter, which is a healthy sensor suite doing what sensor suites do. Then the trend lifts, breaches the 2.0 sigma tolerance around day seventy, and keeps climbing. No individual frame in that record would have triggered anything. The population did.

Anchoring the horizontal axis to the repair order rather than the calendar is what makes drift attributable. Drift that begins at day zero is a repair problem. Drift that begins at day sixty is a road problem, and the two deserve different responses.

Cross-sensor residual since repair Seventy days of healthy scatter near 0.7σ, then a trend that breaches the 2.0σ tolerance at day seventy and keeps climbing.

Thresholds

A threshold that survives being wrong.

Where you put the threshold decides what the system is worth. Set it tight and every pothole becomes a work order. Set it loose and the drift you exist to catch arrives as an incident report instead.

The residual distribution is close to half-normal, so a flag at 2.4 sigma leaves 2.9 percent of samples beyond the line — a rate a maintenance team can actually absorb. But a per-axis threshold has a specific blind spot. Yaw and range residuals correlate at 0.58 on this fleet, so a vehicle can sit inside the limit on each axis independently while sitting well outside the joint distribution. Mahalanobis distance uses the covariance rather than the margins and catches the corner cases that per-axis limits wave through: eight of 260 here, every one of which two independent thresholds would have passed.

Residual distribution A half-normal fit with the flag at 2.4σ. 2.9% of samples fall beyond the line — a recheck rate a bay can absorb.
Joint gate, not per-axis Yaw and range residuals correlate at 0.58. Eight of 260 sit beyond the 3σ ellipse while passing both axis limits independently.

Confidence

Persistence, and a bound rather than a guess.

One bad window is noise. A Kalman filter over the yaw state separates the two: the filtered track absorbs measurement noise, and the innovation sequence — the part of each measurement the filter did not expect — stays flat while the sensor behaves. When a genuine yaw-rate step arrives, the innovations spike and stay spiked. Ten consecutive excursions is a state change, not a bump in the road.

Point estimates are not enough for evidence. A number handed to an insurer needs a bound with a stated coverage. Split-conformal prediction supplies one without assuming the residuals are Gaussian: fit on one half of the data, calibrate the interval on the other, and the resulting band covers the target fraction of future observations by construction. When empirical coverage runs at 86.9 percent against a 90 percent target, that shortfall is itself reportable — the model is telling you it is less certain than advertised, which is far more useful than a confident number with nothing behind it.

Innovation sequence Filtered yaw state with innovations below. Flat through noise, then ten consecutive excursions at a real yaw-rate step.
Split-conformal coverage A 90% band over yaw residual. Empirical coverage of 86.9% against a 90% target is itself a reportable number.

Output

How good is good enough, and what comes out.

At the shadow operating point this detector runs 82.9 percent recall at an 8.1 percent false positive rate, AUC 0.952. Precision is 40.4 percent, and that is a deliberate choice. At 6.2 percent prevalence a missed post-repair drift is an incident and a false positive is a recheck costing a bay hour. Shadow mode should buy recall with precision, and the operating point is placed accordingly.

What the fleet sees is not a sigma. It is a tier — NOMINAL, CAUTION or CRITICAL — with week-over-week movement attached, because the delta is what tells an operator whether last week's dispatching actually worked. And the pilot itself is scored on one number: MTTFF-RE, the median hours from repair-order close to first flag. Median 76 hours against a 168 hour target, 92.1 percent inside. One metric, pass or fail, agreed before the pilot starts.

Detection ROC 82.9% recall at 8.1% FPR, AUC 0.952. Recall bought with precision, deliberately, because a miss costs more than a recheck.
Fleet tier distribution Tier counts with week-over-week deltas. The movement is what tells an operator whether last week's dispatching worked.
MTTFF-RE Hours from repair-order close to first flag. Median 76 h against a 168 h target, 92.1% inside. The pilot passes or fails here.

Every threshold, model version and coverage figure on this page is written into the evidence bundle alongside the score itself.

A number without the method that produced it is not evidence, and it will not survive the first serious question an auditor asks.

Want this pointed at your fleet?

We are onboarding fleets and repair networks a few at a time. Tell us what you run and we will reach out.

Early access

Join the waitlist

We are onboarding fleets and repair networks a few at a time. Tell us who you are and we will reach out.

No spam. We reply from founders@nadirai.net.