SalvaraSonarTier-2 telemetry intelligence

What your fleet is actually doing.

Explained findings — each one carrying the measured signals behind it, the rule that fired, and what it costs. Not another metrics dashboard.

  • 8

    Findings

    computed

    the number of claim objects the rule engine produced from this capture

  • 5

    Correlation chains

    declared

    DOSSIER §4 chain table — five chains marked REAL. A property of the dossier, not computed from data/.

  • 3

    Real datasets

    computed

    distinct dataset families cited by the evidence: Cisco IOS-XR fabric · Microsoft Philly cluster · CINECA Marconi100

  • 0

    False positives / ~20 quiet days

    0 computed · ~20-day span declared

    The 0 is computed: INS-THM-001 check.quiet_period_false_positives — R6 detections further than 24 h from any Nagios-CRITICAL sample. The "~20 quiet days" span is declared: DOSSIER §5 R6 / FINDINGS F1 — the 3-week slice outside the event window. The 0 beside it is computed (INS-THM-001 quiet_period_false_positives).

2 computed from the insight set · 1 declared from the dossier · 1 both. Open any tile for its source — a number that will not say where it came from is exactly what this product argues against.

CaptureCisco fabric: 2017-09-01 17:08–18:09 UTCPhilly cluster: 2017-10-10 → 2017-10-16 PDTM100 rack 0: 2022-06-21 → 2022-07-11 & May 2021 UTCslices prepared: 2026-08-15captures/2026-08-15_public-sample
Measuredobserved directly in the dataInferredconcluded from co-moving signalsdollar figures assume $0.90/GPU-h · $16.5/kW-month — editable

Findings

Editorial impact order — a reading order, not an exchange rate. Dollar figures sit on different bases (per week, per month, over a job residency), so the list is not sorted by size; maturity is a badge on each card rather than the grouping.

  1. inferred#1 · INS-THM-001capture constant2022-07-04T18:00:00Z

    Site-level cooling/power interruption in rack 0: nodes de-energized to PSU standby, fans stopped, and rack ambient rose +14.8 C above its 7-day baseline and held near 37.6 C

    Co-moving evidence:

    • ambient_avgrack0-node16: 22.8 C baseline (7-day trailing median) -> 34.4 C within 1 h of event start (2022-07-04T18:00:00Z), peak 37.6 C through 2022-07-05T06:30:00Z; rack0-node8 peak 30.2 C; rack0-node9 peak 36.7 C
    • fan0_0..fan3_1 (RPM)all eight fan tachs read 0 RPM on rack0-node8 (51/70 reporting samples in the event window), rack0-node9 (4/20 reporting samples in the event window), rack0-node16 (51/70 reporting samples in the event window) against a per-tach baseline of ~4340-4700 RPM (n=44,024 readings outside the event) — cooling stopped, co-moving with the ambient rise. 0 RPM is a reading; NaN is absence, and NaN samples are excluded rather than read as zero
    • ps0_input_power_avg / total_power_avgrack0-node16 PSU input drops to 10 W standby and node total power reads NaN in 51 samples during the event (node de-energized); ~379 W median idle before and after

    106/107 R6 detections fall inside Nagios CRITICAL windows; 0 detections in the quiet days of the slice; the 1 remaining detection precedes the first CRITICAL label by 15 min on this event (label lag, not a false positive; n=1, not a general capability claim)

    nothing priced on this card

    A cooling/power interruption spanning 17.25 h across a 16-node rack; on this event, the rule tripped 15 minutes before the facility's first Nagios CRITICAL label (n=1 — the defensible claim is 106/107 label agreement and 0 quiet-day false positives)

    rack0-node16 · ambient 22.8 → 37.6 °C · all fans → 0 RPM
    • ambient 20.5–39.4 °C
    • fan median 0–4,889 RPM
  2. inferred#2 · INS-GPU-0012017-10-10T00:00:00-07:00

    Zombie allocation: GPU m63/gpu4 is held by a job in Failed state and has done zero measured work for the entire observation window

    Co-moving evidence:

    • gpu4_utilm63 gpu4: 10,080 consecutive per-minute samples 2017-10-10..2017-10-16, mean 0.0%, max 0.0%
    • job allocationapplication_1504131676014_7699 (status Failed) holds m63/gpu4 from 2017-09-14 22:55 to 2017-10-26 23:40 PDT — 1,008.7 hours
    • gpu1_util/gpu6_util (contrast)other persistently-allocated GPUs on the same box run 77-78% mean over the same window — the box is alive; this slot is not
    $908total over the held period · moves with ASSUMED_gpu_hour_rate

    ~1,009 GPU-hours of schedulable capacity stranded by one dead slot (~$908 at $0.90/GPU-h equivalent); 1.0 idle-GPU-equivalent removed from the fleet as long as it persists

    m63/gpu4 held all window · util 0.0% · gpu1 78.3%
    • sibling GPU 0–100% util
    • gpu4 util 0–100% util
  3. inferred#3 · INS-GPU-0032017-10-10T00:00:00-07:00

    Fully-allocated 8-GPU node m230 shows the same allocation/realization gap (34.5% measured) for the entire window under a single long-running job

    Co-moving evidence:

    • gpu0..7_utilm230 all-GPU mean 34.5% over 2017-10-10..2017-10-16 (per-GPU means 33.0-38.8%)
    • job allocationapplication_1506638472019_5808 (Pass) holds all 8 GPUs of m230 from 2017-10-04 02:11 to 2017-10-29 17:11 PDT (615 h)
    • control: m123 same cluster, same windowm123 runs 91.0% all-GPU mean — ~90% sustained is achievable on this hardware in this cluster
    $2,749total over the job residency · moves with ASSUMED_gpu_hour_rate

    ~5.0 idle-GPU-equivalents; over this job's full 615 h residency that is ~3,050 stranded GPU-hours (~$2,750 at $0.90/GPU-h)

    m230 34.5% vs control m123 91% — same window, same cluster
    • control m123 0–100% util
    • m230 all-GPU mean 0–100% util
  4. inferred#4 · INS-GPU-0022017-10-10T00:00:00-07:00

    Fully-allocated 8-GPU node m139 is realizing about a third of its capacity: allocation says busy, measured utilization says mostly idle

    Co-moving evidence:

    • gpu0..7_utilm139 all-GPU mean 32.3% over 2017-10-10..2017-10-16 (per-GPU means 31.0-35.1%)
    • job allocationapplication_1506638472019_19223 (Pass) holds all 8 GPUs of m139 from 2017-10-10 08:00 to 2017-10-22 11:53 PDT
    • control: m123 same cluster, same windowm123 runs 91.0% all-GPU mean — ~90% sustained is achievable on this hardware in this cluster
    $780per week · moves with ASSUMED_gpu_hour_rate

    ~5.2 idle-GPU-equivalents on one node: ~867 stranded GPU-hours per week (~$780/week at $0.90/GPU-h) while the scheduler reports the node fully busy

    m139 32.3% vs control m123 91% — same window, same cluster
    • control m123 0–100% util
    • m139 all-GPU mean 0–100% util
  5. measured#5 · INS-PWR-0012021-05-24T10:30:00Z

    Demand-charge exposure: rack 0's May 2021 billing peak was set by a ~6.5 h burst — the rack ran at or below 6.58 kW for 95% of the month but its demand peak hit 14.23 kW

    Co-moving evidence:

    • rack_total_power_w (sum of 16 nodes' total_power_avg)2,976 15-min samples, May 2021: p50 6.52 kW, p95 6.58 kW, peak 14.23 kW at 2021-05-24 10:30:00 UTC
    • per-node total_power_avg at the peak12 of 16 nodes at ~1.0-1.1 kW simultaneously (real burst compute), 4 nodes at idle ~372-454 W
    • time above p9537.25 h total above p95 in the month; longest contiguous run 6.5 h
    $126per month, per rack · moves with ASSUMED_demand_rate

    7.65 kW of demand-charge exposure on one rack (~$126/month at ASSUMED $16.5/kW-month): utilities bill the peak even though it was sustained for hours, not the month

    rack peak 14.23 kW vs p95 6.58 kW — reduced by maximum, so the peak survives
    • p95 0–15.4 kW
    • rack kW 0–15.4 kW
  6. measured#6 · INS-NET-0032017-09-01T17:08:57Z

    Fault vs planned-action classification across all 6 labelled events: the state-class rule separates physical faults from admin actions with 6/6 class agreement and zero false positives in this capture

    Co-moving evidence:

    • actual_line_state transitions (up->down only)exactly 6 down-events detected for 6 labelled events; 4 show plain 'down' (physical class); 2 show 'admindown' on the acting end
    • both-end confirmation coverage4 events are both-end confirmed (far-end lag <=2 s, detection within 46 s of the hand-logged label). 2 events are single-end only — the far end does not stream MDT in this capture, so label-time offsets run up to 203 s (the case file logs to the minute)
    • ground-truth labels6/6 class agreement with the injected-event case file

    full case file, all six events; both-end vs single-end confirmation split surfaced per event

    nothing priced on this card

    6 detections, 6 labelled events, 6/6 correct class, 0 false positives on real router telemetry — the credibility check for every inference above; single-end detections carry an explicit lower-confidence note

    6 labelled events (bands) · 6 blind detections (ticks) · 6/6 class agreement
    • leaf3 Hu0/0/0/30 ingress (context) 0–97.51 Gbps
  7. inferred#7 · INS-NET-0012017-09-01T17:14:56Z

    Physical-layer fault (transceiver/optic or fiber) on spine2 HundredGigE0/0/0/8 — not an administrative action

    Co-moving evidence:

    • actual_line_statespine2 Hu0/0/0/8 im-state-up -> im-state-down at 2017-09-01T17:14:56Z; state is 'down', not 'admindown'
    • actual_line_state (far end)leaf3 Hu0/0/0/8 down at 2017-09-01T17:14:58Z — both ends within 2 s, signature of light loss, not config
    • input_data_rate/output_data_ratespine2 Hu0/0/0/8 collapses from ~10.73/11.13 Gbps (5-min pre-event mean, n=55) toward zero (min ~39.5 Mbps) during the down window

    cisco_portflap_events.csv event 1: 'transceiver pull (local port) reinserted after convergence', spine2, labelled 17:15-17:19 UTC — the rule fired blind on telemetry (-4 s vs the hand-logged label) and matches the injected fault

    2,827.7 GbGb not carried · nothing priced on this card

    ~2,828 Gb of transport capacity not carried during the 159 s outage; traffic did not return to 90% of baseline until 242 s after the fault

    spine2 Hu0/0/0/8 · 10.73 → 2.02 Gbps across a 159 s outage
    • ingress rate 0–10.89 Gbps
  8. inferred#8 · INS-NET-0022017-09-01T17:27:59Z

    Planned administrative shutdown on spine2 HundredGigE0/0/0/8 — no hardware fault indicated, no RMA action implied

    Co-moving evidence:

    • actual_line_statespine2 Hu0/0/0/8 -> im-state-admindown at 2017-09-01T17:27:59Z (explicit admin state, unlike INS-NET-001)
    • actual_line_state (far end)leaf3 Hu0/0/0/8 shows im-state-down (it sees carrier loss; only the local end knows it was administrative)

    cisco_portflap_events.csv event 3: 'admin down (local port) then admin up', spine2, labelled 17:28-17:33 UTC — matches

    nothing priced on this card

    Same traffic impact class as a physical fault, but zero hardware-replacement cost — separating these avoids false RMA/dispatch

    spine2 Hu0/0/0/8 · admindown on one end — same rate impact, different state class
    • ingress rate 0–11.02 Gbps

Assumptions — separated from what’s measured, and editable

Every dollar figure names its assumptions apart from its measured inputs. Change one and the findings recompute.

ASSUMED_gpu_hour_rate

on-prem amortized equivalent — a parameter, not a measurement

ASSUMED_demand_rate

typical US commercial demand charges run ~$15–20/kW-month

measured inputs (held-hours, kW exposure, utilization) are fixed — only the assumed rates move

Shown for honesty — not counted among the findings

Illustrative — not real telemetryILL-GPU-001 · not ranked · not counted · not computed from the capture

Utilization-that-isn’t-work: GPU reports ~98% utilization while SM-active is ~12% and NVLink is near saturation — the card is interconnect-bound, burning power while doing little compute

No public dataset in the selected inventory carries DCGM_FI_PROF_SM_ACTIVE together with NVLink/PCIe throughput. The chain is real and its fields are standard DCGM fields (see shape catalog), but these numbers are invented to show the card format. It must render visually distinct from real insights and never enter the real insights list.

  • DCGM_FI_DEV_GPU_UTIL ILLUSTRATIVE
  • DCGM_FI_PROF_SM_ACTIVE ILLUSTRATIVE
  • DCGM_FI_PROF_NVLINK_TX_BYTES + _RX_BYTES ILLUSTRATIVE

No computed values: nothing was computed. There is no drill-down behind this card because there is no telemetry behind it.

Not shown — chains the public data can’t support

Named so nobody has to wonder whether this is the whole set

These need fields no dataset in this capture exposes. They are not failures of the method and they are not findings — on a live feed carrying the fields below, they return unchanged.

  • Thermal-throttle cascade — needs throttle-reason bitmask + clock frequency + throughput
  • Silent power-cap — needs power-cap events
  • ECC failure prediction — needs ECC correctable counters + row-remap counters
  • RoCE PFC training stall — needs PFC pause counters + ECN counters

Capture data notes — flagged, not fixed

Flagged during precompute. Capture files are read-only and anchored; nothing here was edited.

  • philly_machines.csv declares num_gpus=2 for m123, m139, m230, m63, but philly_gpu_util_slice.csv reports 8 reporting GPU columns and philly_job_allocations.csv allocates gpu0..gpu7 on each. DOSSIER §3.7 describes these machines as 8-GPU, which the utilization and job data agree with. Measured width is used everywhere; the declared column is ignored. Flagged for the capture owner — capture files are never edited by the build.