Bearing Run to Failure

Six days of vibration, every sample kept.

Four bearings on a loaded shaft, accelerometers sampled at 20 kHz, one second recorded every ten minutes until bearing 1 fails. All 20 million samples per channel are rows in Arc.

Press play to step through the run. Pause anywhere to pull the raw samples and see the outer race defect frequency rise out of the noise. The query panel shows the SQL behind the trend, the onset and every waveform.

-- aggregating 20 million samples in Arc...
What you just watched

One shaft. 20,152,320 samples per bearing.

A vibration sensor at 20 kHz produces more data in one second than a temperature sensor does in a day. The usual answer is to keep an RMS value per minute and throw the waveform away. Then the one question that matters, what the spectrum looked like the hour before the failure, has no data behind it.

Samples per channel
20,152,320
rows in bearing_monitoring.vibration, four channels each
Snapshots
984
one second every 10 minutes, 0 files rejected
Sampling rate
20,000 Hz
50 microseconds between rows
Run length
163.8 h
2004-02-12 to 2004-02-19, rig local time
Trend query
511 ms
Arc execution time to aggregate every sample into 984 rows of RMS, peak and kurtosis, measured on this page load
Failure onset
117.0 h
bearing 1 first exceeds 2x its day-one baseline (2.15x), 46.9 h before the end of the run
Outer race frequency
236.4 Hz
Rexnord ZA-2115, 16 rollers, 2000 rpm
Imported
2026-09-22
by demos/bearing-monitoring/ingest.py

Raw samples in, aggregates on demand

Each row is one sample: a microsecond timestamp, the snapshot and sample ordinals, and one column per accelerometer. The trend chart is a single time_bucket query over every row, computing RMS, peak and kurtosis per ten-minute snapshot. The waveforms during playback are a per-millisecond min/max envelope, also from Arc. Pause and the browser fetches the 20,480 raw samples of one channel and computes the spectrum itself, labelled as such.

How the time axis works

The file names carry the rig's local clock; sample i of a snapshot is that time plus 50 microseconds times i, stored in UTC. Nothing is resampled. The last two snapshots were recorded after the rig had already been stopped and are kept as they are; the trend shows them dropping to zero.

The onset query
WITH s AS (
  SELECT time_bucket(INTERVAL 10 MINUTE, time) AS bucket,
         min(snapshot) AS snapshot,
         sqrt(avg(b1 * b1)) AS rms
  FROM vibration
  WHERE test_id = 'ims-2'
  GROUP BY bucket
),
base AS (
  SELECT avg(rms) AS baseline
  FROM s
  WHERE bucket < (SELECT min(bucket) FROM s) + INTERVAL 1 DAY
)
SELECT s.snapshot, epoch_ms(s.bucket) AS bucket_ms, s.rms, base.baseline, s.rms / base.baseline AS ratio
FROM s, base
WHERE s.rms > 2 * base.baseline
ORDER BY s.bucket
LIMIT 1

Vibration data: IMS Bearing Data, generated by the NSF I/UCR Center for Intelligent Maintenance Systems, University of Cincinnati, with support from Rexnord Corp., and distributed through the NASA Prognostics Center of Excellence data repository. Test set 2, February 12 to 19, 2004. Reference: Qiu, Lee and Lin, “Wavelet filter-based weak signature detection method and its application on rolling element bearing prognostics”, Journal of Sound and Vibration 289 (2006).

The set was imported once by a one-shot importer that fetches the NASA archive, validates every one-second file, and writes one row per 50 microsecond sample into Arc. Every figure on this page is read back from that import. Defect frequencies are computed from the published bearing geometry at 2000 rpm and are markers, not measurements.

Arc is an open, SQL-native time-series database for high-volume telemetry. A vibration sensor at 20 kHz is the workload that pushes most databases into keeping only aggregates. Arc keeps the samples and computes the aggregates when you ask.