Can a 2B Local Model Follow a Bearing to Failure?

#Arc#local AI#condition monitoring#predictive maintenance#vibration#industrial IoT#Jev-style#NASA IMS
Cover image for Can a 2B Local Model Follow a Bearing to Failure?

A bearing does not fail in one moment. Its vibration changes over hours and days: an impact appears, the spectrum develops structure, RMS rises, and eventually the machine reaches a point where continuing to run is no longer worth the risk.

That makes bearing degradation a better test for a decision model than a handful of isolated alerts. The question is not simply whether the model can name one event. It is whether its judgment moves in the right direction as the physical evidence gets worse.

We tested that question with a 2B local model and a real run-to-failure dataset already stored in Arc.

Arc reduced more than 20 million time-aligned vibration rows into 984 snapshots. We selected 96 points across the bearing's life and gave the model a compact state for each one. The model did not see the timestamp, snapshot number, or time remaining before the test rig stopped.

Its exact maintenance label matched our evaluation policy in 65.6% of the windows. That is not a result we would use to automate maintenance.

The probability-weighted risk score was much more interesting. It had a 0.92 rank correlation with the known failure stage and ordered 96.2% of pairs from different stages correctly. As the bearing degraded, the score generally rose with it.

This is not a story about replacing vibration analysis or maintenance engineers. It is a story about a small local model turning evidence prepared by Arc into a useful ranking signal, and about where that signal may fit in a real condition-monitoring workflow.

The bearing run

The source is NASA's IMS bearing dataset. Four Rexnord bearings ran on a shaft at 2,000 rpm under a 6,000-pound radial load. Accelerometers recorded one second of vibration at 20 kHz every ten minutes. After almost seven days, bearing 1 developed an outer-race failure and the rig stopped.

Our earlier article, Keep the Waveform. A Bearing Fails in Arc, One Sample at a Time., explains how we imported the complete run, how the interactive waveform and spectrum work, and why retaining raw vibration matters.

For this experiment, we used the same live Arc dataset:

Dataset propertyValue
Time-aligned rows in Arc20,152,320
Bearing channels per row4
Snapshots984
Samples per snapshot20,480
Time between snapshots10 minutes
Total run163.8 hours
Known outcomeOuter-race failure on bearing 1

You can explore the run in the interactive demo. Press play to follow the waveforms, pause to inspect the raw samples and spectrum, or jump to the first point where bearing 1 exceeds twice its day-one RMS baseline.

From raw vibration to model evidence

We did not send 20,480 waveform samples to the decision model. Arc first aggregated the entire run into one row per ten-minute snapshot:

arcli query --database bearing_monitoring -o json \
  "SELECT
     time_bucket(INTERVAL 10 MINUTE, time) AS bucket,
     min(snapshot) AS snapshot,
     sqrt(avg(b1 * b1)) AS rms_b1,
     max(abs(b1)) AS peak_b1,
     kurtosis(b1) AS kurtosis_b1,
     sqrt(avg(b2 * b2)) AS rms_b2,
     sqrt(avg(b3 * b3)) AS rms_b3,
     sqrt(avg(b4 * b4)) AS rms_b4
   FROM vibration
   WHERE test_id = 'ims-2'
   GROUP BY bucket
   ORDER BY bucket"

The query scanned the stored run and returned all 984 snapshots in 156 milliseconds on the demos instance.

For each point we wanted to evaluate, we built a compact evidence object containing:

  • bearing 1's current RMS, peak and kurtosis against its day-one baseline;
  • the preceding two hours, split into recent and prior one-hour windows;
  • the change between those two windows;
  • the count of recent snapshots above twice the RMS baseline;
  • the current RMS ratios of bearings 2, 3 and 4 as controls.

The peer bearings matter. If all four channels move together, the cause may be load, speed, mounting, or another shared condition. If one bearing separates from the others, the evidence for a local mechanical problem is stronger.

Raw vibration is stored in Arc, SQL turns it into a compact evidence object, a local 2B decision model scores four ordered risk levels, and policy plus a technician decide what happens next.

This separation is deliberate. Arc owns the raw history and the reproducible feature calculation. The model receives a small, inspectable state. A policy layer and a person retain control of what happens to the machine.

The local decision model

We used Jev-Style-Qwen3.5-2B-Decision, an independent Apache-2.0 checkpoint based on Qwen3.5-2B, loaded locally as jev-style-qwen3.5-2b-decision.

The model follows the Jev-style decision pattern: present a bounded question and a set of criteria, then return a probability for each option. It is not TypeSafe AI's hosted Jev and is not affiliated with or endorsed by TypeSafe.

We asked one question:

Rate the mechanical risk shown by the observed vibration evidence on this ordered scale.

The four options were:

  1. Healthy baseline: vibration is close to baseline and stable.
  2. Emerging deviation: vibration is materially different and deserves closer observation.
  3. Persistent degradation: evidence supports planning maintenance soon.
  4. Severe mechanical risk: vibration is extreme or accelerating enough to justify an immediate stop and inspection.

These were bounded interpretations of the evidence, not permissions to stop equipment.

Building a 96-window replay

Because the rig's final stopping time is known, we divided the run into four evaluation stages:

Evaluation stageTime before the rig stoppedWindows sampled
Continue routine monitoringMore than 48 hours24
Increase monitoring24 to 48 hours24
Schedule maintenance4 to 24 hours24
Stop and inspectLess than 4 hours24

We selected 24 evenly spaced snapshots from each stage. The model saw the vibration evidence for each point, but we deliberately withheld the snapshot number, timestamp, stage, and remaining time.

The stages are an evaluation policy, not a universal maintenance standard. A real threshold depends on the asset, operating state, failure mode, redundancy, safety consequences, and the time required to plan an intervention. Here, the stages give us a known ordering against which to test the model's signal.

Why we kept the probabilities

If we take only the model's most likely option, we force every point into one of four boxes. Two adjacent stages can be nearly tied but appear as a hard disagreement.

Instead, we also calculated the expected value of the probability distribution:

risk score = Σ(level index × probability of that level)

The levels are indexed from 0 to 3. A score near 0 places most probability around the healthy baseline. A score near 3 places it around severe mechanical risk.

This does not create a new physical unit. It gives us a continuous summary of the model's ordered judgment.

The exact labels were uneven

The model's top option matched the evaluation stage in 63 of 96 windows, or 65.6%.

Performance varied sharply by stage:

Evaluation stageMean RMS vs. baselineMean model risk, 0–3Exact top-option accuracy
Continue routine monitoring1.06×0.8975.0%
Increase monitoring1.67×1.7425.0%
Schedule maintenance2.15×1.9083.3%
Stop and inspect4.82×2.1479.2%

The “increase monitoring” stage was the weak point. Many of its windows landed on an adjacent level. That is exactly the kind of result a top-label metric punishes, even when the probability distribution has shifted in the correct direction.

We would not deploy this as a four-class maintenance classifier.

The risk ordering was much stronger

The continuous risk score followed the run more consistently than the exact labels:

Across 96 real vibration windows, mean model risk increased from 0.89 during routine monitoring to 2.14 in the final four hours. The score had a 0.92 Spearman correlation with failure stage and ordered 96.2 percent of cross-stage pairs correctly.

Across all 96 windows:

  • Spearman correlation between model risk and failure stage was 0.92;
  • correlation between model risk and bearing 1's RMS ratio was 0.96;
  • 96.2% of pairs drawn from different stages were ordered in the expected direction;
  • each local decision took 0.49 seconds on average.

The pair result means that when we selected one point from an earlier stage and one from a later stage, the later point received the higher risk score 96.2% of the time.

That is useful for ranking. A fleet view does not always need a perfect diagnosis for every asset. It may need to identify which five machines deserve attention first, which sensor should temporarily sample more often, or where a technician should open the waveform and spectrum.

Why not use an RMS threshold?

For this single run, a deterministic RMS rule is the first baseline to beat. It is cheaper, transparent, and already finds the point where bearing 1 exceeds twice its day-one level. The model's 0.96 correlation with RMS is not evidence that it replaced vibration engineering. It shows that the model reacted strongly to a feature we explicitly provided.

The model becomes interesting when the decision depends on several imperfect signals together: RMS level, trend, peak, kurtosis, persistence, peer sensors, operating context, and asset history. A probability distribution can combine that evidence into a ranking while still exposing uncertainty to the surrounding application.

But “interesting” is not “better.” A production evaluation should compare the model against simple rules and classical condition-monitoring methods on multiple assets and multiple failure runs. If the simpler method performs as well, use the simpler method.

What the experiment actually shows

The model did not discover a fault from raw vibration. Arc and SQL calculated the features, and the model received those values directly. Its contribution was mapping a compact set of measurements into a smooth, ordered judgment.

The evaluation also has important limits:

  • all 96 windows come from one bearing in one run;
  • nearby windows are correlated, so they are not 96 independent failures;
  • the time bands were defined by us after the run was complete;
  • there was no second failure run reserved as a holdout set;
  • the experiment measured ordering, not whether maintenance at a given point would have improved the outcome.

The result is strong enough to motivate a larger evaluation. It is not strong enough to establish a maintenance policy.

Where this could fit

The first useful deployment would be local, read-only, and in shadow mode.

Arc would continue storing the raw waveform and computing reproducible features. Existing rules would continue to own hard limits. The model would record its probability distribution and risk score without controlling equipment. Technicians could then compare its ranking with inspections, work orders, and eventual outcomes.

If it proves useful across assets, the score could support a few bounded tasks:

  • rank machines for technician review;
  • increase sampling frequency when evidence starts to change;
  • choose which diagnostic query or visualization to open next;
  • add context to an alert without turning that context into an automatic shutdown;
  • run at the plant edge when vibration data cannot leave the site.

The action boundary remains conventional code and human judgment. A local model can suggest where to look. It should not decide by itself when a machine stops.

What comes next

The more interesting next step is not a larger model. It is more failures: different bearings, loads, speeds, fault modes and operating conditions, evaluated against deterministic baselines and held-out runs.

One bearing gave us a promising signal. A fleet would tell us whether it generalizes.

Build on your own data

Arc keeps raw, high-frequency telemetry queryable, so you can revisit old events, calculate new features, and evaluate new models without rebuilding the history first.

Download Arc and start with the data your machines are already producing.

Ready to handle billion-record workloads?

Deploy Arc in minutes. Own your data in open files on your storage. Use for analytics, observability, AI, IoT, or data warehousing.

Get Started ->