TIERA Book demo ->
Technology blog
End-of-lineProduction testLimitsNVHQuality
Production test / 9 min read

Setting Pass/Fail Limits on an End-of-Line Test Without Guessing

The hard part of end-of-line testing is not measuring the unit — it is deciding where the line goes. Set it from a handful of good units and you will fail good product for a year. Set it from the loudest thing that ever shipped and you will pass the defect you built the station to catch.

01

Two different costs, and they pull in opposite directions

Every limit is a trade between two errors. A false reject throws away or reworks a unit that was fine: scrap cost, rework labour, and a line that slows down. A false accept ships a defective unit: warranty, recall, and reputational cost that dwarfs the other side.

Because the costs are so asymmetric, the instinct is to set the limit tight. That instinct is usually wrong, and specifically it is wrong when the limit is tight enough that operators start overriding it — which they will, because a station that fails 8% of good product becomes a station people work around. A limit nobody trusts protects nothing.

So the real target is not "catch everything". It is the tightest limit the process can live with while the false-reject rate stays low enough that the station keeps its authority.

02

You cannot set a limit from good units alone

The most common approach is to measure fifty good units, take mean plus three standard deviations, and call that the limit. It is easy, it is defensible-sounding, and it answers the wrong question.

Mean plus 3σ tells you where the good population ends. It tells you nothing about where the bad population starts. If the two overlap, no limit separates them and you need a different measurement, not a different threshold. If they are widely separated, 3σ is needlessly tight and you are rejecting good product for no diagnostic gain.

You need both distributions. That means deliberately measuring known-bad units — and the honest problem is that at the start of a programme you rarely have enough of them.

The practical routes are: pull genuine defects from warranty returns and field failures and run them through the station; deliberately build defective units at known severities (a bearing with a seeded defect, a deliberately mis-set preload, an out-of-tolerance gear); and where neither is possible, run provisionally wide and tighten as real failures accumulate. What you must not do is invent the bad distribution and then present the limit as if it were derived.

measured level → good units defective units limit false reject false accept Where the two overlap, no threshold separates them — you need a better feature, not a better number.
A limit is a cut through two distributions. Knowing only the left one tells you what you will reject, never what you will miss.
03

Measure the right thing before you argue about the number

If good and bad overlap on your chosen metric, changing the threshold only trades one error for the other. The way out is a better feature, and that is almost always a narrower one.

Overall RMS is the bluntest choice available. It sums everything, so a defect that adds a modest amount of energy in one narrow band gets diluted by all the energy that was legitimately there. It is the right metric for gross faults and the wrong one for anything subtle.

Band-limited levels are usually the fix. Put a band where the defect lives — gear mesh, bearing race frequency, the specific order your failure mode excites — and the good/bad separation often opens up dramatically, because the band excludes the variation that was masking it.

Order-based limits handle speed variation. If your test runs at a nominal speed with real spread, a fixed-Hz band drifts off the feature. Order tracking pins the band to the shaft.

It is worth being blunt about the sequence: choose the feature so the populations separate, then place the limit. Teams routinely do this backwards and spend months tuning a threshold on a metric that could never have worked.

04

The station has to be repeatable before the limit means anything

Before any limit is defensible, you have to know how much of the measured spread is the product and how much is the station. Measure the same unit repeatedly — remounting it each time, as production would — and look at the scatter. Then measure several units.

If station repeatability is a large fraction of the unit-to-unit spread, your limit is mostly measuring your fixture. The usual culprits are sensor mounting that varies between operators, a fixture that loads the unit differently each time, and speed or temperature that is not controlled between tests. Fix those first; they are cheaper than the limit arguments they cause.

Temperature deserves specific mention because it is so often missed: a gearbox tested cold and a gearbox tested after an hour of production run at different viscosities and different clearances, and the measured level moves with it. Either control the warm-up or make the limit temperature-aware, but do not pretend it does not matter.

05

Run it as a living limit

Start deliberately loose, with every result logged. A station that fails almost nothing in month one but records everything gives you the good distribution from real production rather than from a pilot batch, and it does not disrupt the line while you learn.

Tighten in reviewed steps as the data supports it, and every time a defect escapes to the field, run that unit's data back through the station and ask whether the limit would have caught it. That feedback loop is what turns an end-of-line test from a ritual into a control.

Keep the false-reject rate visible on the same dashboard as the false-accept escapes. The moment only one of them is reported, the limit will drift toward whichever error is being counted.

The kit for this job

TIERA instruments that do this work.

PhonoVibe O — 8-Channel IEPE DAQ

PhonoVibe O — 8-Channel IEPE DAQ

Enough simultaneous channels to instrument a unit properly and keep cycle time down.

Channels
8
Resolution
24-bit
TVIB — Sound & Vibration Analysis Software

TVIB — Sound & Vibration Analysis Software

Band-limited and order-based spectral alarms — the features that make good and bad actually separate.

Alarms
Spectral and band-limited
130F20 ICP® Electret Array Microphone

130F20 ICP® Electret Array Microphone

When the customer complaint is noise, the pass/fail has to be measured in sound, not only vibration.

Sensitivity
45 mV/Pa

From the TIERA store

The kit for this job

What we would actually put in front of someone doing the measurement this post describes — not the whole catalogue.

Use cases

Where this shows up in the field

From TIERA

A limit you can defend in a quality audit.

We help OEMs build end-of-line stations where the feature is chosen before the threshold, the station's own repeatability is quantified, and the limit has a documented basis rather than a remembered one.

If you have overlapping populations today, the fix is nearly always a narrower feature — and that is a conversation worth having before you buy anything.

  • Multi-channel simultaneous acquisition sized to your cycle time
  • Band-limited and order-based alarm configuration in TVIB
  • Seeded-defect units on TMFSS to populate the bad distribution
Learn this properly

Where this sits on the TIERA learning ladder.

The theory behind this article is covered free, in full, by the TIERA 101 primers: AI Condition Monitoring 101, Signal Processing 101. They are self-paced, interactive, and end in an exam and a certificate.

The primer covers alarms and severity. Designing a production limit against two measured distributions, with the cost asymmetry made explicit, is Cat II/III programme work.

TIERA 101 is a free introductory primer, not an accredited ISO certification, and its hours do not count towards the formal training ISO 18436 requires.