Skip to content
Mohammad Al-Araidah

Case study 01 · Manufacturing / Quality

High-Volume Manufacturing Scrap Reduction

Recurring scrap on high-volume production lines that could not be resolved by operator adjustment. Structured root-cause analysis and statistical process control cut scrap 15% and the gain was written into the control plan so it held.

Project record

Professional
Project type
Professional project
Domain
High-volume manufacturing
Role
Manufacturing / Industrial Engineering Manager
Methods
Defect Pareto · Measurement system review · Fishbone / 5-Why · SPC · PFMEA and control plan
Date
2024–2026
Data
Production data from the plant's MES and shop-floor reporting; not reproduced here
Confidentiality
Method abstracted. No employer drawings, part numbers, process parameters or records are disclosed.

Context

A high-volume production area running multiple parallel lines, with automated process steps, in-line inspection, and MES-based defect reporting by station and defect code. Scrap was reported daily and reviewed weekly.

The area was not in crisis. That is actually the harder version of this problem: scrap sat at a level everyone had learned to plan around. It appeared in the budget as a fixed rate rather than as a loss with a cause, and the standing explanation was operator technique and material variation — the two explanations that are least actionable and least falsifiable.

Problem

Scrap performance showed recurring defect modes that could not be resolved through operator adjustment alone. Corrective actions had been attempted before and had produced short-lived improvements that decayed within weeks — the signature of containment being mistaken for corrective action.

Stated as a question the work could actually answer: are these defects the output of a stable process running at its natural capability, or the output of an unstable process with assignable causes that have never been isolated?

Constraints

  • Production could not be stopped for experimentation; analysis had to run on live production data and short, scheduled trials.
  • No capital was available at the outset, so the first round of solutions had to be process and standard-work changes.
  • Defect coding was inherited and partially ambiguous — multiple codes described overlapping failure modes.
  • Any change had to survive across three shifts and a rotating operator population.

My role

I owned the improvement agenda for the area and led a nine-person team. I ran the analysis personally, set the trial plan with production leadership, and owned the PFMEA, control plans, work instructions and inspection criteria that the changes were written into. Maintenance, tooling and quality inspection each owned specific actions.

Approach

Four steps, in this order, because each one eliminates work in the next.
Analysis sequence
StepWhat was doneWhy it comes here
Defect ParetoRanked scrap by defect code and by station across a full quarter of production, in units and in dollars.Dollars and units rank differently. Ranking by dollars alone hid a high-frequency, low-cost mode that was driving most of the rework labor.
SegmentationSplit the top modes by line, shift, tool position, material lot and part family.Segmentation is the cheapest hypothesis test available. A defect that appears on one line and one tool position is not a material problem, whatever the standing explanation says.
Measurement system reviewChecked gauge capability and, more importantly, the consistency of defect coding between shifts and inspectors.Two codes were being used interchangeably for distinct failure modes. Until that was fixed, the Pareto was measuring reporting behaviour as much as process behaviour.
Root cause and validationFishbone to narrow the candidate set, 5-Why on the surviving candidates, then a designed check against process data and short production trials.A cause is not established because it is plausible. It is established when the process data moves the way the cause predicts.

Scroll table horizontally →

Analysis

Correcting the defect coding changed the answer. Once the two overlapping codes were separated, the recombined Pareto showed that the largest true contributor was not the mode the area had been working on. It was concentrated on specific tool positions and correlated with time since the last tooling intervention — which pointed away from operator technique and toward a process window that was drifting between maintenance events.

Control charting the associated process characteristic made the mechanism visible: the process was not unstable in the sense of random upsets. It was drifting predictably, spending progressively more time near the specification limit until it crossed. The mean looked acceptable in the weekly summary the whole time, which is exactly why a monthly average review had never detected it. Capability computed over the full interval was materially worse than capability computed just after an intervention.

That reframed the problem from “why do operators produce this defect” to “why is the process allowed to drift to the edge of the window before anything triggers.” The second question has engineering answers.

Risk

Process characteristic drifts toward the specification limit between scheduled tooling interventions; defect escapes are detected at end-of-line inspection rather than prevented at the source.

Severity
6
Occurrence
7
Detection
6
RPN · Priority
252 · High

Ratings are the PFMEA values the team assigned during the review, on the plant's own rating scale. Detection was the weakest of the three: the condition was visible in process data well before it produced scrap, but nothing was watching for it.

Bands: RPN ≥ 150 high · 60–149 medium · < 60 low

Engineering judgment

What decision had to be made?

Whether to treat the dominant defect mode as inherent process capability — and therefore budget for it — or as an assignable cause worth engineering out. Choosing the first would have been defensible on the evidence as originally reported, and would have been wrong.

What evidence mattered?

Three things mattered. First, the corrected defect coding, without which the Pareto ranked the wrong mode. Second, the concentration by tool position, which is incompatible with a material-variation explanation. Third, the control chart showing systematic drift rather than random variation — the distinction that determines whether a problem is a capability problem or a control problem.

What would cause me to change the decision?

If the defect had been distributed evenly across tool positions, lines and shifts, and the control chart had shown a stable process with the specification limit simply inside the natural variation, then this is a capability problem and the honest answer is different: either widen the tolerance with engineering approval, or invest in process capability. Continuing to hunt for an assignable cause in a stable process is how teams spend a year confirming nothing.

Implementation

  • Defect codes separated and the coding standard rewritten, with inspector re-training so the data stayed trustworthy.
  • Process window tightened and the intervention trigger changed from calendar-based to condition-based, using the characteristic that predicted the drift.
  • Control charts placed on that characteristic at the station, with reaction rules written into standard work — what to do, by whom, at which signal.
  • PFMEA updated for the mode, with the detection rating re-scored against the new in-process control rather than end-of-line inspection.
  • Control plan and work instructions revised; inspection criteria aligned so the station-level check and the final check were measuring the same thing.
  • Weekly scrap review changed from reporting the average to reviewing the control charts, which is the only version of that meeting that can catch a drift early.

Result

Scrap fell 15% across the area. A parallel value-stream mapping program run with the same team captured an additional $180K. Gross utilization on the area rose 20% — roughly $250K annually — through time studies, line balancing and standard work, and changeover time fell 18%.

The result I care about more is that the improvement did not decay. The gain was held by a condition-based trigger and a control chart with reaction rules, not by sustained attention from the engineer who found it.

What I learned

Verify the measurement system before trusting the Pareto. The single highest-leverage hour of this project was spent discovering that two defect codes meant the same thing to some inspectors and different things to others. Every analysis downstream of bad categorization is confidently wrong.

Averages hide drift. A mean inside specification is perfectly compatible with a process that is spending half its life at the edge of the window. Reviewing summary statistics instead of control charts is how a problem survives years of attention.

Containment is not corrective action. The previous attempts had all worked — briefly. They worked because someone was watching. The test of a corrective action is whether it still works when nobody is.

Defects attributed to people are usually attributable to process windows. “Operator technique” is rarely a root cause. It is usually a description of what happens when a process window is too narrow to run reliably without exceptional care.

Tools & methods

  • Defect Pareto
  • Data segmentation
  • Gauge / measurement review
  • Fishbone
  • 5-Why
  • Control charts
  • Cp / Cpk
  • PFMEA
  • Control plan
  • Standard work
  • MES data

Next

The companion project on the equipment side of the same role is Production Equipment Run-Off & Acceptance, which covers how requirements were turned into acceptance criteria before new equipment was allowed into production.