Examples / 06

When the result is not a measurement: weld seams with an ordinal and a nominal response

Not every response is a number. A weld seam is looked at and classified: how much spatter lies next to it, and which defect does the cross-section show? The first answer has an order, the second has none. Here 81 plates are welded at three settings, the spatter is analysed with an ordinal model and the weld defect with a multinomial one, and at the end there is the setting at which a good seam is most likely.

Design
3³ full factorial, 3 plates per setting = 81
Models
proportional odds (ordinal), multinomial logit
Factors
current, travel speed, shielding gas
Responses
2 with 4 categories each, confirmed with 24 plates

The judgements are simulated. They come from two models we set ourselves — for every plate a category is drawn from the true probabilities. That makes it possible to check at the end whether the analysis finds what is really in there. Every table and chart was computed and drawn by DoEStat.

Step 1 / The question

One seam, two judgements

For a fillet weld on structural steel, a MAG setting is wanted that gives a defect-free seam with little spatter as often as possible. Three settings are free:

FactorUnitlowcentrehigh
Welding currentA160200240
Travel speedcm/min304560
Shielding gas flowl/min81420

Every plate is judged twice: the spatter by eye on a four-step scale, the weld defect in the cross-section. Both responses have four categories — but not of the same kind:

ResponseKindCategories
Spatterordinal (ranked)none < few < moderate < many
Weld defectnominal (no order)none · porosity · lack of fusion · undercut

“Few” spatter lies between “none” and “moderate” — the steps are ranked, but the distances between them are unknown; “moderate” does not mean “twice as many as few”. The weld defects, by contrast, are only named: porosity is not “more” than lack of fusion, it has a different cause. This distinction decides the model. In DoEStat it is set in the definition of the response: list the levels, tick “ordered” or not.

Putting a regression on these data as if the categories were the numbers 1 to 4 would be wrong in both cases: the weld defect has no order, the spatter no equal distances. What can be modelled is the probability of each category depending on the setting.

Step 2 / Planning

How many plates does a judgement need?

A category carries less information than a measurement. A single plate only says “porosity” or “no porosity” — the probability behind it only emerges from many. A full factorial with three levels per factor has 27 settings; the question is how often each has to be welded.

The table shows how often the analysis is usable when the whole study is drawn a hundred times from the true models. “Usable” means: the model converges, every category occurs, and no estimate runs off to infinity.

Plates per settingPlates in totalSpatter (ordinal): usableWeld defect (nominal): usable
12799%24%
254100%81%
381100%96%
4108100%97%

The ordinal model is frugal — it estimates one slope per factor, the same for all steps. The nominal one is not: it estimates separate slopes for each of the three defects, twelve parameters instead of six. With one plate per setting it is unusable in three studies out of four, with two still in one out of five. From three plates per setting it is 96%, and a fourth adds little. The choice is therefore 81 plates.

What “running off to infinity” means is separation: if a rare defect never occurs for a group of settings, its probability there can be pushed arbitrarily close to zero, and the estimate grows without bound. DoEStat reports this in the analysis (“separation suspected”). The problem grows with every additional term — a squared gas term, which suggests itself in a three-level design, leaves only 45 of 100 nominal analyses usable at 81 plates in the same simulation, instead of 96. The model here therefore has main effects only.

This planning calculation is done by the example itself, by repeating the study from the true models. The power analysis in DoEStat covers numeric and binomial responses, not categorical ones with more than two levels.

Step 3 / Runs

81 plates, judged twice

The 81 plates are welded in random order. Only 22 of them are good in the sense of the question — no weld defect and at most a few spatters. Counted by the levels of each factor, the direction shows even before any model:

Settingnoneporositylack of fusionundercut
Welding current 160 A104130
Welding current 200 A14472
Welding current 240 A105012
Travel speed 30 cm/min16812
Travel speed 45 cm/min12375
Travel speed 60 cm/min62127
Shielding gas flow 8 l/min9963
Shielding gas flow 14 l/min11376
Shielding gas flow 20 l/min14175

Lack of fusion piles up at low current and fast travel — too little heat, the seam does not fuse with the base metal. Undercuts come with high current. Porosity depends on the shielding gas: frequent at 8 l/min, rare at 20 l/min.

Settingnonefewmoderatemany
Welding current 160 A161100
Welding current 200 A8874
Welding current 240 A09711
Travel speed 30 cm/min91125
Travel speed 45 cm/min7965
Travel speed 60 cm/min8865
Shielding gas flow 8 l/min61038
Shielding gas flow 14 l/min8856
Shielding gas flow 20 l/min101061

For the spatter, the current dominates: at 160 A there are never more than a few, at 240 A not a single plate without spatter and two thirds with moderate or many. Little shielding gas makes the arc less steady.

All 81 plates
PlateWelding current [A]Travel speed [cm/min]Shielding gas flow [l/min]SpatterWeld defect
12006014moderatelack of fusion
22403014manyundercut
32406014manyundercut
4240458manyporosity
52406020moderateundercut
61603014fewnone
71606020fewlack of fusion
8240458fewundercut
9200308fewporosity
10240608manyundercut
11240458manyporosity
121606014nonelack of fusion
131603020nonenone
14200608noneporosity
15160608nonelack of fusion
16160308nonenone
171603020nonenone
18240308manynone
191606014nonelack of fusion
202403020fewnone
211606020nonelack of fusion
22160458fewlack of fusion
232006014nonenone
242406020fewundercut
251604520nonenone
262406014fewundercut
272004520manylack of fusion
281604514nonelack of fusion
29200308fewporosity
302406020moderateundercut
312006020fewnone
321604514nonelack of fusion
332404514fewundercut
342003014fewnone
352406014manynone
36240608manyundercut
372404520fewporosity
382404520moderatenone
391606020fewnone
40160308nonenone
411603014noneporosity
421604520nonenone
432003020fewundercut
442004514noneundercut
452004514moderatenone
462403020fewnone
472004520moderatelack of fusion
482404520moderateundercut
491606014nonelack of fusion
502404514manynone
51160458fewlack of fusion
522404514moderateundercut
53160458nonenone
54200308noneporosity
55200458fewnone
562403020moderatenone
57160308fewporosity
58160608fewporosity
592003020nonenone
602004520nonenone
612004514fewnone
62200608manylack of fusion
63200608moderatelack of fusion
642403014manynone
652003014moderatenone
66240308fewnone
672403014fewporosity
682006020fewnone
691604514fewnone
702003014manynone
711603014fewporosity
721604520fewlack of fusion
731603020nonelack of fusion
74160608fewlack of fusion
752006020nonelack of fusion
76240608moderatenone
772006014moderatelack of fusion
78200458moderatenone
79200458manynone
80240308manyporosity
812003020nonenone

Step 4 / Spatter

Ranked steps: the proportional-odds model

For a ranked response DoEStat chooses the cumulative logit model: it describes the probability of reaching at most a given step. One can picture an invisible spatter tendency behind the four steps that the factors shift; three cut-points divide it into the steps. Each factor gets one slope that applies to all three cut-points — hence “proportional odds”.

TermEstimateStd. errorzp-valueTrue value
Welding current1.880.335.73< 0.00011.50
Travel speed0.260.270.970.33420.30
Shielding gas flow−0.720.27−2.630.0085−0.80

A positive slope pushes the seam towards more spatter. The current acts strongly (z = 5.7), the shielding gas dampens (p = 0.008), the travel speed is not detectable (p = 0.33). The column “True value” is the comparison with the model the data come from — the current is somewhat overestimated, the direction and ranking are right.

Cut-pointEstimateStd. errorTrue value
none | few−1.390.32−1.00
few | moderate0.890.290.80
moderate | many2.180.372.40

Does the assumption of equal slopes hold?

The model claims that a factor acts equally strongly at every cut-point. Whether the data support this is checked by the Brant test: it estimates the three cut-points separately and compares the slopes. A small p-value would mean that a factor acts differently between “none” and “few”, say, than between “moderate” and “many” — the multinomial model would then be the more honest choice.

Termχ²dfp-value
Welding current0.0320.9836
Travel speed2.0720.3543
Shielding gas flow2.7020.2594
Overall (omnibus)5.1960.5198

No factor objects (jointly p = 0.52). The ordinal model may stay — and for three factors it needs three slopes instead of nine.

Log-likelihoodwithout factorsLikelihood-ratio test χ² (df)p-valueMcFadden R²AIC
−85.8−108.845.9 (3)< 0.00010.211183.7
observed ↓ / predicted →nonefewmoderatemanyCorrect
none1590015 / 24
few9130613 / 28
moderate010040 / 14
many0401111 / 15
Accuracy48%

The confusion matrix compares the observed step with the most likely step according to the model. 48% hits sound low, but that is the nature of a probability: at a setting with 45% “few” and 35% “none”, “few” is the best prediction and still often wrong. “Moderate” is never the most likely step — it lies between two neighbours that are each more frequent. What the model separates reliably are the ends: no plate with “no” spatter is predicted as “moderate” or “many”, and vice versa.

Step 5 / Weld defect

Named categories: the multinomial logit

Without an order there is no common scale. The multinomial logit therefore compares each defect separately with a reference category — here “none” — and gives each defect its own slopes. That way the current may prevent lack of fusion and at the same time cause undercuts, without the two cancelling each other out.

Class against “none”TermEstimateStd. errorp-valueTrue value
porosityIntercept−1.520.500.0021−1.20
porosityWelding current0.150.470.74350.00
porosityTravel speed−0.470.510.35660.00
porosityShielding gas flow−1.420.530.0075−1.60
lack of fusionIntercept−1.450.510.0044−1.10
lack of fusionWelding current−1.750.620.0045−2.00
lack of fusionTravel speed1.780.530.00091.50
lack of fusionShielding gas flow−0.040.440.93010.00
undercutIntercept−2.200.750.0032−1.70
undercutWelding current2.350.830.00461.80
undercutTravel speed1.050.510.03941.20
undercutShielding gas flow0.130.490.78630.00

Each defect has its own cause, and the analysis finds it: porosity at the shielding gas (−1.42, true −1.6), lack of fusion at low current and fast travel, undercuts at high current and — more weakly — fast travel. All terms that are zero in the true model stay inconspicuous (p > 0.3).

Log-likelihoodwithout factorsLikelihood-ratio test χ² (df)p-valueMcFadden R²AIC
−71.2−105.869.3 (9)< 0.00010.327166.4
observed ↓ / predicted →noneporositylack of fusionundercutCorrect
none1848418 / 34
porosity37217 / 13
lack of fusion3017017 / 20
undercut3101010 / 14
Accuracy64%

The model recognises lack of fusion and undercuts best — they have clear, opposing causes. The defect-free seam is the hardest to predict: it is what remains when none of the three defects occurs, and so it has no cause of its own.

ClassCountAUC95% interval
none340.7310.621 … 0.841
porosity130.7990.660 … 0.938
lack of fusion200.9080.847 … 0.970
undercut140.9000.810 … 0.989

The ROC curves that DoEStat draws for each category against all others show the same: the area under the curve is 0.91 for lack of fusion and 0.73 for the defect-free seam.

ROC curve for lack of fusion against all other categories: the curve rises steeply and reaches the full true-positive rate at about 20 per cent false-positive rate; area 0.908 with interval 0.847 to 0.970.
ROC: lack of fusionLow current and fast travel — a clear cause, easy to recognise.
ROC curve for the defect-free seam against all defects: the curve lies clearly above the diagonal but flatter than for the defects; area 0.731 with interval 0.621 to 0.841.
ROC: no weld defectThe hardest to predict — the defect-free seam has no cause of its own.

Step 6 / Process window

Where a good seam is most likely

The probability profiler shows, for each category, how its probability changes when one factor is moved and the others are held. The four curves in a column add up to one everywhere. In DoEStat the lines are dragged with the mouse and both responses are seen at once.

Probability profiler of the weld defect at the chosen setting, four rows for none, porosity, lack of fusion and undercut over current, travel speed and shielding gas: porosity falls with more gas, lack of fusion rises with faster travel and falls with more current, undercuts rise with the current; at the setting no defect has a probability of 85 per cent.
Profiler: weld defectEach row a category, each column a factor. The dashed lines stand at the chosen setting.
Probability profiler of the spatter at the chosen setting, four rows for none, few, moderate and many over current, travel speed and shielding gas: with rising current the probability moves from none to many, more gas works the other way, the travel speed hardly matters; at the setting none or few spatters have a probability of 94 per cent.
Profiler: spatterOne slope per factor shifts all four steps together — the ordinal model.

The profilers stand at the setting at which a good plate — no weld defect, at most a few spatters — is most likely. It was found as the product of the two probabilities over a grid of the region:

FactorSettingUnitcoded
Welding current180A−0.50
Travel speed30cm/min−1.00
Shielding gas flow20l/min1.00

The current lies in the lower quarter of the range: more current brings spatter and undercuts, less brings lack of fusion. Travel speed and shielding gas sit on the edge of the region — welding slowly with plenty of gas is simply better here. Anyone searching beyond that leaves the region studied; the model knows nothing there.

Step 7 / Confirmation

24 plates at the chosen setting

A probability is not confirmed with one plate. At the chosen setting 24 plates are welded and judged as in the study:

ResponseCategoryModel predictionTrue probabilityConfirmation (24 plates)
Spatternone63%70%18
few31%23%5
moderate4%5%1
many2%1%0
Weld defectnone84%78%19
porosity7%5%2
lack of fusion8%16%3
undercut1%2%0

Result: 18 of the 24 plates are good (75%) — in the study it was 22 of 81 (27%). The model had promised 80%; the truth is 73%. The gap is typical: whoever picks the best of all settings also picks the one at which the model happens to be too optimistic. Most clearly for lack of fusion — the model expects 8%, the truth is 16%, and 3 of 24 were observed.

Summing up

What the right scale achieved

Two responses that cannot be measured, only classified, and for each the model that fits its kind. The ordinal model uses the order of the spatter steps and makes do with three slopes; the Brant test confirmed that this is enough. The multinomial model gives each weld defect its own cause — and that was exactly what was needed, because the current prevents one defect and causes another.

The price is the number of plates. Categorical data need replicates, and the nominal model needs more than the ordinal one. The planning calculation beforehand showed that 54 plates would have led to no usable result in one study out of five, and that an obvious squared term makes matters worse.

Against the truth: all factors that act in the true model are found, with the right sign; all that do not act stay inconspicuous. The estimates deviate from the truth by up to a little over half a logit — with 81 plates and one judgement each, that is the precision one can expect. And the confirmation shows that the probability at the best setting is somewhat overestimated: anyone who has to make a commitment takes the result of the confirmation plates, not the prediction.

Project files

Run it yourself

The example ships with DoEStat as a project: Help ▸ Open sample project ▸ Other ▸ “Categorical responses: weld seam”, once with the data only (with the full quadratic model a three-level design is made for) and once with the finished analysis (main effects). The same files can be downloaded here.

Back: simulation with a Gaussian process All examples

Try DoEStat for free

30 days, the full feature set, no payment details. We send the download link by e-mail, usually on the next business day.

Request the trial