05 — Unbalanced Designs

What happens when the cells have different sizes?

Scientific question

So far, every cell in our design has had the same number of observations.

The real world is rarely so tidy.

Samples get lost. Instruments break. Analysts call in sick.

When cell sizes differ, a new question emerges:

Does the order in which we enter terms into the model matter?

In a balanced design, the answer is no.

In an unbalanced design, the answer is yes — and understanding why is essential.

The dataset

We start from the same experiment, but now some cells have fewer observations.

This is not an accident.

It is a realistic analytical scenario: some method-matrix combinations are harder to run than others.

Method A in Sediment, for example, might have failed more often due to matrix interference.

library(here)
library(data.table)

dat <- fread(
  here("data", "dataset_v01_clean_balanced.csv")
)

dat[, Method := factor(
  Method,
  levels = c("Reference", "Method_A", "Method_B", "Method_C")
)]

dat[, Matrix := factor(
  Matrix,
  levels = c("Fill", "Soil", "Sediment")
)]

# Create unbalanced version by removing observations
set.seed(42)
remove_idx <- c(
  sample(which(dat$Method == "Method_A" & dat$Matrix == "Sediment"), 8),
  sample(which(dat$Method == "Method_C" & dat$Matrix == "Fill"), 6),
  sample(which(dat$Method == "Reference" & dat$Matrix == "Soil"), 4),
  sample(which(dat$Method == "Method_B" & dat$Matrix == "Sediment"), 3),
  sample(which(dat$Method == "Method_A" & dat$Matrix == "Fill"), 2)
)

dat_unbal <- dat[-remove_idx]

dat_unbal[, .N, by = .(Method, Matrix)]
       Method   Matrix     N
       <fctr>   <fctr> <int>
 1: Reference     Fill    20
 2:  Method_A     Fill    18
 3:  Method_B     Fill    20
 4:  Method_C     Fill    14
 5: Reference     Soil    16
 6:  Method_A     Soil    20
 7:  Method_B     Soil    20
 8:  Method_C     Soil    20
 9: Reference Sediment    20
10:  Method_B Sediment    17
11:  Method_C Sediment    20
12:  Method_A Sediment    12

The design is now unbalanced.

Method A in Sediment has only 12 observations.

Method C in Fill has only 14.

The total is 217 observations instead of 240.

Why imbalance matters

In a balanced design, the sums of squares for Method and Matrix are orthogonal.

Each term is tested after adjusting for the grand mean, but not after adjusting for the other factor.

In an unbalanced design, the model terms are no longer orthogonal. The distribution of observations across the levels of one factor depends on the levels of the other, so the amount of variation attributed to one term depends on which other terms have already been accounted for.

The sum of squares assigned to Method depends on whether Matrix is already in the model.

And vice versa.

This leads to three different ways of partitioning the variance:

  • Type I: sequential. Each term is adjusted for the terms that came before it.
  • Type II: each term is adjusted for all other terms of equal or lower order, but not for higher-order interactions.
  • Type III: each term is adjusted for all other terms, including interactions.

These are not three different statistical methods.

They are three different answers to the question:

“Adjusted for what?”

Type I: sequential sums of squares

Type I is the default of R’s anova() function.

The order of terms matters.

fit_unbal <- lm(RelativeBias_pct ~ Method + Matrix, data = dat_unbal)
anova(fit_unbal)
Analysis of Variance Table

Response: RelativeBias_pct
           Df  Sum Sq Mean Sq F value    Pr(>F)    
Method      3 10888.7  3629.6 338.997 < 2.2e-16 ***
Matrix      2   609.8   304.9  28.478 1.125e-11 ***
Residuals 211  2259.1    10.7                      
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

In this table:

  • Method is tested after adjusting only for the intercept.
  • Matrix is tested after adjusting for the intercept and Method.

If we reverse the order:

fit_unbal_rev <- lm(RelativeBias_pct ~ Matrix + Method, data = dat_unbal)
anova(fit_unbal_rev)
Analysis of Variance Table

Response: RelativeBias_pct
           Df  Sum Sq Mean Sq F value    Pr(>F)    
Matrix      2   381.9   191.0  17.836 6.971e-08 ***
Method      3 11116.6  3705.5 346.092 < 2.2e-16 ***
Residuals 211  2259.1    10.7                      
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

The sums of squares for Method and Matrix have changed.

They are no longer the same as in the balanced case.

This is not a bug.

It is a consequence of imbalance.

Type II: adjusting for the other main effect

Type II tests each main effect after adjusting for the other main effect, but not for the interaction.

It is the natural choice when you have no interaction (or when the interaction is not of primary interest).

car::Anova(fit_unbal, type = "II")
Anova Table (Type II tests)

Response: RelativeBias_pct
           Sum Sq  Df F value    Pr(>F)    
Method    11116.6   3 346.092 < 2.2e-16 ***
Matrix      609.8   2  28.478 1.125e-11 ***
Residuals  2259.1 211                      
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Notice that the sums of squares for Method and Matrix are now the same regardless of the order in which the terms were entered into the model.

Type II answers the question:

“What is the effect of Method, adjusting for Matrix?”

and:

“What is the effect of Matrix, adjusting for Method?”

Type III: adjusting for everything

Type III tests each term after adjusting for all other terms, including interactions.

It is the default in many other statistical packages (SAS, SPSS).

car::Anova(fit_unbal, type = "III")
Anova Table (Type III tests)

Response: RelativeBias_pct
             Sum Sq  Df F value    Pr(>F)    
(Intercept)  2040.2   1 190.550 < 2.2e-16 ***
Method      11116.6   3 346.092 < 2.2e-16 ***
Matrix        609.8   2  28.478 1.125e-11 ***
Residuals    2259.1 211                      
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Type III answers the question:

“What is the effect of Method, adjusting for Matrix and the Method:Matrix interaction?”

This sounds appealingly comprehensive.

But it comes with a cost.

In the presence of interactions, Type III main-effect tests can be difficult to interpret because they test the effect at the reference level of the other factor, not averaged across levels.

Type III tests are especially sensitive to how categorical predictors are parameterized. With the default treatment contrasts, a main-effect test can be interpreted relative to the reference level of another factor. With sum-to-zero contrasts, the parameterization instead represents deviations from the corresponding marginal mean.

Therefore, if Type III tests are used, contrast coding should be chosen deliberately rather than left implicit.

With treatment contrasts, the Type III test for a main effect is not what most users think it is.

Which type should you use?

There is no universal answer.

The choice depends on the question:

Type Question it answers Typical use
I What does this term explain after the preceding terms? Sequential or hierarchical hypotheses
II What is the effect of this main effect after accounting for the other main effects? Models without interactions of primary interest
III What remains attributable to this term after accounting for all other terms in the model? Specific hypotheses in models containing interactions

Type II and Type III are not simply different “corrections” for imbalance. They correspond to different hypotheses.

The important point is to understand what each type is adjusting for, and whether that matches your scientific question.

In our unbalanced dataset, all three types agree that Method and Matrix are significant.

The differences are in the sums of squares and F-values, not in the qualitative conclusions.

But in more marginal cases — near the significance threshold, or with stronger imbalance — the choice can change the answer.

What about the interaction?

With imbalance, the interaction model becomes even more delicate.

fit_unbal_int <- lm(RelativeBias_pct ~ Method * Matrix, data = dat_unbal)
car::Anova(fit_unbal_int, type = "III")
Anova Table (Type III tests)

Response: RelativeBias_pct
              Sum Sq  Df  F value    Pr(>F)    
(Intercept)   1108.7   1 107.4696 < 2.2e-16 ***
Method        3290.1   3 106.3068 < 2.2e-16 ***
Matrix         119.5   2   5.7938  0.003567 ** 
Method:Matrix  144.3   6   2.3305  0.033704 *  
Residuals     2114.9 205                       
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

The interaction is still significant.

But notice how the main effects have changed.

In Type III, the main effect of Method is tested after adjusting for the interaction.

In the presence of an interaction, a Type III main-effect test requires particular care. The meaning of the test depends on the parameterization of the factors and may correspond to a comparison at a particular reference level rather than to an overall average across matrices.

When an interaction is present, the scientific interpretation should usually focus on the interaction and on scientifically meaningful simple effects: for example, comparing methods within each matrix. Whether Type II or Type III is appropriate for a particular omnibus test depends on the hypothesis being tested and on the model parameterization.

Take-home message

Imbalance does not break ANOVA.

It forces us to be explicit about what we are adjusting for.

Type I, II and III are not competing statistical philosophies.

They are three ways of answering the same question with different adjustment strategies.

The right choice depends on your design, your questions, and your willingness to interpret the results carefully.

In practice, for unbalanced factorial designs with interactions, a sensible strategy is:

  1. Fit the interaction model.
  2. Test the interaction (Type II or III).
  3. If the interaction is significant, focus on simple effects rather than main effects.
  4. If the interaction is not significant, use Type II for main effects.

Next: ANCOVA

We have now seen what happens with two categorical factors.

But what if a continuous variable — organic matter content — also influences the bias?

And what if the effect of that continuous variable depends on the method?

That leads us to ANCOVA.

The question becomes:

“Does organic matter content explain the method differences we have observed?”