Post 5 — Real datasets are rarely balanced

Balanced experimental designs are statistically convenient.

Real analytical datasets are often not.

Samples disappear. Measurements fail. Some matrices are harder to analyse. Replicates may be missing.

So what happens when the elegant 4 × 3 design starts losing observations?

The important point is that an unbalanced design is not simply a balanced ANOVA with fewer rows.

Once the design becomes unbalanced, the interpretation of sums of squares becomes more complicated.

The order in which factors are entered into the model can matter.

Different definitions of “the effect of a factor after accounting for the others” can lead to different tests.

This is where Type I, Type II and Type III sums of squares enter the discussion.

But I think the deeper lesson is more important than the terminology.

With a balanced design, the experimental design itself gives us a clean separation between factors.

With an unbalanced design, that separation is no longer automatic.

The model has to do more work.

And this is exactly where blindly applying a familiar ANOVA workflow becomes dangerous.

The scientific question has not changed.

We still want to understand Method, Matrix and their interaction.

What has changed is the amount of information available for the different parts of the design.

That is why I prefer to think of unbalanced ANOVA as a modelling problem rather than as a special version of the F-test.

The difficult question is not:

“Which Type of Sum of Squares should I use?”

It is:

“What comparison does this model term actually represent for this particular design?”

Once that is clear, the choice of statistical machinery becomes much easier to justify.


The implementation details and reproducible analysis for this step are available in the GitHub repository at https://github.com/andreabz/analytical-anova/ and on the project page at https://andreabz.github.io/analytical-anova/