02 — Contrasts

Which comparisons actually matter?

Scientific question

The one-way ANOVA told us that the four methods do not share the same mean relative bias.

That answers a global question.

It does not answer the questions that typically matter in practice:

  • Which specific differences are scientifically interesting?
  • Which differences are large enough to be chemically relevant?
  • And how should we estimate and test them?

This is the domain of contrasts.

A contrast is not a post-hoc fishing expedition.

It is a way of encoding a specific scientific question as a linear combination of group means.

The key challenge is not computational.

The key challenge is deciding which questions are worth asking.

The dataset

We continue with the same balanced dataset.

The response is RelativeBias_pct.

The factor is Method, with four levels: Reference, A, B and C.

For this chapter we still ignore Matrix.

We will add it in the next chapter.

Load the data

library(here)
library(data.table)
library(ggplot2)

dat <- fread(
  here("data", "dataset_v01_clean_balanced.csv")
)

dat[, Method := factor(
  Method,
  levels = c("Reference", "Method_A", "Method_B", "Method_C")
)]

dat[, .N, by = Method]
      Method     N
      <fctr> <int>
1: Reference    60
2:  Method_A    60
3:  Method_B    60
4:  Method_C    60

Reminder: what the one-way ANOVA told us

fit_oneway <- lm(RelativeBias_pct ~ Method, data = dat)
summary(fit_oneway)

Call:
lm(formula = RelativeBias_pct ~ Method, data = dat)

Residuals:
    Min      1Q  Median      3Q     Max 
-9.6991 -2.3251  0.0914  2.8272  8.6320 

Coefficients:
               Estimate Std. Error t value Pr(>|t|)    
(Intercept)     -9.5220     0.4772  -19.95   <2e-16 ***
MethodMethod_A  -8.1052     0.6748  -12.01   <2e-16 ***
MethodMethod_B  11.4747     0.6748   17.00   <2e-16 ***
MethodMethod_C   6.9211     0.6748   10.26   <2e-16 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 3.696 on 236 degrees of freedom
Multiple R-squared:  0.8028,    Adjusted R-squared:  0.8003 
F-statistic: 320.3 on 3 and 236 DF,  p-value: < 2.2e-16

The fitted model uses Reference as the baseline.

The intercept is the estimated mean of the Reference method.

The other coefficients are differences from Reference:

\[ \beta_A = \mu_A - \mu_\mathrm{Reference} \]

\[ \beta_B = \mu_B - \mu_\mathrm{Reference} \]

\[ \beta_C = \mu_C - \mu_\mathrm{Reference} \]

These are valid contrasts.

But they are not the only contrasts we might care about.

And the fact that R has chosen Reference as the baseline does not mean that Reference is inherently the most interesting comparison.

The baseline level is a parameterization choice. It is not a scientific decision.

What is a contrast?

A contrast is a linear combination of group means:

\[ C = \sum_i c_i \mu_i \quad \text{with} \quad \sum_i c_i = 0 \]

The constraint that the coefficients sum to zero ensures that the contrast compares groups rather than estimating an overall mean.

Any question of the form “How does group X differ from group Y?” can be written as a contrast.

More complex questions can also be written as contrasts.

For example:

“Is the Reference method more biased than the average of the three alternatives?”

is:

\[ C = \mu_\mathrm{Reference} - \frac{\mu_A + \mu_B + \mu_C}{3} \]

The coefficients are \(1\), \(-\frac{1}{3}\), \(-\frac{1}{3}\), \(-\frac{1}{3}\).

They sum to zero.

It is a valid contrast.

Look at the data before choosing contrasts

method_stats <- dat[, .(
  Mean = mean(RelativeBias_pct),
  SD   = sd(RelativeBias_pct),
  SE   = sd(RelativeBias_pct) / sqrt(.N),
  N    = .N
), by = Method]

method_stats[, `:=`(
  CI_low  = Mean - qt(0.975, N - 1) * SE,
  CI_high = Mean + qt(0.975, N - 1) * SE
)]

method_stats
      Method       Mean       SD        SE     N      CI_low    CI_high
      <fctr>      <num>    <num>     <num> <int>       <num>      <num>
1: Reference  -9.522000 3.226779 0.4165753    60 -10.3555653  -8.688435
2:  Method_A -17.627167 4.060234 0.5241740    60 -18.6760364 -16.578297
3:  Method_B   1.952717 3.775214 0.4873780    60   0.9774755   2.927958
4:  Method_C  -2.600867 3.673626 0.4742631    60  -3.5498649  -1.651868

The estimated means are approximately:

  • Reference: -9.5%
  • Method A: -17.6%
  • Method B: 2%
  • Method C: -2.6%

Method B is closest to zero.

Method C is also close.

The difference between them is the comparison that matters most for method selection.

A visualization focused on the question

ggplot(dat, aes(x = Method, y = RelativeBias_pct)) +
  geom_jitter(width = 0.12, alpha = 0.35, size = 1.5) +
  geom_errorbar(
    data = method_stats,
    aes(y = Mean, ymin = CI_low, ymax = CI_high),
    width = 0.12, linewidth = 0.7
  ) +
  geom_point(data = method_stats, aes(y = Mean), size = 3) +
  geom_hline(yintercept = 0, linetype = "dashed") +
  geom_hline(yintercept = c(-5, 5), linetype = "dotted") +
  labs(
    x = NULL,
    y = "Relative Bias (%)",
    title = "Relative bias by analytical method",
    subtitle = "Points = observations; symbols and intervals = mean ± 95% CI"
  ) +
  annotate("text", x = 4.45, y = 5, label = "+5 pp",
           hjust = 1, vjust = -0.4, size = 3.5) +
  annotate("text", x = 4.45, y = -5, label = "−5 pp",
           hjust = 1, vjust = 1.4, size = 3.5) +
  theme_minimal(base_size = 12)

The dotted lines at ±5 percentage points are the predefined chemical relevance threshold.

Visually, the gap between B and C is smaller than the gap between either of them and the other methods.

Whether that gap is statistically distinguishable from zero is a separate question from whether it exceeds the chemical threshold.

Contrast 1: Is the Reference method itself unbiased?

This is not a comparison between two methods.

It is a one-sample question about the Reference method alone:

\[ H_0: \mu_\mathrm{Reference} = 0 \]

ref_bias <- dat[Method == "Reference", RelativeBias_pct]
t.test(ref_bias, mu = 0)

    One Sample t-test

data:  ref_bias
t = -22.858, df = 59, p-value < 2.2e-16
alternative hypothesis: true mean is not equal to 0
95 percent confidence interval:
 -10.355565  -8.688435
sample estimates:
mean of x 
   -9.522 

The test gives t = -22.86 with p < 0.001.

The Reference method is significantly biased.

This matters scientifically.

If the Reference method were unbiased, there would be less justification for developing alternatives.

The fact that it is biased — and substantially so — is part of the problem.

Contrast 2: Method B versus Method C

In our scenario, the comparison between B and C is particularly relevant because these are the two strongest candidates.

The question is whether the difference between them is large enough to matter.

We formulate this as a contrast:

\[ C = \mu_B - \mu_C \]

with coefficients \(0\), \(0\), \(1\), \(-1\).

Because the design is balanced and the group sizes are equal, the standard error of this contrast is:

\[ SE = \sqrt{\frac{s_B^2}{n_B} + \frac{s_C^2}{n_C}} \]

mean_b <- dat[Method == "Method_B", mean(RelativeBias_pct)]
mean_c <- dat[Method == "Method_C", mean(RelativeBias_pct)]
diff_bc <- mean_b - mean_c

se_bc <- sqrt(
  dat[Method == "Method_B", var(RelativeBias_pct)] / 60 +
  dat[Method == "Method_C", var(RelativeBias_pct)] / 60
)

t_bc <- diff_bc / se_bc
p_bc <- 2 * pt(-abs(t_bc), df = 118)

cat("B - C =", round(diff_bc, 3), "pp\n")
B - C = 4.554 pp
cat("SE      =", round(se_bc, 3), "\n")
SE      = 0.68 
cat("t(118)  =", round(t_bc, 3), "\n")
t(118)  = 6.696 
cat("p       =", format(p_bc, digits = 2, scientific = TRUE), "\n")
p       = 7.7e-10 

The estimated difference is 4.55 percentage points.

The test is highly significant: t = 6.7, p < 0.001.

But the magnitude is 4.55 percentage points.

Our chemical relevance threshold is 5 percentage points.

So this difference is:

  • Statistically significant: the data are not compatible with equal population means.
  • Chemically borderline: the magnitude is close to but below the 5 pp threshold.

This is precisely the kind of situation where statistical significance and practical relevance diverge.

A small p-value does not mean a large effect. It means a precisely estimated effect.

Contrast 3: Reference versus the average of the alternatives

Sometimes the scientific question is not about pairwise differences.

It is about whether the existing method differs from the new methods as a group.

means <- dat[, .(M = mean(RelativeBias_pct)), by = Method]
mu_ref <- means[Method == "Reference", M]
mu_new <- means[Method != "Reference", mean(M)]

c_ref_new <- mu_ref - mu_new

cat("Reference vs mean(A, B, C) =", round(c_ref_new, 3), "pp\n")
Reference vs mean(A, B, C) = -3.43 pp

The Reference method is approximately 3.4 percentage points more negative than the average of the three alternatives.

This contrast is not one of the pairwise comparisons.

It is a different scientific question.

And it can be estimated and tested just as easily.

All pairwise comparisons: Tukey HSD

When we want to compare all pairs while controlling the family-wise error rate, the Tukey HSD is the natural choice.

It adjusts for the fact that we are making multiple comparisons.

fit_aov <- aov(RelativeBias_pct ~ Method, data = dat)
tukey <- TukeyHSD(fit_aov)
tukey
  Tukey multiple comparisons of means
    95% family-wise confidence level

Fit: aov(formula = RelativeBias_pct ~ Method, data = dat)

$Method
                        diff       lwr       upr p adj
Method_A-Reference -8.105167 -9.851208 -6.359125     0
Method_B-Reference 11.474717  9.728675 13.220758     0
Method_C-Reference  6.921133  5.175092  8.667175     0
Method_B-Method_A  19.579883 17.833842 21.325925     0
Method_C-Method_A  15.026300 13.280259 16.772341     0
Method_C-Method_B  -4.553583 -6.299625 -2.807542     0

All six pairwise differences are statistically significant at the 5% level.

But look at the magnitudes, not just the p-values:

Comparison Difference (pp) Chemically relevant?
A vs B -19.6 Yes
A vs C -15 Yes
A vs Ref -8.1 Yes
B vs C -4.6 Borderline
B vs Ref 11.5 Yes
C vs Ref 6.9 Yes

The Tukey table confirms what we already suspected.

Method B and Method C are the two best candidates.

Their difference is real but small.

In a regulatory context, both might be considered “accurate enough”, and the choice could shift to other criteria: precision, cost, throughput, or matrix-specific performance.

Statistical significance versus chemical relevance

This distinction is worth repeating.

The p-value from a contrast tells us whether the observed difference is compatible with a null hypothesis of no difference.

It does not tell us whether the difference matters in the laboratory.

For that, we need a scientific criterion.

In this project, that criterion is:

\[ \boxed{5\text{ percentage points}} \]

A difference of 4.6 pp between B and C is:

  • Precisely estimated (small standard error, large sample).
  • Statistically distinguishable from zero.
  • Below the threshold that would make us prefer one method over the other on accuracy grounds alone.

This is not a failure of statistics.

It is statistics working as intended.

The data have told us that the difference is real but small.

The decision of whether “small” matters is scientific, not statistical.

What assumptions have we made?

The contrasts we have computed rely on the same assumptions as the one-way ANOVA:

  • Independence of observations.
  • Normally distributed residuals.
  • Common residual variance across methods.

Because this is a classical ANOVA contrast, its standard error is based on the common residual variance estimated by the model. This is one reason why the assumptions of the underlying model matter for contrasts as well as for the global F-test.

If the variances differ substantially between methods, this standard error may not be optimal.

We will investigate this in Chapter 07.

For now, we note the assumption and move on.

What the contrasts tell us — and what they do not

What they tell us

  • The Reference method is significantly biased.
  • Method A is substantially worse than all other methods.
  • Method B and Method C are the best candidates.
  • Their difference is statistically significant but chemically borderline.

What they do not tell us

  • Matrix effects: These contrasts average over Fill, Soil and Sediment. If Method B is excellent in one matrix and poor in another, the average is misleading.
  • Precision: We have looked only at mean bias. A method with zero bias but high variance is still problematic.
  • Why the differences exist: These comparisons quantify differences between methods under this experimental setup. They do not, by themselves, identify the physical or chemical mechanism responsible for those differences.

Take-home message

A contrast is a question dressed as mathematics.

The coefficients you choose should reflect the scientific questions you actually need to answer.

R’s default output — comparisons against a baseline — is convenient for software.

It is not necessarily the most interesting comparison for your laboratory.

The Tukey HSD gives you all pairwise answers with proper multiplicity adjustment.

But only you can decide which differences are large enough to act on.

And that decision requires both statistical evidence and a scientific threshold.


Next: Two-way ANOVA

We now know which method differences exist and how large they are.

But all of these contrasts average over Matrix.

What if the ranking of methods changes depending on whether the sample is Fill, Soil or Sediment?

That question cannot be answered by one-way ANOVA.

It requires a second factor.

Next, we move to two-way ANOVA.

The question becomes:

Does Matrix influence the bias, and does it do so independently of Method?