Method, Matrix, and their interaction are fixed effects.
These are the effects we want to estimate and interpret.
Random effects
SampleID is a random effect.
We are not interested in estimating a separate scientific effect for FIL_01, FIL_02, and so on.
We are interested in accounting for the variation among samples.
This distinction is the essence of a mixed model:
Fixed effects describe the systematic effects we want to study. Random effects describe sources of variation that induce dependence or represent a broader population of experimental units.
How much variation is associated with the sample?
The random-effects output gives us two variance components:
cat("Percentage of variance associated with samples:",round(100* icc, 1), "%\n")
Percentage of variance associated with samples: 15.5 %
An ICC of 0.155 means that approximately 15.5% of the modelled variance is associated with differences between samples.
This is not a measure of method performance.
It is a measure of within-sample dependence.
A larger ICC means that measurements from the same sample tend to resemble one another more closely, because they share a common sample-level component.
That is precisely the structure that the random intercept is designed to capture.
In our experimental setup, the four aliquots from the same sample are prepared independently and submitted to the four analytical methods. Independent aliquot preparation can reduce shared preparation errors, but the aliquots still inherit the same underlying sample characteristics.
They are therefore distinct analytical measurements, but they are not four independent samples.
The mixed model explicitly represents this distinction by separating sample-level variation from the remaining observation-level variation.
In our data, approximately 15.5% of the total modelled variance is associated with differences between samples after accounting for Method, Matrix and their interaction.
This suggests that sample-specific heterogeneity remains relevant to the response.
From an analytical-chemistry perspective, this is an important result: the variability between environmental samples is not negligible compared with the remaining measurement-level variability. The model therefore supports treating SampleID as part of the dependence structure rather than treating the 240 measurements as 240 independent experimental units.
Fixed model versus mixed model
It is useful to ask what changes when we acknowledge the sample structure.
The point is not that the mixed model must produce a different scientific conclusion.
The point is that the uncertainty should be estimated under the correct dependence structure.
We can compare the standard error of a method coefficient from the two models.
This would allow the Method effect to vary from sample to sample.
However, this model cannot be estimated with our experimental design.
Each sample has exactly four observations:
Reference
Method_A
Method_B
Method_C
A random intercept plus a separate random Method effect therefore attempts to estimate a very large number of sample-specific parameters from only four observations per sample.
With 60 samples and four method-related random coefficients per sample, the random-effects structure contains 240 random effects for 240 observations.
There is no residual information left to identify the model.
lme4 therefore correctly rejects the model.
This is an important lesson:
A more flexible random-effects structure is not automatically a better model.
The data must contain enough repeated information to estimate it.
For this dataset, the random-intercept model is the appropriate level of complexity.
Batch and Day: another layer of the design
The dataset contains additional variables:
Batch;
Day;
Analyst;
Instrument.
These variables deserve attention, but they do not play the same role as SampleID.
SampleID identifies the physical sample shared by the four method measurements.
Batch and Day describe when and under which analytical conditions a measurement was performed.
This distinction matters.
For example, the four measurements from FIL_01 were not necessarily performed in the same batch or on the same day.
dat[, .(n_batch =uniqueN(Batch),n_day =uniqueN(Day)), by = SampleID][, .(n_samples = .N), by = .(n_batch, n_day)][order(n_batch, n_day)]
If the estimates and their uncertainty remain broadly similar, this provides additional reassurance that the Method conclusions are not simply reflecting differences in Batch or Day.
If they change substantially, that is scientifically important.
It would mean that part of the apparent Method effect is entangled with analytical conditions.
Either outcome is informative.
What about Analyst and Instrument?
The dataset also records the analyst and instrument associated with each measurement.
These are potentially important sources of analytical variation.
However, we should resist the temptation to turn every recorded variable into a random effect.
With only two instruments and three analysts, estimating random-effect variances for these factors would be difficult and would add little to the central lesson of this chapter.
For this dataset, they are better regarded as potential nuisance factors to investigate, rather than automatically adding them to the random-effects structure.
The general principle is:
The presence of a variable in the dataset does not by itself justify including it in the model.
The role of a variable depends on the experimental design and on the scientific question.
Why not make Batch, Day, Analyst and Instrument all random effects?
The ICC tells us how important the sample-level variation is.
Batch and Day represent additional analytical structure. They are useful for sensitivity analysis and can be included as fixed nuisance effects when we want to ask whether the Method conclusions are robust to these sources of systematic variation.
Analyst and Instrument are also relevant variables, but with only 3 and 2 levels respectively, they do not need to become random effects in our main model.
Finally, a more complicated random-effects structure is not necessarily better.
Our attempt to give each sample a random Method effect fails because four observations per sample do not contain enough information to estimate that structure.
The lesson is simple:
Mixed models are not about adding random effects until the model looks sophisticated.
They are about representing the dependence that is actually present in the experiment.
Final reflection
This brings us to an important point in the progression of the series.
We began by treating the dataset as a collection of observations.
Then we asked increasingly specific questions about:
methods;
matrices;
interactions;
unbalanced designs;
covariates;
assumptions;
robustness.
Now we have asked a more fundamental question:
What exactly is an observation?
A measurement made on a sample is not necessarily an independent experimental unit.
Once we recognise that, the statistical model changes.
This is not a technical refinement added at the end of the analysis.
It is part of understanding the experiment itself.
And that is ultimately what applied statistics requires: not simply fitting models to data, but identifying the structure that generated those data.
The complete journey
Chapter
Question
01
Do the four methods have the same mean bias?
02
Which comparisons actually matter?
03
Does Matrix matter independently of Method?
04
Does the method effect depend on the matrix?
05
What happens when the design is unbalanced?
06
Does organic matter explain the pattern?
07
Are the model assumptions defensible?
08
Which conclusions are robust?
09
How repeated measurements should be handled?
10
How does hierarchical structure affect inference?
The statistical journey does not end with the most complicated model.
It ends when the model is appropriate for the question and the structure of the data.