23  Missing plot analysis

In agricultural field experiments the experimenter is often encountered with the situation that the observations of a particular treatment/ plot may be lost or may be affected by some external factors so that it would not be possible to analyse these observations by including it with the normal values. The observation on a treatment may get lost by various reasons like, attack of cattle or birds, manure may be dumped on the side, disease infestation and so on. The data recorded from plots so affected will be omitted and then the analysis is carried out – called as missing plot analysis.

The Analysis of such data may be done by different methods.

  1. the most commonly used method currently adopted is known as ‘method of analysis of non – orthogonal data’ which is highly computer based

  2. Method of substitution by Yates based on minimization of the error sum of squares and

  3. Analysis of the data with missing values by the technique of analysis of covariance, due to Bartlett.

23.1 Covariance method (Bartlett’s approach)

Assume an imaginary covariate X taking values zero for every plot except the missing plot for which it will take the value 1 (or –1) . Now the value of the main variate Y = 0 for the missing plot, and the actual values for the remaining plots. The data will be analysed as per the ANCOVA technique of the respective design used.

In missing plot analysis the degrees of freedom for error and total will be based on the existing number of observations only. By this method also the degrees of freedom of adjusted error sum of squares will be less by one ( when there is one missing value).

23.2 Method of substitution

  1. In this method we will calculate with the aid of a formula an estimate of the missing value. The formula will vary from design to design and actually it will not supply the exact missing value. But the procedure permits the researcher to complete the analysis without resorting to more complex procedures.

  2. Insert the estimated value in the missing position and work out the estimates of treatment means and error sum of squares.

  3. Some additional adjustments to treatment sum of squares are needed. In pair comparison also some changes are made.

  4. An iterative procedure is adopted when more than one observation is missing.

23.2.1 Randomised block design with a single missing value

  1. The missing value in RBD is estimated as

\[x=\frac{rB+vT-G}{(r-1)(v-1)} \tag{23.1}\]

where,
x = estimate of the missing data;
v = number of treatments;
r = number of replications;
B =Total of the observed values of the replication that contain the missing data
T = Total of the observed values of the treatment that contain the missing data
G = Grand total of all existing observations.

  1. This estimate of the missing value is placed in its position and the analysis is carried out based on the procedure of RBD. Subtract one degree of freedom from the total and error degrees of freedom. This method provides a proper estimate of the error variance per plot but there is an inflation in treatment sum of squares; ( the treatment sum of squares is positively biased). If the treatments turn out to be not significant we can ignore the bias and the results are accepted. But if the treatments turn out to be just significant, it may be due to this bias. In that case an actual treatment sum of squares is obtained by a suitable formula.

  2. Estimate of bias in the case of RBD:

\[Bias=\frac{(B+vT-G)^2}{v(v-1)(r-1)^2} \tag{23.2}\]

This bias is subtracted from the treatment sum of squares. Now test this actual treatment mean square against the mean square for error, and draw the conclusion about the significance of treatments.

  1. Pair comparison
  1. For comparing two treatment means in which one of them contains the missing observation:

\[CD=t_{\alpha} \cdot \sqrt{MSE \cdot \left(\frac{2}{r} + \frac{v}{r(r-1)(v-1)}\right)} \tag{23.3}\]

  1. For comparing other pairs of treatment means (in which neither contains a missing observation):

\[CD=t_{\alpha} \cdot \sqrt{\frac{2\,MSE}{r}}\]

where \(t_{\alpha}\) denotes the \(t\) table value at \((r-1)(v-1)-1\) degrees of freedom.

Example 22.1: Data of a RBD with five sources of sulphur (S), 4 blocks chlorophyll content of a pea leaves in mg/g of the fresh weight are represented below. One sample was spoiled due to negligence during analysis. Find the missing value and do further analysis.

Treatment Block 1 Block 2 Block 3 Block 4
\(S_0\) 0.679 0.852 0.513 0.507
\(S_1\) 0.952 1.002 \(x\) 0.621
\(S_2\) 0.899 0.919 0.718 0.679
\(S_3\) 0.986 0.949 0.845 0.780
\(S_4\) 0.911 0.922 0.668 0.746

Solution 22.1

Treatment Block 1 Block 2 Block 3 Block 4 Total
\(S_0\) 0.679 0.852 0.513 0.507 2.551
\(S_1\) 0.952 1.002 \(x\) 0.621 2.575 + \(x\)
\(S_2\) 0.899 0.919 0.718 0.679 3.215
\(S_3\) 0.986 0.949 0.845 0.780 3.56
\(S_4\) 0.911 0.922 0.668 0.746 3.247
Total 4.427 4.644 2.744 + \(x\) 3.333 15.873 + \(x\)

Here, \(r=4\), \(v=5\)

Using Equation 23.1

\[x=\frac{10.976+12.875-15.148}{4\times3}=0.725\]

Correction factor, \(CF=\dfrac{(15.873)^2}{20}=12.597\)

Total SS \(=13.036-CF=13.036-12.597=0.439\)

Treatment SS \(=\dfrac{50.95}{4}-12.597=0.1406\)

Block SS \(=\dfrac{64.307}{5}-12.597=0.264\)

Using Equation 23.2

\[Bias=\frac{(2.744+12.875-15.148)^2}{5\times4\times9}=0.00123\]

Treatment Adj. SS \(=0.1406-0.00123=0.139\)

Total Adj. SS \(=0.439-0.00123=0.437\)

Source df S.S M. S.S \(F_{cal}\) \(F_{tab}\)
Block 3 0.264 0.088 29.33 3.59
Treatment Adj. 4 0.139 0.034 11.003 3.36
Error Adj. 11 0.034 0.003
Total Adj. 18 0.437

Since the \(F_{cal}>F_{tab}\) the treatment pairs are significantly different from each other.

CD for missing data using Equation 23.3

\[CD=2.201\times\sqrt{\frac{0.003}{4}\left(2+\frac{5}{12}\right)}=0.095\]

CD for other data is

\[CD=t_{\alpha}\times\sqrt{\frac{2\,MSE}{r}}=2.201\times\sqrt{\frac{2\times0.003}{4}}=0.086\]

Here, the mean of \(S_3\) is maximum, which was found to be on par with the means of \(S_1\) and \(S_4\). Also, \(S_1\) was on par with \(S_4\) and \(S_2\). The mean of \(S_0\) is minimum.

23.2.2 Latin square design with a single missing value

The same procedure as in RBD is used here also.

  1. The missing value in LSD is estimated as

\[x=\frac{v(R+C+T)-2G}{(v-1)(v-2)} \tag{23.4}\]

where
x = estimate of the missing data
v = number of treatments/ rows or blocks/ columns
R = Total of the observed values of the row that contain the missing data
C = Total of the observed values of the column that contain the missing data
T = Total of the observed values of the treatment that contain the missing data
G = Grand total of all existing observations.

  1. Carry out the analysis similar to the above method after substitution. The estimate of bias in this case is

\[Bias=\frac{[(v-1)T+R+C-G]^2}{[(v-1)(v-2)]^2} \tag{23.5}\]

  1. For pairwise comparison involving a treatment with a missing observation (comparing two treatment means in which one of them contains a missing observation):

\[CD=t_{\alpha} \cdot \sqrt{MSE \cdot \left(\frac{2}{v} + \frac{1}{(v-1)(v-2)}\right)} \tag{23.6}\]

Example 22.2: The data in below table is the grain yield of paddy from a varietal trial in 5x5 LSD with 5 varieties of paddy, one observation was found missing. Estimate the missing value and analyse the data.

Col 1 Col 2 Col 3 Col 4 Col 5
Row 1 E (26) C (42) D (39) B (37) A (24)
Row 2 A (24) D (33) E (21) C (\(x\)) B (38)
Row 3 D (47) B (45) A (31) E (29) C (31)
Row 4 B (38) A (24) C (36) D (41) E (34)
Row 5 C (41) E (24) B (44) A (26) D (30)

Solution 22.2

Here, \(v=5\)

Col 1 Col 2 Col 3 Col 4 Col 5 Row total
Row 1 E (26) C (42) D (39) B (37) A (24) 168
Row 2 A (24) D (33) E (21) C (\(x\)) B (38) 116+\(x\)
Row 3 D (47) B (45) A (31) E (29) C (31) 183
Row 4 B (38) A (24) C (36) D (41) E (34) 173
Row 5 C (41) E (24) B (44) A (26) D (30) 165
Column total 176 168 171 133+\(x\) 157

Using Equation 23.4

\[x=\frac{[5(116+133+150)-(2\times 805)]}{4\times3}=32.083\approx 32\]

Correction factor, \(CF=\dfrac{(837)^2}{25}=28022.76\)

Total SS \(=29399-CF=29399-28022.76=1376.24\)

Treatment SS \(=\dfrac{129^2+202^2+182^2+190^2+134^2}{5}-28022.76=902.24\)

Row SS \(=\dfrac{168^2+148^2+183^2+173^2+165^2}{5}-28022.76=131.44\)

Column SS \(=\dfrac{176^2+168^2+171^2+165^2+157^2}{5}-28022.76=40.24\)

Using Equation 23.5

\[Bias=\frac{[(4\times150)+133+116-805]^2}{(4\times3)^2}=13.44\]

Treatment Adj. SS \(=902.24-13.44=888.8\)

Total Adj. SS \(=1376.24-13.44=1362.8\)

Source df SS MSS \(F_{cal}\) \(F_{tab}\)
Row 4 131.44 32.86 1.19 3.36
Column 4 40.24 10.06 0.36 3.36
Treatment Adj. 4 888.8 222.2 8.08 3.36
Error Adj. 11 302.32 27.48
Total Adj. 23 1362.8

Since the \(F_{cal}>F_{tab}\) the treatment pairs are significantly different from each other.

CD for missing data using Equation 23.6

\[CD=2.201\times\sqrt{27.48\left(\frac{2}{5}+\frac{1}{4\times3}\right)}=8.021\]

CD for other data is

\[CD=t_{\alpha}\times\sqrt{\frac{2\,MSE}{r}}=2.201\times\sqrt{\frac{2\times27.48}{5}}=7.297\]

Here, the mean of B is maximum, which was found to be on par with the means of D and C. The mean of A is minimum, which is on par with E.

Historical Insights

Frank Yates and the art of the missing value

In the field, experiments rarely go exactly to plan. A goat breaks into a plot, a flood washes out a corner, a technician mislabels a sample, and suddenly one of the neatly balanced yields is simply gone. In the early days of agricultural statistics this was a serious problem, because Fisher’s elegant analysis of variance relied on the tidy balance of equal numbers in every treatment and block. A single hole in the data destroyed that balance and, with it, the simple arithmetic of the analysis.

The solution came from Frank Yates, Fisher’s brilliant younger colleague at the Rothamsted Experimental Station. In the early 1930s Yates worked out that a missing observation could be replaced by an estimated value, chosen precisely so as to minimise the error sum of squares, allowing the standard analysis to proceed almost unchanged. His formulae, one for the randomized block design, another for the Latin square, let researchers recover nearly all the information from a damaged experiment with only a small, correctable bias in the treatment sum of squares and the loss of a single degree of freedom. Yates set this out in his 1933 paper on the analysis of replicated experiments with missing values, and the “missing plot technique” quickly became a standard part of every field experimenter’s toolkit. It is a fine example of the statistician’s craft: not discarding hard-won data because of an accident, but finding a principled way to make the most of what remains. (Yates 1933)

Quotes to Inspire

“The best time to plan an experiment is after you’ve done it.” - R. A. Fisher