Материал: [lect] Grubbs - Procedure for Detecting outlying observations in samples (1969)

Внимание! Если размещение файла нарушает Ваши авторские права, то обязательно сообщите нам

DETECTINGOUTLYINGOBSERVATIONSIN SAMPLES

5

s= estimate of the population standard deviation based on the sample data, calculated with n - 1 degrees of freedom as follows:

f= (X=

E-

 

 

1)

V),-x/n(n

 

n -- 1

J

n

(n - 1)

)

If x, rather than xn is the doubtful value, the criterion is as follows:

T- = (x - xI)/s

The critical values for either case, for the 1 per cent and 5 per cent levels of significance, are given in Table 1. Table 1 and the following tables give the "one-sided" significance levels. (In a previous ASTM tentative recommended practice (1961), the tables listed values of significance levels double those in the present practice, since it was considered that the experimenter would test either the lowest or the highest observation (or both) for statistical significance. However, to be consistent with actual practice and in an attempt to avoid further misunderstanding, single-sided significance levels are tabulated here so that both

viewpoints can be represented.)

4.2 The hypothesis that we are testing in every case is that all observations in the sample come from the same normal population. Let us adopt, for example,

a significance level of 0.05. If we are interested

only in outliers that occur on

the high side, we should always use the statistic

Tn = (Xn- x)/s and take as

critical value the 0.05 point of Table 1. On the other hand, if we are interested

only in

outliers

occurring on the

low

side, we would always use the

statistic

T1 = (x

- x,)/s and again

take

as a critical value the 0.05 point of

Table 1. Suppose, however, that we are interested in outliers occurring on either side, but do not believe that outliers can occur on both sides simultaneously. We might, for example, believe that at some time during the experiment something possibly happened to cause an extraneous variation on the high side or on the low side, but that it was very unlikely that two or more such events could have occurred, one being an extraneous variation on the high side and the other an extraneous variation on the low side. With this point of view we should use the statistic T, = (x, - X)/s or the statistic T1 = (x - xl)/s which ever is larger. If in this instance we use the 0.05 point of Table 1 as our critical value,

the true significance level would be twice 0.05 or 0.10. If we wish a significance level of 0.05 and not 0.10, we must in this case use as a critical value the 0.025

point of Table 1. Similar considerations apply to the other tests given below.

Example 1

As an illustration of the use of T, and Table 1, consider the following ten

observations on breaking strength (in pounds) of 0.104-in. hard-drawn copper wire: 568, 570, 570, 570, 572, 572, 572, 578, 584, 596. The doubtful observation is

the high value, x,o =

596. Is the value of 596 significantly high? The mean is

x = 575.2 and the estimated standard deviation is s = 8.70. We compute

 

Tio = (596 - 575.2)/8.70 = 2.39

From Table 1, for n =

10, note that a T,, as large as 2.39 would occur by chance

6 FRANK E. GRUBBS

with probability less than 0.05. In fact, so large a value would occur by chance not much oftener than 1% of the time. Thus, the weight of the evidence is against the doubtful value having come from the same population as the others (assuming the population is normally distributed). Investigation of the doubtful value is therefore indicated.

4.3 An alternative system, the Dixon criteria, based entirely on ratios of

differences between the

observations is described in the literature [5] and may

be used in cases where

it is desirable to avoid calculation of s or where quick

judgment is called for. For the Dixon test, the sample criterion or statistic changes with sample size. Table 2 gives the appropriate statistic to calculate and also gives the critical values of the statistic for the 1%, 5% and 10% levels

of significance.

Example 2

As an illustration of the use of Dixon's test, consider again the observations on breaking strength given in Example 1, and suppose that a large number of such samples had to be screened quickly for outliers and it was judged too time- consuming to compute s. Table 2 indicates use of

ri, = xn

 

~

for a sample size of ten. Thus, for n = 10,

-

X-

Xn

 

X2

 

 

 

 

 

 

 

 

 

xlo

 

-X

 

 

 

 

 

 

XO

 

--

X2

 

 

 

 

Xlo

 

 

2

 

For the measurements of breaking strength above,

 

 

 

596 -

584=

=

.462

 

 

 

596 -

 

570

 

462

 

 

 

 

 

 

 

which is a little less than .477, the 5% critical value for n = 10. Under the Dixon criterion, we should therefore not consider this observation as an outlier at the 5% level of significance. This illustrates how border-line cases may be accepted under one test but rejected under another. It should be remembered, however, that the T-statistic discussed above is the best one to use for the single-outlier case, and final statistical judgment should be based on it. See Ferguson, References [6], [7].

Further examination of the sample observations on breaking strength of hard-drawn copper wire indicates that none of the other values needs testing.

(Note 2.)

4.4 A test equivalent to Tn (or TI) based on the sample sum of squared deviations from the mean for all the observations and the sum of squared de-

viations omitting the "outlier" is given by Grubbs in [8]

4.5 The next type of problem to consider is the case where we have the possi- bility of two outlying observations, the least and the greatest observation, in a

Note 2: With experience we may usually just look at the sample values to observe if an outlier is present. However, strictly speaking the statistical test should be applied to all samples to guarantee the significance levels used. Concerning "multiple" tests on a single sample, we comment on this below.

DETECTINGOUTLYING OBSERVATIONS IN SAMPLES

7

sample. (The problem of testing the two highest or the two lowest observations is considered below.) In testing the least and the greatest observations simultaneously as probable outliers in a sample, we use the ratio of sample range to sample standard deviation test of David, Hartley and Pearson [4]. The signifi- cance levels for this sample criterion are given in Table 3. An example in astronomy follows.

Example 3

There is one rather famous set of observations that a number of writers on the

subject of outlying observations have referred to in applying their various tests for "outliers". This classic set consists of a sample of 15 observations of the

TABLE 2

Dixon Criteria for Testing of Extreme Observation (Single Sample)*

n

3

4

5

6

7

8

9

10

11

12

13

14

15

16

17

18

19

20

21

22

23

24

25

 

 

 

 

 

Criterion

 

X2

 

-

XI

if smallest value

 

 

 

rio

=--=

 

 

-

 

 

Xn

 

Xi

is suspected;

 

 

 

 

 

 

 

Xn

 

-

 

Xn-1

if largest value

 

 

 

 

is suspected.

 

Xn

 

-

XI

rl

=

 

 

-

X

if smallest value

 

Xn-i

is suspected;

 

 

 

 

 

 

 

n -

Xn_l

if largest value

 

 

n -

X2

is suspected.

 

XS

 

-

Xl

if smallest value

 

 

is suspected;

r2l

=-

 

 

 

 

Xn-1

 

-

Xi

 

 

 

 

 

Xn

-

 

Xn-2

if largest value

 

 

 

is suspected.

 

Xn

-

X2

 

 

 

 

 

 

 

X

3-XI

 

if smallest value

 

 

is suspected.

r22 =

 

 

-

Xl

 

Xn-2

 

 

 

 

Xn -

 

xn-2

if largest value

 

X=

 

-

-3is

suspected;

 

Xn

 

X3

 

Significance

Level

10%

5%

1%

.886

.941

.988

.679

.765

.889

.557

.642

.780

.482

.560

.698

.434

.507

.736

.479

.554

.683

.441

.512

.635

.409

.477

.597

.517

.576

.679

.490

.546

.642

.467

.521

.615

.492

.546

.641

.472

.525

.616

.454

.507

.595

.438

.490

.577

.424

.475

.561

.412

.462

.547

.401

.450

.535

.391

.440

.524

.382

.430

.514

.374

.421

.505

.367

.413

.497

.360

.406

.489

* From W. J.

Dixon,

"Processing Data for Outliers", Biometrics, March 1953, Vol. 9,

No. 1, Appendix,

Page 89.

(Reference [5]) xl < x2 _< '- ' < x,

8

FRANK E. GRUBBS

"vertical semi-diameters of Venus made by Lieutenant Herndon in 1846 and given in William Chauvenet's A Manual of Spherical and Practical Astronomy, Vol. II (5th ed., 1876). In the reduction of the observations, Prof. Pierce assumed two unknown quantities and found the following residuals which have been arranged in ascending order of magnitude:

-1.40"

-0.24

-0.05

0.18

0.48

-0.44

-0.22

0.06

0.20

0.63

-0.30

-0.13

0.10

0.39

1.01

TABLE 3

Critical Values for w/s (Ratio of Range to Sample Standard Deviation)*

Number of

5%

1%

0.5%

Observations

Significance

Significance

Significance

n

Level

Level

Level

3

2.00

2.00

2.00

4

2.43

2.44

2.45

5

2.75

2.80

2.81

6

3.01

3.10

3.12

7

3.22

3.34

3.37

8

3.40

3.54

3.58

9

3.55

3.72

3.77

10

3.68

3.88

3.94

11

3.80

4.01

4.08

12

3.91

4.13

4.21

13

4.00

4.24

4.32

14

4.09

4.34

4.43

15

4.17

4.43

4.53

16

4.24

4.51

4.62

17

4.31

4.59

4.69

18

4.38

4.66

4.77

19

4.43

4.73

4.84

20

4.49

4.79

4.91

30

4.89

5.25

5.39

40

5.15

5.54

5.69

50

5.35

5.77

5.91

60

5.50

5.93

6.09

80

5.73

6.18

6.35

100

5.90

6.36

6.54

150

6.18

6.64

6.84

200

6.38

6.85

7.03

500

6.94

7.42

7.60

1000

7.33

7.80

7.99

* Taken from H. A. David, H. O. Hartley and E. S. Pearson, "The Distribution of the Ratio in a Single Sample of Range to Standard Deviation," Biometrika, Vol. 41 (1954),

pp. 482-493. (Reference[4])

 

 

 

 

 

\f

(Xi -

)2

IW

Xn

-

Il

8

=

n

-

1

X

 

 

X1 <

X2

<

" -

< Xn

 

 

 

 

DETECTINGOUTLYINGOBSERVATIONSIN SAMPLES

9

The deviations -1.40 and 1.01 appear to be outliers. Here the suspected observations lie at each end of the sample. Much less work has been accomplished for the case of outliers at both ends of the sample than for the case of one or more outliers at only one end of the sample. This is not necessarily because the "one-sided" case occurs more frequently in practice but because "two-sided" tests are more difficult to deal with. For a high and a low outlier in a single

sample, the procedure below may possess near optimum properties. For optimum procedures when there is at hand an independent estimate, s2 of a2, see "Some Tests for Outliers" by C. P. Quesenberry and H. A. David, Technical Report No. 47, OOR (ARO) project No. 1166, Virginia Polytechnic Institute, Blacks-

burg, Virginia.

4.6 For the observations on the semi-diameters of Venus given above, all the information on the measurement error is contained in the sample of 15 residuals. In cases like this, where no independent estimate of variance is available (i.e. we still have the single sample case), a useful statistic is the ratio of the range of the observations to the sample standard deviation:

w

 

Xn

-

x1X

 

I

 

(Xi

-

X)2

W

=

 

i

s =

 

s

 

s

 

where

 

-i n -

1

 

 

 

 

 

\

If x, is about as far above the mean, x, as xi is below x, and if w/s exceeds some chosen critical value, then one would conclude that boththe doubtful values are outliers. If, however, xl and x, are displaced from the mean by different amounts, some further test would have to be made to decide whether to reject as outlying only the lowest value or only the highest value or both the lowest and highest values.

4.7 For this example the mean of the deviations is x = .018, s = .551, and

/

=

1.01 -

(-1.40)

2.41

4.374

tv/s

 

.551

= --

=

 

 

 

.551

 

From Table 3 for n =

15, we see that the value of w/s = 4.374 falls between

the critical values for the 1% and 5% levels, so if the test were being run at the 5% level of significance, we would conclude that this sample contains one or more outliers. The lowest measurement, -1.40", is 1.418" below the sample mean, and the highest measurement, 1.01", is .992" above the mean. Since

these extremes are not symmetric about the

mean, either both extremes are

outliers or else only -1.40

is an outlier. That -1.40

is an outlier can be verified

by use of the

T, statistic. We have

 

 

T,

= (x - x,)/s

= .018 - -1.40)

=

2.574 and from

 

 

.551

 

 

Table 1 this value is greater than the critical value for the 5% level, so we reject -1.40. Since we have decided that -1.40 should be rejected, we use the remaining 14 observations and test the upper extreme 1.01, either with the criterion

-

n

8

Источник: https://studfile.net/preview/16673013/