Материал: [lect] Grubbs - Procedure for Detecting outlying observations in samples (1969)

Внимание! Если размещение файла нарушает Ваши авторские права, то обязательно сообщите нам

American Society for Quality

Procedures for Detecting Outlying Observations in Samples Author(s): Frank E. Grubbs

Source: Technometrics, Vol. 11, No. 1 (Feb., 1969), pp. 1-21

Published by: American Statistical Association and American Society for Quality Stable URL: http://www.jstor.org/stable/1266761

Accessed: 11/08/2011 11:45

Your use of the JSTOR archive indicates your acceptance of the Terms & Conditions of Use, available at http://www.jstor.org/page/info/about/policies/terms.jsp

JSTOR is a not-for-profit service that helps scholars, researchers, and students discover, use, and build upon a wide range of content in a trusted digital archive. We use information technology and tools to increase productivity and facilitate new forms of scholarship. For more information about JSTOR, please contact support@jstor.org.

American Statistical Association and American Society for Quality are collaborating with JSTOR to digitize, preserve and extend access to Technometrics.

http://www.jstor.org

TECHNOMETRICS

VOL. 11, No. 1

FEBRUARY 1969

Procedures for Detecting Outlying

Observations in Samples

FRANK E. GRUBBS*

UT.S. Army Aberdeen Research and Development Center

Aberdeen Proving Ground, Maryland 21005

Proceduresare given for determiningstatistically whetherthe highest observation, the lowest observation,the highest and lowest observations,the two highest observations, the two lowest observations, or more of the observations in the sample are statistical outliers.Both the statistical formulaeand the applicationof the procedures to examples are given, thus representinga rather complete treatment of tests for outliers in single samples. This paper has been preparedprimarily as an expository and tutorial article on the problem of detecting outlying observations in much experimentalwork. We cover only tests of significancein this paper.

1.SCOPE OF PAPER

1.1This is an expository and tutorial type of paper which deals with the problem of outlying observations in samples and how to test the statistical

significance of them. An outlying observation, or "outlier," is one that appears to deviate markedly from other members of the sample in which it occurs. In this connection, the following two alternatives are of interest:

1.1.1 An outlying observation may be merely an extreme manifestation of the random variability inherent in the data. If this is true, the values should be retained and processed in the same manner as the other observations in the sample.

1.1.2 On the other hand, an outlying observation may be the result of gross deviation from prescribed experimental procedure or an error in calculating or recording the numerical value. In such cases, it may be desirable to institute an investigation to ascertain the reason for the aberrant value. The observation

may even eventually be rejected as a result of the investigation, though not necessarily so. At any rate, in subsequent data analysis the outlier or outliers will be recognized as probably being from a different population than that of the sample values.

1.2 It is our purpose here to provide statistical rules that will lead the experi-

Received December 1967; revised April 1968.

* Member, Committee E-ll on Statistical Methods, The American Society for Testing Materials (ASTM). This work in a slightly different form was prepared primarily for the AmericanSociety for Testing Materials and representsa ratherextensive revisionof an earlier

Tentative RecommendedPractice which was drafted by Dr. R. J. Hader and others in 1960.

The authoris indebtedto W. E. Deming, AchesonJ. Duncan, E. V. Harrington,Helen J. Coon and others for commentsleading to the present paper. Permissionhas been obtained from the

AmericanSociety for Testing Materials to publish this paper in Technometrics.

1

of the data.
2.3 Thus, for purposes of orientation relative to the overall problem of experimentation, our position on the matter of screening samples for outlying observations is precisely the following:
Physical Reason Known or Discoveredfor Outlier(s)
values and indicate to what extent

2

FRANK E. GRUBBS

menter almost unerringly to look for causes of outliers when they really exist, and hence to decide whether alternative 1.1.1 above is not the more plausible hypothesis to accept as compared to alternative 1.1.2 in order that the most appropriate action in further data analysis may be taken. The procedures covered herein apply primarily to the simplest kind of experimental data, i.e., replicate measurements of some property of a given material, or observations in a supposedly single random sample. Nevertheless, the tests suggested do cover a wide enough range of cases in practice to have rather broad utility.

2.GENERAL

2.1When the skilled experimenter is clearly aware that a gross deviation from prescribed experimental procedure has taken place, the resultant observations should be discarded, whether or not it agrees with the rest of the data

and without recourse to statistical tests for outliers. If a reliable correction procedure, for example, for temperature, is available, the observation may some- times be corrected and retained.

2.2 In many cases evidence for deviation from prescribed procedure will consist primarily of the discordant value itself. In such cases it is advisable to adopt a cautious attitude. Use of one of the criteria discussed below will sometimes permit a clear-cut decision to be made. In doubtful cases the experimenter's judgment will have considerable influence. When the experimenter cannot identify abnormal conditions, he should at least report the discordant

they have been used in the analysis

(i)Reject observation(s)

(ii)Correct observation(s) on physical grounds

(iii)Reject it (them) and possibly take additional observation(s)

Physical Reason Unknown-Use Statistical Test

(i)Reject observation(s)

(ii)Correct observation(s) statistically

(iii)Reject it (them) and possibly take additional observation(s)

(iv)Employ truncated sample theory for censored observations

2.4The statistical test may always be used to lend support to a judgment

that a physical reason does actually exist for an outlier, or the statistical criterion may be used routinely as a basis to initiate action to find a physical cause.

3.BASIS OF STATISTICAL CRITERIA FOR OUTLIERS

3.1There are a number of criteria for testing outliers. In all of these the doubtful observation is included in the calculation of the numerical value of a

DETECTINGOUTLYING OBSERVATIONS IN SAMPLES

3

sample criterion (or statistic), which is then compared with

a critical value

based on the theory of random sampling to determine whether the doubtful observation is to be retained or rejected. The critical value is that value of the sample criterion which would be exceeded by chance with some specified (small) probability on the assumption that all the observations did indeed constitute a random sample from a common system of causes, a single parent population, distribution or universe. The specified small probability is called the "significance levels" or "percentage point" and can be thought of as the risk of erroneously rejecting a good observation. It becomes clear, therefore, that if there exists a real shift or change in the value of an observation that arises from non-random causes (human error, loss of calibration of instrument, change of measuring instrument, or even change of time of measurements, etc.), then the observed value of the sample criterion used would exceed the "critical value" based on random sampling theory. Tables of critical values are usually given for several different significance levels, for example, 5%, 1%. For statistical tests of outlying observations, it is generally recommended that a low significance level, such as 1%, be used and that significance levels greater than 5% should not be common practice. (Note 1).

3.2 It should be pointed out that almost all criteria for outliers are based on an assumed underlying normal (Gaussian) population or distribution. When

the data are not normally or approximately normally distributed, the probabilities associated with these tests will be different. Until such time as criteria not

sensitive to the normality assumption are developed, the experimenter is cautioned against interpreting the probabilities too literally when normality of the data is not assured.

3.3 Although our primary interest here is that of detecting outlying observations, we remark that the statistical criteria used also test the hypothesis that the random sample taken did indeed come from a normal or Gaussian popula- tion. The end result is for all practical purposes the same, i.e., we really want to know once and for all whether we have in hand a sample of homogeneous observations.

4.RECOMMENDEDCRITERIAFORSINGLESAMPLES

4.1Let the sample of n observations be denoted in order of increasing mag-

nitude by x1 < x2

< xa3< .**

largest value. The

test criterion,

is as follows:

 

< Xn. Let xn be the doubtful

value, i.e. the

Tn , recommended here for a

single outlier

Tn = (xn - X)/s

where

x = arithmetic average of all n values, and

Note 1: In this paper, we will usually illustrate the use of the 5% significancelevel. Proper choice of level in probability depends on the particular problem and just what may be involved, along with the risk that one is willing to take in rejecting a good observation, i.e., if the null-

hypothesis stating "all observations in the sample come from the same normal population" may be assumed.

4

FRANK E. GRUBBS

 

TABLE 1

Table of Critical Values for

T (One-sided Test) When Standard Deviation

is Calculated from the Same Sample

Number of

5%

Observations

Significance

n

Level

3

1.15

4

1.46

5

1.67

6

1.82

7

1.94

8

2.03

9

2.11

10

2.18

11

2.23

12

2.29

13

2.33

14

2.37

15

2.41

16

2.44

17

2.47

18

2.50

19

2.53

20

2.56

21

2.58

22

2.60

23

2.62

24

2.64

25

2.66

30

2.75

35

2.82

40

2.87

45

2.92

50

2.96

60

3.03

70

3.09

80

3.14

90

3.18

100

3.21

 

Xn -X

 

I n-1

 

 

S

 

;

T,

--

X1 <

X2 <

*'

< Xn

 

 

s

2.5%

1%

Significance

Significance

Level

Level

1.15

1.15

1.48

1.49

1.71

1.75

1.89

1.94

2.02

2.10

2.13

2.22

2.21

2.32

2.29

2.41

2.36

2.48

2.41

2.55

2.46

2.61

2.51

2.66

2.55

2.71

2.59

2.75

2.62

2.79

2.65

2.82

2.68

2.85

2.71

2.88

2.73

2.91

2.76

2.94

2.78

2.96

2.80

2.99

2.82

3.01

2.91

 

2.98

 

3.04

 

3.09

 

3.13

 

3.20

 

3.26

 

3.31

 

3.35

 

3.38

 

n

-

-

1)

n

 

 

n(n

- 1)

Note: Values of T for n < 25 are based on those given in Reference [8]. For n > 25, the values of T are approximated. All values have been adjusted for division by n - 1 instead of n in calculating s.

Источник: https://studfile.net/preview/16673013/