Jarque–Bera test

In statistics, the Jarque–Bera test is a goodness-of-fit test of whether sample data have the skewness and kurtosis matching a normal distribution. The test is named after Carlos Jarque and Anil K. Bera. The test statistic JB is defined as

{\mathit {JB}}={\frac {n-k+1}{6}}\left(S^{2}+{\frac 14}(C-3)^{2}\right)

where n is the number of observations (or degrees of freedom in general); S is the sample skewness, C is the sample kurtosis, and k is the number of regressors:

S={\frac {{\hat {\mu }}_{3}}{{\hat {\sigma }}^{3}}}={\frac {{\frac 1n}\sum _{{i=1}}^{n}(x_{i}-{\bar {x}})^{3}}{\left({\frac 1n}\sum _{{i=1}}^{n}(x_{i}-{\bar {x}})^{2}\right)^{{3/2}}}},

C={\frac {{\hat {\mu }}_{4}}{{\hat {\sigma }}^{4}}}={\frac {{\frac 1n}\sum _{{i=1}}^{n}(x_{i}-{\bar {x}})^{4}}{\left({\frac 1n}\sum _{{i=1}}^{n}(x_{i}-{\bar {x}})^{2}\right)^{{2}}}},

where ${\hat {\mu }}_{3}$ and ${\hat {\mu }}_{4}$ are the estimates of third and fourth central moments, respectively, ${\bar {x}}$ is the sample mean, and ${\hat {\sigma }}^{2}$ is the estimate of the second central moment, the variance.

If the data comes from a normal distribution, the JB statistic asymptotically has a chi-squared distribution with two degrees of freedom, so the statistic can be used to test the hypothesis that the data are from a normal distribution. The null hypothesis is a joint hypothesis of the skewness being zero and the excess kurtosis being zero. Samples from a normal distribution have an expected skewness of 0 and an expected excess kurtosis of 0 (which is the same as a kurtosis of 3). As the definition of JB shows, any deviation from this increases the JB statistic.

For small samples the chi-squared approximation is overly sensitive, often rejecting the null hypothesis when it is true. Furthermore, the distribution of p-values departs from a uniform distribution and becomes a right-skewed uni-modal distribution, especially for small p-values. This leads to a large Type I error rate. The table below shows some p-values approximated by a chi-squared distribution that differ from their true alpha levels for small samples.

Calculated p-values equivalents to true alpha levels at given sample sizes
True α level	20	30	50	70	100
0.1	0.307	0.252	0.201	0.183	0.1560
0.05	0.1461	0.109	0.079	0.067	0.062
0.025	0.051	0.0303	0.020	0.016	0.0168
0.01	0.0064	0.0033	0.0015	0.0012	0.002

(These values have been approximated by using Monte Carlo simulation in Matlab)

In MATLAB's implementation, the chi-squared approximation for the JB statistic's distribution is only used for large sample sizes (> 2000). For smaller samples, it uses a table derived from Monte Carlo simulations in order to interpolate p-values.^[1]

History

Considering normal sampling, and √β₁ and β₂ contours, Bowman & Shenton (1975) noticed that the statistic JB will be asymptotically χ²(2)-distributed; however they also noted that “large sample sizes would doubtless be required for the χ² approximation to hold”. Bowman and Shelton did not study the properties any further, preferring D’Agostino’s K-squared test.

Jarque–Bera test in regression analysis

According to Robert Hall, David Lilien, et al. (1995) when using this test along with multiple regression analysis the right estimate is:

\mathit{JB} = \frac{n-k}{6} \left( S^2 + \frac14 (C-3)^2 \right)

where n is the number of observations and k is the number of regressors when examining residuals to an equation.

Implementations

ALGLIB includes implementation of the Jarque–Bera test in C++, C#, Delphi, Visual Basic, etc.
gretl includes an implementation of the Jarque–Bera test
R includes implementations of the Jarque–Bera test: jarque.bera.test in package tseries,^[2] for example, and jarque.test in package moments.^[3]
MATLAB includes implementation of the Jarque–Bera test, the function "jbtest".
Python statsmodels includes implementation of the Jarque–Bera test, "statsmodels.stats.stattools.py".

References

↑ "Analysis of the JB-Test in MATLAB". MathWorks. Retrieved May 24, 2009.
↑ "tseries: Time Series Analysis and Computational Finance". R Project.
↑ "moments: Moments, cumulants, skewness, kurtosis and related tests". R Project.

Bowman, K.O.; Shenton, L.R. (1975). "Omnibus contours for departures from normality based on √b₁ and b₂". Biometrika. 62 (2): 243–250. doi:10.1093/biomet/62.2.243. JSTOR 2335355.
Jarque, Carlos M.; Bera, Anil K. (1980). "Efficient tests for normality, homoscedasticity and serial independence of regression residuals". Economics Letters. 6 (3): 255–259. doi:10.1016/0165-1765(80)90024-5.
Jarque, Carlos M.; Bera, Anil K. (1981). "Efficient tests for normality, homoscedasticity and serial independence of regression residuals: Monte Carlo evidence". Economics Letters. 7 (4): 313–318. doi:10.1016/0165-1765(81)90035-5.
Jarque, Carlos M.; Bera, Anil K. (1987). "A test for normality of observations and regression residuals". International Statistical Review. 55 (2): 163–172. JSTOR 1403192.
Judge; et al. (1988). Introduction and the theory and practice of econometrics (3rd ed.). pp. 890–892.
Hall, Robert E.; Lilien, David M.; et al. (1995). EViews User Guide. p. 141.

Statistics

Descriptive statistics

Continuous data

Center	Mean arithmetic geometric harmonic Median Mode

Dispersion	Variance Standard deviation Coefficient of variation Percentile Range Interquartile range

Shape	Moments Skewness Kurtosis L-moments

Count data

Index of dispersion

Summary tables

Dependence

Graphics

Data collection

Study design	Population Statistic Effect size Statistical power Sample size determination Missing data

Survey methodology	Sampling Standard error stratified cluster Opinion poll Questionnaire

Controlled experiments	Design control optimal Controlled trial Randomized Random assignment Replication Blocking Interaction Factorial experiment

Uncontrolled studies	Observational study Natural experiment Quasi-experiment

Statistical inference

Statistical theory

Frequentist inference

Point estimation	Estimating equations Maximum likelihood Method of moments M-estimator Minimum distance Unbiased estimators Mean-unbiased minimum-variance Rao–Blackwellization Lehmann–Scheffé theorem Median unbiased Plug-in

Interval estimation	Confidence interval Pivot Likelihood interval Prediction interval Tolerance interval Resampling Bootstrap Jackknife

Testing hypotheses	1- & 2-tails Power Uniformly most powerful test Permutation test Randomization test Multiple comparisons

Parametric tests	Likelihood-ratio Wald Score

Specific tests

Z (normal) Student's t-test F

Goodness of fit	Chi-squared Kolmogorov–Smirnov Anderson–Darling Normality (Shapiro–Wilk) Likelihood-ratio test Model selection Cross validation AIC BIC

Rank statistics	Sign Sample median Signed rank (Wilcoxon) Hodges–Lehmann estimator Rank sum (Mann–Whitney) Nonparametric anova 1-way (Kruskal–Wallis) 2-way (Friedman) Ordered alternative (Jonckheere–Terpstra)

Bayesian inference

Correlation	Pearson product–moment Partial correlation Confounding variable Coefficient of determination

Regression analysis	Errors and residuals Regression model validation Mixed effects models Simultaneous equations models Multivariate adaptive regression splines (MARS)

Linear regression	Simple linear regression Ordinary least squares General linear model Bayesian regression

Non-standard predictors	Nonlinear regression Nonparametric Semiparametric Isotonic Robust Heteroscedasticity Homoscedasticity

Generalized linear model	Exponential families Logistic (Bernoulli) / Binomial / Poisson regressions

Partition of variance	Analysis of variance (ANOVA, anova) Analysis of covariance Multivariate ANOVA Degrees of freedom

Categorical / Multivariate / Time-series / Survival analysis

Categorical

Multivariate

Time-series

General	Decomposition Trend Stationarity Seasonal adjustment Exponential smoothing Cointegration Structural break Granger causality

Specific tests	Dickey–Fuller Johansen Q-statistic (Ljung–Box) Durbin–Watson Breusch–Godfrey

Time domain	Autocorrelation (ACF) partial (PACF) Cross-correlation (XCF) ARMA model ARIMA model (Box–Jenkins) Autoregressive conditional heteroskedasticity (ARCH) Vector autoregression (VAR)

Frequency domain	Spectral density estimation Fourier analysis Wavelet

Survival

Survival function	Kaplan–Meier estimator (product limit) Proportional hazards models Accelerated failure time (AFT) model First hitting time

Hazard function	Nelson–Aalen estimator

Test	Log-rank test

Applications

Biostatistics	Bioinformatics Clinical trials / studies Epidemiology Medical statistics

Engineering statistics	Chemometrics Methods engineering Probabilistic design Process / quality control Reliability System identification

Social statistics	Actuarial science Census Crime statistics Demography Econometrics National accounts Official statistics Population statistics Psychometrics

Spatial statistics	Cartography Environmental statistics Geographic information system Geostatistics Kriging

Category
Portal
Commons
WikiProject

This article is issued from Wikipedia - version of the 3/22/2016. The text is available under the Creative Commons Attribution/Share Alike but additional terms may apply for the media files.

Jarque–Bera test

History

Jarque–Bera test in regression analysis

Implementations

References

Further reading