CNTSELECT Procedure

Negative Binomial Regression

The Poisson regression model can be generalized by introducing an unobserved heterogeneity term for observation i. Thus, the individuals are assumed to differ randomly in a manner that is not fully accounted for by the observed covariates. This is formulated as

upper E left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline comma tau Subscript i Baseline right-parenthesis equals mu Subscript i Baseline tau Subscript i Baseline equals e Superscript bold x prime Super Subscript i Superscript bold-italic beta plus epsilon Super Subscript i

where the unobserved heterogeneity term tau Subscript i Baseline equals e Superscript epsilon Super Subscript i is independent of the vector of regressors bold x Subscript i. Then the distribution of y Subscript i conditional on bold x Subscript i and tau Subscript i is Poisson with conditional mean and conditional variance mu Subscript i Baseline tau Subscript i:

f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline comma tau Subscript i Baseline right-parenthesis equals StartFraction exp left-parenthesis minus mu Subscript i Baseline tau Subscript i Baseline right-parenthesis left-parenthesis mu Subscript i Baseline tau Subscript i Baseline right-parenthesis Superscript y Super Subscript i Superscript Baseline Over y Subscript i Baseline factorial EndFraction

Let g left-parenthesis tau Subscript i Baseline right-parenthesis be the probability density function of tau Subscript i. Then, the distribution f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis (no longer conditional on tau Subscript i) is obtained by integrating f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline comma tau Subscript i Baseline right-parenthesis with respect to tau Subscript i:

f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis equals integral Subscript 0 Superscript normal infinity Baseline f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline comma tau Subscript i Baseline right-parenthesis g left-parenthesis tau Subscript i Baseline right-parenthesis d tau Subscript i Baseline

An analytical solution to this integral exists when tau Subscript i is assumed to follow a gamma distribution. This solution is the negative binomial distribution. If the model contains a constant term, then in order to identify the mean of the distribution, it is necessary to assume that upper E left-parenthesis e Superscript epsilon Super Subscript i Superscript Baseline right-parenthesis equals upper E left-parenthesis tau Subscript i Baseline right-parenthesis equals 1. Thus, it is assumed that tau Subscript i follows a gamma(theta comma theta) distribution with upper E left-parenthesis tau Subscript i Baseline right-parenthesis equals 1 and upper V left-parenthesis tau Subscript i Baseline right-parenthesis equals 1 slash theta,

g left-parenthesis tau Subscript i Baseline right-parenthesis equals StartFraction theta Superscript theta Baseline Over normal upper Gamma left-parenthesis theta right-parenthesis EndFraction tau Subscript i Superscript theta minus 1 Baseline exp left-parenthesis minus theta tau Subscript i Baseline right-parenthesis

where normal upper Gamma left-parenthesis x right-parenthesis equals integral Subscript 0 Superscript normal infinity Baseline z Superscript x minus 1 Baseline exp left-parenthesis negative z right-parenthesis d z is the gamma function and theta is a positive parameter. Then, the density of y Subscript i given bold x Subscript i is derived as

StartLayout 1st Row 1st Column f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis 2nd Column equals 3rd Column integral Subscript 0 Superscript normal infinity Baseline f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline comma tau Subscript i Baseline right-parenthesis g left-parenthesis tau Subscript i Baseline right-parenthesis d tau Subscript i 2nd Row 1st Column Blank 2nd Column equals 3rd Column StartFraction theta Superscript theta Baseline mu Subscript i Superscript y Super Subscript i Superscript Baseline Over y Subscript i Baseline factorial normal upper Gamma left-parenthesis theta right-parenthesis EndFraction integral Subscript 0 Superscript normal infinity Baseline e Superscript minus left-parenthesis mu Super Subscript i Superscript plus theta right-parenthesis tau Super Subscript i Baseline tau Subscript i Superscript theta plus y Super Subscript i minus 1 Baseline d tau Subscript i 3rd Row 1st Column Blank 2nd Column equals 3rd Column StartFraction theta Superscript theta Baseline mu Subscript i Superscript y Super Subscript i Superscript Baseline normal upper Gamma left-parenthesis y Subscript i Baseline plus theta right-parenthesis Over y Subscript i Baseline factorial normal upper Gamma left-parenthesis theta right-parenthesis left-parenthesis theta plus mu Subscript i Baseline right-parenthesis Superscript theta plus y Super Subscript i Superscript Baseline EndFraction 4th Row 1st Column Blank 2nd Column equals 3rd Column StartFraction normal upper Gamma left-parenthesis y Subscript i Baseline plus theta right-parenthesis Over y Subscript i Baseline factorial normal upper Gamma left-parenthesis theta right-parenthesis EndFraction left-parenthesis StartFraction theta Over theta plus mu Subscript i Baseline EndFraction right-parenthesis Superscript theta Baseline left-parenthesis StartFraction mu Subscript i Baseline Over theta plus mu Subscript i Baseline EndFraction right-parenthesis Superscript y Super Subscript i EndLayout

If you make the substitution alpha equals StartFraction 1 Over theta EndFraction (alpha greater-than 0), the negative binomial distribution can then be rewritten as

f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis equals StartFraction normal upper Gamma left-parenthesis y Subscript i Baseline plus alpha Superscript negative 1 Baseline right-parenthesis Over y Subscript i Baseline factorial normal upper Gamma left-parenthesis alpha Superscript negative 1 Baseline right-parenthesis EndFraction left-parenthesis StartFraction alpha Superscript negative 1 Baseline Over alpha Superscript negative 1 Baseline plus mu Subscript i Baseline EndFraction right-parenthesis Superscript alpha Super Superscript negative 1 Superscript Baseline left-parenthesis StartFraction mu Subscript i Baseline Over alpha Superscript negative 1 Baseline plus mu Subscript i Baseline EndFraction right-parenthesis Superscript y Super Subscript i Superscript Baseline comma y Subscript i Baseline equals 0 comma 1 comma 2 comma ellipsis

Thus, the negative binomial distribution is derived as a gamma mixture of Poisson random variables. It has the conditional mean

upper E left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis equals mu Subscript i Baseline equals e Superscript bold x prime Super Subscript i Superscript bold-italic beta

and the conditional variance

upper V left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis equals mu Subscript i Baseline left-bracket 1 plus StartFraction 1 Over theta EndFraction mu Subscript i Baseline right-bracket equals mu Subscript i Baseline left-bracket 1 plus alpha mu Subscript i Baseline right-bracket greater-than upper E left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis

The conditional variance of the negative binomial distribution exceeds the conditional mean. Overdispersion results from neglected unobserved heterogeneity. The negative binomial model with variance function upper V left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis equals mu Subscript i Baseline plus alpha mu Subscript i Superscript 2, which is quadratic in the mean, is referred to as the NEGBIN2 model Cameron and Trivedi (1986). To estimate this model, specify DIST=NEGBIN(P=2) in the MODEL statement. The Poisson distribution is a special case of the negative binomial distribution where alpha equals 0. A test of the Poisson distribution can be carried out by testing the hypothesis that alpha equals StartFraction 1 Over theta Subscript i Baseline EndFraction equals 0. A Wald test of this hypothesis is provided (it is the reported t statistic for the estimated alpha in the negative binomial model).

The log-likelihood function of the negative binomial regression model (NEGBIN2) is given by

StartLayout 1st Row 1st Column script upper L 2nd Column equals 3rd Column sigma-summation Underscript i equals 1 Overscript upper N Endscripts left-brace sigma-summation Underscript j equals 0 Overscript y Subscript i Baseline minus 1 Endscripts ln left-parenthesis j plus alpha Superscript negative 1 Baseline right-parenthesis minus ln left-parenthesis y Subscript i Baseline factorial right-parenthesis 2nd Row 1st Column Blank 2nd Column Blank 3rd Column minus left-parenthesis y Subscript i Baseline plus alpha Superscript negative 1 Baseline right-parenthesis ln left-parenthesis 1 plus alpha exp left-parenthesis bold x prime Subscript i Baseline bold-italic beta right-parenthesis right-parenthesis plus y Subscript i Baseline ln left-parenthesis alpha right-parenthesis plus y Subscript i Baseline bold x prime Subscript i Baseline bold-italic beta right-brace EndLayout

where use of the following fact is made if y is an integer:

normal upper Gamma left-parenthesis y plus a right-parenthesis slash normal upper Gamma left-parenthesis a right-parenthesis equals product Underscript j equals 0 Overscript y minus 1 Endscripts left-parenthesis j plus a right-parenthesis

Cameron and Trivedi (1986) consider a general class of negative binomial models that have mean mu Subscript i and variance function mu Subscript i Baseline plus alpha mu Subscript i Superscript p. The NEGBIN2 model, with p equals 2, is the standard formulation of the negative binomial model. Models that have other values of p, negative normal infinity less-than p less-than normal infinity, have the same density f left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis, except that alpha Superscript negative 1 is replaced everywhere by alpha Superscript negative 1 Baseline mu Superscript 2 minus p. The negative binomial model NEGBIN1, which sets p equals 1, has the variance function upper V left-parenthesis y Subscript i Baseline vertical-bar bold x Subscript i Baseline right-parenthesis equals mu Subscript i Baseline plus alpha mu Subscript i, which is linear in the mean. To estimate this model, specify DIST=NEGBIN(P=1) in the MODEL statement.

The log-likelihood function of the NEGBIN1 regression model is given by

StartLayout 1st Row 1st Column script upper L 2nd Column equals 3rd Column sigma-summation Underscript i equals 1 Overscript upper N Endscripts left-brace sigma-summation Underscript j equals 0 Overscript y Subscript i Baseline minus 1 Endscripts ln left-parenthesis j plus alpha Superscript negative 1 Baseline exp left-parenthesis bold x prime Subscript i Baseline bold-italic beta right-parenthesis right-parenthesis 2nd Row 1st Column Blank 2nd Column Blank 3rd Column minus ln left-parenthesis y Subscript i Baseline factorial right-parenthesis minus left-parenthesis y Subscript i Baseline plus alpha Superscript negative 1 Baseline exp left-parenthesis bold x prime Subscript i Baseline bold-italic beta right-parenthesis right-parenthesis ln left-parenthesis 1 plus alpha right-parenthesis plus y Subscript i Baseline ln left-parenthesis alpha right-parenthesis right-brace EndLayout

For more information about the negative binomial regression model, see SAS/ETS User's Guide.

Last updated: July 09, 2026