TEXTPROFILE Procedure

_STATS Data Table

The _STATS data table is one of the output tables in the library that is specified by the OUTTOLIBRARY= option that contain overall summary statistics on information complexity, information density, vocabulary diversity, and domain specificity. The actual name of this table is constructed by appending _STATS to the DATA= option. For example, if you have the options DATA=ABC and OUTTOLIBRARY=DEF, then the table is created as DEF.ABC_STATS. Table 16 shows the variables in this data table.

Table 16: Variables in the _STATS Data Table

Variable Description
_CORPUS_NAME_ Data set name
_TOTAL_SENTENCES_ Total number of sentences in the data set
_AVG_SENTENCES_DOC_ Average number of sentences per document
_MAX_SENTENCES_DOC_ Number of sentences in the longest document by sentence count
_AVG_TOKENS_SENTENCE_ Average number of tokens per sentence
_MAX_TOKENS_SENTENCE_ Number of tokens in the longest sentence by token count
_TOTAL_TOKENS_ Total number of tokens in the data set
_AVG_TOKEN_LEN_ Average number of characters (or bytes for some languages) per token (all non-unique tokens counted)
_MAX_TOKEN_LEN_ Number of characters or bytes in the longest token
_TOTAL_FORMS_ Number of unique tokens in the data set
_FORM_80_PERCENT_ Number of forms (unique tokens) to account for 80% of the data
_PERCENT_CONTENT_TOKENS_ Percentage of tokens that are content (does not include numbers, stop words, or punctuation marks)
_PERCENT_STOP_TOKENS_ Percentage of tokens that are stop words
_PERCENT_NUM_TOKENS_ Percentage of tokens that contain a number or digit
_PERCENT_PUNCT_TOKENS_ Percentage of tokens that are punctuation marks


If you omit the NOREFERENCE option in the OUTPUT statement, reference statistics are included in the output.[6]



[6] These reference statistics are available in the Chinese, Dutch, English, French, Japanese, and Polish languages.

Last updated: January 14, 2026