TEXTPROFILE Procedure
_STATS Data Table
The _STATS data table is one of the output tables in the library that is specified by the OUTTOLIBRARY= option that contain overall summary statistics on information complexity, information density, vocabulary diversity, and domain specificity. The actual name of this table is constructed by appending _STATS to the DATA= option. For example, if you have the options DATA=ABC and OUTTOLIBRARY=DEF, then the table is created as DEF.ABC_STATS. Table 16 shows the variables in this data table.
Table 16: Variables in the _STATS Data Table
| Variable | Description |
|---|---|
_CORPUS_NAME_ | Data set name |
_TOTAL_SENTENCES_ | Total number of sentences in the data set |
_AVG_SENTENCES_DOC_ | Average number of sentences per document |
_MAX_SENTENCES_DOC_ | Number of sentences in the longest document by sentence count |
_AVG_TOKENS_SENTENCE_ | Average number of tokens per sentence |
_MAX_TOKENS_SENTENCE_ | Number of tokens in the longest sentence by token count |
_TOTAL_TOKENS_ | Total number of tokens in the data set |
_AVG_TOKEN_LEN_ | Average number of characters (or bytes for some languages) per token (all non-unique tokens counted) |
_MAX_TOKEN_LEN_ | Number of characters or bytes in the longest token |
_TOTAL_FORMS_ | Number of unique tokens in the data set |
_FORM_80_PERCENT_ | Number of forms (unique tokens) to account for 80% of the data |
_PERCENT_CONTENT_TOKENS_ | Percentage of tokens that are content (does not include numbers, stop words, or punctuation marks) |
_PERCENT_STOP_TOKENS_ | Percentage of tokens that are stop words |
_PERCENT_NUM_TOKENS_ | Percentage of tokens that contain a number or digit |
_PERCENT_PUNCT_TOKENS_ | Percentage of tokens that are punctuation marks |
If you omit the NOREFERENCE option in the OUTPUT statement, reference statistics are included in the output.[6]
[6] These reference statistics are available in the Chinese, Dutch, English, French, Japanese, and Polish languages.