CARDINALITY Procedure
Details: CARDINALITY Procedure
The CARDINALITY procedure extracts from a specified data table the distinct values of its variables up to a specified maximum number of levels. The maximum number of levels that you specify cannot exceed 254. The number of distinct levels of a variable is called its cardinality.
You need this information in a data exploration phase in order to understand the contents of big data and the possible role that a variable can play in subsequent analyses.
A limited cardinality study is sufficient to determine the variable roles—class, interval, or ID (a record identifier)—on the basis of the distinct values that are observed from the data. A full cardinality study might be computationally prohibitive, because it requires a lot of memory and resources. You can arrive at the same conclusion by running a limited cardinality analysis on the data.
The limited cardinality study reads and processes data until the number of levels reaches the value that you specify in the MAXLEVELS= option. All the levels that are higher than the value of this option are grouped together into one level that has a missing value (.) in the _INDEX_ column in the output details data table. All variables whose visibility is 100 are identified as having a class role. Any nonnumeric variable whose number of levels is greater than the MAXLEVELS= option value is identified as having an ID role. A numeric variable whose number of levels is greater than the MAXLEVELS= option value is identified as having either an interval or ID role, depending on whether all the processed data of the variable are unique or not. If all the processed data of the variable are unique, then the variable is identified as having an ID role; otherwise it is identified as having an interval role.