The SVD statement specifies the options for calculating a truncated singular value decomposition (SVD) of the large, sparse term-by-document matrix that is created during the parsing phase of PROC TEXTMINE. Table 4 summarizes the svd-options in the statement by function. The svd-options are then described fully in alphabetical order.
Specifies the data table to contain the topics that have been discovered
You can specify the following svd-options:
COL=variable
specifies the variable that contains the column indices of the term-by-document matrix. You must specify this option when you run PROC TEXTMINE in SVD-only mode (that is, when you specify the SVD statement but not the PARSE statement).
ENTRY=variable
specifies the variable that contains the entries of the term-by-document matrix. You must specify this option when you run PROC TEXTMINE in SVD-only mode (that is, when you specify the SVD statement but not the PARSE statement).
EXACTWEIGHT
requests that the weights aggregated during topic derivation not be rounded. By default, the calculated weights are rounded to the nearest 0.001.
IN_TERMS=CAS-libref.data-table
specifies the input data table that contains information about the terms in the document collection. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The data table should have the variables that are described in Table 11. The terms are required to generate topic names in the OUTTOPICS= data table. This option is only for topic discovery in SVD-only mode. This option conflicts with the PARSE statement, and only one of the two can be specified. If you want to run SVD-only mode without topic discovery, then you do not need to specify this option.
K=k
specifies the number of columns in the matrices , , and . This value is the number of dimensions of the data table after SVD is performed. If the value of k is too large, then the TEXTMINE procedure runs for an unnecessarily long time. This option takes precedence over the MAX_K= option. This option also controls the number of topics that are extracted from the text corpus when the ROTATION= option is specified.
MAX_K=n
specifies the maximum value that the TEXTMINE procedure should return as the recommended value of k (the number of columns in the matrices , , and ) when the RESOLUTION= option is specified (to recommend the value of k). The TEXTMINE procedure attempts to calculate k dimensions (as opposed to recommending it) when it performs SVD. This option is ignored if the K= option has been specified. This option also controls the number of topics that are extracted from the text corpus when the ROTATION= option is specified.
NOCUTOFFS
uses all weights in the matrix to form the document projections. When topics are requested, weights below the term cutoff (as calculated in the OUTTOPICS= data table) are set to 0 before the projection is formed.
NUMLABELS=n
specifies the number of terms to use in the descriptive label for each topic. The descriptive label provides a quick synopsis of the discovered topics. The labels are stored in the OUTTOPICS= data table. By default, NUMLABELS=5.
specifies the output data table to contain the projections of the columns of the term-by-document matrix onto the columns of . CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. Because each column of the term-by-document matrix corresponds to a document, the output forms a new representation of the input documents in a space that has much lower dimensionality.
You can copy the variables from the data table that is specified in the DATA= option in the PROC TEXTMINE statement to the data table that is specified in this option. You can specify the following suboptions:
KEEPVARIABLES=variable-list
attaches the content of the variables that are specified in the variable-list to the output. These variables must appear in the data table that is specified in the DATA= option in the PROC TEXTMINE statement.
NONORMDOC
suppresses normalization of the columns that contain the projections of documents to have a unit norm.
OUTTOPICS=CAS-libref.data-table
specifies the output data table to contain the topics that are discovered. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
RESOLUTION=LOW | MED | HIGH
RES=LOW | MED | HIGH
specifies how to calculate the recommended number of dimensions (resolution) for the singular value decomposition. If you specify this option, you must also specify the MAX_K= option. A low-resolution singular value decomposition returns fewer dimensions than a high-resolution singular value decomposition. This option recommends the value of k (the number of columns in the matrices , , and ) heuristically based on the value specified in the MAX_K= option. Assume that the MAX_K= option is set to n and a singular value decomposition that has n dimensions accounts for of the total variance. You can specify the following values:
HIGH
always recommends the maximum number of dimensions; that is, n.
MED
recommends a k that explains of the total variance.
LOW
recommends a k that explains of the total variance.
By default, RESOLUTION=HIGH.
ROTATION=VARIMAX | PROMAX
specifies the type of rotation to be used in order to maximize the explanatory power of each topic. You can specify the following values:
PROMAX
does an oblique rotation on the original left singular vectors and generates topics that might be correlated.
VARIMAX
does an orthogonal rotation on the original left singular vectors and generates uncorrelated topics.
By default, ROTATION=VARIMAX.
ROW=variable
specifies the variable that contains the row indices of the term-by-document matrix. You must specify this option when you run PROC TEXTMINE in SVD-only mode (that is, when you specify the SVD statement but not the PARSE statement).
SVDS=CAS-libref.data-table
specifies the output data table to contain the calculated singular values. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
SVDU=CAS-libref.data-table
specifies the data table to contain the calculated left singular vectors. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
SVDV=CAS-libref.data-table
specifies the data table to contain the calculated right singular vectors. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.
TOL=
specifies the maximum allowable tolerance for the singular value. Let be a matrix. Suppose is the ith singular value of and is the corresponding right singular vector. The SVD computation terminates when for all , and satisfy . The default value of is , which is more than adequate for most text mining problems.