The TEXTMINE Procedure

SVD Statement

  • SVD <svd-options>;

The SVD statement specifies the options for calculating a truncated singular value decomposition (SVD) of the large, sparse term-by-document matrix that is created during the parsing phase of PROC TEXTMINE. Table 4 summarizes the svd-options in the statement by function. The svd-options are then described fully in alphabetical order.

Table 4: SVD Statement Options

svd-option Description
Input Options
COL= Specifies the column variable, which contains the column indices of the term-by-document matrix, which is stored in coordinate list (COO) format
ROW= Specifies the row variable, which contains the row indices of the term-by-document matrix, which is stored in COO format
ENTRY= Specifies the entry variable, which contains the entries of the term-by-document matrix, which is stored in COO format
SVD Computation Options
K= Specifies the number of dimensions to be extracted
MAX_K= Specifies the maximum number of dimensions to be extracted
TOL= Specifies the maximum allowable tolerance for the singular value
RESOLUTION | RES= Specifies the recommended number of dimensions (resolution) to be extracted by SVD, when the MAX_K= option is specified
Topic Discovery Options
NUMLABELS= Specifies the number of terms to be used in the descriptive label for each topic
ROTATION= Specifies the type of rotation to be used for topic discovery
IN_TERMS= Specifies the data table that contains the terms for topic discovery in SVD-only mode
EXACTWEIGHT Prevents rounding of the topic weights
NOCUTOFFS Prevents setting term weights to 0 when they are below the threshold
Output Options
SVDU= Specifies the U matrix, which contains the left singular vectors
SVDV= Specifies the V matrix, which contains the right singular vectors
SVDS= Specifies the S matrix, whose diagonal elements are the singular values
OUTDOCPRO= Specifies the data table to contain the projections of the documents
OUTTOPICS= Specifies the data table to contain the topics that have been discovered


You can specify the following svd-options:

COL=variable

specifies the variable that contains the column indices of the term-by-document matrix. You must specify this option when you run PROC TEXTMINE in SVD-only mode (that is, when you specify the SVD statement but not the PARSE statement).

ENTRY=variable

specifies the variable that contains the entries of the term-by-document matrix. You must specify this option when you run PROC TEXTMINE in SVD-only mode (that is, when you specify the SVD statement but not the PARSE statement).

EXACTWEIGHT

requests that the weights aggregated during topic derivation not be rounded. By default, the calculated weights are rounded to the nearest 0.001.

IN_TERMS=CAS-libref.data-table

specifies the input data table that contains information about the terms in the document collection. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The data table should have the variables that are described in Table 11. The terms are required to generate topic names in the OUTTOPICS= data table. This option is only for topic discovery in SVD-only mode. This option conflicts with the PARSE statement, and only one of the two can be specified. If you want to run SVD-only mode without topic discovery, then you do not need to specify this option.

K=k

specifies the number of columns in the matrices bold upper U, bold upper V, and bold upper S. This value is the number of dimensions of the data table after SVD is performed. If the value of k is too large, then the TEXTMINE procedure runs for an unnecessarily long time. This option takes precedence over the MAX_K= option. This option also controls the number of topics that are extracted from the text corpus when the ROTATION= option is specified.

MAX_K=n

specifies the maximum value that the TEXTMINE procedure should return as the recommended value of k (the number of columns in the matrices bold upper U, bold upper V, and bold upper S) when the RESOLUTION= option is specified (to recommend the value of k). The TEXTMINE procedure attempts to calculate k dimensions (as opposed to recommending it) when it performs SVD. This option is ignored if the K= option has been specified. This option also controls the number of topics that are extracted from the text corpus when the ROTATION= option is specified.

NOCUTOFFS

uses all weights in the bold upper U matrix to form the document projections. When topics are requested, weights below the term cutoff (as calculated in the OUTTOPICS= data table) are set to 0 before the projection is formed.

NUMLABELS=n

specifies the number of terms to use in the descriptive label for each topic. The descriptive label provides a quick synopsis of the discovered topics. The labels are stored in the OUTTOPICS= data table. By default, NUMLABELS=5.

OUTDOCPRO=CAS-libref.data-table <KEEPVARIABLES=variable-list><NONORMDOC>
OUTDOCPRO=CAS-libref.data-table <KEEPVARS=variable-list><NONORMDOC>

specifies the output data table to contain the projections of the columns of the term-by-document matrix onto the columns of bold upper U. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. Because each column of the term-by-document matrix corresponds to a document, the output forms a new representation of the input documents in a space that has much lower dimensionality.

You can copy the variables from the data table that is specified in the DATA= option in the PROC TEXTMINE statement to the data table that is specified in this option. You can specify the following suboptions:

KEEPVARIABLES=variable-list

attaches the content of the variables that are specified in the variable-list to the output. These variables must appear in the data table that is specified in the DATA= option in the PROC TEXTMINE statement.

NONORMDOC

suppresses normalization of the columns that contain the projections of documents to have a unit norm.

OUTTOPICS=CAS-libref.data-table

specifies the output data table to contain the topics that are discovered. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

RESOLUTION=LOW | MED | HIGH
RES=LOW | MED | HIGH

specifies how to calculate the recommended number of dimensions (resolution) for the singular value decomposition. If you specify this option, you must also specify the MAX_K= option. A low-resolution singular value decomposition returns fewer dimensions than a high-resolution singular value decomposition. This option recommends the value of k (the number of columns in the matrices bold upper U, bold upper V, and bold upper S) heuristically based on the value specified in the MAX_K= option. Assume that the MAX_K= option is set to n and a singular value decomposition that has n dimensions accounts for t percent-sign of the total variance. You can specify the following values:

HIGH

always recommends the maximum number of dimensions; that is, k equalsn.

MED

recommends a k that explains left-parenthesis 5 slash 6 right-parenthesis asterisk t percent-sign of the total variance.

LOW

recommends a k that explains left-parenthesis 2 slash 3 right-parenthesis asterisk t percent-sign of the total variance.

By default, RESOLUTION=HIGH.

ROTATION=VARIMAX | PROMAX

specifies the type of rotation to be used in order to maximize the explanatory power of each topic. You can specify the following values:

PROMAX

does an oblique rotation on the original left singular vectors and generates topics that might be correlated.

VARIMAX

does an orthogonal rotation on the original left singular vectors and generates uncorrelated topics.

By default, ROTATION=VARIMAX.

ROW=variable

specifies the variable that contains the row indices of the term-by-document matrix. You must specify this option when you run PROC TEXTMINE in SVD-only mode (that is, when you specify the SVD statement but not the PARSE statement).

SVDS=CAS-libref.data-table

specifies the output data table to contain the calculated singular values. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

SVDU=CAS-libref.data-table

specifies the data table to contain the calculated left singular vectors. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

SVDV=CAS-libref.data-table

specifies the data table to contain the calculated right singular vectors. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs.

TOL=epsilon

specifies the maximum allowable tolerance for the singular value. Let bold upper A be a matrix. Suppose lamda Subscript i is the ith singular value of bold upper A and xi Subscript i is the corresponding right singular vector. The SVD computation terminates when for all i element-of StartSet 1 comma ellipsis comma k EndSet, lamda Subscript i and xi Subscript i satisfy double-vertical-bar bold upper A Superscript down-tack Baseline bold upper A xi minus lamda xi double-vertical-bar squared less-than-or-equal-to epsilon. The default value of epsilon is 10 Superscript negative 6, which is more than adequate for most text mining problems.

Last updated: November 11, 2020