The PROC TMSCORE statement invokes the procedure. Table 1 summarizes the options in the statement by function. The options are then described fully in alphabetical order.
Specifies the data table that contains the term-by-document frequency matrix that is used to model the document collection. In this matrix, the child terms are not represented and child terms’ frequencies are attributed to their corresponding parents.
names the input data table for PROC TMSCORE to use. CAS-libref.data-table is a two-level name, where
CAS-libref
refers to a collection of information that is defined in the LIBNAME statement and includes the caslib, which includes a path to the data, and a session identifier, which defaults to the active session but which can be explicitly defined in the LIBNAME statement. For more information about CAS-libref, see the section Using CAS Sessions and CAS Engine Librefs.
data-table
specifies the name of the input data table.
The input data table contains documents for PROC TMSCORE to score. Each row of the input data table must contain one text variable and one ID variable, which correspond to the text and the unique ID of a document, respectively.
You can also specify the following options:
CONFIG=CAS-libref.data-table
specifies the input data table that contains configuration information for PROC TMSCORE. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. Specify the table that was generated by the OUTCONFIG= option in the PARSE statement of the TEXTMINE procedure during training. For more information about this data table, see the section The OUTCONFIG= Data Table of Chapter 24, The TEXTMINE Procedure.
OUTPARENT=CAS-libref.data-table
specifies the output data table to contain a compressed representation of the sparse term-by-document frequency matrix. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The data table contains only the kept representative terms, and the child frequencies are attributed to the corresponding parent. For more information about the compressed representation of the sparse term-by-document frequency matrix, see the section The OUTPARENT= Data Table of Chapter 24, The TEXTMINE Procedure.
SVDDOCPRO=CAS-libref.data-table
specifies the output data table to contain the reduced dimensional projections for each document. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the output data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The contents of this data table are formed by multiplying the term-by-document frequency matrix by the input data table that is specified in the SVDU= option and then normalizing the result.
SVDU=CAS-libref.data-table
specifies the input data table that contains the matrix, which is created during training by PROC TEXTMINE. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. The data table contains the information that is needed to project each document into the reduced dimensional space. For more information about the contents of this data table, see the SVDU= option in Chapter 24, The TEXTMINE Procedure.
TERMS=CAS-libref.data-table
specifies the input data table of terms to be used by PROC TMSCORE. CAS-libref.data-table is a two-level name, where CAS-libref refers to the caslib and session identifier, and data-table specifies the name of the input data table. For more information about this two-level name, see the DATA= option and the section Using CAS Sessions and CAS Engine Librefs. Specify the table that was generated by the OUTTERMS= option in the PARSE statement of the TEXTMINE procedure during training. This data table conveys to PROC TMSCORE which terms should be used in the analysis and whether they should be mapped to a parent. The data table also assigns to each term a key that corresponds to the key that is used in the input data table that is specified by the SVDU= option. For more information about this data table, see the section The OUTTERMS= Data Table of Chapter 24, The TEXTMINE Procedure.