Text Mining Action Set: Syntax

Provides actions for mining textual data

tmSvd Action

Computes the SVD factorization and generates topics. This action requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

CASL Syntax

textMining.tmSvd <result=results> <status=rc> /
config={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
}
count="variable-name"
docId="variable-name"
docPro={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
exactDocPro=TRUE | FALSE
exactWeight=TRUE | FALSE
k=integer
legacyNames=TRUE | FALSE
maxK=integer
nThreads=integer
norm="ALL" | "DOC" | "NONE" | "WORD"
numLabels=integer
required parameter parent={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
}
resolution="HIGH" | "LOW" | "MED"
rotate="PROMAX" | "VARIMAX"
rowPivot=double
s={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
scoreConfig={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
termId="variable-name"
termTopics={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
terms={
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=TRUE | FALSE,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
}
tolerance=double
topicDecision=TRUE | FALSE
topics={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
u={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
v={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
wordPro={
caslib="string",
compress=TRUE | FALSE,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
;

Parameter Descriptions

config={castable}

specifies the name of the input CAS table that contains parsing configuration information

AliasparseConfig
Long formconfig={name="table-name"}
Shortcut formconfig="table-name"
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

count="variable-name"

specifies the variable that contains the, possibly weighted, term count. The values in this variable must be numeric. There can be no missing values in this variable.

Default"_COUNT_"

docId="variable-name"

specifies the variable that contains the document ID. The type of this variable can either be numeric or a string. There can be no missing values in this variable.

Default"_DOCUMENT_"

docPro={casouttable}

specifies the name of the table to contain the SVD projections of the documents.

For more information about specifying the docPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

docStdMultiple=double

Specifies how many standard deviations above the mean to set the document cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

exactDocPro=TRUE | FALSE

Specifies if the exact document projection values should be output. This parameter requires a SAS Visual Text Analytics license.

DefaultTRUE

exactWeight=TRUE | FALSE

AliasexactWeights
DefaultFALSE

k=integer

specifies the number of dimensions to be extracted (also the number of derived topics). If the input data is too small for the requested number of dimensions, this value is adjusted to complete the calculation.

AliasnumTopics
Range1–1000

legacyNames=TRUE | FALSE

specifies whether to use the legacy variable names on tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

DefaultFALSE

maxK=integer

specifies the maximum number of dimensions to be extracted. The maxK option can be used in conjunction with the resolution option to dynamically select the recommended number of dimensions. If you wish to use a specific number of dimensions use maxK and set the resolution to high, or use the k parameter.

Default10
Range1–1000

nThreads=integer

specifies number of threads to be used per node. If not set, or if a value of 0 is specified, all available threads will be used.

Default8
Range0–64

norm="ALL" | "DOC" | "NONE" | "WORD"

indicates whether the document projections, term projections, or both are normalized. The normalization converts the representation from depending on angles between vectors to one based on Euclidean distances between vectors.

DefaultALL

numLabels=integer

specifies the number of terms to use in the descriptive label for each topic.

Default5
Range1–500

* parent={castable}

specifies the input CAS table that contains the term-by-document matrix in transaction form. The table must have at last three variables, one containing the document id, a second containing the term id, and the third containing the value in the cell corresponding to that particular term and document.

Long formparent={name="table-name"}
Shortcut formparent="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

resolution="HIGH" | "LOW" | "MED"

specifies the desired resolution level for the recommended number of dimensions to be extracted by the SVD.

DefaultHIGH

rotate="PROMAX" | "VARIMAX"

specifies the type of rotation used to maximize the explanatory power of each topic. A VARIMAX rotation produces uncorrelated topics and a PROMAX rotation produces correlated topics.

DefaultVARIMAX

rowPivot=double

Specifies the row pivot weight for document normalization of the parent table before the SVD. A negative value turns off the rowPivot process. When topics are requested a rowPivot=1 value is used by default. This parameter requires a SAS Visual Text Analytics license.

Default-1
Range-1–1

s={casouttable}

specifies the S matrix, which is a diagonal matrix that is output in compressed form, with two variables and k rows. The variable _ID_ indicates the row and column of the entry and the variable S contains the singular values.

For more information about specifying the s parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scoreConfig={casouttable}

Specifies the output scoring config file.

For more information about specifying the scoreConfig parameter, see the common casouttable parameter (Appendix A: Common Parameters).

termId="variable-name"

specifies the variable that contains the term ID. The contents of this variable must be an integer greater than or equal to 1. There can be no missing values in this variable.

Default"_TERMNUM_"

termStdMultiple=double

Specifies how many standard deviations above the mean to set the term cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

termTopics={casouttable}

specifies the name of the output CAS table to contain the term-by-topic sparse matrix information.

For more information about specifying the termTopics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

terms={castable}

specifies the name of the input table that contains information about the terms in the document collection. The table is used to determine which terms to use in the topic calculation.

Long formterms={name="table-name"}
Shortcut formterms="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the stopping threshold for the iterative factorization algorithm. If 0 is specified the default value is used.

Default1e-06
Range0–1

topicDecision=TRUE | FALSE

Specifies to include topic membership decisions and document cutoffs in the output tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

DefaultFALSE

topics={casouttable}

specifies the output CAS table to contain the topics that are discovered.

For more information about specifying the topics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

u={casouttable}

specifies the U matrix, which contains the left singular vectors. The matrix U is number of terms by k+1.

For more information about specifying the u parameter, see the common casouttable parameter (Appendix A: Common Parameters).

v={casouttable}

specifies the transpose of the matrix containing the right singular vectors. The matrix V is number of documents by k+1.

For more information about specifying the v parameter, see the common casouttable parameter (Appendix A: Common Parameters).

wordPro={casouttable}

specifies the table to contain the projections of the terms. If k dimensions of the SVD are found and the input data set contains n terms, this table will have n rows and k+1 columns.

For more information about specifying the wordPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

tmSvd Action

Computes the SVD factorization and generates topics. This action requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

Lua Syntax

results, info = s:textMining_tmSvd{
config={
caslib="string",
computedOnDemand=true | false,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
},
count="variable-name",
docId="variable-name",
docPro={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
exactDocPro=true | false,
exactWeight=true | false,
k=integer,
legacyNames=true | false,
maxK=integer,
nThreads=integer,
norm="ALL" | "DOC" | "NONE" | "WORD",
numLabels=integer,
required parameter parent={
caslib="string",
computedOnDemand=true | false,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
},
resolution="HIGH" | "LOW" | "MED",
rotate="PROMAX" | "VARIMAX",
rowPivot=double,
s={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
scoreConfig={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
termId="variable-name",
termTopics={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
terms={
caslib="string",
computedOnDemand=true | false,
computedVars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>},
groupBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter name="table-name",
onDemand=true | false,
orderBy={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
singlePass=true | false,
vars={{
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
}, {...}},
where="where-expression"
},
tolerance=double,
topicDecision=true | false,
topics={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
u={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
v={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
},
wordPro={
caslib="string",
compress=true | false,
indexVars={"variable-name-1" <, "variable-name-2", ...>},
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=true | false,
promote=true | false,
replace=true | false,
replication=integer,
timeStamp="string",
where={"string-1" <, "string-2", ...>}
}
}

Parameter Descriptions

config={castable}

specifies the name of the input CAS table that contains parsing configuration information

AliasparseConfig
Long formconfig={name="table-name"}
Shortcut formconfig="table-name"
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=true | false

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
Defaultfalse
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=true | false

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

Defaultfalse
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

count="variable-name"

specifies the variable that contains the, possibly weighted, term count. The values in this variable must be numeric. There can be no missing values in this variable.

Default"_COUNT_"

docId="variable-name"

specifies the variable that contains the document ID. The type of this variable can either be numeric or a string. There can be no missing values in this variable.

Default"_DOCUMENT_"

docPro={casouttable}

specifies the name of the table to contain the SVD projections of the documents.

For more information about specifying the docPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

docStdMultiple=double

Specifies how many standard deviations above the mean to set the document cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

exactDocPro=true | false

Specifies if the exact document projection values should be output. This parameter requires a SAS Visual Text Analytics license.

Defaulttrue

exactWeight=true | false

AliasexactWeights
Defaultfalse

k=integer

specifies the number of dimensions to be extracted (also the number of derived topics). If the input data is too small for the requested number of dimensions, this value is adjusted to complete the calculation.

AliasnumTopics
Range1–1000

legacyNames=true | false

specifies whether to use the legacy variable names on tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

Defaultfalse

maxK=integer

specifies the maximum number of dimensions to be extracted. The maxK option can be used in conjunction with the resolution option to dynamically select the recommended number of dimensions. If you wish to use a specific number of dimensions use maxK and set the resolution to high, or use the k parameter.

Default10
Range1–1000

nThreads=integer

specifies number of threads to be used per node. If not set, or if a value of 0 is specified, all available threads will be used.

Default8
Range0–64

norm="ALL" | "DOC" | "NONE" | "WORD"

indicates whether the document projections, term projections, or both are normalized. The normalization converts the representation from depending on angles between vectors to one based on Euclidean distances between vectors.

DefaultALL

numLabels=integer

specifies the number of terms to use in the descriptive label for each topic.

Default5
Range1–500

* parent={castable}

specifies the input CAS table that contains the term-by-document matrix in transaction form. The table must have at last three variables, one containing the document id, a second containing the term id, and the third containing the value in the cell corresponding to that particular term and document.

Long formparent={name="table-name"}
Shortcut formparent="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=true | false

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
Defaultfalse
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=true | false

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

Defaultfalse
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

resolution="HIGH" | "LOW" | "MED"

specifies the desired resolution level for the recommended number of dimensions to be extracted by the SVD.

DefaultHIGH

rotate="PROMAX" | "VARIMAX"

specifies the type of rotation used to maximize the explanatory power of each topic. A VARIMAX rotation produces uncorrelated topics and a PROMAX rotation produces correlated topics.

DefaultVARIMAX

rowPivot=double

Specifies the row pivot weight for document normalization of the parent table before the SVD. A negative value turns off the rowPivot process. When topics are requested a rowPivot=1 value is used by default. This parameter requires a SAS Visual Text Analytics license.

Default-1
Range-1–1

s={casouttable}

specifies the S matrix, which is a diagonal matrix that is output in compressed form, with two variables and k rows. The variable _ID_ indicates the row and column of the entry and the variable S contains the singular values.

For more information about specifying the s parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scoreConfig={casouttable}

Specifies the output scoring config file.

For more information about specifying the scoreConfig parameter, see the common casouttable parameter (Appendix A: Common Parameters).

termId="variable-name"

specifies the variable that contains the term ID. The contents of this variable must be an integer greater than or equal to 1. There can be no missing values in this variable.

Default"_TERMNUM_"

termStdMultiple=double

Specifies how many standard deviations above the mean to set the term cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

termTopics={casouttable}

specifies the name of the output CAS table to contain the term-by-topic sparse matrix information.

For more information about specifying the termTopics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

terms={castable}

specifies the name of the input table that contains information about the terms in the document collection. The table is used to determine which terms to use in the topic calculation.

Long formterms={name="table-name"}
Shortcut formterms="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=true | false

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
Defaultfalse
computedVars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions={key-1=any-list-or-data-type-1 <, key-2=any-list-or-data-type-2, ...>}

specifies data source options.

Aliasoptions, dataSource
groupBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions={fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=true | false

This parameter is deprecated.

Defaulttrue
orderBy={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=true | false

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

Defaultfalse
vars={{casinvardesc-1} <, {casinvardesc-2}, ...>}

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the stopping threshold for the iterative factorization algorithm. If 0 is specified the default value is used.

Default1e-06
Range0–1

topicDecision=true | false

Specifies to include topic membership decisions and document cutoffs in the output tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

Defaultfalse

topics={casouttable}

specifies the output CAS table to contain the topics that are discovered.

For more information about specifying the topics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

u={casouttable}

specifies the U matrix, which contains the left singular vectors. The matrix U is number of terms by k+1.

For more information about specifying the u parameter, see the common casouttable parameter (Appendix A: Common Parameters).

v={casouttable}

specifies the transpose of the matrix containing the right singular vectors. The matrix V is number of documents by k+1.

For more information about specifying the v parameter, see the common casouttable parameter (Appendix A: Common Parameters).

wordPro={casouttable}

specifies the table to contain the projections of the terms. If k dimensions of the SVD are found and the input data set contains n terms, this table will have n rows and k+1 columns.

For more information about specifying the wordPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

tmSvd Action

Computes the SVD factorization and generates topics. This action requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

Python Syntax

results= s.textMining.tmSvd(
config={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"importOptions":{"fileType":"AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression"
},
count="variable-name",
docId="variable-name",
docPro={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
exactDocPro=True | False,
exactWeight=True | False,
k=integer,
legacyNames=True | False,
maxK=integer,
nThreads=integer,
norm="ALL" | "DOC" | "NONE" | "WORD",
numLabels=integer,
required parameter parent={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"importOptions":{"fileType":"AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression"
},
resolution="HIGH" | "LOW" | "MED",
rotate="PROMAX" | "VARIMAX",
rowPivot=double,
s={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
scoreConfig={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
termId="variable-name",
termTopics={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
terms={
"caslib":"string",
"computedOnDemand":True | False,
"computedVars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"computedVarsProgram":"string",
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>},
"groupBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"importOptions":{"fileType":"AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters},
required parameter "name":"table-name",
"onDemand":True | False,
"orderBy":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"singlePass":True | False,
"vars":[{
"format":"string",
"formattedLength":integer,
"label":"string",
required parameter "name":"variable-name",
"nfd":integer,
"nfl":integer
}<, {...}>],
"where":"where-expression"
},
tolerance=double,
topicDecision=True | False,
topics={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
u={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
v={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
},
wordPro={
"caslib":"string",
"compress":True | False,
"indexVars":["variable-name-1" <, "variable-name-2", ...>],
"label":"string",
"maxMemSize":64-bit-integer,
"name":"table-name",
"onDemand":True | False,
"promote":True | False,
"replace":True | False,
"replication":integer,
"timeStamp":"string",
"where":["string-1" <, "string-2", ...>]
}
)

Parameter Descriptions

config={castable}

specifies the name of the input CAS table that contains parsing configuration information

AliasparseConfig
Long formconfig={"name":"table-name"}
Shortcut formconfig="table-name"
"caslib":"string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

"computedOnDemand":True | False

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFalse
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"computedVarsProgram":"string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}

specifies data source options.

Aliasoptions, dataSource
"groupBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the variables to use for grouping results.

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"groupByMode":"NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

"importOptions":{fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the table to use.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"orderBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"singlePass":True | False

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFalse
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"where":"where-expression"

specifies an expression for subsetting the input data.

count="variable-name"

specifies the variable that contains the, possibly weighted, term count. The values in this variable must be numeric. There can be no missing values in this variable.

Default"_COUNT_"

docId="variable-name"

specifies the variable that contains the document ID. The type of this variable can either be numeric or a string. There can be no missing values in this variable.

Default"_DOCUMENT_"

docPro={casouttable}

specifies the name of the table to contain the SVD projections of the documents.

For more information about specifying the docPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

docStdMultiple=double

Specifies how many standard deviations above the mean to set the document cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

exactDocPro=True | False

Specifies if the exact document projection values should be output. This parameter requires a SAS Visual Text Analytics license.

DefaultTrue

exactWeight=True | False

AliasexactWeights
DefaultFalse

k=integer

specifies the number of dimensions to be extracted (also the number of derived topics). If the input data is too small for the requested number of dimensions, this value is adjusted to complete the calculation.

AliasnumTopics
Range1–1000

legacyNames=True | False

specifies whether to use the legacy variable names on tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

DefaultFalse

maxK=integer

specifies the maximum number of dimensions to be extracted. The maxK option can be used in conjunction with the resolution option to dynamically select the recommended number of dimensions. If you wish to use a specific number of dimensions use maxK and set the resolution to high, or use the k parameter.

Default10
Range1–1000

nThreads=integer

specifies number of threads to be used per node. If not set, or if a value of 0 is specified, all available threads will be used.

Default8
Range0–64

norm="ALL" | "DOC" | "NONE" | "WORD"

indicates whether the document projections, term projections, or both are normalized. The normalization converts the representation from depending on angles between vectors to one based on Euclidean distances between vectors.

DefaultALL

numLabels=integer

specifies the number of terms to use in the descriptive label for each topic.

Default5
Range1–500

* parent={castable}

specifies the input CAS table that contains the term-by-document matrix in transaction form. The table must have at last three variables, one containing the document id, a second containing the term id, and the third containing the value in the cell corresponding to that particular term and document.

Long formparent={"name":"table-name"}
Shortcut formparent="table-name"
The castable value can be one or more of the following:
"caslib":"string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

"computedOnDemand":True | False

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFalse
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"computedVarsProgram":"string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}

specifies data source options.

Aliasoptions, dataSource
"groupBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the variables to use for grouping results.

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"groupByMode":"NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

"importOptions":{fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the table to use.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"orderBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"singlePass":True | False

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFalse
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"where":"where-expression"

specifies an expression for subsetting the input data.

resolution="HIGH" | "LOW" | "MED"

specifies the desired resolution level for the recommended number of dimensions to be extracted by the SVD.

DefaultHIGH

rotate="PROMAX" | "VARIMAX"

specifies the type of rotation used to maximize the explanatory power of each topic. A VARIMAX rotation produces uncorrelated topics and a PROMAX rotation produces correlated topics.

DefaultVARIMAX

rowPivot=double

Specifies the row pivot weight for document normalization of the parent table before the SVD. A negative value turns off the rowPivot process. When topics are requested a rowPivot=1 value is used by default. This parameter requires a SAS Visual Text Analytics license.

Default-1
Range-1–1

s={casouttable}

specifies the S matrix, which is a diagonal matrix that is output in compressed form, with two variables and k rows. The variable _ID_ indicates the row and column of the entry and the variable S contains the singular values.

For more information about specifying the s parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scoreConfig={casouttable}

Specifies the output scoring config file.

For more information about specifying the scoreConfig parameter, see the common casouttable parameter (Appendix A: Common Parameters).

termId="variable-name"

specifies the variable that contains the term ID. The contents of this variable must be an integer greater than or equal to 1. There can be no missing values in this variable.

Default"_TERMNUM_"

termStdMultiple=double

Specifies how many standard deviations above the mean to set the term cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

termTopics={casouttable}

specifies the name of the output CAS table to contain the term-by-topic sparse matrix information.

For more information about specifying the termTopics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

terms={castable}

specifies the name of the input table that contains information about the terms in the document collection. The table is used to determine which terms to use in the topic calculation.

Long formterms={"name":"table-name"}
Shortcut formterms="table-name"
The castable value can be one or more of the following:
"caslib":"string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

"computedOnDemand":True | False

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFalse
"computedVars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"computedVarsProgram":"string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
"dataSourceOptions":{"key-1":{any-list-or-data-type-1} <, "key-2":{any-list-or-data-type-2}, ...>}

specifies data source options.

Aliasoptions, dataSource
"groupBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the names of the variables to use for grouping results.

"format":"string"

specifies the format to apply to the variable.

"formattedLength":integer

specifies the length of format field plus the format precision.

"label":"string"

specifies the descriptive label for the variable.

* "name":"variable-name"

specifies the name for the variable.

"nfd":integer

specifies the length of the format precision.

"nfl":integer

specifies the length of the format field.

"groupByMode":"NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

"importOptions":{fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters}

specifies the settings for reading a table from a data source.

Aliasimport_

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* "name":"table-name"

specifies the name of the table to use.

"onDemand":True | False

This parameter is deprecated.

DefaultTrue
"orderBy":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"singlePass":True | False

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFalse
"vars":[{casinvardesc-1} <, {casinvardesc-2}, ...>]

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

"where":"where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the stopping threshold for the iterative factorization algorithm. If 0 is specified the default value is used.

Default1e-06
Range0–1

topicDecision=True | False

Specifies to include topic membership decisions and document cutoffs in the output tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

DefaultFalse

topics={casouttable}

specifies the output CAS table to contain the topics that are discovered.

For more information about specifying the topics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

u={casouttable}

specifies the U matrix, which contains the left singular vectors. The matrix U is number of terms by k+1.

For more information about specifying the u parameter, see the common casouttable parameter (Appendix A: Common Parameters).

v={casouttable}

specifies the transpose of the matrix containing the right singular vectors. The matrix V is number of documents by k+1.

For more information about specifying the v parameter, see the common casouttable parameter (Appendix A: Common Parameters).

wordPro={casouttable}

specifies the table to contain the projections of the terms. If k dimensions of the SVD are found and the input data set contains n terms, this table will have n rows and k+1 columns.

For more information about specifying the wordPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

tmSvd Action

Computes the SVD factorization and generates topics. This action requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

R Syntax

results <– cas.textMining.tmSvd(s,
config=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression"
),
count="variable-name",
docId="variable-name",
docPro=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
exactDocPro=TRUE | FALSE,
exactWeight=TRUE | FALSE,
k=integer,
legacyNames=TRUE | FALSE,
maxK=integer,
nThreads=integer,
norm="ALL" | "DOC" | "NONE" | "WORD",
numLabels=integer,
required parameter parent=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression"
),
resolution="HIGH" | "LOW" | "MED",
rotate="PROMAX" | "VARIMAX",
rowPivot=double,
s=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
scoreConfig=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
termId="variable-name",
termTopics=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
terms=list(
caslib="string",
computedOnDemand=TRUE | FALSE,
computedVars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>),
groupBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters),
required parameter name="table-name",
onDemand=TRUE | FALSE,
orderBy=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
singlePass=TRUE | FALSE,
vars=list( list(
format="string",
formattedLength=integer,
label="string",
required parameter name="variable-name",
nfd=integer,
nfl=integer
) <, list(...)>),
where="where-expression"
),
tolerance=double,
topicDecision=TRUE | FALSE,
topics=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
u=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
v=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
),
wordPro=list(
caslib="string",
compress=TRUE | FALSE,
indexVars=list("variable-name-1" <, "variable-name-2", ...>),
label="string",
maxMemSize=64-bit-integer,
name="table-name",
onDemand=TRUE | FALSE,
promote=TRUE | FALSE,
replace=TRUE | FALSE,
replication=integer,
timeStamp="string",
where=list("string-1" <, "string-2", ...>)
)
)

Parameter Descriptions

config=list(castable)

specifies the name of the input CAS table that contains parsing configuration information

AliasparseConfig
Long formconfig=list(name="table-name")
Shortcut formconfig="table-name"
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)

specifies data source options.

Aliasoptions, dataSource
groupBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

count="variable-name"

specifies the variable that contains the, possibly weighted, term count. The values in this variable must be numeric. There can be no missing values in this variable.

Default"_COUNT_"

docId="variable-name"

specifies the variable that contains the document ID. The type of this variable can either be numeric or a string. There can be no missing values in this variable.

Default"_DOCUMENT_"

docPro=list(casouttable)

specifies the name of the table to contain the SVD projections of the documents.

For more information about specifying the docPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

docStdMultiple=double

Specifies how many standard deviations above the mean to set the document cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

exactDocPro=TRUE | FALSE

Specifies if the exact document projection values should be output. This parameter requires a SAS Visual Text Analytics license.

DefaultTRUE

exactWeight=TRUE | FALSE

AliasexactWeights
DefaultFALSE

k=integer

specifies the number of dimensions to be extracted (also the number of derived topics). If the input data is too small for the requested number of dimensions, this value is adjusted to complete the calculation.

AliasnumTopics
Range1–1000

legacyNames=TRUE | FALSE

specifies whether to use the legacy variable names on tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

DefaultFALSE

maxK=integer

specifies the maximum number of dimensions to be extracted. The maxK option can be used in conjunction with the resolution option to dynamically select the recommended number of dimensions. If you wish to use a specific number of dimensions use maxK and set the resolution to high, or use the k parameter.

Default10
Range1–1000

nThreads=integer

specifies number of threads to be used per node. If not set, or if a value of 0 is specified, all available threads will be used.

Default8
Range0–64

norm="ALL" | "DOC" | "NONE" | "WORD"

indicates whether the document projections, term projections, or both are normalized. The normalization converts the representation from depending on angles between vectors to one based on Euclidean distances between vectors.

DefaultALL

numLabels=integer

specifies the number of terms to use in the descriptive label for each topic.

Default5
Range1–500

* parent=list(castable)

specifies the input CAS table that contains the term-by-document matrix in transaction form. The table must have at last three variables, one containing the document id, a second containing the term id, and the third containing the value in the cell corresponding to that particular term and document.

Long formparent=list(name="table-name")
Shortcut formparent="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)

specifies data source options.

Aliasoptions, dataSource
groupBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

resolution="HIGH" | "LOW" | "MED"

specifies the desired resolution level for the recommended number of dimensions to be extracted by the SVD.

DefaultHIGH

rotate="PROMAX" | "VARIMAX"

specifies the type of rotation used to maximize the explanatory power of each topic. A VARIMAX rotation produces uncorrelated topics and a PROMAX rotation produces correlated topics.

DefaultVARIMAX

rowPivot=double

Specifies the row pivot weight for document normalization of the parent table before the SVD. A negative value turns off the rowPivot process. When topics are requested a rowPivot=1 value is used by default. This parameter requires a SAS Visual Text Analytics license.

Default-1
Range-1–1

s=list(casouttable)

specifies the S matrix, which is a diagonal matrix that is output in compressed form, with two variables and k rows. The variable _ID_ indicates the row and column of the entry and the variable S contains the singular values.

For more information about specifying the s parameter, see the common casouttable parameter (Appendix A: Common Parameters).

scoreConfig=list(casouttable)

Specifies the output scoring config file.

For more information about specifying the scoreConfig parameter, see the common casouttable parameter (Appendix A: Common Parameters).

termId="variable-name"

specifies the variable that contains the term ID. The contents of this variable must be an integer greater than or equal to 1. There can be no missing values in this variable.

Default"_TERMNUM_"

termStdMultiple=double

Specifies how many standard deviations above the mean to set the term cutoff. This parameter requires a SAS Visual Text Analytics license.

Default1
Range0–10

termTopics=list(casouttable)

specifies the name of the output CAS table to contain the term-by-topic sparse matrix information.

For more information about specifying the termTopics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

terms=list(castable)

specifies the name of the input table that contains information about the terms in the document collection. The table is used to determine which terms to use in the topic calculation.

Long formterms=list(name="table-name")
Shortcut formterms="table-name"
The castable value can be one or more of the following:
caslib="string"

specifies the caslib that contains the table that you want to use with the action. By default, the active caslib is used. Specify a value only if you need to access a table from a different caslib.

computedOnDemand=TRUE | FALSE

when set to True, creates the computed variables when the table is loaded instead of when the action begins.

AliascompOnDemand
DefaultFALSE
computedVars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the computed variables to create. Specify an expression for each variable in the computedVarsProgram parameter.

AliascompVars
format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

computedVarsProgram="string"

specifies an expression for each computed variable that you include in the computedVars parameter.

AliascompPgm
dataSourceOptions=list(key-1=list(any-list-or-data-type-1) <, key-2=list(any-list-or-data-type-2), ...>)

specifies data source options.

Aliasoptions, dataSource
groupBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the names of the variables to use for grouping results.

format="string"

specifies the format to apply to the variable.

formattedLength=integer

specifies the length of format field plus the format precision.

label="string"

specifies the descriptive label for the variable.

* name="variable-name"

specifies the name for the variable.

nfd=integer

specifies the length of the format precision.

nfl=integer

specifies the length of the format field.

groupByMode="NOSORT" | "REDISTRIBUTE"

specifies how to create groups.

DefaultNOSORT
NOSORT

group the data without sorting on each machine, and then group the data again on the controller.

REDISTRIBUTE

transfer rows between nodes to guarantee ordering within groups. This method is slower.

importOptions=list(fileType="AUTO" | "BASESAS" | "CSV" | "DOCUMENT" | "DTA" | "ESP" | "EXCEL" | "FMT" | "HDAT" | "JMP" | "LASR" | "SPSS" | "XLS", fileType-specific-parameters)

specifies the settings for reading a table from a data source.

Aliasimport

The value that you specify for fileType determines the other parameters that apply. For more information about this common parameter, see importOptions (Appendix A: Common Parameters).

* name="table-name"

specifies the name of the table to use.

onDemand=TRUE | FALSE

This parameter is deprecated.

DefaultTRUE
orderBy=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use for ordering observations within partitions. This parameter applies to partitioned tables, or it can be combined with variables that are specified in the groupBy parameter when the value of the groupByMode parameter is set to REDISTRIBUTE.

For more information about specifying the orderBy parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

singlePass=TRUE | FALSE

when set to True, does not create a transient table on the server. Setting this parameter to True can be efficient, but the data might not have stable ordering upon repeated runs.

DefaultFALSE
vars=list( list(casinvardesc-1) <, list(casinvardesc-2), ...>)

specifies the variables to use in the action.

For more information about specifying the vars parameter, see the common casinvardesc parameter (Appendix A: Common Parameters).

where="where-expression"

specifies an expression for subsetting the input data.

tolerance=double

specifies the stopping threshold for the iterative factorization algorithm. If 0 is specified the default value is used.

Default1e-06
Range0–1

topicDecision=TRUE | FALSE

Specifies to include topic membership decisions and document cutoffs in the output tables. This parameter requires a SAS Visual Text Analytics license or a SAS Visual Data Mining and Machine Learning license.

DefaultFALSE

topics=list(casouttable)

specifies the output CAS table to contain the topics that are discovered.

For more information about specifying the topics parameter, see the common casouttable parameter (Appendix A: Common Parameters).

u=list(casouttable)

specifies the U matrix, which contains the left singular vectors. The matrix U is number of terms by k+1.

For more information about specifying the u parameter, see the common casouttable parameter (Appendix A: Common Parameters).

v=list(casouttable)

specifies the transpose of the matrix containing the right singular vectors. The matrix V is number of documents by k+1.

For more information about specifying the v parameter, see the common casouttable parameter (Appendix A: Common Parameters).

wordPro=list(casouttable)

specifies the table to contain the projections of the terms. If k dimensions of the SVD are found and the input data set contains n terms, this table will have n rows and k+1 columns.

For more information about specifying the wordPro parameter, see the common casouttable parameter (Appendix A: Common Parameters).

Last updated: June 07, 2018