Text Management Action Set: Syntax
Provides utility actions for managing the properties of the text
profileText Action
Profiles text data and generates descriptive statistics. This action requires a SAS Visual Text Analytics license.
| See: | Profile Text Data |
|---|---|
| Example: | Profile Text Data Using the profileText Action |
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the input table that contains reference mutual information scores. | |
|
required parametertable |
— |
specifies the input CAS table that contains the text data. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the output table that contains extracted bigrams. | |
|
— |
specifies the output table that contains information complexity, information density and vocabulary diversity statistics. | |
|
— |
specifies the output table that contains document-level information complexity statistics. | |
|
— |
specifies the output table that contains the token count of each sentence in each document. | |
|
— |
specifies the output table that contains bigram information. | |
|
— |
specifies the output table that contains document-sentence information. | |
|
— |
specifies the output table that contains document length information. | |
|
— |
specifies the output table that contains sentence token information. | |
|
— |
specifies the output table that contains sentence information. | |
|
— |
specifies the output table that contains sentence length information. | |
|
— |
specifies the output table that contains statistical information. | |
|
— |
specifies the output table that contains token information. | |
|
— |
specifies the output table that contains token length information. | |
|
— |
specifies the output table that contains token type information. | |
|
— |
specifies the output table that contains sentence-level information complexity statistics based on token count. | |
|
— |
specifies the output table that contains word-level information density statistics based on unique tokens. |
Parameter Descriptions
bigram={casouttable}
specifies the output table that contains extracted bigrams.
For more information about specifying the bigram parameter, see the common casouttable parameter.
bigramRatio=double
specifies the sample ratio of the input data set for the token bigram analysis. By default, 0.2 (for example, 20%) is used for the input data set.
| Default | 0.2 |
|---|---|
| Range | 0.01–1 |
casOut={casouttable}
specifies the output table that contains information complexity, information density and vocabulary diversity statistics.
For more information about specifying the casOut parameter, see the common casouttable parameter.
documentId="string"
specifies the CAS table variable name that contains the document IDs.
| Default | "_DOCUMENT_ID_" |
|---|
documentOut={casouttable}
specifies the output table that contains document-level information complexity statistics.
For more information about specifying the documentOut parameter, see the common casouttable parameter.
intermediateOut={casouttable}
specifies the output table that contains the token count of each sentence in each document.
For more information about specifying the intermediateOut parameter, see the common casouttable parameter.
language="string"
specifies the language used in the text variable of the input table.
| Default | "ENGLISH" |
|---|
outLibBigram={casouttable}
specifies the output table that contains bigram information.
For more information about specifying the outLibBigram parameter, see the common casouttable parameter.
outLibDoc={casouttable}
specifies the output table that contains document-sentence information.
For more information about specifying the outLibDoc parameter, see the common casouttable parameter.
outLibDocLen={casouttable}
specifies the output table that contains document length information.
For more information about specifying the outLibDocLen parameter, see the common casouttable parameter.
outLibInterm={casouttable}
specifies the output table that contains sentence token information.
For more information about specifying the outLibInterm parameter, see the common casouttable parameter.
outLibSent={casouttable}
specifies the output table that contains sentence information.
For more information about specifying the outLibSent parameter, see the common casouttable parameter.
outLibSentLen={casouttable}
specifies the output table that contains sentence length information.
For more information about specifying the outLibSentLen parameter, see the common casouttable parameter.
outLibStats={casouttable}
specifies the output table that contains statistical information.
For more information about specifying the outLibStats parameter, see the common casouttable parameter.
outLibTok={casouttable}
specifies the output table that contains token information.
For more information about specifying the outLibTok parameter, see the common casouttable parameter.
outLibTokLen={casouttable}
specifies the output table that contains token length information.
For more information about specifying the outLibTokLen parameter, see the common casouttable parameter.
outLibTokType={casouttable}
specifies the output table that contains token type information.
For more information about specifying the outLibTokType parameter, see the common casouttable parameter.
referenceData=TRUE | FALSE
Whether to output reference statistics.
| Default | FALSE |
|---|
refMIScore={castable}
specifies the input table that contains reference mutual information scores.
For more information about specifying the refMIScore parameter, see the common castable (Form 1) parameter.
sentenceOut={casouttable}
specifies the output table that contains sentence-level information complexity statistics based on token count.
For more information about specifying the sentenceOut parameter, see the common casouttable parameter.
* table={castable}
specifies the input CAS table that contains the text data.
For more information about specifying the table parameter, see the common castable (Form 1) parameter.
text="string"
specifies the input variable to profile. The default value is "_text_".
| Default | "_TEXT_" |
|---|
tokenOut={casouttable}
specifies the output table that contains word-level information density statistics based on unique tokens.
For more information about specifying the tokenOut parameter, see the common casouttable parameter.
profileText Action
Profiles text data and generates descriptive statistics. This action requires a SAS Visual Text Analytics license.
| See: | Profile Text Data |
|---|---|
| Example: | Profile Text Data Using the profileText Action |
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the input table that contains reference mutual information scores. | |
|
required parametertable |
— |
specifies the input CAS table that contains the text data. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the output table that contains extracted bigrams. | |
|
— |
specifies the output table that contains information complexity, information density and vocabulary diversity statistics. | |
|
— |
specifies the output table that contains document-level information complexity statistics. | |
|
— |
specifies the output table that contains the token count of each sentence in each document. | |
|
— |
specifies the output table that contains bigram information. | |
|
— |
specifies the output table that contains document-sentence information. | |
|
— |
specifies the output table that contains document length information. | |
|
— |
specifies the output table that contains sentence token information. | |
|
— |
specifies the output table that contains sentence information. | |
|
— |
specifies the output table that contains sentence length information. | |
|
— |
specifies the output table that contains statistical information. | |
|
— |
specifies the output table that contains token information. | |
|
— |
specifies the output table that contains token length information. | |
|
— |
specifies the output table that contains token type information. | |
|
— |
specifies the output table that contains sentence-level information complexity statistics based on token count. | |
|
— |
specifies the output table that contains word-level information density statistics based on unique tokens. |
Parameter Descriptions
bigram={casouttable}
specifies the output table that contains extracted bigrams.
For more information about specifying the bigram parameter, see the common casouttable parameter.
bigramRatio=double
specifies the sample ratio of the input data set for the token bigram analysis. By default, 0.2 (for example, 20%) is used for the input data set.
| Default | 0.2 |
|---|---|
| Range | 0.01–1 |
casOut={casouttable}
specifies the output table that contains information complexity, information density and vocabulary diversity statistics.
For more information about specifying the casOut parameter, see the common casouttable parameter.
documentId="string"
specifies the CAS table variable name that contains the document IDs.
| Default | "_DOCUMENT_ID_" |
|---|
documentOut={casouttable}
specifies the output table that contains document-level information complexity statistics.
For more information about specifying the documentOut parameter, see the common casouttable parameter.
intermediateOut={casouttable}
specifies the output table that contains the token count of each sentence in each document.
For more information about specifying the intermediateOut parameter, see the common casouttable parameter.
language="string"
specifies the language used in the text variable of the input table.
| Default | "ENGLISH" |
|---|
outLibBigram={casouttable}
specifies the output table that contains bigram information.
For more information about specifying the outLibBigram parameter, see the common casouttable parameter.
outLibDoc={casouttable}
specifies the output table that contains document-sentence information.
For more information about specifying the outLibDoc parameter, see the common casouttable parameter.
outLibDocLen={casouttable}
specifies the output table that contains document length information.
For more information about specifying the outLibDocLen parameter, see the common casouttable parameter.
outLibInterm={casouttable}
specifies the output table that contains sentence token information.
For more information about specifying the outLibInterm parameter, see the common casouttable parameter.
outLibSent={casouttable}
specifies the output table that contains sentence information.
For more information about specifying the outLibSent parameter, see the common casouttable parameter.
outLibSentLen={casouttable}
specifies the output table that contains sentence length information.
For more information about specifying the outLibSentLen parameter, see the common casouttable parameter.
outLibStats={casouttable}
specifies the output table that contains statistical information.
For more information about specifying the outLibStats parameter, see the common casouttable parameter.
outLibTok={casouttable}
specifies the output table that contains token information.
For more information about specifying the outLibTok parameter, see the common casouttable parameter.
outLibTokLen={casouttable}
specifies the output table that contains token length information.
For more information about specifying the outLibTokLen parameter, see the common casouttable parameter.
outLibTokType={casouttable}
specifies the output table that contains token type information.
For more information about specifying the outLibTokType parameter, see the common casouttable parameter.
referenceData=true | false
Whether to output reference statistics.
| Default | false |
|---|
refMIScore={castable}
specifies the input table that contains reference mutual information scores.
For more information about specifying the refMIScore parameter, see the common castable (Form 1) parameter.
sentenceOut={casouttable}
specifies the output table that contains sentence-level information complexity statistics based on token count.
For more information about specifying the sentenceOut parameter, see the common casouttable parameter.
* table={castable}
specifies the input CAS table that contains the text data.
For more information about specifying the table parameter, see the common castable (Form 1) parameter.
text="string"
specifies the input variable to profile. The default value is "_text_".
| Default | "_TEXT_" |
|---|
tokenOut={casouttable}
specifies the output table that contains word-level information density statistics based on unique tokens.
For more information about specifying the tokenOut parameter, see the common casouttable parameter.
profileText Action
Profiles text data and generates descriptive statistics. This action requires a SAS Visual Text Analytics license.
| See: | Profile Text Data |
|---|---|
| Example: | Profile Text Data Using the profileText Action |
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the input table that contains reference mutual information scores. | |
|
required parametertable |
— |
specifies the input CAS table that contains the text data. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the output table that contains extracted bigrams. | |
|
— |
specifies the output table that contains information complexity, information density and vocabulary diversity statistics. | |
|
— |
specifies the output table that contains document-level information complexity statistics. | |
|
— |
specifies the output table that contains the token count of each sentence in each document. | |
|
— |
specifies the output table that contains bigram information. | |
|
— |
specifies the output table that contains document-sentence information. | |
|
— |
specifies the output table that contains document length information. | |
|
— |
specifies the output table that contains sentence token information. | |
|
— |
specifies the output table that contains sentence information. | |
|
— |
specifies the output table that contains sentence length information. | |
|
— |
specifies the output table that contains statistical information. | |
|
— |
specifies the output table that contains token information. | |
|
— |
specifies the output table that contains token length information. | |
|
— |
specifies the output table that contains token type information. | |
|
— |
specifies the output table that contains sentence-level information complexity statistics based on token count. | |
|
— |
specifies the output table that contains word-level information density statistics based on unique tokens. |
Parameter Descriptions
bigram={casouttable}
specifies the output table that contains extracted bigrams.
For more information about specifying the bigram parameter, see the common casouttable parameter.
bigramRatio=double
specifies the sample ratio of the input data set for the token bigram analysis. By default, 0.2 (for example, 20%) is used for the input data set.
| Default | 0.2 |
|---|---|
| Range | 0.01–1 |
casOut={casouttable}
specifies the output table that contains information complexity, information density and vocabulary diversity statistics.
For more information about specifying the casOut parameter, see the common casouttable parameter.
documentId="string"
specifies the CAS table variable name that contains the document IDs.
| Default | "_DOCUMENT_ID_" |
|---|
documentOut={casouttable}
specifies the output table that contains document-level information complexity statistics.
For more information about specifying the documentOut parameter, see the common casouttable parameter.
intermediateOut={casouttable}
specifies the output table that contains the token count of each sentence in each document.
For more information about specifying the intermediateOut parameter, see the common casouttable parameter.
language="string"
specifies the language used in the text variable of the input table.
| Default | "ENGLISH" |
|---|
outLibBigram={casouttable}
specifies the output table that contains bigram information.
For more information about specifying the outLibBigram parameter, see the common casouttable parameter.
outLibDoc={casouttable}
specifies the output table that contains document-sentence information.
For more information about specifying the outLibDoc parameter, see the common casouttable parameter.
outLibDocLen={casouttable}
specifies the output table that contains document length information.
For more information about specifying the outLibDocLen parameter, see the common casouttable parameter.
outLibInterm={casouttable}
specifies the output table that contains sentence token information.
For more information about specifying the outLibInterm parameter, see the common casouttable parameter.
outLibSent={casouttable}
specifies the output table that contains sentence information.
For more information about specifying the outLibSent parameter, see the common casouttable parameter.
outLibSentLen={casouttable}
specifies the output table that contains sentence length information.
For more information about specifying the outLibSentLen parameter, see the common casouttable parameter.
outLibStats={casouttable}
specifies the output table that contains statistical information.
For more information about specifying the outLibStats parameter, see the common casouttable parameter.
outLibTok={casouttable}
specifies the output table that contains token information.
For more information about specifying the outLibTok parameter, see the common casouttable parameter.
outLibTokLen={casouttable}
specifies the output table that contains token length information.
For more information about specifying the outLibTokLen parameter, see the common casouttable parameter.
outLibTokType={casouttable}
specifies the output table that contains token type information.
For more information about specifying the outLibTokType parameter, see the common casouttable parameter.
referenceData=True | False
Whether to output reference statistics.
| Default | False |
|---|
refMIScore={castable}
specifies the input table that contains reference mutual information scores.
For more information about specifying the refMIScore parameter, see the common castable (Form 1) parameter.
sentenceOut={casouttable}
specifies the output table that contains sentence-level information complexity statistics based on token count.
For more information about specifying the sentenceOut parameter, see the common casouttable parameter.
* table={castable}
specifies the input CAS table that contains the text data.
For more information about specifying the table parameter, see the common castable (Form 1) parameter.
text="string"
specifies the input variable to profile. The default value is "_text_".
| Default | "_TEXT_" |
|---|
tokenOut={casouttable}
specifies the output table that contains word-level information density statistics based on unique tokens.
For more information about specifying the tokenOut parameter, see the common casouttable parameter.
profileText Action
Profiles text data and generates descriptive statistics. This action requires a SAS Visual Text Analytics license.
| See: | Profile Text Data |
|---|---|
| Example: | Profile Text Data Using the profileText Action |
Summary: Input and Output Tables
If a row includes a subparameter, you can specify the name, caslib, and so on in the subparameter. Otherwise, you can specify the name, caslib, and so on in the parameter.
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the input table that contains reference mutual information scores. | |
|
required parametertable |
— |
specifies the input CAS table that contains the text data. |
|
Parameter |
Subparameter |
Description |
|---|---|---|
|
— |
specifies the output table that contains extracted bigrams. | |
|
— |
specifies the output table that contains information complexity, information density and vocabulary diversity statistics. | |
|
— |
specifies the output table that contains document-level information complexity statistics. | |
|
— |
specifies the output table that contains the token count of each sentence in each document. | |
|
— |
specifies the output table that contains bigram information. | |
|
— |
specifies the output table that contains document-sentence information. | |
|
— |
specifies the output table that contains document length information. | |
|
— |
specifies the output table that contains sentence token information. | |
|
— |
specifies the output table that contains sentence information. | |
|
— |
specifies the output table that contains sentence length information. | |
|
— |
specifies the output table that contains statistical information. | |
|
— |
specifies the output table that contains token information. | |
|
— |
specifies the output table that contains token length information. | |
|
— |
specifies the output table that contains token type information. | |
|
— |
specifies the output table that contains sentence-level information complexity statistics based on token count. | |
|
— |
specifies the output table that contains word-level information density statistics based on unique tokens. |
Parameter Descriptions
bigram=list(casouttable)
specifies the output table that contains extracted bigrams.
For more information about specifying the bigram parameter, see the common casouttable parameter.
bigramRatio=double
specifies the sample ratio of the input data set for the token bigram analysis. By default, 0.2 (for example, 20%) is used for the input data set.
| Default | 0.2 |
|---|---|
| Range | 0.01–1 |
casOut=list(casouttable)
specifies the output table that contains information complexity, information density and vocabulary diversity statistics.
For more information about specifying the casOut parameter, see the common casouttable parameter.
documentId="string"
specifies the CAS table variable name that contains the document IDs.
| Default | "_DOCUMENT_ID_" |
|---|
documentOut=list(casouttable)
specifies the output table that contains document-level information complexity statistics.
For more information about specifying the documentOut parameter, see the common casouttable parameter.
intermediateOut=list(casouttable)
specifies the output table that contains the token count of each sentence in each document.
For more information about specifying the intermediateOut parameter, see the common casouttable parameter.
language="string"
specifies the language used in the text variable of the input table.
| Default | "ENGLISH" |
|---|
outLibBigram=list(casouttable)
specifies the output table that contains bigram information.
For more information about specifying the outLibBigram parameter, see the common casouttable parameter.
outLibDoc=list(casouttable)
specifies the output table that contains document-sentence information.
For more information about specifying the outLibDoc parameter, see the common casouttable parameter.
outLibDocLen=list(casouttable)
specifies the output table that contains document length information.
For more information about specifying the outLibDocLen parameter, see the common casouttable parameter.
outLibInterm=list(casouttable)
specifies the output table that contains sentence token information.
For more information about specifying the outLibInterm parameter, see the common casouttable parameter.
outLibSent=list(casouttable)
specifies the output table that contains sentence information.
For more information about specifying the outLibSent parameter, see the common casouttable parameter.
outLibSentLen=list(casouttable)
specifies the output table that contains sentence length information.
For more information about specifying the outLibSentLen parameter, see the common casouttable parameter.
outLibStats=list(casouttable)
specifies the output table that contains statistical information.
For more information about specifying the outLibStats parameter, see the common casouttable parameter.
outLibTok=list(casouttable)
specifies the output table that contains token information.
For more information about specifying the outLibTok parameter, see the common casouttable parameter.
outLibTokLen=list(casouttable)
specifies the output table that contains token length information.
For more information about specifying the outLibTokLen parameter, see the common casouttable parameter.
outLibTokType=list(casouttable)
specifies the output table that contains token type information.
For more information about specifying the outLibTokType parameter, see the common casouttable parameter.
referenceData=TRUE | FALSE
Whether to output reference statistics.
| Default | FALSE |
|---|
refMIScore=list(castable)
specifies the input table that contains reference mutual information scores.
For more information about specifying the refMIScore parameter, see the common castable (Form 1) parameter.
sentenceOut=list(casouttable)
specifies the output table that contains sentence-level information complexity statistics based on token count.
For more information about specifying the sentenceOut parameter, see the common casouttable parameter.
* table=list(castable)
specifies the input CAS table that contains the text data.
For more information about specifying the table parameter, see the common castable (Form 1) parameter.
text="string"
specifies the input variable to profile. The default value is "_text_".
| Default | "_TEXT_" |
|---|
tokenOut=list(casouttable)
specifies the output table that contains word-level information density statistics based on unique tokens.
For more information about specifying the tokenOut parameter, see the common casouttable parameter.