Text Miner Node Properties

Contents

Note: The Text Miner node is not available on the Text Mining tab in SAS Text Miner 14.3. The Text Miner node has now been replaced by the functionality in other SAS Text Miner nodes. You can import diagrams from a previous release of SAS Text Miner that had a Text Miner node in the process flow diagram. However, new Text Miner nodes can no longer be created, and property values cannot be changed in imported Text Miner nodes. For more information, see Converting SAS Text Miner Diagrams from a Previous Version.

Text Miner Node General Properties

These are the general properties that are available on the Text Miner node:
  • Node ID — displays the ID that SAS Enterprise Miner automatically assigns the node. Node IDs are especially useful for distinguishing between two or more nodes of the same type in a process flow diagram. For example, the first Text Miner node that is added to a diagram will have the node ID TEXT, and the second Text Miner node that is added will have the node ID TEXT2.
  • Imported Data — accesses a list of the data sets that are imported by the node and the ports that provide them. Click the ellipsis button to open the Imported Data window, which displays this list. If data exists for an imported data set, then you can select a row in the list and do any of the following:
    • browse the data set
    • explore (sample and plot) the data in a data set
    • view the table and variable properties of a data set
  • Exported Data — accesses a list of the data sets exported by the node and the ports to which they are provided. Click the ellipsis button to open the Exported Data window, which displays this list. If data exists for an exported data set, then you can select a row in the list and do any of the following:
    • browse the data set
    • explore (sample and plot) the data in a data set
    • view the table and variable properties of a data set
  • Notes — accesses a window that you can use to store notes of interest, such as data or configuration information. Click the ellipsis button to open the Notes window.

Text Miner Node Train Properties

General Train Properties

These are the training properties that are available on the Text Miner node:
  • Variables — accesses a list of variables and associated properties in the data source. Only variables with the roles Text, Text Location, or Web Address are displayed. Click the ellipsis button to open the Variables window.
  • Interactive — This property is not available in SAS Text Miner 14.3.
  • Force Run — specifies whether to rerun the node, even if it has already been successfully run. This function is useful if the underlying input data has changed.

Parse Properties

  • Parse Variable — (value is populated after the node has run) displays the name of the variable in the input data source that was used for parsing. Depending on the structure of the data source, this variable contains either the entire text of each document in the document collection or it contains paths to plain text or HMTL files that contain that text.
  • Language — specifies the language to use when parsing text. Only supported languages that are licensed to you are available for selection. For a list of supported languages, see About SAS Text Miner.
  • Stop List — accesses a window in which you can select a SAS data set that contains terms to exclude from parsing. If you include a stop list, then the terms therein are not included in the node results. Default data sets are provided for several languages. You can edit these data sets or create your own. Click the ellipsis button to open the Select a SAS Table window.
  • Start List — accesses a window in which you can select a SAS data set that contains the terms to parse. If you include a start list, then terms other than those therein are not included in the node results. Click the ellipsis button to open the Select a SAS Table window.
  • Stem Terms — specifies whether to stem terms.
  • Terms in Single Document — specifies whether to parse terms that appear in only one document.
  • Punctuation — specifies whether to parse punctuation marks as terms.
  • Numbers — specifies whether to parse numbers as terms.
  • Different Parts of Speech — specifies whether to identify the parts of speech of parsed terms. If the value of this property is Yes, then same terms with different parts of speech are treated as different terms.
  • Ignore Parts of Speech — accesses a window in which you can select one or more parts of speech to ignore when parsing. Click the ellipsis button to open the Ignore Parts of Speech window. Terms with the selected parts of speech are not parsed and do not appear in node results. This property is not used if the value of Different Parts of Speech is No.
  • Noun Groups — specifies whether to identify noun groups. If stemming is turned on, then noun group elements are also stemmed.
  • Synonyms — specifies a SAS data set that contains synonyms to be treated as equivalent. Default data sets are provided for several languages. You can edit these data sets or create your own. Click the ellipsis button to open the Select a SAS Table window.
  • Find Entities — specifies whether to identify the entities of parsed terms.
  • Types of Entities — accesses a window in which you can select one or more entity classifications to parse. Click the ellipsis button to open the Select Entity Types window.
    Note: If the value of Find Entities is Yes but you have not selected any entity types in the Select Entity Types window, then all entity types are parsed. In other words, deselecting all entity types has the same effect as selecting all of them.

Transform Properties

  • Compute SVD — specifies whether to compute the singular-value decomposition (SVD) of the term-by-document frequency matrix.
  • SVD Resolution — specifies the resolution to use to generate the SVD dimensions.
  • Max SVD Dimensions — specifies the maximum number, greater than or equal to 2, of SVD dimensions to generate.
  • Scale SVD Dimensions — specifies whether to scale the SVD dimensions by the inverse of the singular values in order to generate equal variances.
  • Frequency Weighting — specifies the frequency weighting method to use.
  • Term Weight — specifies the term weighting method to use.
  • Roll up Terms — specifies whether to first sort parsed terms in descending order of the value of the term weight multiplied by the square root of the number of documents and then roll up these selected terms as variables on the Document data set.
  • No. of Rolled-up Terms — specifies the number of terms to use to roll up lower-weighted terms. This property is used if the value of Roll up Terms is Yes.
  • Drop Other Terms — specifies whether to drop terms if they are not rolled up in the Document data set. This property is used if the value of Roll up Terms is Yes.

Cluster Properties

  • Automatically Cluster — specifies whether to perform clustering analysis. If you select No, then the remaining Cluster Properties are not used.
  • Exact or Maximum Number — specifies whether to find an exact number of clusters or any number less than or equal to a maximum number of clusters.
  • Number of Clusters — specifies the number of clusters. This is the exact number if the value of Exact or Maximum Number is Exact, and it is the maximum number if the value of Exact or Maximum Number is Maximum.
  • Cluster Algorithm — specifies the clustering algorithm.
  • Ignore Outliers — specifies whether to ignore outliers in the clustering analysis. If the value of this property is No, then do one of the following:
    • for Hierarchical clustering, outliers are removed
    • for Expectation-Maximization clustering, outliers are placed in a single cluster
  • Hierarchy Levels — (for Hierarchical clustering) specifies the number of levels in the cluster hierarchy. To specify the maximum depth (all levels), enter a period (.).
  • Descriptive Terms — specifies the number of descriptive terms to display in each cluster.
  • What to Cluster — specifies whether to cluster roll-up terms or the SVD dimensions.

Text Miner Node Status Properties

These are the status properties that are displayed on the Text Miner node:
  • Create Time — time at which the node was created.
  • Run ID — identifier of the run of the node. A new identifier is assigned every time the node is run.
  • Last Error — error message, if any, from the last run.
  • Last Status — last reported status of the node.
  • Last Run Time — time at which the node was last run.
  • Run Duration — length of time required to complete the last node run.
  • Grid Host — grid host, if any, that was used for computation.
  • User-Added Node — denotes whether the node was created by a user as a SAS Enterprise Miner Extension node. The value of this property is always No for the Text Miner node.
Last updated: September 15, 2017