Text Parsing Node Properties
Contents
Text Parsing Node General Properties
These are the general
properties that are available on the Text Parsing node:
-
Node ID — displays the ID that is assigned to the node. Node IDs are especially useful for distinguishing between two or more nodes of the same type in a process flow diagram. For example, the first Text Parsing node that is added to a diagram will have the Node ID
TextParsing. The second Text Parsing node that is added will have the Node IDTextParsing2. -
Imported Data — accesses a list of the data sets that are imported by the node and the ports that provide them. Click the ellipsis button to open the Imported Data window, which displays this list. If data exists for an imported data set, then you can select a row in the list and do any of the following:
-
browse the data set
-
explore (sample and plot) the data in a data set
-
view the table and variable properties of a data set
-
-
Exported Data — accesses a list of the data sets that are exported by the node and the ports to which they are provided. Click the ellipsis button to open the Exported Data window, which displays this list. If data exists for an exported data set, then you can select a row in the list and do any of the following:
-
browse the data set
-
explore (sample and plot) the data in a data set
-
view the table and variable properties of a data set
-
-
Notes — accesses a window that you can use to store notes of interest, such as data or configuration information. Click the ellipsis button to open the Notes window.
Text Parsing Node Train Properties
General Train Properties
These are the training
properties that are available on the Text Parsing node:
-
Variables — accesses a list of variables and associated properties in the data source. Click the ellipsis button to open the Variables window. For more information, see Text Parsing Node Input Data.
Parse Properties
-
Parse Variable — (value is populated after the node is run) displays the name of the variable in the input data source that was used for parsing. Depending on the structure of the data source, this variable contains either the entire text of each document in the document collection or the paths to plain text or HMTL files that contain that text.
-
Language — accesses a window in which you can select the language to use when parsing. Click the ellipsis button to open the Languages window. Only supported languages that are licensed to you are available for selection. For a list of supported languages, see About SAS Text Miner.
Detect Properties
-
Different Parts of Speech — specifies whether to identify the parts of speech of parsed terms. If the value of this property is Yes, then same terms with different parts of speech are treated as different terms. For more information, see Parts of Speech in SAS Text Miner.
-
Noun Groups — specifies whether to identify noun groups. If stemming is turned on, then noun group elements are also stemmed. For more information, see Noun Groups in SAS Text Miner.
-
Multi-word Terms — (for all supported languages except Chinese, Japanese, and Korean) specifies a SAS data set that contains multi-word terms. For more information, see Multi-Term Lists. Default data sets are provided for several languages. For more information, see SAS Text Miner Sample Data Sets. You can edit these data sets or create your own. Click the ellipsis button to open a window in which you can do the following:
-
Click Replace Table to replace the currently selected table.
-
Click Add Table to add to the currently selected table.
-
(If a multi-word term data set is selected) Add, delete, and edit terms in the multi-term list.
-
-
Find Entities — specifies whether to identify the entities that are contained in the documents. Entity detection relies on linguistic rules and lists that are provided for many entity types; these are known as standard entities. You can create custom entity types by defining linguistic rules and referencing a custom file within SAS Text Miner. If you have SAS Concept Creation for SAS Text Miner software installed, you can create a file with custom entity types. For more information about how to create a custom entity type, see your SAS representative. For more information about entities, see Entities in SAS Text Miner.
-
Noneidentifies neither standard nor custom entities. -
Standardidentifies standard entities, but not custom entities. -
Customidentifies custom entities, but not standard entities. -
Allidentifies both standard and custom entities.
-
-
Custom Entities — specifies the path (relative to the SAS Text Miner server) to a file that contains compiled custom entities. Valid files have the extension
.li. No custom entity should have the same name as a standard entity. For more information, see Entities in SAS Text Miner.
Ignore Properties
-
Ignore Parts of Speech — accesses a window in which you can select one or more parts of speech. Terms that are assigned these parts of speech will be ignored when parsing. Click the ellipsis button to open the Ignore Parts of Speech dialog box. Use the SHIFT and Ctrl keys to make multiple selections. Terms with the selected parts of speech are not parsed and do not appear in node results.
-
Ignore Types of Entities — (if the value of Find Entities is
StandardorAll) accesses a dialog box in which you can select one or more standard entities to ignore when parsing. For more information, see Entities in SAS Text Miner. Click the ellipsis button to open the Ignore Types of Entities window. Use the SHIFT and Ctrl keys to make multiple selections. Terms with the selected entity types are not parsed and do not appear in node results. -
Ignore Types of Attributes — accesses a window in which you can select one or more attributes to ignore when parsing. For more information, see Attributes in SAS Text Miner . Click the ellipsis button to open the Ignore Types of Attributes dialog box. Use the SHIFT and Ctrl keys to make multiple selections. Terms with the selected attribute types are not parsed and do not appear in node results.
Synonyms Properties
-
Stem Terms — specifies whether to treat different terms with the same root as equivalent. For more information see Term Stemming.
-
Synonyms — specifies a SAS data set that contains synonyms to be treated as equivalent. For more information, see Synonym Lists. Default data sets are provided for several languages. For more information, see SAS Text Miner Sample Data Sets.Note: Multi-word term lists are not supported in Chinese, Japanese, and Korean.You can edit these data sets or create your own. Click the ellipsis button to open a window in which you can do the following:
-
Click Replace Table to replace the currently selected table.
-
Click Add Table to add to the currently selected table.
-
(If a synonym data set is selected) Add, delete, and edit terms in the synonym list.
-
Filter Properties
-
Start List — specifies a SAS data set that contains the terms to parse. If you include a start list, then the terms that are not included in the start list appear in the results Term table with a Keep status of
N. For more information, see Start Lists and Stop Lists. Click the ellipsis button to open a window in which you can do the following:-
Click Replace Table to replace the currently selected table.
-
Click Add Table to add to the currently selected table.
-
(If a start list is selected) Add, delete, and edit terms in the start list.
-
-
Stop List — specifies a SAS data set that contains terms to exclude from parsing. If you include a stop list, then the terms that are included in the stop list appear in the results Term table with a Keep status of
N. For more information, see Start Lists and Stop Lists. Default data sets are provided for several languages. For more information, see SAS Text Miner Sample Data Sets. You can edit these data sets or create your own. Click the ellipsis button to open a window in which you can do the following:-
Click Replace Table to replace the currently selected table.
-
Click Add Table to add to the currently selected table.
-
(If a stop list is selected) Add, delete, and edit terms in the stop list.
-
-
Select Languages — specifies languages to keep in the document collection. If no languages are selected, all documents will be processed. To process documents with the selected languages, the input data set must include a Language variable.
Text Parsing Node Report Properties
This is the report property
that is available on the Text Parsing node:
-
Number of Terms to Display — indicates the maximum number of terms to be displayed in the Results viewer. Terms are first sorted by the number of documents in which they appear, and then the list is truncated to the maximum number. If the value of this property is
All, then all terms are displayed.
Text Parsing Node Status Properties
These are the status
properties that are displayed on the Text Parsing node:
-
Create Time — time that the node was created.
-
Run ID — identifier of the run of the node. A new identifier is assigned every time the node is run.
-
Last Error — error message, if any, from the last run.
-
Last Status — last reported status of the node.
-
Last Run Time — time at which the node was last run.
-
Run Duration — length of time required to complete the last node run.
-
Grid Host — grid host, if any, that was used for computation.
-
User-Added Node — denotes whether the node was created by a user as a SAS Enterprise Miner extension node. The value of this property is always
Nofor the Text Parsing node.
Copyright © SAS Institute Inc. All Rights Reserved.
Last updated: September 15, 2017