The TEXTMINE Procedure
Getting Started: TEXTMINE Procedure
The input data must be a table on your CAS server, and a CAS session must be set up. For more information, see the sections Using CAS Sessions and CAS Engine Librefs and Loading a SAS Data Set onto a CAS Server in Chapter 2, Shared Concepts.
The following DATA step creates the getstart data table, which contains 16 observations that have two variables, in your CAS session. The text variable contains the input documents, and the did variable contains the ID of the documents. Each row in the data table represents a document for analysis.
data mycas.getstart;
infile datalines delimiter='|' missover;
length text $150;
input text$ did;
datalines;
Reduces the cost of maintenance. Improves revenue forecast. | 1
Analytics holds the key to unlocking big data. | 2
The cost of updates between different environments is eliminated. | 3
Ensures easy deployment in the cloud or on-site. | 4
Organizations are turning to SAS for business analytics. | 5
This removes concerns about maintenance and hidden costs. | 6
Service-oriented and cloud-ready for many cloud infrastructures. | 7
Easily apply machine learning and data mining techniques to data. | 8
SAS Viya will address data analysis, modeling and learning. | 9
Helps customers reduce cost and make better decisions faster. | 10
Simple, powerful architecture ensures easy deployment in the cloud.| 11
SAS is helping industries glean insights from data. | 12
Solve complex business problems faster than ever. | 13
Shatter the barriers associated with data volume with SAS Viya. | 14
Casual business users, data scientists and application developers. | 15
Serves as the basis for innovation causing revenue growth. | 16
run;
These statements assume that your CAS engine libref is named mycas, but you can substitute any appropriately defined CAS engine libref.
The following DATA step uses the default stop list to eliminate noisy, noninformative terms:
proc cas;
loadtable caslib="ReferenceData" path="en_stoplist.sashdat";
run;
quit;
The following statements parse the input collection and use singular value decomposition followed by a rotation to discover topics that exist in the sample collection. The statements specify that all terms in the document collection, except for those on the stop list, are to be kept for generating the term-by-document matrix. The summary information about the terms in the document collection is stored in a data table named mycas.terms. The SVD statement requests that the first three singular values and singular vectors be computed. The topic assignments of the documents are stored in a data table named mycas.docpro, and the descriptive terms that define each topic are stored in a data table named mycas.topics.
proc textmine data=mycas.getstart;
doc_id did;
variables text;
parse
outterms = mycas.terms
reducef = 1
stop = mycas.en_stoplist;
svd
k = 3
outdocpro = mycas.docpro
outtopics = mycas.topics;
savestate rstore = mycas.aStoreTab;
run;
The output from this analysis is presented in Figure 2, Figure 3 and Figure 4.
Figure 1 shows the SAS log that is generated by PROC TEXTMINE; the log provides information about the default configurations used by the procedure and about the input and output files including the number of observations in each of the output tables. The mycas.terms data table lists the discovered terms. The mycas.docpro data table contains four variables: the first variable is the document ID, and the remaining three variables are obtained by projecting the original document onto the three left-singular vectors that have been rotated with the default orthogonal (varimax) rotation. The mycas.topics data table has 3 variables containing summary information of the discovered topics. Finally, the mycas.astoretab table contains a binary representation of a scoring model.
Figure 1: SAS Log
| NOTE: Stemming will be used in parsing. |
| NOTE: Tagging will be used in parsing. |
| NOTE: Noun groups will be used in parsing. |
| NOTE: No TERMWGT option is specified. TERMWGT=ENTROPY will be run by default. |
| NOTE: No CELLWGT option is specified. CELLWGT=LOG will be run by default. |
| NOTE: No ENTITIES option is specified. ENTITIES=NONE will be run by default. |
| NOTE: Topics have been requested so the document unit normalization will not |
| occur unless requested. |
| NOTE: The dense SVD solver was used for this calculation. |
| NOTE: Wrote 12532 bytes to the savestate file ASTORETAB. |
| NOTE: The Cloud Analytic Services server processed the request in 1.670414 |
| seconds. |
| NOTE: The data set MYCAS.TERMS has 134 observations and 11 variables. |
| NOTE: The data set MYCAS.DOCPRO has 16 observations and 4 variables. |
| NOTE: The data set MYCAS.TOPICS has 3 observations and 3 variables. |
| NOTE: The data set MYCAS.ASTORETAB has 1 observations and 2 variables. |
The following statements use PROC PRINT in Base SAS to show the contents of the first 10 rows of the sorted mycas.docpro data table that is generated by the TEXTMINE procedure:
data docpro;
set mycas.docpro;
run;
proc sort data=docpro;
by did;
run;
proc print data = docpro (obs=10);
run;
Figure 2 shows the output of PROC PRINT. For information about the output of the OUTDOCPRO= option, see the section The OUTDOCPRO= Data Table.
Figure 2: The mycas.docpro Data Table
| Obs | did | COL1 | COL2 | COL3 |
|---|---|---|---|---|
| 1 | 1 | 0 | 0 | 0.7460570931 |
| 2 | 2 | 0 | 0.1111856451 | 0 |
| 3 | 3 | 0 | 0 | 0.0964494952 |
| 4 | 4 | 0.8688770161 | 0 | 0 |
| 5 | 5 | 0 | 0.4742893251 | 0 |
| 6 | 6 | 0 | 0 | 0.6276285113 |
| 7 | 7 | 0.0901933118 | 0 | 0 |
| 8 | 8 | 0 | 0.0626896657 | 0 |
| 9 | 9 | 0 | 0.5236329356 | 0 |
| 10 | 10 | 0 | 0.0478786576 | 0.0703302315 |
The following statements use a DATA step and PROC PRINT to show the contents of the mycas.topics data table that is generated by the TEXTMINE procedure:
data topics; set mycas.topics; run;
proc print data = topics;
run;
Figure 3 shows the output of PROC PRINT. The three discovered topics are listed with four descriptive terms to characterize each topic.
Figure 3: The mycas.topics Data Table
| Obs | _topicid | _name | _termCutOff |
|---|---|---|---|
| 1 | 1 | easy deployment, deployment, +ensure, easy, cloud | 0.135 |
| 2 | 2 | sas, data, viya, analytics, +industry | 0.149 |
| 3 | 3 | +cost, maintenance, revenue forecast, forecast, +improve | 0.146 |
The following statements use a DATA step and the SORT and PRINT procedures to show the first 10 observations of the mycas.terms data table that is generated by the TEXTMINE procedure:
data terms; set mycas.terms; run;
proc sort data = terms; by key; run;
proc print data = terms (obs=10);
var term role freq numdocs key parent;
run;
Figure 4 shows the output of PROC PRINT, which provides details about the terms that are identified by the TEXTMINE procedure. Only the values of the variables term, role, freq, numdocs, key, and parent are displayed. For information about the output of the OUTTERMS= option, see the section The OUTTERMS= Data Table.
Figure 4: The mycas.terms Data Table
| Obs | Term | Role | Freq | numdocs | Key | Parent |
|---|---|---|---|---|---|---|
| 1 | simple | A | 1 | 1 | 1 | . |
| 2 | revenue forecast | nlpNounGroup | 1 | 1 | 2 | . |
| 3 | technique | N | 1 | 1 | 3 | . |
| 4 | different environment | nlpNounGroup | 1 | 1 | 4 | . |
| 5 | decision | N | 1 | 1 | 5 | . |
| 6 | cloud infrastructure | nlpNounGroup | 1 | 1 | 6 | . |
| 7 | hold | V | 1 | 1 | 7 | . |
| 8 | application developer | nlpNounGroup | 1 | 1 | 8 | . |
| 9 | analysis | N | 1 | 1 | 9 | . |
| 10 | analytics | N | 2 | 2 | 10 | . |
The following DATA step and statements create data and then score that data with PROC ASTORE.
data mycas.scoreData;
infile datalines delimiter='|' missover;
length text $150;
input text$ id;
datalines;
Deployment in the cloud or on-site. | 1
SAS for business analytics. | 2
Maintenance and hidden costs. | 3
run;
proc astore;
score rstore=mycas.aStoreTab
data=mycas.scoreData
out= mycas.scoreResults
copyVars= id;
run;
proc sort data=mycas.scoreResults out=scoreResults;
by id;
run;
proc print data = scoreResults;
run;
Figure 5 shows the output of PROC PRINT, which provides the topic score for the documents processed by the ASTORE PROCEDURE.
Figure 5: The mycas.scoreResults Data Table
| Obs | COL1 | COL2 | COL3 | id |
|---|---|---|---|---|
| 1 | 0.56920 | 0.00000 | 0.00000 | 1 |
| 2 | 0.00000 | 0.41840 | 0.00000 | 2 |
| 3 | 0.00000 | 0.00000 | 0.55244 | 3 |