The HPTMSCORE Procedure
Getting Started: HPTMSCORE Procedure
The following DATA steps generate two data sets. The getstart data set contains 36 observations, and the getstart_score data set contains 31 observations. Both data sets have two variables. The text variable contains the input documents, and the did variable contains the ID of the documents. Each row in the data set represents a "document" for analysis.
data getstart;
infile datalines delimiter='|' missover;
length text $150;
input text$ did$;
datalines;
High-performance analytics hold the key to |d01
unlocking the unprecedented business value of big data.|d02
Organizations looking for optimal ways to gain insights|d03
from big data in shorter reporting windows are turning to SAS.|d04
As the gold-standard leader in business analytics |d05
for more than 36 years,|d06
SAS frees enterprises from the limitations of |d07
traditional computing and enables them |d08
to draw instant benefits from big data.|d09
Faster Time to Insight.|d10
From banking to retail to health care to insurance, |d11
SAS is helping industries glean insights from data |d12
that once took days or weeks in just hours, minutes or seconds.|d13
It's all about getting to and analyzing relevant data faster.|d14
Revealing previously unseen patterns, sentiments and relationships.|d15
Identifying unknown risks.|d16
And speeding the time to insights.|d17
High-Performance Analytics from SAS Combining industry-leading |d18
analytics software with high-performance computing technologies|d19
produces fast and precise answers to unsolvable problems|d20
and enables our customers to gain greater competitive advantage.|d21
SAS In-Memory Analytics eliminate the need for disk-based processing|d22
allowing for much faster analysis.|d23
SAS In-Database executes analytic logic into the database itself |d24
for improved agility and governance.|d25
SAS Grid Computing creates a centrally managed,|d26
shared environment for processing large jobs|d27
and supporting a growing number of users efficiently.|d28
Together, the components of this integrated, |d29
supercharged platform are changing the decision-making landscape|d30
and redefining how the world solves big data business problems.|d31
Big data is a popular term used to describe the exponential growth,|d32
availability and use of information,|d33
both structured and unstructured.|d34
Much has been written on the big data trend and how it can |d35
serve as the basis for innovation, differentiation and growth.|d36
run;
data getstart_score;
infile datalines delimiter='|' missover;
length text $150;
input text$ did$;
datalines;
Big data according to SAS|d1
At SAS, we consider two other dimensions|d2
when thinking about big data:|d3
Variability. In addition to the|d4
increasing velocities and varieties of data, data|d5
flows can be highly inconsistent with periodic peaks.|d6
Is something big trending in the social media?|d7
Perhaps there is a high-profile IPO looming.|d8
Maybe swimming with pigs in the Bahamas is suddenly|d9
the must-do vacation activity. Daily, seasonal and|d10
event-triggered peak data loads can be challenging|d11
to manage - especially with social media involved.|d12
Complexity. When you deal with huge volumes of data,|d13
it comes from multiple sources. It is quite an|d14
undertaking to link, match, cleanse and|d15
transform data across systems. However,|d16
it is necessary to connect and correlate|d17
relationships, hierarchies and multiple data|d18
linkages or your data can quickly spiral out of|d19
control. Data governance can help you determine|d20
how disparate data relates to common definitions|d21
and how to systematically integrate structured|d22
and unstructured data assets to produce|d23
high-quality information that is useful,|d24
appropriate and up-to-date.|d25
Ultimately, regardless of the factors involved,|d26
we believe that the term big data is relative|d27
it applies (per Gartner's assessment)|d28
whenever an organization's ability|d29
to handle, store and analyze data|d30
exceeds its current capacity.|d31
run;
The following statements use PROC HPTMINE for processing the input text data set getstart and create three data sets, outconfig, outterms, and svdu, which can be used in PROC HPTMSCORE for scoring. The statements then use PROC HPTMSCORE to score the input text data set getstart_score. The statements take the three data sets that are generated by PROC HPTMINE as input and create a SAS data set named docpro, which contains the projection of the documents in the input data set getstart_score.
proc hptmine data = getstart;
doc_id did;
variables text;
parse
outterms = outterms
outconfig = outconfig
reducef = 2;
svd
k = 10
svdu = svdu;
performance details;
run;
proc hptmscore
data = getstart_score
terms = outterms
config = outconfig
svdu = svdu
svddocpro = docpro;
doc_id did;
variable text;
performance details;
run;
The output from this analysis is presented in Figure 6.1 through Figure 6.4.
Figure 6.1 shows the "Performance Information" table, which indicates that PROC HPTMSCORE executes in single-machine mode. That is, PROC HPTMSCORE runs on the machine where the SAS system is running. The table also shows that four threads are used for computing.
Figure 6.1: Performance Information
Figure 6.2 shows the "Procedure Task Timing" table, which provides details about how much time is used by each processing step.
Figure 6.2: Procedure Task Timing
Figure 6.3 shows the "Data Access Information" table, which provides the information about the data sets that the HPTMSCORE procedure has accessed and generated.
Figure 6.3: Data Access Information
The following statements use PROC PRINT to show the content of the first 10 rows of the docpro data set that is generated by the HPTMSCORE procedure.
proc print data = docpro (obs=10) round; run;
Figure 6.4 shows the output of PROC PRINT. For more information about the output of the OUTDOCPRO= option, see the description of the SVDDOCPRO= option.
Figure 6.4: The DOCPRO Data Set
| Obs | did | COL1 | COL2 | COL3 | COL4 | COL5 | COL6 | COL7 | COL8 | COL9 | COL10 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | d1 | 0.87 | -0.04 | -0.07 | -0.34 | -0.02 | -0.17 | -0.02 | 0.06 | 0.02 | -0.28 |
| 2 | d2 | 0.31 | 0.63 | -0.26 | -0.58 | 0.14 | 0.11 | 0.13 | 0.01 | 0.23 | 0.02 |
| 3 | d3 | 0.82 | -0.37 | -0.21 | -0.20 | -0.01 | -0.16 | -0.01 | 0.15 | -0.10 | -0.21 |
| 4 | d4 | 0.50 | 0.30 | -0.08 | 0.20 | -0.13 | -0.15 | -0.50 | -0.40 | -0.11 | -0.39 |
| 5 | d5 | 0.69 | -0.19 | -0.03 | -0.01 | 0.06 | 0.05 | 0.02 | 0.33 | 0.56 | 0.24 |
| 6 | d6 | 0.34 | -0.22 | 0.43 | 0.08 | 0.42 | -0.54 | 0.34 | -0.16 | -0.16 | -0.02 |
| 7 | d7 | 0.73 | -0.13 | -0.09 | 0.18 | 0.23 | -0.38 | -0.10 | -0.26 | -0.23 | -0.28 |
| 8 | d8 | 0.33 | -0.17 | 0.32 | 0.04 | 0.56 | -0.61 | 0.22 | 0.07 | 0.14 | 0.08 |
| 9 | d9 | 0.57 | 0.01 | 0.02 | 0.26 | 0.30 | -0.46 | -0.12 | -0.41 | -0.29 | -0.22 |
| 10 | d10 | 0.52 | 0.10 | 0.14 | 0.46 | 0.15 | 0.34 | 0.02 | 0.25 | 0.22 | 0.49 |