The HPTMSCORE Procedure

Getting Started: HPTMSCORE Procedure

The following DATA steps generate two data sets. The getstart data set contains 36 observations, and the getstart_score data set contains 31 observations. Both data sets have two variables. The text variable contains the input documents, and the did variable contains the ID of the documents. Each row in the data set represents a "document" for analysis.

data getstart;
    infile datalines delimiter='|' missover;
    length text $150;
    input text$ did$;
    datalines;
        High-performance analytics hold the key to |d01
        unlocking the unprecedented business value of big data.|d02
        Organizations looking for optimal ways to gain insights|d03
        from big data in shorter reporting windows are turning to SAS.|d04
        As the gold-standard leader in business analytics |d05
        for more than 36 years,|d06
        SAS frees enterprises from the limitations of |d07
        traditional computing and enables them |d08
        to draw instant benefits from big data.|d09
        Faster Time to Insight.|d10
        From banking to retail to health care to insurance, |d11
        SAS is helping industries glean insights from data |d12
        that once took days or weeks in just hours, minutes or seconds.|d13
        It's all about getting to and analyzing relevant data faster.|d14
        Revealing previously unseen patterns, sentiments and relationships.|d15
        Identifying unknown risks.|d16
        And speeding the time to insights.|d17
        High-Performance Analytics from SAS Combining industry-leading |d18
        analytics software with high-performance computing technologies|d19
        produces fast and precise answers to unsolvable problems|d20
        and enables our customers to gain greater competitive advantage.|d21
        SAS In-Memory Analytics eliminate the need for disk-based processing|d22
        allowing for much faster analysis.|d23
        SAS In-Database executes analytic logic into the database itself |d24
        for improved agility and governance.|d25
        SAS Grid Computing creates a centrally managed,|d26
        shared environment for processing large jobs|d27
        and supporting a growing number of users efficiently.|d28
        Together, the components of this integrated, |d29
        supercharged platform are changing the decision-making landscape|d30
        and redefining how the world solves big data business problems.|d31
        Big data is a popular term used to describe the exponential growth,|d32
        availability and use of information,|d33
        both structured and unstructured.|d34
        Much has been written on the big data trend and how it can |d35
        serve as the basis for innovation, differentiation and growth.|d36
run;
data getstart_score;
    infile datalines delimiter='|' missover;
    length text $150;
    input text$ did$;
    datalines;
        Big data according to SAS|d1
        At SAS, we consider two other dimensions|d2
        when thinking about big data:|d3
        Variability. In addition to the|d4
        increasing velocities and varieties of data, data|d5
        flows can be highly inconsistent with periodic peaks.|d6
        Is something big trending in the social media?|d7
        Perhaps there is a high-profile IPO looming.|d8
        Maybe swimming with pigs in the Bahamas is suddenly|d9
        the must-do vacation activity. Daily, seasonal and|d10
        event-triggered peak data loads can be challenging|d11
        to manage - especially with social media involved.|d12
        Complexity. When you deal with huge volumes of data,|d13
        it comes from multiple sources. It is quite an|d14
        undertaking to link, match, cleanse and|d15
        transform data across systems. However,|d16
        it is necessary to connect and correlate|d17
        relationships, hierarchies and multiple data|d18
        linkages or your data can quickly spiral out of|d19
        control. Data governance can help you determine|d20
        how disparate data relates to common definitions|d21
        and how to systematically integrate structured|d22
        and unstructured data assets to produce|d23
        high-quality information that is useful,|d24
        appropriate and up-to-date.|d25
        Ultimately, regardless of the factors involved,|d26
        we believe that the term big data is relative|d27
        it applies (per Gartner's assessment)|d28
        whenever an organization's ability|d29
        to handle, store and analyze data|d30
        exceeds its current capacity.|d31
run;

The following statements use PROC HPTMINE for processing the input text data set getstart and create three data sets, outconfig, outterms, and svdu, which can be used in PROC HPTMSCORE for scoring. The statements then use PROC HPTMSCORE to score the input text data set getstart_score. The statements take the three data sets that are generated by PROC HPTMINE as input and create a SAS data set named docpro, which contains the projection of the documents in the input data set getstart_score.

 proc hptmine data     = getstart;
 doc_id      did;
 variables   text;
 parse
             outterms  = outterms
             outconfig = outconfig
             reducef   = 2;
 svd
             k         = 10
             svdu      = svdu;
 performance details;
 run;

 proc hptmscore
             data      = getstart_score
             terms     = outterms
             config    = outconfig
             svdu      = svdu
             svddocpro = docpro;
 doc_id      did;
 variable    text;
 performance details;
 run;

The output from this analysis is presented in Figure 6.1 through Figure 6.4.

Figure 6.1 shows the "Performance Information" table, which indicates that PROC HPTMSCORE executes in single-machine mode. That is, PROC HPTMSCORE runs on the machine where the SAS system is running. The table also shows that four threads are used for computing.

Figure 6.1: Performance Information

The HPTMSCORE Procedure

Performance Information
Execution ModeSingle-Machine
Number of Threads16


Figure 6.2 shows the "Procedure Task Timing" table, which provides details about how much time is used by each processing step.

Figure 6.2: Procedure Task Timing

Procedure Task Timing
TaskSecondsPercent
Parse documents1.0792.81%
Generate term-document matrix0.032.90%
Compute SVD0.032.93%
Output OUTDOCPRO table0.021.36%


Figure 6.3 shows the "Data Access Information" table, which provides the information about the data sets that the HPTMSCORE procedure has accessed and generated.

Figure 6.3: Data Access Information

Data Access Information
DataEngineRolePath
WORK.GETSTART_SCOREV9InputOn Client
WORK.DOCPROV9OutputOn Client


The following statements use PROC PRINT to show the content of the first 10 rows of the docpro data set that is generated by the HPTMSCORE procedure.

 proc print data = docpro (obs=10) round; run;

Figure 6.4 shows the output of PROC PRINT. For more information about the output of the OUTDOCPRO= option, see the description of the SVDDOCPRO= option.

Figure 6.4: The DOCPRO Data Set

ObsdidCOL1COL2COL3COL4COL5COL6COL7COL8COL9COL10
1d10.87-0.04-0.07-0.34-0.02-0.17-0.020.060.02-0.28
2d20.310.63-0.26-0.580.140.110.130.010.230.02
3d30.82-0.37-0.21-0.20-0.01-0.16-0.010.15-0.10-0.21
4d40.500.30-0.080.20-0.13-0.15-0.50-0.40-0.11-0.39
5d50.69-0.19-0.03-0.010.060.050.020.330.560.24
6d60.34-0.220.430.080.42-0.540.34-0.16-0.16-0.02
7d70.73-0.13-0.090.180.23-0.38-0.10-0.26-0.23-0.28
8d80.33-0.170.320.040.56-0.610.220.070.140.08
9d90.570.010.020.260.30-0.46-0.12-0.41-0.29-0.22
10d100.520.100.140.460.150.340.020.250.220.49


Last updated: October 30, 2018