Examples: Use a SAS Engine to Process SAS Data

Example: Assign the V9 Engine in a LIBNAME Statement

Example Code

This library assignment specifies the V9 engine.


libname myfiles v9 'library-path'; /*1*/
data myfiles.myclass;              /*2*/
   set sashelp.class;
run;
  1. The LIBNAME statement assigns the myfiles libref and the V9 engine to the location of a library. Substitute the location of your library for library-path. The location must already exist and must be accessible by the SAS Compute Server.

  2. The DATA step creates the myclass data set in the myfiles library by copying the class data set from the sashelp library.

SAS Log Showing a Successful V9 Library Assignment
1    libname myfiles v9 'library-path';
NOTE: Libref MYFILES was successfully assigned as follows:
      Engine:        V9
      Physical Name: library-path
2  !                                     
3    data myfiles.myclass;               
4       set sashelp.class;
5    run;

NOTE: There were 19 observations read from the data set SASHELP.CLASS.
NOTE: The data set MYFILES.MYCLASS has 19 observations and 5 variables.

Key Ideas

  • The LIBNAME statement is a common way to programmatically assign a library. A library assignment consists of a libref, an engine, a physical location, and options that are specific to the engine or environment.
  • When you assign a library in SAS® Viya®, the location must already exist and must be accessible by the Compute Server.
  • The shipped default Base SAS engine is BASE, which is an alias for the V9 engine.
  • If you do not specify an engine name when you create a new library, and if you have not specified the ENGINE system option, then the V9 engine is automatically selected.
  • If the library location already contains SAS files, then SAS might be able to assign the correct engine based on those files. For example, if the location contains V9 data sets only, then SAS assigns the V9 engine. However, if a library location contains a mix of different engine files, then SAS might not assign the engine you want. Therefore, specifying the engine is a best practice.

Example: Assign the SPD Engine in a LIBNAME Statement

Example Code

The following LIBNAME statement for the SPD Engine is very similar to a LIBNAME statement for the V9 engine.


libname mylib spde 'library-path'        /*1*/
   datapath=('path-for-data-partitions') /*2*/
   indexpath=('path-for-indexes');       /*3*/
  1. This portion of the LIBNAME statement assigns the mylib libref and the SPD Engine to a primary path name. The first (and usually only) metadata file for a data set is always stored in the library’s primary path.

  2. Optionally, you can assign one or more path names in the DATAPATH= option to store data partitions. Otherwise, the data partition files are stored in the primary path.

  3. Optionally, you can assign one or more path names in the INDEXPATH= option to store index files. Otherwise, the index files are stored in the primary path.

SAS Log Showing a Successful SPD Engine Library Assignment
1    libname mylib spde 'library-path' 
2       datapath=('path-for-data-partitions') 
3       indexpath=('path-for-indexes');
NOTE: Libref MYLIB was successfully assigned as follows:
      Engine:        SPDE
      Physical Name: library-path
3  !                                   

Key Ideas

  • The SPD Engine is an alternative Base SAS engine.
  • The SPD Engine is designed for high-speed processing of very large tables. The engine uses threads to read data very rapidly and in parallel, executing on multiple CPUs. Contributing to this performance is the partitioned file format, which can take advantage of distributed environments.
  • Although SPD Engine stores a data set in multiple files, you can process an SPD Engine data set very similarly to a V9 engine data set. Most of the Base SAS language works very well with an SPD Engine data set. However, the engine supports some language elements that are specific to its processing and storage optimizations. For differences from V9 engine capabilities, see the documentation.

Example: Read and Write SAS Data in Hadoop by Using the SPD Engine

Example Code

The following example assigns a Base SAS library to a Hadoop cluster.


options set=SAS_HADOOP_CONFIG_PATH='/myconfigpath';         /*1*/
options set=SAS_HADOOP_JAR_PATH='/myjarpath';

libname mydata spde '/data/abcdef' hdfs=yes accelwhere=yes; /*2*/
  1. The SET= system option defines environment variables for Hadoop. If these environment variables are already set (for example, during configuration), do not submit these lines of code. If these environment variables are not correctly set, then the LIBNAME statement produces errors in the SAS log.

  2. The LIBNAME statement assigns the mydata libref to the SPD Engine and a directory in the Hadoop cluster. The HDFS=YES argument specifies to connect to the Hadoop cluster that is defined in the Hadoop cluster configuration files. The ACCELWHERE=YES option requests that data subsetting be performed by a MapReduce program in the Hadoop cluster.

Key Ideas

  • The SPD Engine is an alternative Base SAS engine that can read and write SAS data on a traditional file system or on Hadoop. The engine does not require you to configure additional SAS products such as SAS/ACCESS, but you must be running a supported Hadoop distribution.
  • Customers often choose Hadoop for low-cost storage of very large data. The distributed storage and processing of the SPD Engine works well with the Hadoop file system (HDFS). In addition, the engine can optimize most WHERE expressions by automatically submitting a MapReduce program in the Hadoop cluster.
  • The engine can read SPD Engine data sets on Hadoop. After you use the engine to store a data set on Hadoop, you can use most of the Base SAS language for processing the data. However, the engine supports some language elements that are specific to its processing and storage on Hadoop.

Example: Avoid Truncation by Using the CVP Engine with the V9 Engine

Example Code

To run this example, first create a data set named myclass as in Example: Assign the V9 Engine in a LIBNAME Statement. Run PROC CONTENTS to see the length of the variables:

libname myfiles v9 'library-path-1';
proc contents data=myfiles.myclass;
run;

In the PROC CONTENTS output, notice the two character variables. Name has a length of 8, and Sex has a length of 1.

PROC CONTENTS Showing Variable Lengths before Expansion
Portion of PROC CONTENTS Before CVP Expansion

The example below uses the CVP engine with the V9 engine to expand the size of character variables. The CVP engine can help you avoid truncation if you copy a data set to an encoding that uses more bytes to represent the characters.


libname srclib cvp 'library-path-1' cvpengine=v9 cvpmult=2.5; /*1*/
libname target v9 'library-path-2';                           /*2*/
proc copy in=srclib out=target;                               /*3*/
   select myclass;
run;

proc contents data=target.myclass;                            /*4*/
run; 
  1. This LIBNAME statement assigns the srclib library to the CVP engine and the location of the data that you want to copy. The CVPENGINE= option specifies the V9 engine as the underlying engine to process the data. The CVPMULT= option specifies a multiplication factor of 2.5 to expand all character variables. If this option is not specified, the CVP engine automatically chooses a multiplier value.

  2. This LIBNAME statement assigns the target library to contain the copied data.

  3. The COPY procedure with SELECT statement copies the myclass data set to the target library. During the copy, the CVP engine expands the character variable lengths 2.5 times larger.

  4. The CONTENTS procedure shows that the lengths of the character variables have been multiplied by 2.5:

    For Name, 8 × 2.5 = 20.

    For Sex, 1 × 2.5 = 2.5, which is 3 when rounded up to a whole number.

PROC CONTENTS Showing Variable Lengths after Expansion
Portion of PROC CONTENTS After CVP Expansion

Key Ideas

  • When you copy a data set to an encoding that uses more bytes to represent the characters, truncation might occur if the column length does not accommodate the larger character size. For example, a character might be represented in wlatin1 encoding as one byte but in UTF-8 as two bytes.
  • If an error in the log states character data was lost during transcoding, it usually indicates that truncation has occurred. You can troubleshoot the error by using the CVP engine to expand the length of character variables.
  • By default, the CVP engine automatically chooses a multiplier value. The automatic value is usually sufficient to avoid truncation. You can also use CVPMULTIPLIER= to specify it yourself.
  • The default CVPFORMATWIDTH=YES option expands the length for formats but does not affect user-defined formats. For user-defined formats, see Example: Avoid Truncation in Formats When Using PROC FORMAT in SAS V9 LIBNAME Engine: Reference.
  • If you copy data sets to a different operating environment, or to a different character encoding, you probably want to specify options such as NOCLONE for PROC COPY. NOCLONE does not copy certain data set attributes, such as data representation and character encoding, so that the library is compatible in the target environment.

    As an alternative to the NOCLONE option, you can use the OVERRIDE= option to specify ENCODING= and OUTREP= options. This method keeps other data set attributes from the source library, so make sure that you want those attributes.

  • PROC COPY does not preserve an audit trail. Under CEDA processing, PROC COPY does not copy indexes or integrity constraints. If you are migrating, then PROC MIGRATE could be a better choice. See Example: Avoid Truncation When Migrating a SAS Library by Using the CVP Engine in SAS V9 LIBNAME Engine: Reference.

See Also

Example: Load a SAS Data Set to a CAS Server

Example Code

The following example uses the DATA step to load a SAS data set into memory as a SAS Cloud Analytic Services (CAS) table.


cas casauto host="cloud.example.com" port=5570; /*1*/

libname mycas cas;                              /*2*/
data mycas.cars (promote=yes);                  /*3*/
   set sashelp.cars; 
run;
proc contents data=mycas.cars;                  /*4*/
run;
  1. The CAS statement starts a CAS session and specifies casauto as the CAS session name. Use your connection information in the HOST= and PORT= options.

  2. The LIBNAME statement assigns the mycas libref to the CAS engine. The SESSREF= LIBNAME option is not specified, so the engine uses the casauto session.

  3. The DATA step copies the SAS data set sashelp.cars to the CAS session. The PROMOTE=YES data set option promotes the table with global scope.

  4. PROC CONTENTS shows the mycas.cars table is available on the CAS server for the duration of the session. After data is loaded into memory, subsequent steps can process the data in memory. Loading and processing are done in separate steps.

Portion of PROC CONTENTS Output for mycas.cars
PROC CONTENTS of mycas.cars

Key Ideas

  • You can submit a LIBNAME statement that uses the CAS engine to connect your SAS Compute Server session to a CAS session. You must have access to a CAS server and an existing CAS session.
  • The CAS LIBNAME engine with the DATA step is one way of loading SAS data to the CAS server as an in-memory table. Other methods might be more efficient for large tables.
  • After you load data to the CAS server, you can execute SAS procedures or the DATA step from your SAS session by referencing the SAS libref and table name. You do not process the table in memory in the same DATA step that you use to load the table into memory. Loading and processing are done in separate steps.
  • Tables are not automatically saved when they are loaded to a caslib. You can use the CASUTIL procedure to save tables. Native CAS tables have the file extension .sashdat.

See Also

Last updated: August 14, 2026