GPREG Procedure

Optimization

PROC GPREG supports the stochastic gradient descent (SGD) and Adam optimizations and uses options that are specified in the OPTIMIZATION statement. For the widely used SGD optimization (which is the default), the procedure supports momentum, adaptive learning rate, and adaptive decay. These options correspond to the momentum and ADADELTA variant of SGD. The validation error for one iteration is the previous iteration’s loss. For more information about Adam optimization, see Kingma and Ba (2015). For more information about ADADELTA, see Zeiler (2012).

Last updated: August 06, 2026