SAS Cloud Analytic Services: Troubleshooting
- Insufficient Disk Space (CAS_DISK_CACHE)
- Cannot Access Permstore
- CAS Pods in Pending Status
- Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error
- CAS Pods Restart in a Kubernetes 1.28 Cluster That Is Configured with Linux cgroup v2
- CAS Resources Are Defined Even Though CAS Servers Are Removed from the SAS Viya Programming-Only Deployment
- Permstore Item Cannot be Processed During Initial CAS Server Start Up
Insufficient Disk Space (CAS_DISK_CACHE)
Issue: An
upload fails with a message such as Failed to
open temporary file for upload (80BFE801): /tmp/cascache1/_f_43d6c87c_7f5d854996e8.sas7bdat
Explanation: Insufficient disk space in CAS_DISK_CACHE on the CAS controller.
Resolution: Add more disk space.
Cannot Access Permstore
Issue: The
following message is displayed. A CAS server
is currently running and has exclusive access to the access control
storage location.
Explanation: This error message is displayed when a CAS server fails to access the permstore and terminates because the permstore is already being accessed by another CAS server.
CAS Pods in Pending Status
Issue: One
or more CAS pods are stuck in Pending status.
Explanation: The CAS pods that are in a pending state cannot currently be scheduled in the Kubernetes cluster. This state can occur if the available nodes have insufficient resources (memory or CPU), are missing the appropriate taints or labels, or the pod or node affinity rules are not met.
Resolution: If you describe the pod, the scheduler provides a reason for the pending state.
kubectl describe pods podName
Inadequate resources can be fixed by adjusting resource requests, deleting pods, or adding new nodes to your cluster. Ensure that nodes have the appropriate taints and labels and that all the affinity rules are met.
Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error
Issue: The following message
is displayed in the server log. Child terminated by signal: PID nnn,
signal 9, status 0x00000009. The syslog for the Kubernetes node(s)
where CAS pods are scheduled, also records out-of-memory (OOM) messages that varies
by the
provider.
Explanation: This error message indicates that a session was killed. One possible cause is that the OOM killer terminated a session because of high memory use.
Tips for reducing OOM error issues follow:
TIP 1:
Reducing the percentage of node memory that can be used for writeback (lazy writes) can reduce chances of cgroup OOM killer events for workloads that have heavy disk write activity to CAS node local storage like cascache.
You can modify the following two OS configurations:
- Reduce
vm.dirty_ratiofrom 20 to 10. - Reduce
vm.dirty_background_ratiofrom 10 to 5.
Refer the documentation of your respective
environment on how to modify these settings on CAS nodes. If you have access to the
nodes
where the CAS pods are running, then you can run ssh to each CAS
node, sudo to root, and run sysctl to modify
the suggested values.
TIP 2:
Reducing container memory request = limit to a
percentage (for example, 90%) of the Remaining Allocatable memory on the node can
help
mitigate issues that are due to a host OOM error. You can see the values of Allocatable
and
Allocated Memory by running
kubectl -n name-of-namespace describe node.
Note that by reducing this value, you might not have enough space for your in-memory
CAS
tables.
TIP 3:
To prevent OOM failures, you can also set a backing store to support memory allocations. For more information, see Backing Store for CAS Memory Allocations.
- All the tips provided also apply to the compute node if any of the SAS processes are being OOM killed.
- Kubernetes 1.32 has introduced a
configuration flag - singleprocessOOMKill to disable the OOM
group kill for
cgroup v2. For more information, see Kubelet Configuration. Contact your Kubernetes vendor for instructions on modifying kublet configuration flags.
CAS Pods Restart in a Kubernetes 1.28 Cluster That Is Configured with Linux cgroup v2
Explanation: Kubernetes
1.28 implemented a change in behavior that impacts CAS if your cluster is configured
with
Linux cgroup v2 and encounters a Linux out-of-memory (OOM) kill on
CAS sessions.
To determine whether your site is experiencing Linux OOM kill situations, see Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error.
Resolution:
- For information about how to tune CAS to avoid Linux OOM failures, see Session Terminated by cgroup OOM Error or CAS Pod Evicted by Host OOM Error.
- To prevent OOM failures, you can also set a backing store to support memory allocations. For more information, see Backing Store for CAS Memory Allocations.
- Kubernetes 1.32 has introduced a configuration
flag - singleprocessOOMKill to disable the OOM group kill for
cgroup v2. For more information, see Kubelet Configuration. Contact your Kubernetes vendor for instructions about modifying kubelet configuration flags.
CAS Resources Are Defined Even Though CAS Servers Are Removed from the SAS Viya Programming-Only Deployment
Explanation: If CAS Servers are not defined in your SAS Viya Programming-Only deployment, the following three CAS resources are still deployed:
sas-cas-controlpodsas-cas-operatorpodcasdeploymentCustom Resource Definition (CRD)
Resolution: These resources are used by other services in your SAS Viya Programming-Only deployment and must remain deployed. The pods require minimal resources and do not require a dedicated CAS NodePool.
Permstore Item Cannot be Processed During Initial CAS Server Start Up
Explanation: When starting a CAS server, if there are any issues while processing a caslib item store, that caslib is removed from the in-memory data structure and does not appear in the list of caslibs after server initialization.
CAS server changes the
.casitm suffix of the file for this caslib in the permstore to the
.problem or .bad suffix. The server start-up
continues to process any additional caslibs in the permstore.
You might see one of the following error messages in the log:
Error processing item store: name
Error processing caslib AutoData: /path/permstore/UUID.casitm, s:table_addCaslib{….
Moving damaged item store from /path/permstore/UUID.casitm to /path/permstore/UUID.casitm.problem.
Resolution: You can resolve this error in either of these ways:
- Re-create the missing caslib.
- Rename the file with the
.problemor.badsuffix back to.casitmand restart the CAS server to attempt processing this caslib item store again.