Description

Voyager is a heterogeneous system designed to support complex deep learning AI workflows. The system features 42 Intel Habana Gaudi training nodes, each with 8 training processors (336 in total). Each training node has 512GB of memory and 6.4TB of node local NVMe storage. The Gaudi training processors feature specialized hardware units for AI, HBM2, and on-chip high-speed Ethernet. The on-chip ethernet ports are used in a non-blocking all-to-all network between processors on a node and the remaining ports are aggregated into 6 400G connections on each node that are plugged into a 400G Arista switch to provide scale out of network. Voyager also has two first-generation inference nodes, each with 8 inference processors (16 in total). In addition to the custom AI hardware, the system also has 36 Intel x86 processors compute nodes for general purpose computing and data processing. Voyager features 3PB of storage currently deployed as a Ceph filesystem.

Resource ID
2088
Global Resource ID
voyager.sdsc.access-ci.org
Resource Type
Compute
Latest Status
production
Latest Status Begin
Project Affiliation
ACCESS
Organization Name
San Diego Supercomputer Center
RP Description

Voyager is an innovative AI testbed system consisting of Habana Gaudi training and first-generation Habana inference processors, along with a high-performance, low latency 400 gigabit-per-second interconnect from Arista. Voyager is particularly well suited for research in science and engineering that is increasingly dependent upon artificial intelligence and deep learning as a critical element in the experimental and/or computational work.

MFA Required
Off
File Transfer Methods
Transfer Method
Globus (Coming Soon)
Storage Text

See [Storage] for information. 

Jobs Information

Voyager runs Kubernetes, an open-source platform for managing containerized workloads and services. A Kubernetes cluster consists of a set of worker machines, called nodes, that run containerized applications. The application workloads are executed by placing containers into Pods to run on nodes. The resources required by the Pods are specified in YAML files. 

For computer, inference, or gaudi examples: Basic Jobs

Storage Filesystems
Directory
home
File System Path
/home/username
Quota (Deprecated)
200GB
Quota size
200
Backup Policy
Not backed up
Notes
The home directory is limited in space and should be used only for source code storage. User will have access to 200GB in /home. Users should keep usage on $HOME under 200GB.
Directory
projects
File System Path
/voyager/projects/project/username
Quota (Deprecated)
153TB
Quota size
153000
Backup Policy
Not backed up
Notes
NSF mounted project space
Directory
scratch
File System Path
emptyDir
Quota (Deprecated)
-
Purge Policy
Purged when the pod is removed from the node.
Backup Policy
Not backed up
Notes
Needs to be mounted as an emptyDir. Only exists while the Pod is running on the node.
Directory
ceph
File System Path
/voyager/ceph/users/username
Quota (Deprecated)
2,000,000 files
Backup Policy
Not backed up
Notes
The SDSC Ceph file system ( /voyager/ceph/users/$USER) IS NOT an archival file
Queue Specifications
Queue Name
gaudi
Purpose
Designed for high-performance AI training using Intel Habana Gaudi training processors.
CPU Type
Intel Xeon Gold 6336
GPU Type
Supermicro X12 Gaudi Training Nodes
RAM (Deprecated)
512 GB DDR4 DRAM
GPU Count
8
Node RAM
512
Queue Name
goya
Purpose
Dedicated for Habana Gaudi model inference, utilizing 2 first-generation nodes.
CPU Type
Xeon Gold 6240
RAM (Deprecated)
384 GB DDR4 DRAM
CPU Count
40
Node RAM
384
Queue Name
compute
Purpose
General-purpose pre/post-data processing
CPU Type
Intelx86
RAM (Deprecated)
384 GB
CPU Count
2
Node RAM
384