Document navigation

Overview

ZStack Cloud supports the Graphics Processing Unit (GPU) passthrough feature. A physical GPU (pGPU), along with all its peripherals (including the GPU graphics card, GPU sound card, and other small devices on the GPU), can be passed through to a VM instance as a group. This allows the VM instance to leverage the powerful parallel computing capabilities of the pGPU. This feature applies to 3D rendering, high-definition transcoding and decoding, as well as High-Performance Computing (HPC) scenarios that demand high computational intensity.

Through passthrough, ZStack Cloud can monitor the load status of GPU devices in real time and sound alarm in abnormal conditions, providing you with a complete and convenient solution for GPU operation and maintenance.

The GPU passthrough feature applies to the pGPUs of the following models.
Vendor Model
NVIDIA Nvidia RTX 6000Ada, Nvidia RTX A6000
GeForce RTX 5090, Geforce RTX 4090, Nvidia RTX 3090
Quadro RTX 8000, Quadro RTX 6000
M4000, P2000
GTX 1650/1660, GTX 1060ti
H100, H200, H800, H20
Note: Supported only by ZStack Cloud H84R ISO.
Nvidia L40, Nvidia L20, Nvidia L4
Nvidia A100, Nvidia A30
Nvidia A40, Nvidia A16, Nvidia A10, Nvidia T4
Tesla V100, Tesla P4/6/40/100, M6/10/60
K6000
AMD Radeon v620, Radeon RX5700
RRO W7800
Note: Passthrough to VM instances running Windows operating system is not currently supported.
FirePro S7150, FirePro S7150X2
Huawei Atlas 300i pro
Note: Only GPUs of this model installed on ARM hosts can be passthroughed.
910B3/4
Note: Only GPUs of this model installed on ARM hosts can be passthroughed.
Hygon Z100, Z100L
K100-AI
Enflame S60
Iluvatar CoreX Zhikai MR-V100
Tiangai
Vastai SV100, SG100
Note: We do not recommend passthrough SG100 directly. It is recommended to slice it into vGPUs before use.
KUNLUNXIN P800
Alibaba PPU PPU-ZW810E
Other MetaX N100, MOORE THREADS, Cambricon, and others

Application Scenarios

3D Rendering

Both Pre-Rendering (or Offline Rendering) and Real-Time Rendering (or Online Rendering) for 3D computer graphics are time-consuming. Pre-Rendering, a technology often used in film-making, demands high computational intensity that needs to be provided by a substantial number of servers. Real-Time Rendering, a technology commonly used in 3D video games, relies on GPUs to complete this process.

Nowadays, with the rapid development of the GPU technology, considerable 3D rendering cases are accomplished in GPU server clusters. With the GPU passthrough feature provided by ZStack Cloud, you can perform centralized cluster management efficiently with a low GPU performance loss (less than 5%). Combined with intelligent monitoring software and the billing service of ZStack Cloud, the feature provides you with a comprehensive, convenient, and efficient rendering farm solution.

Figure 1. 3D Rendering


Artificial Intelligence

GPU supports deep learning tasks due to its compute capability. Since the launch of TensorFlow, a tool for creating neural networks, by Google, GPUs have been gaining favour among research institutes and companies, and adopted as infrastructures.

Take the NVIDIA P100 graphic card as an example, when passed through to a VM instance on ZStack Cloud, test results show that its performance is nearly identical to specifications, fully meeting the infrastructure requirements for large-scale model training.

Figure 2. Artificial Intelligence


Cloud Gaming

With the expansion of broadband networks and the widespread usage of mobile devices, a new trend in gaming has emerged where the computational load is shifted to the cloud, while the client merely handles display and control. In this model, cloud servers undertake the task of 3D game rendering. Combined with GPUs, these cloud servers can encode each frame instantly and stream it to any devices connected to the network.

Leveraging the GPU and server CPU capabilities, this cloud gaming model uses ZStack Cloud’s GPU passthrough feature to create a virtual gaming environment with enhanced isolation and smooth computation and rendering, providing you with an improved gaming experience.

Figure 3. Cloud Gaming


Virtual Desktop Infrastructure (VDI)

GPU has always been an integral part of Virtual Desktop Infrastructure (VDI), for it not only enhances visual experience, but also takes on a primary computational role in special applications, making it a good substitute for traditional PC graphics workstation. Also, GPU enables you to perform 3D designs in a safer environment.

Through the GPU passthrough feature provided by ZStack Cloud, coupled with protocols such as RDP (Remote Desktop Protocol) or PCoIP (PC over IP), you can leverage the GPU to ensure the smooth operation of 3D designs or games and gain the VDI experience as great as that in a physical environment.

Figure 4. VDI


Considerations

To use the GPU passthrough feature, note the following considerations:
  • A single VM instance can attach multiple pGPUs simultaneously but does not support attaching both pGPUs and vGPUs at the same time.
  • After the GPU is passed through to a VM instance, functions such as changing host, changing host and primary storage, or HA may not work well.
  • We recommend that you stop the VM instance before detaching GPUs. Otherwise, a blue screen or a suspension may occur.
  • For GPU passthrough to a Windows VM instance, you need to install the operating system using the UEFI boot mode.
  • Enter the global settings and find PCI Hot Plugging. The default is true. If a hardware incompatibility error occurs during hot plugging or a hardware device does not support hot plugging, you can set this parameter to false.
  • To use the GPU passthrough feature and obtain real-time GPU load monitoring data, you need to install GPU drivers and GuestToos on both the host and the VM instance. The recommended versions for GPU drivers are as follows.
    GPU Recommended Version for Host Recommended Version for VM Instance
    NVIDIA NVIDIA-Linux-x86_64-510.47.03-grid.run Latest version recommended by NVIDIA. For more information, see NVIDIA Official Documentation
    AMD rocm-smi 6.1.2 and later versions rocm-smi 6.1.2 and later versions
    Note: If the VM instance uses the OS of RHEL7 series, make sure that the VM kernel is of 4.18.0 or later versions.
    Hygon

    rock-5.2.0-5.16.29-V01.13.run

    After Hygon uses the GPU passthrough feature, it does not have access to load monitoring data via GPU drivers.
    Huawei

    Ascend-hdk-310p-npu-driver_24.1.rc1_linux-aarch64.run

    Ascend-hdk-310p-npu-driver_24.1.rc1_linux-aarch64.run

    Iluvatar CoreX
    • x86: corex-installer-linux64-4.0.1_x86_64_10.2.run
    • ARM: corex-installer-linux64-4.0.1_arm64_10.2.run
    • x86: corex-installer-linux64-4.0.1_x86_64_10.2.run
    • ARM: corex-installer-linux64-4.0.1_arm64_10.2.run

Preparations

To utilize the GPU passthrough feature on ZStack Cloud, the following preparations are necessary:
  • Enter the BIOS settings of your host and enable the Intel VT-d feature or the AMD IOMMU feature. Verify that the host kernel has IOMMU support enabled.
  • For NVIDIA A series users, you have to ensure that your host kernel is upgraded to 4.18 version, and GCC to 8.3.1 version.
  • Make sure that the version of your driver matches that of your GPU device.
    Note: For more information about driver service or installation methods, please contact GPU device supplier. For recommended driver versions, see Considerations.
  • Configure the global settings: Click Settings > Platform Setting > Global Setting. The following global settings are related to the pGPU passthrough feature, you can adjust them as needed:
    • Hide KVM Virtualization Flag: If you use NVIDIA graphics cards for VM instances, you need to set this parameter to true.

      This setting specifies whether to hide the KVM virtualization flag. The default is false. If set to true, <hidden state='on'> is inserted in the <kvm> field of the XML file defined for a newly started VM instance.

    • VM Hyper-V Virtualization: If you use NVIDIA graphics cards for VM instances, you need to set this parameter to true.

      This setting specifies whether to enable Hyper-V emulation for a VM instance. The default is false.

    • PCI Hot Plugging: Configure it as needed.

      This setting specifies whether to enable hot plugging of PCI devices for a VM instance. The default is true. If a hardware incompatibility error occurs during hot plugging or a hardware device does not support hot plugging, you can set this parameter to false.

    • vGPU Default Quota: Configure it as needed.

      This setting is used to set the quota of GPU devices (including pGPUs and vGPUs) that an account/project can use. The default is 20.

Typical Scenarios

About this task

To pass through a GPU device:
  1. Enable IOMMU in the host BIOS.
  2. Set ROM (Optional).
  3. Attach pGPUs to the VM instance.
  4. Install GPU drivers on the VM instance.
  5. Install GuestTools.

Before using the GPU passthrough feature, ensure that requirements in Preparations have been met completely and correctly. Below is detailed operational procedures for setting up GPU passthrough:

Procedure

  1. Enable IOMMU in the host BIOS.
    Make sure that Intel VT-d or AMD IOMMU is enabled in the host BIOS before you enable the IOMMU option on ZStack Cloud.
    • For adding host: Choose Resource Center > Hardware > Computing Facility > Host > Add Host. Then set Scan Host IOMMU Setting to true to enable IOMMU.
      Figure 5. Add Host and Enable IOMMU


    • For added host: Select one added host and set IOMMU State to true on its details page. Reboot the host and the IOMMU setting will take effect.
      Figure 6. Enable IOMMU for Added Host

    Note: After enabling IOMMU in the host, you have to make sure that IOMMU Status on the same page is available. Otherwise, the GPU passthrough feature cannot work as expected. If IOMMU State is enabled, yet IOMMU Status is unavailable, the reasons could be as follows:
    • The IOMMU setting is enabled but the host is not rebooted. Just reboot the host.
    • If a host configuration error occurs, please enter the host BIOS and enable Intel VT-d or AMD IOMMU.
  2. Set ROM (Optional).

    ROM is a configuration file used for passing through pGPUs. The ROM file you upload is updated to the pGPU that the specification specifies.

    ZStack Cloud provides built-in basic ROM files, which can satisfy the majority of passthrough needs. Additionally, you can obtain other ROM files you need on the official site of the GPU supplier and then upload them.

    On the main menu of ZStack Cloud, choose Resource Center > Resource Pool > Compute Configuration > GPU Specification. Select the pGPU that you need and click Actions > Set ROM. A pop-up menu appears and you can upload ROM file herein.

    Figure 7. Set ROM


    Note: When uploading ROM file, note the following cases:
    • The upoaded ROM file must match the target pGPU's specifications and versions. Otherwise, the passed-through pGPU cannot work properly.
    • The latest ROM file that you upload overwrites the previous ROM file.
  3. Attach pGPUs to the VM instance.
    This step enables the pGPU to be directly passed through to the VM instance. On the ZStack Cloud, you can use the following methods to attach pGPUs to the VM instance:
    • Method One: Create a VM Instance and Attach pGPUs to It
      To create a VM instance, you need to choose Resource Center > Resource Pool > Virtual Resource > VM Instance > Create VM Instance. After you complete Basic Configuration, you come to next stage, that is, Resource Configurations. We support two GPU attachment policies including attaching GPU specification and attaching GPU device. Set the following parameters as you need:
      • Attach GPU Specification: Select a GPU specification and the system allocates GPU device(s) to the VM instance according to this specification. You can choose whether to make these GPU device(s) automatically detached when the VM is stopped. If you set Auto Detach to true, these GPU device(s) would be automatically detached when the VM is stopped. When the VM restarts, the system re-allocates GPU device(s) to it according to the GPU specification. If you set Auto Detach to false, the VM would keep these GPU devices attached and continue using them when it restarts.
        Figure 8. Attach GPU Specification


      • Attach GPU Device: Select a GPU device and attach it directly to the VM instance.
        Figure 9. Attach GPU Device


      After completing the configurations, click OK. Then you'll get a VM instance attached to a pGPU.
    • Method Two: Attach pGPUs to an Existing VM Instance

      In the management interface of VM Instance, select the name of one existing VM instance and enter its details page. Choose Configuration info on the top row. Find pGPU Device on this page and click Attach.

      Figure 10. Attach a pGPU Device


      • A single VM instance can attach multiple pGPUs simultaneously but does not support attaching both pGPUs and vGPUs at the same time.
      • If you want to detach a GPU device, select it and click Actions > Detach.
        Note: If you detach a pGPU from a running VM instance, a blue screen or suspension may occur. We recommend stopping the VM instance before performing detach operation.
    • Method Three: Attach pGPUs to Existing VM Instances

      Select one or more stopped VM instances in the management interface of VM Instance, and click Bulk Action > System Configurations > Set GPU Policy. Then you have two options to choose, that is, attach GPU specification or attach GPU device.

      Figure 11. Attach pGPUs in Bulk


  4. Install GPU drivers on the VM instance.
    After attaching a GPU device to the VM instance, you need to install corresponding GPU drivers. Download paths for AMD or NVIDIA drivers are as follows:
    • Linux OS supports AMD GPU drivers (computing, gaming or professional series included). Linux has built-in community driver, providing you with services such as compute acceleration, displaying acceleration, and checking GPU monitoring. Click here to install the official driver.
    • Linux OS supports NVIDA GPU drivers (computing, gaming or professional series included). Linux has built-in community driver, providing you with services such as compute acceleration, displaying acceleration, and checking GPU monitoring. Click here to install the official driver.
    • Windows OS supports AMD GPU drivers (computing, gaming or professional series included). Click here to download the proper driver that matches the type of GPU and the version of OS.
    • Window OS only supports the computing series of NVIDA GPU drivers. Click here to download the proper driver that matches the type of GPU and the version of OS.
    Note: If you use AMD Firepro S7150 X2 that features two physical GPUs (pGPUs), and want the two pGPUs to be passed through to different VMs, please install the drivers of the same versions on your VMs, so that GPU monitoring data can be obtained normally.
    The procedures of installing the GPU driver vary depending on the version of your GPU device. For more information, you can contact GPU suppliers for assistance. This chapter takes installing a NVIDA GPU on the Linux VM instance as an example. You can refer to the operational procedures below:
    1. Obtain the required driver installation packages:

      Obtain the driver and CUDA toolkit compatible with the GPU device.

    2. Disable the Nouveau kernel driver:
      If NVIDIA drivers conflict with the Nouveau kernel driver, you can run the command lsmod | grep nouveau to check whether the Nouveau driver has been installed. If the output data suggests the Nouveau driver has been installed, you can perform the following operations to disable it. If no output is displayed, just skip this procedure.
      # touch  /etc/modprobe.d/nvidia-installer-disable-nouveau.conf  # Create a file and save the two lines below into it
      blacklist nouveau
      options nouveau modeset=0
    3. Install the gcc, kernel-devel, and kernel-headers files:
      Run the following commands to install the gcc, kernel-devel, and kernel-headers files and ensure that these kernel source files are of the same version. We recommend using the same version of ISO to configure local installations.
      # yum install gcc kernel-devel-$(uname -r)  kernel-headers-$(uname -r)     # Reconstruct initramfs image
      # cp /boot/initramfs-$(uname -r).img /boot/initramfs-$(uname -r).img.bak
      # dracut /boot/initramfs-$(uname -r).img $(uname -r) --force       # Only reboot the VM in the text mode
      # systemctl set-default multi-user.target
      # init 3
      # reboot
      # lsmod | grep nouveau    # After the VM instance is rebooted, check whether the nouveau driver is used or not
    4. Install an NVIDIA GPU driver:
      Upload the downloaded package to the VM instance and run the following commands to install the driver.
      # chmod +x NVIDIA-Linux-x86_64-346.47.run    # Configure executable permissions
      # ./NVIDIA-Linux-x86_64-346.47.run      # Execute the driver script 
      After you run the commands, the driver package will begin to unpack and you can follow the installation instructions. During the installation, some warnings may appear. Confirm these warnings in sequence as they do not have any real impact. If some errors occur, please refer to the table below to check the environment.
      Error Message Solution

      ERROR: Unable to find the kernel source tree for the currently running kernel. Please make sure you have installed the kernel source files for your kernel and that they are properly configured; on Red Hat Linux systems, for example, be sure you have the 'kernel-source' or 'kernel-devel' RPM installed. If you know the correct kernel source files are installed, you may specify the kernel source path with the '--kernel-source-path' command line option.

      You need to have all of the kernel source files (including kernel, kernel-headers, and kernel-devel) installed and ensure that they are of the same version

      ERROR: The Nouveau kernel driver is currently in use by your system. This driver is incompatible with the NVIDIA driver, and must be disabled before proceeding. Please consult the ow to correctly disable the Nouveau kernel driver.

      You have to disable the Nouveau kernel driver

      ERROR: Failed to find dkms on the system!

      ERROR: Failed to install the kernel module through DKMS. No kernel module was installed; please try installing again without DKMS, or check the DKMS logs for more information.

      You need to install DKMS, which helps maintain out-of-tree drivers by automatically regenerating new modules when the kernel version changes

      ERROR: Unable to load the kernel module 'nvidia.ko'. This happens most frequently when this kernel module was built against the wrong or improperly configured kernel sources, with a version of gcc that differs from the one used to build the target kernel, or if a driver such as rivafb, nvidiafb, or nouveau is present and prevents the NVIDIA kernel module from obtaining ownership of the NVIDIA graphics device(s), or no NVIDIA GPU installed in this system is supported by this NVIDIA Linux graphics driver release.

      Just run the commands ./NVIDIA-Linux-x86_64-384.98.run --kernel-source-path=/usr/src/kernels/3.10.0-XXX.x86_64/ -k $(uname -r)
    5. Check whether the installation is successful:
      Respectively run the following two commands to check whether the installation is successful. If GPU information such as model is displayed in the command output, the driver has been installed successfully.
      # lspci |grep NVIDIA
      # nvidia-smi
    6. Install the CUDA Toolkit:
      Download CUDA Toolkit installation package and upload this package to the VM system. Run the following commands to execute the driver script:
      # chmod +x cuda_8.0.61_375.26_linux.run      # Configure executable permissions
      # ./cuda_8.0.61_375.26_linux.run     # Execute the driver script 
      During the installation, please set the following parameters:
      Figure 12. Install the CUDA Toolkit


    7. Configure the environment variables:
      Run the vim /root/.bashrc command and save the content below to the same file:
      #gpu driver
      export CUDA_HOME=/usr/local/cuda-8.0
      export PATH=/usr/local/cuda-8.0/bin:$PATH
      export LD_LIBRARY_PATH=/usr/local/cuda-8.0/lib64:$LD_LIBRARY_PATH
      export LD_LIBRARY_PATH="/usr/local/cuda-8.0/lib:${LD_LIBRARY_PATH}"
      Environment variables will take effect once added. To verify the effect, you can run the following commands:
      # source ~/.bashrc
      # cd /usr/local/cuda-8.0/samples/1_Utilities/deviceQuery
      # make
      # ./deviceQuery
  5. Install GuestTools.
    To capture real-time data on GPU load monitor, GuestTools are required for the VM instance. The procedures of installing GuestTools vary depending on the OS of your VM instance.
    • For Linux VM Instance
      1. Enter the VM instance details page and find GuestTools on the top row.
      2. Attach ISO.
      3. Launch VM console and run the following commands:
        # Create a mount point.
        mkdir /mnt/cdrom
        # Mount the CD-ROM image.
        mount /dev/cdrom /mnt/cdrom
        # Install GuestTools.
        cd /mnt/cdrom/
        bash ./zs-tools-install.sh
        # Unmount the CD-ROM image(Optional)
        cd ~
        umount /mnt/cdrom
        Note:
        • The commands above can be directly copied to VM console.
        • Before you install GuestTools, ensure that you have installed Linux command-line tools, for example, tar, wget, curl.
        • If you install GuestTools for OpenEuler VM, you need to disable selinux. Otherwise, the QGA feature may be affected.
        Figure 13. Install GuestTools | Linux VM Instance




    • For Windows VM Instance
      1. Enter the VM instance detail page and find GuestTools on the top row.
      2. Install ISO.
      3. Launch VM console and follow the steps to install GuestTools.
        Figure 14. Install GuestTools | Windows VM Instance


Glossary

Instance

An instance is a virtual machine or server that runs the images of operating systems in Cloud, such as VM instance and elastic baremetal instance.

VM Instance

A VM instance is a virtual machine instance running on a host. A VM instance has its own IP address and can access public networks and run application services.

Volume

A volume provides storage space for a VM instance. Volumes are categorized into root volumes and data volumes.

Root Volume

A root volume provides support for the system operations of a VM instance.

Data Volume

A data volume provides extended storage space for a VM instance.

Image

An image is a template file used to create a VM instance or volume. Images are categorized into system images and volume images.

Instance Offering

An instance offering defines the number of vCPU cores, memory size, network bandwidth, and other configuration settings of VM instances.

Disk Offering

A disk offering defines the capacity and other configuration settings of volumes.

GPU Specification

A GPU specification defines the frame per second (FPS), video memory, resolution, and other configuration settings of a physical or virtual GPU. GPU specifications are categorized into physical GPU specifications and virtual GPU specifications.

vNUMA Configuration

vNUMA uses CPU pinning to passthrough the topology of associated host physical NUMA (pNUMA) nodes to a VM instance, generating a topology of virtual NUMA (vNUMA) nodes for the VM instance. This topology enables a vCPU on a vNUMA node to primarily access the local memory and thus improves VM performance.

NUMA (Non-Uniform Memory Access)

Non-uniform memory access (NUMA) is a computer memory design where the memory access time depends on the memory location relative to the CPU. Under NUMA, a processor can access its own local memory faster than non-local memory and thus improves VM performance.

pNUMA Node (physical NUMA Node)

A pNUMA node (physical NUMA node) is a host NUMA node predefined based on the host NUMA architecture. It is used to manage the CPUs and memory of the host.

pNUMA Topology (physical NUMA Topology)

A pNUMA topology (physical NUMA topology) is the topology of the host NUMA nodes predefined by the CPU vendor based on the host NUMA architecture.

vNUMA Node (virtual NUMA Node)

A vNUMA node (virtual NUMA node) is generated by passing-through associated pNUMA nodes via CPU pinning. It is used to manage the CPUs and memory of a VM instance.

vNUMA Topology (virtual NUMA Topology)

A vNUMA topology (virtual NUMA topology) is the topology of VM NUMA nodes generated by passing-through associated pNUMA nodes via CPU pinning.

Local Memory

Local memory is the memory that a CPU (pCPU or vCPU) accesses through the Uncore iMC (Integrated Memory Controller) of the same NUMA (pNUMA or vNUMA) node. Compared with accessing non-local memory, accessing local memory has lower latencies.

CPU Pinning

CPU pinning assigns the virtual CPUs (vCPUs) of a VM instance to specific physical CPUs (pCPUs) of the host, which improves VM performance.

EmulatorPin Configuration

EmulatorPin assigns all other threads than virtual CPU (vCPU) threads and IO threads of a VM instance to physical CPUs (pCPUs) of the host so that these threads run on assigned pCPUs.

Auto-Scaling Group

An auto-scaling group is a group of VM instances that are used for the same scenarios. An auto-scaling group can automatically scale out or in based on application workloads or health status of VM instances in the group.

Snapshot

A snapshot is a point-in-time capture of data status in a volume.

Affinity Group

A VM scheduling policy is a resource orchestration policy based on which VM instances are assigned hosts to achieve the high performance and high availability of businesses.

Zone

A zone is a logical group of resources such as clusters, L2 networks, and primary storage. Zone is the largest resource scope defined in the Cloud.

Cluster

A cluster is a logical group of hosts (compute nodes).

Host

A host provides compute, network, and storage resources for VM instances.

Primary Storage

A primary storage is one or more servers that store volume files of VM instances. These files include root volume snapshots, data volume snapshots, image caches, root volumes, and data volumes.

Image Storage

An image storage is a storage server that stores VM image templates, including ISO image files.

iSCSI Storage

iSCSI storage is an SAN storage that uses the iSCSI protocol for data transmission. You can add an iSCSI SAN block as a Shared Block primary storage or pass through the block to a VM instance.

FC Storage

FC storage is an SAN storage that uses the FC technology for data transmission. You can add an FC SAN block as a Shared Block primary storage or pass through the block to a VM instance.

NVMe Storage

A type of storage implemented via the NVMe-oF (NVMe over fabrics) protocol. You can add a block device configured from an NVMe storage as SharedBlock primary storage.

L2 Network

An L2 network is a layer 2 broadcast domain used for layer 2 isolation. Generally, L2 networks are identified by names of devices on the physical network.

VXLAN Pool

A VXLAN pool is a collection of VXLAN networks established based on VXLAN Tunnel Endpoints (VTEPs). The VNI of each VXLAN network in a VXLAN pool must be unique.

L3 Network

An L3 network includes IP ranges, gateway, DNS, and other network configurations that are used by VM instances.

Public Network

Generally, a public network is a logical network that is connected to the Internet. However, in an environment that has no access to the Internet, you can also create a public network.

Flat Network

A flat network is connected to the network where the host is located and has direct access to the Internet. VM instances in a flat network can access public networks by using elastic IP addresses.

VPC Network

A VPC network is a private network where VM instances can be created. A VM instance in a VPC network can access the Internet through a VPC vRouter.

Management Network

A management network is used to manage physical resources in the Cloud. For example, you can create a management network to manage access to hosts, primary storage, image storage, and VPC vRouters.

Flow Network

A flow network is a dedicated network for port mirror transmission. You can use a flow network to transmit the mirrors of data packets of NIC ports to the target ports.

VPC vRouter

A VPC vRouter is a dedicated VM instance that provides multiple network services.

VPC vRouter HA Group

A VPC vRouter HA group consists of two VPC vRouters. Either VPC vRouter can be a primary or secondary VPC vRouter for the group. If the primary VPC vRouter does not work as expected, the VPC vRouter becomes the secondary VPC vRouter in the group to ensure high availability of business.

vRouter Image

A vRouter image encapsulates network services and can be used to create VPC vRouters.

Dedicated-Performance LB Image

A dedicated-performance load balancer (LB) image encapsulates dedicated-performance load-balancing services and can be used to create load balancer instances. However, a dedicated-performance load balancer image cannot be used to create VM instances.

vRouter Offering

A vRouter offering defines the number of vCPU cores, memory size, image, management network, and public network configuration settings of VPC vRouters. You can use a vRouter offering to create VPC vRouters that can provide network services for public networks and VPC networks.

LB Instance Offering

A load balancer (LB) instance offering defines the CPU, memory, image, and management network configuration settings used to create LB instances. LB instances provide load balancing services for the public network, flat network, and VPC network.

SDN Controller

The SDN controller is the core of the SDN architecture, responsible for centralized management and control of network devices.

SDN Cluster

A cluster of dedicated VM instances designed to provide highly available SDN capabilities.

SDN Instance

A dedicated VM instance designed to provide SDN network capabilities.

SDN Image

An SDN image encapsulates an SDN software and can be used to create SDN instances.

SDN Instance Offering

An SDN instance offering defines the CPU, memory, SDN image, and management network configuration used for creating SDN instances.

Security Group

A security group provides security control services for VM NICs. It filters the ingress or egress TCP, UDP, and ICMP packets of VM NICs based on the specified security rules.

VIP

In bridged network environments, a virtual IP address (VIP) provides network services such as serving as an elastic IP address (EIP), port forwarding, load balancing, IPsec tunneling. When a VIP provides the preceding network services, packets are sent to the VIP and then routed to the destination network where VM instances are located.

EIP

An elastic IP address (EIP) functions based on the NAT technology. IP addresses in a private network are translated into an EIP that is in another network. This way, private networks can be accessed from other networks by using EIPs.

Port Forwarding

Port forwarding functions based on the layer-3 forwarding service of VPC vRouters. This service forwards traffic flows of the specified IP addresses and ports in a public network to specified ports of VM instances by using the specified protocol. If your public IP addresses are insufficient, you can configure port forwarding for multiple VM instances by using one public IP address and port.

Load Balancer

A load balancer distributes traffic flows of a virtual IP address to backend servers. It automatically inspects the availability of backend servers and isolates unavailable servers during traffic distribution. This way, the load balancer improves the availability and service capability of your business.

Listener

A listener monitors the frontend requests of a load balancer and distributes the requests to a backend server based on the specified policy. In addition, the listener performs health checks on backend servers.

Forwarding Rule

A forwarding rule forwards the requests from different domain names or URLs to different backend server groups.

Backend Server Group

A backend server group is a group of backend servers that handles requests distributed by load balancers. It is the basic unit for traffic distribution by load balancer instances.

Backend Server

A backend server handles requests distributed by a load balancer. You can add a VM instance on the Cloud or a server on a third-party cloud as a backend server.

Frontend Network

A frontend network is a type of network that is associated with a load balancer. Requests from the network are distributed by the load balancer to backend servers based on a specified policy.

Backend Network

A backend network is a type of network that is associated with a load balancer. Requests from frontend networks are distributed by the load balancer to servers in the backend network.

Load Balancer Instance

A load balancer instance is a custom VM instance used to provide load balancing services.

Certificate

If you select HTTPS for a listener, associate it with a certificate to make the listener take effect. You can upload either a certificate or certificate chain.

Firewall

A firewall is an access control policy that monitors ingress and egress traffic of VPC vRouters and decides whether to allow or block specific traffic based on the associated rule sets and rules.

Firewall Rule Set

A firewall rule set is a set of rules that a firewall uses to defend against network attacks. You need to associate a rule set with the egress or ingress flow direction of VPC vRouter NICs to make the rule set take effect.

Firewall Rule

A firewall rule is an access control entry associated with the egress or ingress flow direction of VPC vRouter NICs to defend against network attacks. A firewall rule includes rule priority, match condition, and behavior.

Rule Template

A rule template is a template that you can select when you add rules to a rule set or a firewall.

IP/Port Set

An IP or port set is a set of IP addresses or ports that you can select when you add rules to a rule set or a firewall.

IPsec Tunnel

An IPSec tunnel encrypts and verifies IP packets that transmit over a virtual private network (VPN) from one site to another.

OSPF Area

An Open Shortest Path First (OSPF) area is divided from an autonomous system based on the OSPF protocol. This simplifies the hierarchical management of vRouters.

NetFlow

A NetFlow monitors the ingress and egress traffic of the NICs of VPC vRouters. The supported versions of data flows are V5 and V9.

Port Mirroring

Port mirroring mirrors the traffic data of VM NICs and sends the traffic data to the target ports. This allows for the analysis of data packets of ports and simplifies the monitoring and management of data traffic and makes it easier to locate network errors and exceptions.

Route Table

A route table contains information about various routes that you configure. Route entries in a route table must include the destination network, next hop, and route priority.

CloudFormation

CloudFormation is a service that simplifies the management of cloud resources and automates deployment and O&S. You can create a stack template to configure cloud resources and their dependencies. This way, resources can be automatically configured and deployed in batches. CloudFormation provides easy management of the lifecycle of cloud resources and integrates automatic O&S into API and SDK.

Resource Stack

A resource stack is a stack of resources that are configured by using a stack template. The resources in the stack have dependencies with each other. You can manage resources in the stack by managing the resource stack.

Stack Template

A stack template is a UTF8-encoded file based on which you can create resource stacks. The stack template defines the resources that you want, the dependencies between the resources, and the configuration settings of the resources. When you use a stack template to create a resource stack, CloudFormation parses the template and the resources are automatically created and configured.

Sample Template

A sample template is a commonly used resource stack. You can use a sample template provide by the Cloud to create resource stacks.

Designer

A designer is a CloudFormation tool that allows you to orchestrate cloud resources. You can drag and drop resources on a canvas and use lines to establish dependencies between the resources.

Baremetal Cluster

A baremetal cluster consists of baremetal chassis. You can manage baremetal chassis by managing a baremetal cluster where the chassis reside.

Deployment Server

A deployment server is a server that provides PXE service and console proxy service for baremetal chassis.

Baremetal Chassis

A baremetal chassis is used to create a baremetal instance and is identified based on the BMC interface and IPMI configuration setting.

Preconfigured Template

A preconfigured template is used to create a preconfigured file that allows for unattended batch installation of an operating system for baremetal instances.

Baremetal Instance

A baremetal instance is an instantiated baremetal chassis.

Elastic Baremetal Management

Elastic Baremetal Management provides dedicated physical servers for your applications to ensure high performance and stability. In addition, this feature allows elastic scaling. You can apply for and scale resources based on your needs.

Provision Network

A provision network is a dedicated network for PXE boot and image downloads while creating elastic baremetal instances in a gateway proxy cluster.

Elastic Baremetal Cluster

Provides a separated cluster to manage baremetal nodes.

Gateway Node

A gateway node is a node where the ingress and egress traffic of the Cloud and elastic baremetal instances in gateway proxy clusters is forwarded.

Baremetal Node

A baremetal node is used to create a baremetal instance and is identified based on the BMC interface and IPMI configuration setting.

Elastic Baremetal Instance

An elastic baremetal instance has the same performance as physical servers and allows elastic scaling. You can apply for and scale resources based on your needs.

Elastic Baremetal Offering

An elastic baremetal offering defines the number of vCPU cores, memory size, CPU architecture, CPU model, and other configuration settings of elastic baremetal instances.

vCenter

The Cloud allows you to take over vCenter and manage resources on the vCenter.

VM Instance

A VM instance is an ESXi virtual machine instance running on a host. A VM instance has its own IP address to access public networks and can run application services.

Network

A vCenter network defines the network settings of VM instances on vCenter, such as IP range, gateway, DNS, and network services.

Volume

A volume provides storage space for a VM instance on vCenter. A volume attached to a VM instance can be used as a root volume or data volume. A root volume provides support for the system operations of a VM instance. A data volume provides extended storage space for a VM instance.

Image

An image is a template file used to create a VM instance or volume on vCenter. Images are categorized into system images and volume images.

Event Message

Event Message displays event alarm messages of vCenter that is took over by the Cloud. This feature allows you to locate errors and exceptions efficiently.

Network Topology

A network topology visualizes the network architecture of the Cloud. It allows for efficient planning, management, and improvement of network architecture. Network topologies can be categorized into global topologies and custom topologies.

Performance Analysis

Performance Analysis displays the performance metrics of key resources monitored externally or internally in the Cloud. You can view the performance analysis or export the analysis report as needed to improve the O&M efficiency.

Capacity Management

Capacity Management visualizes the capacities and usages of key resources in the Cloud. You can use this feature to improve O&S efficiency.

MN Monitoring

Management Node (MN) monitoring allows you to view the health status of each management node when you use multiple management nodes to achieve high availability.

Alarm

An alarm is used to monitor the status of time-series data and events and respond to the status change. Alarms can be categorized into resource alarm, event alarm, and extended alarm.

One-Click Alarm

A one-click alarm integrates multiple metrics of a resource. You can create one-click alarms for multiple resources to monitor these resources.

Alarm Template

An alarm template is a template of alarm rules. If you associate an alarm template with a resource group, an alarm is created to monitor the resources in the group.

Resource Group

A resource group consists of resources grouped based on your business needs. If you associate an alarm template with a resource group, the alarm rules specified by the template take effect on all the resources in the group.

Message Template

A message template specifies the text template of a resource alarm message or event alarm message sent to an SNS system.

Message Source

A message source is used to take over extended alarm messages. If you configure alarms for message sources, extended alarm messages can be sent to various endpoints.

Endpoint

An endpoint is a method that users obtain subscribed messages. Endpoints are categorized into system endpoints, email, DingTalk, HTTP application, short message service, and Microsoft Teams.

Alarm Message

An alarm message is a message sent the time when an alarm is triggered.

Current Task

A current task is an ongoing operation performed in the Cloud. You can perform centralized management over ongoing operations.

Operation Log

An operation log is a chronological record of operations on the specified objects and their operation results.

Audit

Audit monitors and records all activities on the Cloud. You can use this feature to implement operation tracking, cybersecurity classified protection compliance, security analysis, troubleshooting, and automatic O&M.

Log Collection

Allows you to collect with one click the log data from the Cloud and various nodes on the Cloud generated in the specified time period and download the log data.

One-Click Inspection

Comprehensively inspects the health status of key resources and services of the Cloud and scores their healthiness based on the inspection results. In addition, the one-click inspection service provides O&M suggestions and inspection reports.

Backup Management

Backup management integrates multiple disaster recovery technologies such as incremental backup and full backup that are suitable for multiple business scenarios. You can implement local backup and remote backup based on your business needs.

Backup Job

You can create a backup job to back up local VM instances, volumes, or databases to a specified storage server on a regular basis.

Local Backup Data

Local backup data of VM instances, volumes, and databases is stored in the local backup server.

Local Backup Server

A local backup server is located at the local data center and is used to store local backup data.

Remote Backup Server

A remote backup server is located at a remote data center or a public cloud and is used to store remote backup data.

Continuous Data Protection (CDP)

Continuous Data Protection (CDP) provides second-level and fine-grained continuous backups for important business systems in VM instances, allowing users to restore VM data to a specific time state, and retrieve files without restoring the system.

CDP Task

You can create a CDP task to continuously back up your VM data to a specified backup server to achieve continuous data protection and recovery.

CDP Data

The backup data generated from continuous data protection on VM instances is stored in local backup servers.

Recovery Point

A recovery point is a data point generated during continuous data protection. A recovery point corresponds to a data record within the recovery point interval specified by the user.

Locked Recovery Point

You can lock or unlock a recovery point as needed. After a recovery point is locked, data of the recovery point will not be automatically cleared or deleted.

Recovery Task

A recovery task helps you quickly restore data by specifying a CDP task and recovery point, and allows you to view the recovery progress and logs in a more friendly way.

Cryptography Security Compliance

The Cryptography Security Compliance service provides applications with cloud security capabilities based on commercial cryptography, meeting the requirements of commercial cryptography application security assessments.

HSM Pool

An HSM pool is a logical group of hardware security modules (HSMs) and is used to provide unified cryptography services such as signature validation and encryption.

HSM

A hardware security module (HSM) is a dedicated device that encrypts, decrypts, and authenticates information by using the cryptographic technology.

Platform Cryptography Security Compliance

Enables the Cloud to meet the requirements of Cryptography Security Compliance through the cryptography capabilities provided by HSM pools.

Certificate Login

Authenticates the identity of a user by using a UKey device.

Data Protection

Protects important data on the Cloud to ensure the data confidentiality and integrity.

Scheduled Job

A scheduled job defines that a specific action be implemented at a specified time based on a scheduler.

Scheduler

A scheduler is used to schedule jobs. It is suitable for business scenarios that last for a long time.

Tag

A tag is used to mark resources. You can use a tag to search for and aggregate resources.

Migration Service

The Cloud provides V2V migration service that allows you to migrate VM instances and data from other virtualized platform to the current cloud platform.

ZMigrate Migration Service

A migration service installed from Application Market that migrates VM instances and their data from VMware environments to the current cloud platform.

V2V Migration

V2V Migration allows you to migrate VM instances from the VMware or KVM platform to the current cloud platform.

V2V Conversion Host

A V2V conversion host is a host in the destination cluster that you need to specify during V2V migration to cache VM instances and data when you implement V2V migration. After the VM instances and data are cached in the V2Vconversion host, they are migrated to the destination primary storage.

User

A user is a natural person that constructs the most basic unit in Tenant Management.

User Group

A user group is a collection of natural persons or a collection of project members. You can use a user group to grant permissions.

Role

A role is a collection of permissions that can be granted to users. A user that assumes a role can call API operations based on the permissions specified by the role. Roles are categorized into platform roles and project roles.

Single Sign-On

The Single Sign-On service provided by the Cloud. It supports seamless access to SSO systems. Through the service, related users can directly log in to the Cloud and manage cloud resources.

Project

A project is a task that needs to be accomplished by specific personnel at a specified time. In Tenant Management, you can plan resources at the project granularity and allocate an independent resource pool to a project. The word Tenant in Tenant Management mainly refers to projects. A project is a tenant.

Project Member

A project member is a member in a project who is granted permissions on specific project resources and can use the resources to accomplish tasks. Project members include the project admin, project managers, and normal project members.

Process Management

Process management is part of ticket management that manages the processes related to the resources of projects. Processes can be categorized into default processes and custom processes.

My Approvals

In the Cloud, only the administrator and project administrators are granted approval permissions. the administrator and project administrators can approve or reject a ticket. If a ticket is approved, resources are automatically deployed and allocated to the specified project.

Bills

A bill is the expense of resources totaled at a specified time period. Billing is accurate to the second. Bills can be categorized into project bills, department bills, and account bills.

Pricing List

A pricing list is a list of unit prices of different resources. The unit price of a resource is set based on the specification and usage time of the resource.

Console Proxy

Console proxy allows you to log in to a VM instance by using the IP address of a proxy.

AccessKey Management

An AccessKey pair is a security credential that one party authorizes another party to call API operations and access its resources in the Cloud. AccessKey pairs shall be kept confidential.

IP Allowlist/Blocklist

An IP allowlist or blocklist identifies and filters IP addresses that access the Cloud. You can create an IP allowlist or blocklist to improve access control of the Cloud.

Application Center

Application Market allows you to add applications to the Cloud and then access the applications with one click. It extends the functionality of the Cloud. You can add default applications through the built-in installation package or add more applications through URLs.

Sub-Account Management

A sub-account can be created by the admin or synced from an SSO authentication system and is managed by the admin. Resources created under a sub-account are managed by the sub-account.

Theme and Appearance

You can customize the theme and appearance of the Cloud.

Email Server

If you select Email as the endpoint of an alarm, you need to set an email server. Then alarm messages are sent to the email server.

Log Server

A log server is used to collect management node logs or the platform operation logs. You can add a log server to the cloud and use the collected logs for operation trace or troubleshooting. This makes your O&M more efficient.

Global Setting

Global Setting allows you to configure settings that take effect on the whole platform.

Scenario Template

Scenario Template provides multiple templates that encapsulate scenario-based global settings. You can apply a template globally with one click based on your business needs. This improves your O&M efficiency.

HA Policy

HA Policy is a mechanism that ensures sustained and stable running of the business if VM instances are unexpectedly stopped or are errored because of errors occurring to compute, network, or storage resources associated with the VM instances. By enabling this feature, you can customize VM HA policies to ensure your business continuity and stability.

Time Management

Manages the Cloud system time and allows you to configure time servers for the Cloud. After you configure NTP time servers for the Cloud, the clock of the time servers is synced with all nodes of the Cloud.

GPU Device

A GPU device is a powerful microprocessor with high computational capabilities. You can use a GPU device to handle intricate graphics rendering and parallel computing jobs, thus improving the efficiency of businesses such as graphic production, video processing, and machine learning.

Script Library

The script library stores and manages script files centrally. By executing scripts on VM instances, you can complete complex O&M operations and automated jobs.

XML Hook

An XML Hook is a script that can flexibly insert or modify parameters in XML files of VM instances. By attaching an XML Hook to a VM instance, you can customize VM configurations and enable specialized functionalities.

Container Service

A simple and user-friendly container management service, providing features like GPU management & scheduling, multi-tenancy, multi-cluster, quota configuration, CI/CD. and microservice. The service reduces the container using complexity and aligns well with traditional user's habits, helping you easily manage and deploy your container cluster, and enjoy the benefits of cloud-native technologies in a quick and convenient way.

Advanced Monitoring Server

An advanced monitoring server is a dedicated VM instance used to receive advanced monitoring data of load balancers and other resources.

Advanced Monitoring Server Image

An advanced monitoring server image encapsulates the advanced monitoring service and can be used to create advanced monitoring server.

Advanced Monitoring Server Offering

An advanced monitoring server offering defines the CPU cores, memory size, image, management network, and public network configurations of advanced monitoring server. You can use an advanced monitoring server offering to create advanced monitoring servers.

Plugin Management

You can package extended resources or tools into standardized plugins for quick installation and integration, expanding the Cloud capabilities.

Region Management

A region is a self-contained cloud environment with independent management node(s), networks, hardware, and cloud resources. ZStack IAM enabled user synchronization and SSO across multiple regions.
GPU Passthrough Tutorial | 5.5.38 | ZStack Cloud · ZCF | ZStack Resource Center