Skip to content

Inconsistent GPU Memory Usage Reporting Between dpctl and xpu-smi #1761

Description

@avimanyu786

Description

When using dpctl to report GPU memory usage on Intel GPUs, the reported free and total memory values appear to be incorrect when compared to the output from xpu-smi. Specifically, dpctl reports 0 bytes of used memory, while xpu-smi correctly reports the used memory as 17 MiB.

Steps to Reproduce

  1. Set up an environment with dpctl and xpu-smi installed.
  2. Use the following Python script to get GPU memory information using dpctl:
import os
import dpctl
from dpctl.utils import intel_device_info

def get_intel_gpu_memory_info():
    try:
        # Set the environment variable ZES_ENABLE_SYSMAN to 1
        os.environ["ZES_ENABLE_SYSMAN"] = "1"
        
        # Get the list of GPU devices
        devices = dpctl.get_devices(device_type=dpctl.device_type.gpu)
        for device in devices:
            # Get Intel GPU device info
            device_info = intel_device_info(device)
            if device_info:
                free_memory = device_info.get('free_memory', None)
                if free_memory is not None:
                    free_memory_mib = free_memory / (1024 * 1024)
                    print(f"Free Memory: {free_memory_mib:.2f} MiB")

                # Get the total global memory size
                try:
                    global_mem_size = device.get_info(dpctl.device_info.global_mem_size)
                except AttributeError:
                    global_mem_size = device.global_mem_size

                global_mem_size_mib = global_mem_size / (1024 * 1024)
                print(f"Total Memory: {global_mem_size_mib:.2f} MiB")

                # Calculate and display used memory
                if free_memory is not None and global_mem_size is not None:
                    used_memory = global_mem_size - free_memory
                    used_memory_mib = used_memory / (1024 * 1024)
                    print(f"Used Memory: {used_memory_mib:.2f} MiB")
                else:
                    print("Unable to calculate used memory due to missing information.")

                return
        print("No Intel GPU devices found or no information available.")
    except Exception as e:
        print(f"An error occurred: {e}")

if __name__ == "__main__":
    get_intel_gpu_memory_info()
  1. Compare the output with the results of running xpu-smi stats -d 0:
xpu-smi stats -d 0

Observed Behavior

  • Output from the Python script using dpctl:
Free Memory: 15473.60 MiB
Total Memory: 15473.60 MiB
Used Memory: 0.00 MiB

Also, python -c "import torch; import intel_extension_for_pytorch as ipex; print(torch.__version__); print(ipex.__version__); [print(f'[{i}]: {torch.xpu.get_device_properties(i)}') for i in range(torch.xpu.device_count())];" shows the following output which matches the total memory:

2.1.0.post2+cxx11.abi
2.1.30+xpu
[0]: _DeviceProperties(name='Intel(R) Arc(TM) A770 Graphics', platform_name='Intel(R) Level-Zero', dev_type='gpu', driver_version='1.3.27642', has_fp64=0, total_memory=15473MB, max_compute_units=512, gpu_eu_count=512)
  • Output from xpu-smi:
+-----------------------------+--------------------------------------------------------------------+
| Device ID                   | 0                                                                  |
+-----------------------------+--------------------------------------------------------------------+
| GPU Memory Used (MiB)       | 17                                                                 |
| GPU Memory Util (%)         | 0                                                                  |
+-----------------------------+--------------------------------------------------------------------+

Expected Behavior

The used memory reported by dpctl should match the used memory reported by xpu-smi.

Environment

  • dpctl version: 0.17.0
  • xpu-smi version: 1.2.38.20240718
  • OS: HiveOS [Based on Ubuntu 20.04]
  • Docker version: 24.0.7, build 24.0.7-0ubuntu2~20.04.1
  • Docker image: intel/intel-extension-for-pytorch:2.1.30-xpu
  • Python version: 3.10.12
  • GPU: Intel(R) Arc(TM) A770 Graphics

Additional Information

Setting the environment variable ZES_ENABLE_SYSMAN to 1 was necessary as mentioned in the documentation, to report the free_memory. The discrepancy in reported values suggests a potential issue within the dpctl library or its interaction with the GPU drivers.

Further information on OS:

# uname -r
6.1.0-hiveos
# lsb_release -a
No LSB modules are available.
Distributor ID:	Ubuntu
Description:	Ubuntu 20.04.6 LTS
Release:	20.04
Codename:	focal

The container was launched through the following command:

docker run -ti --cap-add=PERFMON --device /dev/dri intel/intel-extension-for-pytorch:2.1.30-xpu bash

The intel-basekit (provides the necessary SYCL runtime and development tools for dpctl) and xpu-smi packages were installed with the following commands before the testing the issue inside the container:

wget -qO - https://repositories.intel.com/gpu/intel-graphics.key | gpg --yes --dearmor --output /usr/share/keyrings/intel-graphics.gpg \
wget https://github.com/intel/xpumanager/releases/download/V1.2.38/xpu-smi_1.2.38_20240718.060204.0db09695+deb10u1_amd64.deb \
apt update && apt install -y ./xpu-smi_1.2.38_20240718.060204.0db09695+deb10u1_amd64.deb intel-basekit \
source /opt/intel/oneapi/setvars.sh

Proposed Solution

Investigate and resolve the inconsistency in GPU memory reporting between dpctl and xpu-smi. Ensure that dpctl accurately reflects the actual GPU memory usage.


Thank you for looking into this issue. Please let me know if further information or testing is required.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions