Description
When using dpctl to report GPU memory usage on Intel GPUs, the reported free and total memory values appear to be incorrect when compared to the output from xpu-smi. Specifically, dpctl reports 0 bytes of used memory, while xpu-smi correctly reports the used memory as 17 MiB.
Steps to Reproduce
- Set up an environment with
dpctl and xpu-smi installed.
- Use the following Python script to get GPU memory information using
dpctl:
import os
import dpctl
from dpctl.utils import intel_device_info
def get_intel_gpu_memory_info():
try:
# Set the environment variable ZES_ENABLE_SYSMAN to 1
os.environ["ZES_ENABLE_SYSMAN"] = "1"
# Get the list of GPU devices
devices = dpctl.get_devices(device_type=dpctl.device_type.gpu)
for device in devices:
# Get Intel GPU device info
device_info = intel_device_info(device)
if device_info:
free_memory = device_info.get('free_memory', None)
if free_memory is not None:
free_memory_mib = free_memory / (1024 * 1024)
print(f"Free Memory: {free_memory_mib:.2f} MiB")
# Get the total global memory size
try:
global_mem_size = device.get_info(dpctl.device_info.global_mem_size)
except AttributeError:
global_mem_size = device.global_mem_size
global_mem_size_mib = global_mem_size / (1024 * 1024)
print(f"Total Memory: {global_mem_size_mib:.2f} MiB")
# Calculate and display used memory
if free_memory is not None and global_mem_size is not None:
used_memory = global_mem_size - free_memory
used_memory_mib = used_memory / (1024 * 1024)
print(f"Used Memory: {used_memory_mib:.2f} MiB")
else:
print("Unable to calculate used memory due to missing information.")
return
print("No Intel GPU devices found or no information available.")
except Exception as e:
print(f"An error occurred: {e}")
if __name__ == "__main__":
get_intel_gpu_memory_info()
- Compare the output with the results of running
xpu-smi stats -d 0:
Observed Behavior
- Output from the Python script using
dpctl:
Free Memory: 15473.60 MiB
Total Memory: 15473.60 MiB
Used Memory: 0.00 MiB
Also, python -c "import torch; import intel_extension_for_pytorch as ipex; print(torch.__version__); print(ipex.__version__); [print(f'[{i}]: {torch.xpu.get_device_properties(i)}') for i in range(torch.xpu.device_count())];" shows the following output which matches the total memory:
2.1.0.post2+cxx11.abi
2.1.30+xpu
[0]: _DeviceProperties(name='Intel(R) Arc(TM) A770 Graphics', platform_name='Intel(R) Level-Zero', dev_type='gpu', driver_version='1.3.27642', has_fp64=0, total_memory=15473MB, max_compute_units=512, gpu_eu_count=512)
+-----------------------------+--------------------------------------------------------------------+
| Device ID | 0 |
+-----------------------------+--------------------------------------------------------------------+
| GPU Memory Used (MiB) | 17 |
| GPU Memory Util (%) | 0 |
+-----------------------------+--------------------------------------------------------------------+
Expected Behavior
The used memory reported by dpctl should match the used memory reported by xpu-smi.
Environment
- dpctl version: 0.17.0
- xpu-smi version: 1.2.38.20240718
- OS: HiveOS [Based on Ubuntu 20.04]
- Docker version: 24.0.7, build 24.0.7-0ubuntu2~20.04.1
- Docker image: intel/intel-extension-for-pytorch:2.1.30-xpu
- Python version: 3.10.12
- GPU: Intel(R) Arc(TM) A770 Graphics
Additional Information
Setting the environment variable ZES_ENABLE_SYSMAN to 1 was necessary as mentioned in the documentation, to report the free_memory. The discrepancy in reported values suggests a potential issue within the dpctl library or its interaction with the GPU drivers.
Further information on OS:
# uname -r
6.1.0-hiveos
# lsb_release -a
No LSB modules are available.
Distributor ID: Ubuntu
Description: Ubuntu 20.04.6 LTS
Release: 20.04
Codename: focal
The container was launched through the following command:
docker run -ti --cap-add=PERFMON --device /dev/dri intel/intel-extension-for-pytorch:2.1.30-xpu bash
The intel-basekit (provides the necessary SYCL runtime and development tools for dpctl) and xpu-smi packages were installed with the following commands before the testing the issue inside the container:
wget -qO - https://repositories.intel.com/gpu/intel-graphics.key | gpg --yes --dearmor --output /usr/share/keyrings/intel-graphics.gpg \
wget https://github.com/intel/xpumanager/releases/download/V1.2.38/xpu-smi_1.2.38_20240718.060204.0db09695+deb10u1_amd64.deb \
apt update && apt install -y ./xpu-smi_1.2.38_20240718.060204.0db09695+deb10u1_amd64.deb intel-basekit \
source /opt/intel/oneapi/setvars.sh
Proposed Solution
Investigate and resolve the inconsistency in GPU memory reporting between dpctl and xpu-smi. Ensure that dpctl accurately reflects the actual GPU memory usage.
Thank you for looking into this issue. Please let me know if further information or testing is required.
Description
When using
dpctlto report GPU memory usage on Intel GPUs, the reported free and total memory values appear to be incorrect when compared to the output fromxpu-smi. Specifically,dpctlreports 0 bytes of used memory, whilexpu-smicorrectly reports the used memory as 17 MiB.Steps to Reproduce
dpctlandxpu-smiinstalled.dpctl:xpu-smi stats -d 0:Observed Behavior
dpctl:Also,
python -c "import torch; import intel_extension_for_pytorch as ipex; print(torch.__version__); print(ipex.__version__); [print(f'[{i}]: {torch.xpu.get_device_properties(i)}') for i in range(torch.xpu.device_count())];"shows the following output which matches the total memory:xpu-smi:Expected Behavior
The used memory reported by
dpctlshould match the used memory reported byxpu-smi.Environment
Additional Information
Setting the environment variable
ZES_ENABLE_SYSMANto1was necessary as mentioned in the documentation, to report thefree_memory. The discrepancy in reported values suggests a potential issue within thedpctllibrary or its interaction with the GPU drivers.Further information on OS:
The container was launched through the following command:
docker run -ti --cap-add=PERFMON --device /dev/dri intel/intel-extension-for-pytorch:2.1.30-xpu bashThe
intel-basekit(provides the necessary SYCL runtime and development tools for dpctl) andxpu-smipackages were installed with the following commands before the testing the issue inside the container:Proposed Solution
Investigate and resolve the inconsistency in GPU memory reporting between
dpctlandxpu-smi. Ensure thatdpctlaccurately reflects the actual GPU memory usage.Thank you for looking into this issue. Please let me know if further information or testing is required.