User Tools

Site Tools


unite_python_cupy

Using CuPy for GPU-Accelerated Computing in Python

This guide explains how to load the required Python and CUDA environment modules to use CuPy (a NumPy-compatible array library for GPU-accelerated computing) and submit jobs to the unite partition.

IMPORTANT: Execution & Environment Policy

WARNING: Do not run interactive Python code or initialization directly on the Login Node!
CuPy dynamically binds with local hardware drivers and initializes CUDA runtime kernels upon execution. This overhead can heavily lock processor threads. Always request an interactive compute node session on the unite partition using `srun` before running or testing your Python files.

Available Modules

To use CuPy, you must load the specific Python package ecosystem paired with the correct CUDA version (CUDA 12.8).

To see the available Python module configurations, run:

$ module avail unite/python

Step-by-Step Guide

1. Request a GPU-Enabled Interactive Node

Switch to a safe environment on a compute node under the unite partition. Since CuPy requires a graphics card to run, ensure your resource allocation includes access to a GPU:

$ srun --partition=unite --gres=gpu:1 --cpus-per-task=4 --time=00:30:00 --pty bash

Once your shell prompt updates, you are active on a GPU compute node.

2. Load the Required Modules

Load the base Python 3.14 environment module together with the CuPy application layer (which implicitly connects with CUDA 12.8):

$module load unite/python/3.14/base
$ module load unite/python/3.14/cupy

To verify your environment's Python path and check if the modules are recognized, run:

$ python3 -c "import cupy; print('CuPy version:', cupy.__version__)"

3. Create a Sample CuPy Script

Create a parallel matrix operation script named `cupy_test.py`. This example allocates multi-dimensional arrays directly inside the GPU memory architecture, performs matrix multiplication, and moves the output result back to standard system memory space.

cupy_test.py
import cupy as cp
import time
 
# Verify available GPU hardware allocation
try:
    device = cp.cuda.Device(0)
    print(f"Running on Device: {device}")
except Exception as e:
    print(f"CUDA Error: {e}. Make sure you are running on a GPU compute node!")
    exit(1)
 
# Initialize large matrices directly on the GPU memory space
print("Allocating multi-dimensional arrays on the GPU...")
size = 5000
x_gpu = cp.ones((size, size), dtype=cp.float32)
y_gpu = cp.ones((size, size), dtype=cp.float32)
 
# Perform matrix multiplication on the GPU cores
print(f"Performing matrix multiplication ({size}x{size})...")
start_time = time.time()
z_gpu = cp.dot(x_gpu, y_gpu)
 
# Force CUDA synchronization to get accurate benchmarks
cp.cuda.Stream.null.synchronize()
end_time = time.time()
 
print(f"Operation completed successfully in: {end_time - start_time:.4f} seconds")
 
# Pull calculation matrix value back to CPU host memory for confirmation
result_cpu = cp.asnumpy(z_gpu)
print(f"Verification check value: {result_cpu[0, 0]} (Expected: {float(size)})")

4. Testing the Script

You can safely test your execution parameters right here on this interactive node session by running:

$ python3 cupy_test.py

Exit your interactive session to return to the login node when the validation completes:

$ exit

Running CuPy Jobs via Slurm Batch

To launch headless Python jobs using your CuPy environment via Slurm, create a batch script named `submit_cupy.sh`:

submit_cupy.sh
#!/bin/bash
#SBATCH --job-name=cupy_gpu_job
#SBATCH --partition=unite           # Target partition for the UNITE cluster
#SBATCH --output=cupy_job_%j.out
#SBATCH --error=cupy_job_%j.err
#SBATCH --nodes=1                   # Run on a single node
#SBATCH --ntasks=1                  # Run a single task instance
#SBATCH --gres=gpu:1                # Request 1 GPU allocation
#SBATCH --cpus-per-task=4           # Allocate CPU cores for background thread pooling
#SBATCH --time=00:15:00             # Format: HH:MM:SS
 
# Clean out inherited environment parameters
module purge
 
# Crucial: Load the exact same Python and CuPy dependency module matrix
module load unite/python/3.14/python-3.14.0
module load unite/python/3.14/cupy
 
# Launch your Python script
python3 cupy_test.py

Submit the parallel run to the scheduler queue:

$ sbatch submit_cupy.sh

Troubleshooting

* Error: `ModuleNotFoundError: No module named 'cupy'`
Solution: You forgot to load the module layer or your interactive session timed out. Run `module load unite/python/3.14/cupy` inside your active session before launching your application. * Error: `cupy.cuda.compiler.CompileException / CUDA_ERROR_NO_DEVICE`
Solution: You tried running the Python script directly on the login node or forgot to include the `#SBATCH –gres=gpu:1` flag inside your Slurm script configuration. CuPy requires an accessible NVIDIA GPU architecture to successfully build or load runtime mathematical kernels.

unite_python_cupy.txt · Last modified: 2026/06/13 15:06 by nshegunov

Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki