====== Using CuPy for GPU-Accelerated Computing in Python ======
This guide explains how to load the required Python and CUDA environment modules to use CuPy (a NumPy-compatible array library for GPU-accelerated computing) and submit jobs to the **unite** partition.
===== IMPORTANT: Execution & Environment Policy =====
^ WARNING: Do not run interactive Python code or initialization directly on the Login Node! ^
| CuPy dynamically binds with local hardware drivers and initializes CUDA runtime kernels upon execution. This overhead can heavily lock processor threads. Always request an interactive compute node session on the **unite** partition using **`srun`** before running or testing your Python files. |
===== Available Modules =====
To use CuPy, you must load the specific Python package ecosystem paired with the correct CUDA version (CUDA 12.8).
To see the available Python module configurations, run:
$ module avail unite/python
===== Step-by-Step Guide =====
==== 1. Request a GPU-Enabled Interactive Node ====
Switch to a safe environment on a compute node under the **unite** partition. Since CuPy requires a graphics card to run, ensure your resource allocation includes access to a GPU:
$ srun --partition=unite --gres=gpu:1 --cpus-per-task=4 --time=00:30:00 --pty bash
Once your shell prompt updates, you are active on a GPU compute node.
==== 2. Load the Required Modules ====
Load the base Python 3.14 environment module together with the CuPy application layer (which implicitly connects with CUDA 12.8):
$module load unite/python/3.14/base
$ module load unite/python/3.14/cupy
To verify your environment's Python path and check if the modules are recognized, run:
$ python3 -c "import cupy; print('CuPy version:', cupy.__version__)"
==== 3. Create a Sample CuPy Script ====
Create a parallel matrix operation script named `cupy_test.py`. This example allocates multi-dimensional arrays directly inside the GPU memory architecture, performs matrix multiplication, and moves the output result back to standard system memory space.
import cupy as cp
import time
# Verify available GPU hardware allocation
try:
device = cp.cuda.Device(0)
print(f"Running on Device: {device}")
except Exception as e:
print(f"CUDA Error: {e}. Make sure you are running on a GPU compute node!")
exit(1)
# Initialize large matrices directly on the GPU memory space
print("Allocating multi-dimensional arrays on the GPU...")
size = 5000
x_gpu = cp.ones((size, size), dtype=cp.float32)
y_gpu = cp.ones((size, size), dtype=cp.float32)
# Perform matrix multiplication on the GPU cores
print(f"Performing matrix multiplication ({size}x{size})...")
start_time = time.time()
z_gpu = cp.dot(x_gpu, y_gpu)
# Force CUDA synchronization to get accurate benchmarks
cp.cuda.Stream.null.synchronize()
end_time = time.time()
print(f"Operation completed successfully in: {end_time - start_time:.4f} seconds")
# Pull calculation matrix value back to CPU host memory for confirmation
result_cpu = cp.asnumpy(z_gpu)
print(f"Verification check value: {result_cpu[0, 0]} (Expected: {float(size)})")
==== 4. Testing the Script ====
You can safely test your execution parameters right here on this interactive node session by running:
$ python3 cupy_test.py
Exit your interactive session to return to the login node when the validation completes:
$ exit
===== Running CuPy Jobs via Slurm Batch =====
To launch headless Python jobs using your CuPy environment via Slurm, create a batch script named `submit_cupy.sh`:
#!/bin/bash
#SBATCH --job-name=cupy_gpu_job
#SBATCH --partition=unite # Target partition for the UNITE cluster
#SBATCH --output=cupy_job_%j.out
#SBATCH --error=cupy_job_%j.err
#SBATCH --nodes=1 # Run on a single node
#SBATCH --ntasks=1 # Run a single task instance
#SBATCH --gres=gpu:1 # Request 1 GPU allocation
#SBATCH --cpus-per-task=4 # Allocate CPU cores for background thread pooling
#SBATCH --time=00:15:00 # Format: HH:MM:SS
# Clean out inherited environment parameters
module purge
# Crucial: Load the exact same Python and CuPy dependency module matrix
module load unite/python/3.14/python-3.14.0
module load unite/python/3.14/cupy
# Launch your Python script
python3 cupy_test.py
Submit the parallel run to the scheduler queue:
$ sbatch submit_cupy.sh
===== Troubleshooting =====
* **Error:** `ModuleNotFoundError: No module named 'cupy'`\\ **Solution:** You forgot to load the module layer or your interactive session timed out. Run `module load unite/python/3.14/cupy` inside your active session before launching your application.
* **Error:** `cupy.cuda.compiler.CompileException / CUDA_ERROR_NO_DEVICE`\\ **Solution:** You tried running the Python script directly on the login node or forgot to include the `#SBATCH --gres=gpu:1` flag inside your Slurm script configuration. CuPy requires an accessible NVIDIA GPU architecture to successfully build or load runtime mathematical kernels.