Skip to content

Repository files navigation

Phase-based Frequency Scaling for Energy-efficient Heterogeneous Computing

Artifact for the IPDPS 2025 paper "Phase-based Frequency Scaling for Energy-efficient Heterogeneous Computing" by L. Carpentieri, A. De Caro, M. Salimi Beni, K. Fan and B. Cosenza (University of Salerno).

📄 Paper (preprint)

Frequency scaling on GPUs is usually applied either per application (coarse-grained, one frequency for the whole run) or per kernel (fine-grained, one frequency per kernel launch). Neither is ideal: a single frequency ignores that kernels have different energy characteristics, while per-kernel scaling pays a frequency-change cost of roughly 0.3–0.6 ms on every launch, which dominates for applications made of many short kernels.

This repository implements a third option. Kernels with similar energy behaviour are grouped into phases, and one frequency is set per phase. For MPI applications, the frequency change is additionally overlapped with communication, hiding its latency behind non-blocking collectives and stencil exchanges.

Reported results: up to 37 % energy saving and 1.87× speedup on single-GPU benchmarks, and 68 % energy saving and 3.63× speedup on two multi-GPU applications, across AMD MI100, Intel Max 1100 and NVIDIA V100S/A30 GPUs.

The work extends SYnergy (SC '23, paper), which provides the portable frequency-scaling and energy-profiling API used here. SYnergy is included as a git submodule.


Contents


Repository layout

Path Contents
SYnergy/ Submodule: the SYnergy frequency-scaling and energy-profiling library
app/ Single-GPU SYCL benchmarks: ace, aop, bh, lulesh, matMul, median, metropolis, mnist, srad
mpi-freq-change/ MPI + SYCL applications used for the communication-overlap experiments
script_mpi_freq/ CMake configure scripts for the four MPI build variants (sync/async × hiding/no-hiding)
config/ Frequency configuration files, one set per policy: app/, kernel/, phase/
kernel_freq_info/phases.txt Notes on the kernel sequence and the phases detected for each benchmark
include/, utils/ Shared headers; utils/map_reader.hpp implements the FreqManager that applies the policy at runtime
scripts/ Parsing and plotting scripts, plus scripts/aurora/ for the PBS-based cluster workflow
run_*.sh, process_phase.sh Top-level drivers for profiling, running and post-processing

Requirements

  • A SYCL 2020 compiler — the experiments used Intel DPC++ / oneAPI DPC++ (icpx), with the CUDA, ROCm or Level Zero backend as appropriate.
  • CMake ≥ 3.5, a C++17 toolchain.
  • An MPI implementation (Intel MPI in the paper's setup). MPI is required by the top-level CMakeLists.txt even for single-GPU builds.
  • Vendor energy/frequency interface matching the target GPU:
    • NVIDIA — NVML (nvidia-smi)
    • AMD — ROCm SMI (rocm-smi)
    • Intel — Level Zero (intel_gpu_frequency), or GEOPM on systems where Level Zero is not exposed
  • Python 3 with pandas, numpy, matplotlib and seaborn for the parsing and plotting scripts.
  • Permission to set the GPU core clock (on most clusters this requires administrator support).

Building

git clone --recurse-submodules https://github.com/unisa-hpc/PhaseAwareFrequencyScaling.git
cd PhaseAwareFrequencyScaling
mkdir build && cd build

If the repository was cloned without --recurse-submodules:

git submodule update --init --recursive

CMake options

Option Default Meaning
ENABLED_SYNERGY OFF Build and link SYnergy. Required for frequency scaling and energy profiling.
ENABLED_TIME_EVENT_PROFILING OFF Enable SYCL queue profiling for per-kernel timings.
WITH_MPI_ASYNCH OFF Use non-blocking MPI calls (MPI_Ibcast, MPI_Ireduce, …) in the MPI applications.
ENABLE_FREQ_CHANGE_MPI_HIDING OFF Issue the frequency change during the communication instead of before the following kernel.
WITH_PROCESS_FREQ_CHANGE OFF Perform the frequency change from a separate MPI process.
DPCPP_WITH_CUDA_BACKEND + CUDA_ARCH — NVIDIA target, e.g. -DCUDA_ARCH=sm_70.
DPCPP_WITH_ROCM_BACKEND + ROCM_ARCH — AMD target, e.g. -DROCM_ARCH=gfx908.
DPCPP_WITH_LZ_BACKEND + LZ_ARCH — Intel target, e.g. -DLZ_ARCH=pvc.

SYnergy contributes its own options (SYNERGY_LZ_SUPPORT, SYNERGY_GEOPM_SUPPORT, SYNERGY_DEVICE_PROFILING, SYNERGY_HOST_PROFILING, SYNERGY_KERNEL_PROFILING, SYNERGY_USE_PROFILING_ENERGY, SYNERGY_SYCL_IMPL); see the submodule's documentation.

Example: NVIDIA

cmake -DCMAKE_CXX_COMPILER=clang++ \
      -DDPCPP_WITH_CUDA_BACKEND=ON -DCUDA_ARCH=sm_70 \
      -DENABLED_SYNERGY=ON \
      -DENABLED_TIME_EVENT_PROFILING=ON \
      -DSYNERGY_DEVICE_PROFILING=ON \
      -DSYNERGY_USE_PROFILING_ENERGY=ON \
      ..
make -j

Example: Intel GPU (Level Zero)

export SYSMAN=
cmake -DCMAKE_CXX_COMPILER=icpx \
      -DCMAKE_CXX_FLAGS="-O3 -Wno-deprecated" \
      -DDPCPP_WITH_LZ_BACKEND=ON -DLZ_ARCH=pvc \
      -DENABLED_SYNERGY=ON -DSYNERGY_LZ_SUPPORT=ON \
      -DENABLED_TIME_EVENT_PROFILING=ON \
      -DSYNERGY_DEVICE_PROFILING=ON \
      -DSYNERGY_USE_PROFILING_ENERGY=ON \
      ..
make -j

The ready-made variants under script_mpi_freq/ follow the same pattern and additionally set the MPI flags; note that they hard-code the oneAPI compiler path, so adjust -DCMAKE_CXX_COMPILER for your system.

Example: AMD

cmake -DCMAKE_CXX_COMPILER=clang++ \
      -DDPCPP_WITH_ROCM_BACKEND=ON -DROCM_ARCH=gfx908 \
      -DENABLED_SYNERGY=ON \
      -DENABLED_TIME_EVENT_PROFILING=ON \
      -DSYNERGY_DEVICE_PROFILING=ON \
      -DSYNERGY_USE_PROFILING_ENERGY=ON \
      ..
make -j

Configuration file format

Each benchmark reads its frequency configuration from standard input, which is why the run scripts pipe a .conf file into the executable:

cat config/phase/ace.conf | ./ace_main 1

If nothing is available on stdin, frequency scaling is disabled and the application runs at the default clock — this is how the profiling runs are performed.

A configuration file starts with the policy name, followed by one line per kernel:

APP|PHASE|KERNEL|NONE
<kernel_name> <freq_MHz> <KEEP|NO_KEEP>
...

The policy determines how FreqManager::getAndSetFreq() (in utils/map_reader.hpp) behaves:

Policy Behaviour
APP Coarse-grained baseline. The first non-zero frequency is applied once and then cleared for all kernels, so the whole application runs at a single clock.
KERNEL Fine-grained baseline (SYnergy). Every kernel gets its own frequency on every launch.
PHASE This work. A frequency is set only at the kernels that begin a phase; the remaining kernels in the phase carry a frequency of 0, meaning "leave the clock as it is".
NONE Frequency scaling disabled; used for the profiling runs.

KEEP / NO_KEEP control what happens on the next invocation of the same kernel and are only meaningful under PHASE. NO_KEEP (the default) clears the entry after the first call, so a kernel inside a loop sets the frequency once rather than on every iteration. KEEP re-applies the frequency at each invocation — needed when a preceding kernel in the loop body changes the clock, so that the phase boundary is re-established on the next iteration.

Example (config/phase/ace.conf):

PHASE
calculateForce 1035 NO_KEEP
allenCahn 0 KEEP
boundaryConditionsPhi 607 NO_KEEP
thermalEquation 202 NO_KEEP
boundaryConditionsU 0 NO_KEEP
swapGrid_1 0 NO_KEEP
swapGrid_2 0 NO_KEEP

Three phases are set here — at calculateForce (1035 MHz), boundaryConditionsPhi (607 MHz) and thermalEquation (202 MHz) — and the remaining kernels inherit the clock of the phase they belong to. kernel_freq_info/phases.txt records the kernel sequence and phase reasoning for each benchmark.

The order of lines matters, since it mirrors the order in which kernels are submitted to the in-order SYCL queue and therefore defines where phase boundaries fall.

Single-GPU workflow

The phase-based configuration is produced by profiling the application once per available frequency, picking the per-kernel optimum, and then grouping kernels into phases.

1. Profile across frequencies

./run_profiling.sh --arch cuda -o data/logs-profiling --benchmarks=ace,aop,metropolis,mnist,srad --sampling=5

--arch is one of cuda, rocm, lz, geopm; it selects the vendor tool used to set the clock. --sampling=N runs every N-th supported frequency. The script writes one CSV per benchmark and frequency into the output directory and resets the clock when it finishes.

2. Derive the optimal frequencies and configurations

scripts/extract_freqs.py and scripts/find_freq.py process the profiling CSVs to obtain the per-kernel optimum for a chosen target (Min Energy, Min EDP, Max Perf). The scripts/aurora/parse-csv/ scripts do the same in a more automated form and additionally emit the APP, KERNEL and PHASE configuration files — see the next section.

Phase grouping is a per-application step: starting from the per-kernel optima and the runtime share of each kernel, adjacent kernels are merged into a phase when the saving from a separate frequency does not repay the change overhead. The resulting files, as used for the paper, are committed under config/.

3. Run the three policies

./run_phase.sh -o data/logs-phase --benchmarks=ace,aop,metropolis,mnist,srad --num-runs=5

For each benchmark this executes the app, phase and kernel configurations in turn and appends the results to <out>/<bench>/<bench>_{app,phase,kernel}.dat, with per-run stderr logs alongside. The executables are expected in ./build.

4. Parse and plot

./process_phase.sh --parse

This runs scripts/parse_phase_logs.py over logs/phase/native, writing parsed/phase/phase_results.csv, and then scripts/plot_phase_results.py to produce the figures in plot/. Omit --parse to re-plot from an existing CSV. Adjust the paths inside the script if your log directory differs. scripts/plot_per_app.py and scripts/kernels_info.py produce the per-application breakdowns.

Cluster workflow (PBS / Aurora)

scripts/aurora/ contains a fully scripted version of the workflow for PBS-managed systems. It was written for Aurora but is portable: replace the submission scripts in bash-pbs/ and the frequency list for your hardware.

Step Command
1. Profile at all frequencies python3 scripts/aurora/run/freq_scaling_profiling.py --app-dir=$(pwd)/build/ --benchmarks aop ace metropolis mnist srad --log-dir=$(pwd)/data/logs-profiling/ --pbs-path=$(pwd)/scripts/aurora/bash-pbs/ --config-dir=$(pwd)/config/none/
2. Aggregate into CSV python3 scripts/aurora/parse-csv/parse_profiling.py --log-dir=... --output-dir=$(pwd)/data/profiling-csv/ --benchmarks ...
3. Extract optimal frequencies python3 scripts/aurora/parse-csv/extract_opt_freq.py --csv-dir=... --output-dir=$(pwd)/data/opt-freq-csv/ --benchmarks ...
4. Generate configurations python3 scripts/aurora/parse-csv/generate_config.py --profiling-csv-dir=... --out-config-dir=$(pwd)/data/config/ --benchmarks ...
5. Adjust the PHASE files by hand see below
6. Run all three policies python3 scripts/aurora/run/opt_freq_scaling.py --app-dir=$(pwd)/build/ --log-dir=$(pwd)/data/final-logs/ --pbs-path=$(pwd)/scripts/aurora/bash-pbs/ --config-dir=$(pwd)/data/config/ --benchmarks ...
7. Parse and plot as in the single-GPU workflow

Step 5 is the one manual step. generate_config.py emits per-kernel optima; turning them into phases means choosing the phase boundaries and, where the frequency must be re-applied at the start of each loop iteration, adding the corresponding KEEP entries. scripts/aurora/RUN.md and scripts/aurora/STRUCTURE.md document each script's arguments in detail.

Step 1 expects a NONE configuration directory (frequency scaling disabled) that is not committed; create one containing a <bench>.conf per benchmark whose first line is NONE.

MPI experiments

The applications in mpi-freq-change/ measure the effect of overlapping the frequency change with MPI communication. Four build variants are provided:

Script WITH_MPI_ASYNCH ENABLE_FREQ_CHANGE_MPI_HIDING
compile_synch_no_hiding.sh OFF OFF
compile_synch_hiding.sh OFF ON
compile_asynch_no_hiding.sh ON OFF
compile_asynch_hiding.sh ON ON

With hiding disabled, the frequency is changed immediately before the kernel that follows the communication. With hiding enabled and non-blocking MPI, the change is issued between the MPI_I* call and the matching MPI_Wait/MPI_Waitall, so its latency is absorbed by the transfer. WITH_PROCESS_FREQ_CHANGE moves the change to a separate MPI process; see mpi-freq-change/README.md for the variants that were explored.

To reproduce the comparison:

./run_mpi_freq_bench.sh     # builds both async variants and runs each 100× on 4 ranks
./run_parse_mpi.sh          # summarises logs/*.log via parse_mpi_out.py

The script rebuilds from scratch between variants, so build/ is wiped twice; run it from a checkout where that is acceptable. Adjust NUM_RUNS, the rank count and the mpirun invocation for your system. The two commented-out blocks run the synchronous variants.

The real-world MPI applications used in the paper — CloverLeaf and miniWeather — are not part of this repository; they are the SYCL+MPI ports of the upstream codes, instrumented with the same FreqManager mechanism.

Output format

Each benchmark prints a CSV header on stdout followed by one row per kernel:

kernel_name,host_energy[j],memory_freq [MHz],core_freq [MHz],times[ms],kernel_energy[j],
total_real_time[ms],sum_kernel_times[ms],total_device_energy[j],sum_kernel_energy[j]

Diagnostics — including the frequency retrieved for each kernel — go to stderr, which the run scripts capture into .log files. scripts/parse_phase_logs.py reads the totals (Total time [ms], Host energy [J], Device energy [J]) from these logs.

Host energy is measured through the Linux powercap interface; device energy comes from SYnergy's per-vendor backends.

Known issues

  • CMakeLists.txt lists two sources that are not present in the tree: lulesh-sycl/lulesh_main.cc (the file lives at app/lulesh-sycl/lulesh_main.cc) and mpi-freq-change/freq_change_overhead.cpp. Remove or fix these entries before configuring.
  • The build unconditionally requires MPI and applies the MPI-related compile definitions to every target, including the single-GPU benchmarks.
  • run_phase.sh reassigns num_runs inside the loop body when running ace, which overrides the --num-runs argument for subsequent iterations.
  • process_phase.sh uses hard-coded logs/phase/native, parsed/ and plot/ paths.
  • The AMD frequency list in run_profiling.sh is hard-coded for the MI100 (see the TODO in get_core_frequencies).
  • metropolis and mnist do not fit in MI100 memory at the input sizes used in the paper.

Citation

@inproceedings{carpentieri2025phase,
  title     = {Phase-based Frequency Scaling for Energy-efficient Heterogeneous Computing},
  author    = {Carpentieri, Lorenzo and De Caro, Antonio and Salimi Beni, Majid and
               Fan, Kaijie and Cosenza, Biagio},
  booktitle = {2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS)},
  year      = {2025},
  publisher = {IEEE}
}

If you use the underlying frequency-scaling library, please also cite SYnergy:

@inproceedings{fan2023synergy,
  title     = {{SYnergy}: Fine-grained Energy-efficient Heterogeneous Computing for
               Scalable Energy Saving},
  author    = {Fan, Kaijie and D'Antonio, Marco and Carpentieri, Lorenzo and Cosenza, Biagio and
               Ficarelli, Federico and Cesarini, Daniele},
  booktitle = {Proceedings of the International Conference for High Performance Computing,
               Networking, Storage and Analysis (SC)},
  year      = {2023}
}

The benchmarks under app/ are derived from HeCBench and retain their original licences where included.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages