Artifact for the IPDPS 2025 paper "Phase-based Frequency Scaling for Energy-efficient Heterogeneous Computing" by L. Carpentieri, A. De Caro, M. Salimi Beni, K. Fan and B. Cosenza (University of Salerno).
Frequency scaling on GPUs is usually applied either per application (coarse-grained, one frequency for the whole run) or per kernel (fine-grained, one frequency per kernel launch). Neither is ideal: a single frequency ignores that kernels have different energy characteristics, while per-kernel scaling pays a frequency-change cost of roughly 0.3–0.6 ms on every launch, which dominates for applications made of many short kernels.
This repository implements a third option. Kernels with similar energy behaviour are grouped into phases, and one frequency is set per phase. For MPI applications, the frequency change is additionally overlapped with communication, hiding its latency behind non-blocking collectives and stencil exchanges.
Reported results: up to 37 % energy saving and 1.87× speedup on single-GPU benchmarks, and 68 % energy saving and 3.63× speedup on two multi-GPU applications, across AMD MI100, Intel Max 1100 and NVIDIA V100S/A30 GPUs.
The work extends SYnergy (SC '23, paper), which provides the portable frequency-scaling and energy-profiling API used here. SYnergy is included as a git submodule.
- Repository layout
- Requirements
- Building
- Configuration file format
- Single-GPU workflow
- Cluster workflow (PBS / Aurora)
- MPI experiments
- Output format
- Known issues
- Citation
| Path | Contents |
|---|---|
SYnergy/ |
Submodule: the SYnergy frequency-scaling and energy-profiling library |
app/ |
Single-GPU SYCL benchmarks: ace, aop, bh, lulesh, matMul, median, metropolis, mnist, srad |
mpi-freq-change/ |
MPI + SYCL applications used for the communication-overlap experiments |
script_mpi_freq/ |
CMake configure scripts for the four MPI build variants (sync/async × hiding/no-hiding) |
config/ |
Frequency configuration files, one set per policy: app/, kernel/, phase/ |
kernel_freq_info/phases.txt |
Notes on the kernel sequence and the phases detected for each benchmark |
include/, utils/ |
Shared headers; utils/map_reader.hpp implements the FreqManager that applies the policy at runtime |
scripts/ |
Parsing and plotting scripts, plus scripts/aurora/ for the PBS-based cluster workflow |
run_*.sh, process_phase.sh |
Top-level drivers for profiling, running and post-processing |
- A SYCL 2020 compiler — the experiments used Intel DPC++ / oneAPI DPC++ (
icpx), with the CUDA, ROCm or Level Zero backend as appropriate. - CMake ≥ 3.5, a C++17 toolchain.
- An MPI implementation (Intel MPI in the paper's setup). MPI is required by the top-level
CMakeLists.txteven for single-GPU builds. - Vendor energy/frequency interface matching the target GPU:
- NVIDIA — NVML (
nvidia-smi) - AMD — ROCm SMI (
rocm-smi) - Intel — Level Zero (
intel_gpu_frequency), or GEOPM on systems where Level Zero is not exposed
- NVIDIA — NVML (
- Python 3 with
pandas,numpy,matplotlibandseabornfor the parsing and plotting scripts. - Permission to set the GPU core clock (on most clusters this requires administrator support).
git clone --recurse-submodules https://github.com/unisa-hpc/PhaseAwareFrequencyScaling.git
cd PhaseAwareFrequencyScaling
mkdir build && cd buildIf the repository was cloned without --recurse-submodules:
git submodule update --init --recursive| Option | Default | Meaning |
|---|---|---|
ENABLED_SYNERGY |
OFF |
Build and link SYnergy. Required for frequency scaling and energy profiling. |
ENABLED_TIME_EVENT_PROFILING |
OFF |
Enable SYCL queue profiling for per-kernel timings. |
WITH_MPI_ASYNCH |
OFF |
Use non-blocking MPI calls (MPI_Ibcast, MPI_Ireduce, …) in the MPI applications. |
ENABLE_FREQ_CHANGE_MPI_HIDING |
OFF |
Issue the frequency change during the communication instead of before the following kernel. |
WITH_PROCESS_FREQ_CHANGE |
OFF |
Perform the frequency change from a separate MPI process. |
DPCPP_WITH_CUDA_BACKEND + CUDA_ARCH |
— | NVIDIA target, e.g. -DCUDA_ARCH=sm_70. |
DPCPP_WITH_ROCM_BACKEND + ROCM_ARCH |
— | AMD target, e.g. -DROCM_ARCH=gfx908. |
DPCPP_WITH_LZ_BACKEND + LZ_ARCH |
— | Intel target, e.g. -DLZ_ARCH=pvc. |
SYnergy contributes its own options (SYNERGY_LZ_SUPPORT, SYNERGY_GEOPM_SUPPORT,
SYNERGY_DEVICE_PROFILING, SYNERGY_HOST_PROFILING, SYNERGY_KERNEL_PROFILING,
SYNERGY_USE_PROFILING_ENERGY, SYNERGY_SYCL_IMPL); see the submodule's documentation.
cmake -DCMAKE_CXX_COMPILER=clang++ \
-DDPCPP_WITH_CUDA_BACKEND=ON -DCUDA_ARCH=sm_70 \
-DENABLED_SYNERGY=ON \
-DENABLED_TIME_EVENT_PROFILING=ON \
-DSYNERGY_DEVICE_PROFILING=ON \
-DSYNERGY_USE_PROFILING_ENERGY=ON \
..
make -jexport SYSMAN=
cmake -DCMAKE_CXX_COMPILER=icpx \
-DCMAKE_CXX_FLAGS="-O3 -Wno-deprecated" \
-DDPCPP_WITH_LZ_BACKEND=ON -DLZ_ARCH=pvc \
-DENABLED_SYNERGY=ON -DSYNERGY_LZ_SUPPORT=ON \
-DENABLED_TIME_EVENT_PROFILING=ON \
-DSYNERGY_DEVICE_PROFILING=ON \
-DSYNERGY_USE_PROFILING_ENERGY=ON \
..
make -jThe ready-made variants under script_mpi_freq/ follow the same pattern and additionally set the
MPI flags; note that they hard-code the oneAPI compiler path, so adjust
-DCMAKE_CXX_COMPILER for your system.
cmake -DCMAKE_CXX_COMPILER=clang++ \
-DDPCPP_WITH_ROCM_BACKEND=ON -DROCM_ARCH=gfx908 \
-DENABLED_SYNERGY=ON \
-DENABLED_TIME_EVENT_PROFILING=ON \
-DSYNERGY_DEVICE_PROFILING=ON \
-DSYNERGY_USE_PROFILING_ENERGY=ON \
..
make -jEach benchmark reads its frequency configuration from standard input, which is why the run
scripts pipe a .conf file into the executable:
cat config/phase/ace.conf | ./ace_main 1If nothing is available on stdin, frequency scaling is disabled and the application runs at the default clock — this is how the profiling runs are performed.
A configuration file starts with the policy name, followed by one line per kernel:
APP|PHASE|KERNEL|NONE
<kernel_name> <freq_MHz> <KEEP|NO_KEEP>
...
The policy determines how FreqManager::getAndSetFreq() (in utils/map_reader.hpp) behaves:
| Policy | Behaviour |
|---|---|
APP |
Coarse-grained baseline. The first non-zero frequency is applied once and then cleared for all kernels, so the whole application runs at a single clock. |
KERNEL |
Fine-grained baseline (SYnergy). Every kernel gets its own frequency on every launch. |
PHASE |
This work. A frequency is set only at the kernels that begin a phase; the remaining kernels in the phase carry a frequency of 0, meaning "leave the clock as it is". |
NONE |
Frequency scaling disabled; used for the profiling runs. |
KEEP / NO_KEEP control what happens on the next invocation of the same kernel and are only
meaningful under PHASE. NO_KEEP (the default) clears the entry after the first call, so a kernel
inside a loop sets the frequency once rather than on every iteration. KEEP re-applies the
frequency at each invocation — needed when a preceding kernel in the loop body changes the clock, so
that the phase boundary is re-established on the next iteration.
Example (config/phase/ace.conf):
PHASE
calculateForce 1035 NO_KEEP
allenCahn 0 KEEP
boundaryConditionsPhi 607 NO_KEEP
thermalEquation 202 NO_KEEP
boundaryConditionsU 0 NO_KEEP
swapGrid_1 0 NO_KEEP
swapGrid_2 0 NO_KEEP
Three phases are set here — at calculateForce (1035 MHz), boundaryConditionsPhi (607 MHz) and
thermalEquation (202 MHz) — and the remaining kernels inherit the clock of the phase they belong
to. kernel_freq_info/phases.txt records the kernel sequence and phase reasoning for each
benchmark.
The order of lines matters, since it mirrors the order in which kernels are submitted to the in-order SYCL queue and therefore defines where phase boundaries fall.
The phase-based configuration is produced by profiling the application once per available frequency, picking the per-kernel optimum, and then grouping kernels into phases.
./run_profiling.sh --arch cuda -o data/logs-profiling --benchmarks=ace,aop,metropolis,mnist,srad --sampling=5--arch is one of cuda, rocm, lz, geopm; it selects the vendor tool used to set the clock.
--sampling=N runs every N-th supported frequency. The script writes one CSV per benchmark and
frequency into the output directory and resets the clock when it finishes.
scripts/extract_freqs.py and scripts/find_freq.py process the profiling CSVs to obtain the
per-kernel optimum for a chosen target (Min Energy, Min EDP, Max Perf). The
scripts/aurora/parse-csv/ scripts do the same in a more automated form and additionally emit the
APP, KERNEL and PHASE configuration files — see the next section.
Phase grouping is a per-application step: starting from the per-kernel optima and the runtime share
of each kernel, adjacent kernels are merged into a phase when the saving from a separate frequency
does not repay the change overhead. The resulting files, as used for the paper, are committed under
config/.
./run_phase.sh -o data/logs-phase --benchmarks=ace,aop,metropolis,mnist,srad --num-runs=5For each benchmark this executes the app, phase and kernel configurations in turn and appends
the results to <out>/<bench>/<bench>_{app,phase,kernel}.dat, with per-run stderr logs alongside.
The executables are expected in ./build.
./process_phase.sh --parseThis runs scripts/parse_phase_logs.py over logs/phase/native, writing
parsed/phase/phase_results.csv, and then scripts/plot_phase_results.py to produce the figures in
plot/. Omit --parse to re-plot from an existing CSV. Adjust the paths inside the script if your
log directory differs. scripts/plot_per_app.py and scripts/kernels_info.py produce the
per-application breakdowns.
scripts/aurora/ contains a fully scripted version of the workflow for PBS-managed systems. It was
written for Aurora but is portable: replace the submission scripts in bash-pbs/ and the frequency
list for your hardware.
| Step | Command |
|---|---|
| 1. Profile at all frequencies | python3 scripts/aurora/run/freq_scaling_profiling.py --app-dir=$(pwd)/build/ --benchmarks aop ace metropolis mnist srad --log-dir=$(pwd)/data/logs-profiling/ --pbs-path=$(pwd)/scripts/aurora/bash-pbs/ --config-dir=$(pwd)/config/none/ |
| 2. Aggregate into CSV | python3 scripts/aurora/parse-csv/parse_profiling.py --log-dir=... --output-dir=$(pwd)/data/profiling-csv/ --benchmarks ... |
| 3. Extract optimal frequencies | python3 scripts/aurora/parse-csv/extract_opt_freq.py --csv-dir=... --output-dir=$(pwd)/data/opt-freq-csv/ --benchmarks ... |
| 4. Generate configurations | python3 scripts/aurora/parse-csv/generate_config.py --profiling-csv-dir=... --out-config-dir=$(pwd)/data/config/ --benchmarks ... |
5. Adjust the PHASE files by hand |
see below |
| 6. Run all three policies | python3 scripts/aurora/run/opt_freq_scaling.py --app-dir=$(pwd)/build/ --log-dir=$(pwd)/data/final-logs/ --pbs-path=$(pwd)/scripts/aurora/bash-pbs/ --config-dir=$(pwd)/data/config/ --benchmarks ... |
| 7. Parse and plot | as in the single-GPU workflow |
Step 5 is the one manual step. generate_config.py emits per-kernel optima; turning them into phases
means choosing the phase boundaries and, where the frequency must be re-applied at the start of each
loop iteration, adding the corresponding KEEP entries. scripts/aurora/RUN.md and
scripts/aurora/STRUCTURE.md document each script's arguments in detail.
Step 1 expects a NONE configuration directory (frequency scaling disabled) that is not committed;
create one containing a <bench>.conf per benchmark whose first line is NONE.
The applications in mpi-freq-change/ measure the effect of overlapping the frequency change with
MPI communication. Four build variants are provided:
| Script | WITH_MPI_ASYNCH |
ENABLE_FREQ_CHANGE_MPI_HIDING |
|---|---|---|
compile_synch_no_hiding.sh |
OFF | OFF |
compile_synch_hiding.sh |
OFF | ON |
compile_asynch_no_hiding.sh |
ON | OFF |
compile_asynch_hiding.sh |
ON | ON |
With hiding disabled, the frequency is changed immediately before the kernel that follows the
communication. With hiding enabled and non-blocking MPI, the change is issued between the
MPI_I* call and the matching MPI_Wait/MPI_Waitall, so its latency is absorbed by the transfer.
WITH_PROCESS_FREQ_CHANGE moves the change to a separate MPI process; see mpi-freq-change/README.md
for the variants that were explored.
To reproduce the comparison:
./run_mpi_freq_bench.sh # builds both async variants and runs each 100× on 4 ranks
./run_parse_mpi.sh # summarises logs/*.log via parse_mpi_out.pyThe script rebuilds from scratch between variants, so build/ is wiped twice; run it from a checkout
where that is acceptable. Adjust NUM_RUNS, the rank count and the mpirun invocation for your
system. The two commented-out blocks run the synchronous variants.
The real-world MPI applications used in the paper — CloverLeaf and miniWeather — are not part
of this repository; they are the SYCL+MPI ports of the upstream codes, instrumented with the same
FreqManager mechanism.
Each benchmark prints a CSV header on stdout followed by one row per kernel:
kernel_name,host_energy[j],memory_freq [MHz],core_freq [MHz],times[ms],kernel_energy[j],
total_real_time[ms],sum_kernel_times[ms],total_device_energy[j],sum_kernel_energy[j]
Diagnostics — including the frequency retrieved for each kernel — go to stderr, which the run scripts
capture into .log files. scripts/parse_phase_logs.py reads the totals (Total time [ms],
Host energy [J], Device energy [J]) from these logs.
Host energy is measured through the Linux powercap interface; device energy comes from SYnergy's per-vendor backends.
CMakeLists.txtlists two sources that are not present in the tree:lulesh-sycl/lulesh_main.cc(the file lives atapp/lulesh-sycl/lulesh_main.cc) andmpi-freq-change/freq_change_overhead.cpp. Remove or fix these entries before configuring.- The build unconditionally requires MPI and applies the MPI-related compile definitions to every target, including the single-GPU benchmarks.
run_phase.shreassignsnum_runsinside the loop body when runningace, which overrides the--num-runsargument for subsequent iterations.process_phase.shuses hard-codedlogs/phase/native,parsed/andplot/paths.- The AMD frequency list in
run_profiling.shis hard-coded for the MI100 (see theTODOinget_core_frequencies). metropolisandmnistdo not fit in MI100 memory at the input sizes used in the paper.
@inproceedings{carpentieri2025phase,
title = {Phase-based Frequency Scaling for Energy-efficient Heterogeneous Computing},
author = {Carpentieri, Lorenzo and De Caro, Antonio and Salimi Beni, Majid and
Fan, Kaijie and Cosenza, Biagio},
booktitle = {2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS)},
year = {2025},
publisher = {IEEE}
}If you use the underlying frequency-scaling library, please also cite SYnergy:
@inproceedings{fan2023synergy,
title = {{SYnergy}: Fine-grained Energy-efficient Heterogeneous Computing for
Scalable Energy Saving},
author = {Fan, Kaijie and D'Antonio, Marco and Carpentieri, Lorenzo and Cosenza, Biagio and
Ficarelli, Federico and Cesarini, Daniele},
booktitle = {Proceedings of the International Conference for High Performance Computing,
Networking, Storage and Analysis (SC)},
year = {2023}
}The benchmarks under app/ are derived from
HeCBench and retain their original licences where included.