2024
Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning Techniques
HARUTYUNYAN, Suren; Eduardo CÉSAR; Anna SIKORA; Jiří FILIPOVIČ; Akash DUTTA et al.Basic information
Original name
Efficient Code Region Characterization Through Automatic Performance Counters Reduction Using Machine Learning Techniques
Authors
HARUTYUNYAN, Suren; Eduardo CÉSAR; Anna SIKORA; Jiří FILIPOVIČ; Akash DUTTA; Ali JANNESARI and Jordi ALCARAZ
Edition
Madrid, Spain, European Conference on Parallel Processing, p. 18-32, 15 pp. 2024
Publisher
Springer Nature Switzerland
Other information
Language
English
Type of outcome
Proceedings paper
Field of Study
10201 Computer sciences, information science, bioinformatics
Country of publisher
Spain
Confidentiality degree
is not subject to a state or trade secret
Publication form
electronic version available online
References:
Impact factor
Impact factor: 0.402 in 2005
Marked to be transferred to RIV
Yes
RIV identification code
RIV/00216224:14610/24:00137325
Organization unit
Institute of Computer Science
ISBN
978-3-031-69576-6
ISSN
UT WoS
EID Scopus
Keywords in English
Performance counters; Automatic dimension reduction; machine learning ensambles; parallel region classification
Tags
International impact, Reviewed
Changed: 4/4/2025 13:13, Mgr. Eva Špillingová
Abstract
In the original language
Leveraging hardware performance counters provides valuable insights into system resource utilization, aiding performance analysis and tuning for parallel applications. The available counters vary with architecture and are collected at execution time. Their abundance and the limited number of registers for measurement make gathering laborious and costly. Efficient characterization of parallel regions necessitates a dimension reduction strategy. While recent efforts have focused on manually reducing the number of counters for specific architectures, this paper introduces a novel approach: an automatic dimension reduction technique for efficiently characterizing parallel code regions across diverse architectures. The methodology is based on Machine Learning ensembles because of their precision and ability at capturing different relationships between the input features and the target variables. Evaluation results show that ensembles can successfully reduce the number of hardware performance counters that characterize a code region. We validate our approach on CPUs using a comprehensive dataset of OpenMP regions, showing that any region can be accurately characterized by 8 relevant hardware performance counters. In addition, we also apply the proposed methodology on GPUs using a reduced set of kernels, demonstrating its effectiveness across various hardware configurations and workloads.
Links
| LM2023054, research and development project |
|