Investigation of the applicability of autotuners in supercomputer co-design problems using MLKAPS
Authors
-
Tatiana E. Zagoruiko
-
Alexander S. Antonov
Keywords:
autotuner
supercomputer co-design
Kernel Tuner
KTT
GPTune
HYPERF
VibeCodeHPC
MLKAPS
HPL
optimization
Abstract
This paper provides an overview of a subset of modern autotuners, focusing on their functionality and demonstrating their applicability to solving problems within the framework of supercomputer co-design concept. The results of an experimental study of the MLKAPS autotuner’s effectiveness in relation to the problem of optimizing HPL benchmark parameters are presented. A comparative analysis of the MLKAPS autotuner with the full iteration method is conducted. It is demonstrated that MLKAPS, given a larger search space of all possible configurations, can significantly reduce the search time for the optimal configuration. The obtained results demonstrate the practical applicability of autotuners for automating the tuning of parallel applications in conditions of availability of limited resources of supercomputer complexes.
Section
Parallel software tools and technologies
References
- Vl. V. Voevodin, A. S. Antonov, D. A. Nikitenko, et al., “Supercomputer Lomonosov-2: Large Scale, Deep Monitoring and Fine Analytics for the User Community,” Supercomput. Front. Innov. 6 (2), 4–11 (2019).
doi 10.14529/jsfi190201
- A. S. Antonov, R. V. Maier, D. A. Nikitenko, and V. V. Voevodin, “An Approach to Solving the Problem of Supercomputer Co-design,” Lobachevskii J. Math. 45 (7), 2965–2973 (2024).
doi 10.1134/S1995080224603680
- J. A. Ang, “High Performance Computing Co-Design Strategies,” in MEMSYS’15: Proceedings of the 2015 International Symposium on Memory Systems, USA, Washington, October 5–8, 2015 (ACM, New York, 2015), pp. 51–52.
doi 10.1145/2818950.2818959
- H. Kasim, V. March, R. Zhang, and S. See, “Survey on Parallel Programming Model,” in Network and Parallel Computing IFIP International Conference (NPC 2008), China, Shanghai, October 18–20, 2008 (Springer, Berlin, 2008), pp. 266–275.
doi 10.1007/978-3-540-88140-7_24
- A. A. Romanenko, Features of Software Adaptation for GPU Using OpenACC Technology (RIC NGU, Novosibirsk, 2016) [in Russian].
- A. V. Boreskov, A. A. Kharlamov, N. D. Markovsky, et al., Parallel computing on the GPU. CUDA Architecture and Software model (Izdat. Moskow. Universiteta, Moscow, 2015) [in Russian].
- X. Wu, M. Kruse, P. Balaprakash, et al., “Autotuning PolyBench Benchmarks with LLVM Clang/Polly Loop Optimization Pragmas Using Bayesian Optimization,” Concurrency and Computation: Practice and Experience 34 (20), Article Number e6683 (2022).
doi 10.1002/cpe.6683
- X. Wu, P. Balaprakash, M. Kruse, et al., “ytopt: Autotuning Scientific Applications for Energy Efficiency at Large Scales,” Concurrency and Computation: Practice and Experience 37 (1), Article Number e8322 (2025).
doi 10.1002/cpe.8322
- GitHub – ytopt-team/ytopt: ytopt: machine-learning-based autotuning and hyperparameter optimization framework using Bayesian Optimization cdot GitHub.
https://github.com/ytopt-team/ytopt Cited July 9, 2026.
- J. Park, Y. Shin, J. Lee, et al., “HYPERF: End-to-End Autotuning Framework for High-Performance Computing,” in HPDC’25: Proceedings of the 34th International Symposium on High-Performance Parallel and Distributed Computing, USA, July 20–23, 2025 (ACM, New York, 2025), pp. 1–14.
doi 10.1145/3731545.3731588
- GitHub – SNU-CODElab/HYPERF cdot GitHub.
https://github.com/SNU-CODElab/HYPERF Cited July 9, 2026.
- M. E. Jam, E. Petit, P. de Oliveira Castro, et al., “MLKAPS: Machine Learning and Adaptive Sampling for HPC Kernel Auto-tuning,” ACM Trans. Archit. Code Optim. 22 (4), 1–22 (2025).
doi 10.1145/3774418
- GitHub – MLCGO/MLKAPS: MLKAPS: Machine Learning for Kernel Accuracy and Performance Studies cdot GitHub.
https://github.com/MLCGO/MLKAPS Cited July 9, 2026.
- C. Nugteren and V. Codreanu, “CLTune: A Generic Auto-Tuner for OpenCL Kernels,” in 2015 IEEE 9th International Symposium on Embedded Multicore/Many-core Systems-on-Chip, Italy, Turin, September 23–25, 2015 (IEEE Press, 2015), pp. 195–202.
doi 10.1109/MCSoC.2015.10
- GitHub – CNugteren/CLTune: CLTune: An automatic OpenCL & CUDA kernel tuner cdot GitHub.
https://github.com/CNugteren/CLTune Cited July 9, 2026.
- B. van Werkhoven, “Kernel Tuner: A search-optimizing GPU code auto-tuner,” Future Generation Computer Systems 90, 347–358 (2019).
doi 10.1016/j.future.2018.08.004
- Design documentation – Kernel Tuner 1.3.1 documentation.
https://kerneltuner.github.io/kernel_tuner/stable/design.html Cited July 13, 2026.
- GitHub – KernelTuner/kernel_tuner: Kernel Tuner cdot GitHub.
https://github.com/KernelTuner/kernel_tuner Cited July 13, 2026.
- F. Petrovič, D. Střelák, J. Hozzová, et al., “A Benchmark Set of Highly-Efficient CUDA and OpenCL Kernels and its Dynamic Autotuning with Kernel Tuning Toolkit,” Future Generation Computer Systems 108, 161–177 (2020).
doi 10.1016/j.future.2020.02.069
- GitHub – HiPerCoRe/KTT: Kernel Tuning Toolkit cdot GitHub.
https://github.com/HiPerCoRe/KTT Cited July 13, 2026.
- KTT – Kernel Tuning Toolkit.
https://hipercore.github.io/KTT/index.html Cited July 13, 2026.
- R. C. Whaley and J. J. Dongarra, “Automatically Tuned Linear Algebra Software,” in SC’98: Proceedings of the 1998 ACM/IEEE Conference on Supercomputing, USA, Orlando, November 7–13, 1998 (IEEE Press, 1998), p. 38.
doi 10.1109/SC.1998.10004
- Automatically Tuned Linear Algebra Software (ATLAS).
https://math-atlas.sourceforge.net/ Cited July 13, 2026.
- A. B. Hayes, L. Li, D. Chavarr’ia-Miranda, et al., “Orion: A Framework for GPU Occupancy Tuning,” in Middleware’16: Proceedings of the 17th International Middleware Conference, Italy, Trento, December 12–16, 2016 (ACM, New York, 2016), pp. 1–13.
doi 10.1145/2988336.2988355
- D. Mustafa, “A Survey of Performance Tuning Techniques and Tools for Parallel Applications,” IEEE Access 10, 15036–15055 (2022).
doi 10.1109/ACCESS.2022.3147846
- A. H. Ashouri, W. Killian, J. Cavazos, et al., “A Survey on Compiler Autotuning Using Machine Learning,” ACM Computing Surveys (CSUR). 51 (5), 1–42 (2018).
doi 10.1145/3197978
- P. Balaprakash, J. Dongarra, T. Gamblin, et al., “Autotuning in High-Performance Computing Applications,” Proceedings of the IEEE. 106 (11), 2068–2083 (2018).
doi 10.1109/JPROC.2018.2841200
- G. Fursin, R. Miceli, A. Lokhmotov, et al., “Collective mind: Towards practical and collaborative auto-tuning,” Scientific Programming 22 (4), 309–329 (2014).
doi 10.3233/SPR-140396
- GitHub – gfursin/cm-ctuning-code-source: Collective Mind Repository (cTuning; Benchmarks) cdot GitHub.
https://github.com/gfursin/cm-ctuning-code-source Cited July 13, 2026.
- F.-J. Willemsen, R. Schoonhoven, J. Filipovič, et al., “A methodology for comparing optimization algorithms for auto-tuning,” Future Generation Computer Systems 159, 489–504 (2024).
doi 10.1016/j.future.2024.05.021
- J. W. Romein and B. Veenboer, “PowerSensor 2: A Fast Power Measurement Tool,” in 2018 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), UK, Belfast, April 2–4, 2018 (IEEE Press, 2018), pp. 111–113.
doi 10.1109/ISPASS.2018.00020
- NVIDIA Management Library (NVML) | NVIDIA Developer.
https://developer.nvidia.com/management-library-nvml Cited July 14, 2026.
- Y. Liu, W. M. Sid-Lakhdar, O. Marques, et al., “GPTune: Multitask Learning for Autotuning Exascale Applications,” in PPoPP’21: Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming, Republic of Korea, February 27, 2021 (ACM, New York, 2021), pp. 234–246.
doi 10.1145/3437801.3441621
- Y. Cho, J. W. Demmel, G. Dinh, et al., “GPTune user guide,”
https://gptune.lbl.gov/documentation/gptune-user-guide/ Cited July 16, 2026.
- GitHub – gptune/GPTune cdot GitHub.
https://github.com/gptune/GPTune Cited July 14, 2026.
- S. Hayashi, K. Morita, D. Mukunoki, et al., “VibeCodeHPC: An Agent-Based Iterative Prompting Auto-Tuner for HPC Code Generation Using LLMs,” (2025).
doi 10.48550/arXiv.2510.00031
- GitHub – Katagiri-Hoshino-Lab/VibeCodeHPC: CLI-based multi-agents for auto-tuning (e.g. HPC code optimazation loops) supporting Local LLMs cdot GitHub.
https://github.com/Katagiri-Hoshino-Lab/VibeCodeHPC Cited July 14, 2026.
- GitHub – anthropics/claude-code: Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands cdot GitHub.
https://github.com/anthropics/claude-code Cited July 14, 2026.
- GitHub – wonderwhy-er/DesktopCommanderMCP: This is MCP server for Claude that gives it terminal control, file system search and diff file editing capabilities cdot GitHub.
https://github.com/wonderwhy-er/DesktopCommanderMCP Cited July 14, 2026.