
We’re excited to welcome ClusterShell to the High Performance Software Foundation!
ClusterShell is an event-driven, open source Python library designed to run local or remote commands in parallel on server farms and large Linux clusters, together with a set of command-line tools built on top of it. It handles the daily work of HPC cluster administration: operating on groups of nodes, executing distributed commands with optimized algorithms, and gathering and merging output across thousands of nodes.
ClusterShell leverages remote shell facilities already present on your systems, such as SSH, and its primary goal is to improve the administration of high-performance clusters by offering a lightweight but scalable Python API for developers, as well as convenient command-line tools for shell scripts and interactive use.
A Project Born in Production
ClusterShell was created at CEA, the French Alternative Energies and Atomic Energy Commission — an HPSF member organization — where it was first developed to administer some of the largest supercomputers in Europe. CEA provided the resources and the demanding production environment that shaped the project from its earliest days, and CEA engineers have contributed to its design and code throughout its history.
The project’s lead maintainer, Stéphane Thiell, later continued its development at Stanford Research Computing, which supported ClusterShell for over a decade — providing development hardware and the time to keep the project growing while it ran in production on Stanford’s research clusters, including Sherlock, the university’s flagship shared HPC cluster serving thousands of researchers. Today, ClusterShell is used at major HPC sites and research institutions worldwide, scaling from small clusters to systems with tens of thousands of nodes.
“ClusterShell exists because two great institutions gave it room to grow: CEA, where it was born to manage some of Europe’s largest supercomputers, and Stanford Research Computing, which supported its development for more than a decade,” said Thiell. “Joining the High Performance Software Foundation is the natural next step — a neutral, sustainable home where the whole HPC community can help shape what comes next.”
What Makes ClusterShell Different
What sets ClusterShell apart is its combination of a clean Python API and battle-tested command-line tools. The clush command lets admins run commands on arbitrarily complex node sets, with output gathering built right in: identical output is merged and formatted on the fly, and an integrated diff mode highlights the nodes that don’t match, often the fastest way to find the one misbehaving node in ten thousand. clush also doubles as an interactive parallel shell and can copy files to entire node sets or gather them back.
At the heart of ClusterShell is its node set engine, one of the project’s most beloved features, used as much through the Python API as on the command line. It lets you do math on your cluster: the nodeset command (also available as cluset) folds ten thousand hostnames into a single expression like node[0001-9999], expands it back in an instant, and combines sets with unions, intersections, and differences. Named node groups like @gpu or @rack12 add your cluster’s own vocabulary, defined in simple files or pulled live from sources like Slurm. A question that once meant a morning of scripting, “everything in @rack12 except the nodes down for maintenance”, collapses into one command that answers in milliseconds. It works so well that other HPC tools embed the ClusterShell library just for this.
ClusterShell supports multiple execution backends (SSH, pdsh, rsh, and local execution). For the largest systems, its tree propagation mode routes commands through gateway nodes, scaling parallel execution to tens of thousands of nodes.
ClusterShell is a community-driven open source project maintained on GitHub, with contributors spanning multiple HPC sites and institutions. Co-maintainer Aurélien Degrémont has helped steer the project’s design since its earliest days at CEA and continues to guide its evolution today. The project is governed by a Technical Charter and operates under LF Europe stewardship as an HPSF Established project.
“Every significant change in ClusterShell gets debated, reviewed, and weighed against years of production experience before it ships,” said Degrémont. “That discipline is why administrators trust it on their largest systems, and joining HPSF is how we make sure that trust outlives any single maintainer.”
What’s Next
As part of the High Performance Software Foundation, ClusterShell will keep its focus where it has always been: reliability at scale. Work is under way on significant performance improvements for operations on very large node sets — more on that soon — along with continued scalability work and making it easier for new contributors to get involved with the project.
“The open source community plays a critical role in advancing high performance computing,” said Todd Gamblin, HPSF Governing Board Chair. “Projects like ClusterShell help system administrators manage large clusters more efficiently and lower the barrier to parallel command execution at scale.”
Learn More
To learn more about ClusterShell, visit the project’s documentation at https://clustershell.readthedocs.io/ or explore the GitHub repository at https://github.com/clustershell/clustershell.
ClusterShell is packaged in major Linux distributions — including Fedora, EPEL, Debian, Ubuntu, and openSUSE — and is available on PyPI (pip install ClusterShell).
Get Involved
The ClusterShell community welcomes contributors, users, and collaborators from across the HPC ecosystem. Whether you’re an HPC system administrator, a developer building cluster tooling, or a researcher who relies on parallel command execution, there’s a place for you: report issues or start a discussion at https://github.com/clustershell/clustershell/issues, or find the community in the HPSF Slack.
About the CEA
The CEA is a public research organization that supports public policy decision-making and equips French and European businesses and communities with the scientific and technological means to better navigate four major societal transitions: energy transition, digital transition, future healthcare, and national/global security. Its mission is to ensure France and Europe maintain scientific, technological, and industrial leadership, contributing to a more secure and controlled present and future for all. The CEA is guided by three core values: curiosity, cooperation, and a strong sense of responsibility.
Learn more at: www.cea.fr/english
About Stanford Research Computing
Stanford Research Computing, a joint effort of the Stanford Dean of Research and University IT, delivers and supports comprehensive programs that advance computational and data-intensive research across Stanford University. Its team engineers, manages, and supports high-performance computing systems and services, including shared compute clusters and storage systems for modeling, simulation, and data analysis, hosted in an award-winning, energy-efficient research data center. Research Computing also provides training, consultation,
and support to help researchers explore data and answer research questions at a scale not possible on desktops or departmental servers.
Learn more at: srcc.stanford.edu