Storm™ for HPC and AI

Keep hardware.
Simplify cluster management.

Bring your existing HPC, scientific-computing, or AI cluster under Storm™ management. Keep compatible servers in place while the controller rebuilds their software configuration and manages the changes that follow.

Available for deployment · Supplied controller or installation image

An existing cluster, a new way to operate it

What changes. What carries forward.

Moving to Storm™ replaces the cluster's software configuration. Compatible physical hardware can stay, with controller cabling and other changes identified during planning.

  1. Scope what carries forward.

    Review the servers, storage, fabrics, applications, and workload requirements. Custom libraries and tunings may need initial integration; once integrated, their configuration can be reused. Keeping existing storage and datasets in place depends on the migration and service scope.

  2. Connect the controller and rebuild.

    Use a supplied controller or install a controller image on your own compatible node. Storm™ connects through the servers' out-of-band BMC management interfaces, configures SAN-backed root filesystems, and re-images selected bare-metal nodes with new configurations.

  3. Move workloads onto Storm™.

    Choose a staged migration or a full rebuild, with maintenance windows planned around your workloads. NetThunder, your team, or both can carry out the integration and transition, according to the services and support you need.

After deployment

Add capacity. Add applications.

Cluster changes should start with what you need to run. Storm™ calculates the configuration work from the requested change, the cluster's state, and its service dependencies.

Add a compute node.

  1. Connect a compatible node. It appears in inventory as unallocated.
  2. Select the node and the cluster it should join.
  3. Submit the configuration. Storm™ applies the node and cluster-service instructions.

Add a packaged application.

  1. Choose the application in the marketplace.
  2. Select Run.
  3. Storm™ calculates and applies the configuration and dependencies.
Manual administration and changes that need a restart

You can install software manually. Changes to managed dependencies need coordination, and edits made while the controller is disconnected may not be reflected in its recorded state.

Storm™ can restart bare-metal nodes through BMC control. Plan compute-interrupting changes around maintenance windows; workload recovery depends on application checkpointing or rerunning jobs. The operating design defines which lifecycle actions are automatic, approval-gated, or operator-driven.

Scheduler agnostic

Choose your software stack.

Storm™ supports user-defined software stacks. Choose the scheduler, applications, and services your workloads need.

NetThunder autoconfiguration

Package it. Deploy it.

Software that can be installed manually can be packaged with its service dependencies and the services they require. NetThunder autoconfiguration uses those definitions to calculate and apply the installation automatically.

Available package examples include Kubernetes, vLLM, Open WebUI, PyTorch, and agentic AI tools, with storage options including WEKA as a root filesystem. Versions, hardware, and integration scope follow the selected deployment.

Declarative bare-metal control

From request to running cluster.

Rules, templates, dependencies, and recorded state guide the controller's work. Compute accesses its storage directly, with the management path outside workload execution.

Authorize the request.

Define the workload and the resources Storm™ is permitted to use.

Approved cluster definition

Storm™ cluster management

  1. Resolve the dependencies.

    Rules, templates, service dependencies, and required state guide the work.

    Infrastructure control path
  2. Realize the cluster.

    Coordinate supported bare-metal compute, accelerators, storage, fabrics, and cluster services.

    Workloads execute directly on bare metal
  3. Validate and record.

    Evaluate workload-defined acceptance criteria and record intended, observed, and qualified state.

Recorded state informs the next approved change.

Storm™'s out-of-band control stays outside the workload execution path.

Inspect the detailed architecture
A left-to-right architecture diagram showing an authorized cluster request passing through declarative Storm™ control, direct realization on supported bare metal, qualification, and managed cluster state.
Read left to right: authorized request, resolved cluster intent, direct realization on supported bare metal, qualification, and managed state.On a narrow screen, swipe or scroll horizontally to inspect the full diagram.
Read the architecture sequence
  1. Authorize resources: The selected operating environment grants an authorized resource allocation and cluster definition.
  2. Define the cluster: Storm™ resolves service definitions, rules, templates, requirements, dependencies, and required state.
  3. Realize bare metal and services: Storm™ coordinates supported compute, accelerators, storage, networks, fabrics, and packaged services.
  4. Start and validate: Storm™ initializes the cluster, evaluates workload-defined acceptance criteria, and records intended, observed, and qualified state.
  5. Operate: Storm™ carries approved lifecycle changes forward as stateful cluster work.

Related products

Choose the scope you need.

Storm™ is the cluster-management platform. It can run your scientific-computing and AI applications without requiring Storm AI Factory™.

Building a cloud service?

CloudSpawn™ adds tenant-environment orchestration and isolation. Storm™ and CloudSpawn™ are separate platforms with distinct jobs. When paired, Storm™ manages clusters within the tenant boundary.

Explore CloudSpawn™

Want a complete AI appliance?

Storm AI Factory™ packages an AI operations module with Storm™ infrastructure software in an all-in-one inference or training appliance for on-premises or colocation deployment.

Explore Storm AI Factory™

Live virtual demo

See Storm™ build a cluster.

Watch Storm™ take unallocated nodes through configuration to a cluster running a workload. Then discuss how deployment would fit your existing infrastructure.

Email NetThunder +1 847.477.7676