ConsultancyHPC and parallelisation

Computing power that scales.

We speed up simulations, data pipelines and AI training: from performance analysis to parallelisation across thousands of cores and GPUs. Synchronous where it must be, asynchronous where it can be.

Synchronous or asynchronous: where the time goes.

In a bulk-synchronous program every processor waits at each barrier for the slowest one. With asynchronous execution, a task starts as soon as its own neighbours are done. The larger the variation in compute time, the bigger the difference.

The same tasks, two execution models. Each task depends on its neighbours in the previous time step.
  • Compute
  • Waiting
  • Barrier
0Synchronous runtime, as index
0Asynchronous runtime, synchronous = 100
0Utilisation, synchronous and asynchronous
0Speedup from asynchronous execution

Amdahl’s law is relentless.

If 5% of a program stays serial, the maximum speedup is 20 times, no matter how many cores you add. That is why we always start by measuring: where is the serial part, and how do we get rid of it?

Gustafson shows the other side: if you also solve bigger problems with more computing power, you scale much further. Which law matters for you depends on your question.

Speedup versus number of cores, logarithmic axes
  • Amdahl
  • Gustafson
  • Ideal
0Speedup according to Amdahl
0Speedup according to Gustafson
0Maximum with infinitely many cores

From a single node to the whole cluster.

We work at every level of the stack: from the vector instructions in an inner loop to the workload manager that distributes thousands of jobs.

Distributed memory

Domain decomposition, non-blocking and one-sided communication, and topology-aware process placement.

  • MPI
  • UCX
  • Halo exchange
  • RMA

Shared memory

Threading, NUMA-aware data layout and vectorisation, so every core and every SIMD lane contributes.

  • OpenMP
  • oneTBB
  • AVX-512
  • NEON and SVE

GPU acceleration

Rewriting and tuning kernels, hiding memory transfers and scaling across multiple GPUs and nodes.

  • CUDA
  • HIP and ROCm
  • SYCL
  • NCCL

Asynchronous runtimes

Task-based parallelism, futures and event loops that overlap computation, communication and I/O.

  • HPX
  • Taskflow
  • asyncio
  • Ray
  • Dask

Performance engineering

Profiling, roofline analysis and measuring memory bandwidth and cache behaviour, before we change anything.

  • Nsight
  • VTune
  • perf
  • LIKWID

AI at scale

Distributed training with data, model and pipeline parallelism, and high-throughput inference.

  • PyTorch DDP and FSDP
  • DeepSpeed
  • Triton

Infrastructure

Clusters, cloud bursting and containers, set up so researchers and engineers can carry on by themselves.

  • Slurm
  • Kubernetes
  • Apptainer
  • Spack

Our research: parallel simulation of tumour growth.

In this chapter we analyse the parallel efficiency of a framework that simulates the growth of malignant pleural mesothelioma: a Cellular Potts Model coupled with PDEs for oxygen, nutrients and cytokines, in a three-dimensional domain built from CT data.

A dynamic bounding box shrinks the domain on which the PDEs are solved to the region around the tumour. The PDEs are solved with the finite volume method and an implicit Euler scheme; parallelisation uses mpi4py with PETSc and a GMRES solver. The result: shorter solve times than serial computation, more efficient memory use and better load balancing across the cores.

Title
Multiscale Parallel Simulation of Malignant Pleural Mesothelioma via Adaptive Domain Partitioning – An Efficiency Analysis Study
Authors
Anton Dolganov, Valeria Krzhizhanovskaya, Stefano Trebeschi, Vivek M. Sheraton
Published in
Computational Science – ICCS 2025 Workshops, Lecture Notes in Computer Science, vol. 15911, Springer, 2025, pp. 20–32
DOI
10.1007/978-3-031-97570-7_3
Adaptive computational domainSchematic
0Tumour cells in this view
0%Share of the cross-section being computed
Schematic illustration of the approach, not results from the study. Only the boxed region is computed; the dashed lines split it across four processes by cell count.

Hiding communication behind computation.

A small difference in code, a big difference at scale. By starting the exchange with neighbours before computing the interior of the domain, only the boundary still waits for data.

Blocking

computation waits for communication
for (int it = 0; it < iters; ++it) {
    exchange_halos(u);          /* MPI_Sendrecv */
    compute_interior(u, unew);
    compute_boundary(u, unew);
    swap(&u, &unew);
}

Overlapping

computing while data is in flight
for (int it = 0; it < iters; ++it) {
    MPI_Request req[8];
    post_halo_exchange(u, req);  /* MPI_Irecv, MPI_Isend */
    compute_interior(u, unew);   /* overlaps with transfer */
    MPI_Waitall(8, req, MPI_STATUSES_IGNORE);
    compute_boundary(u, unew);   /* only the boundary waits */
    swap(&u, &unew);
}

Measure first, then accelerate.

Every optimisation starts with a reproducible measurement and ends with a team that can hold on to the gains.

  1. 01

    Measure

    A reproducible baseline and profiling on your own hardware and datasets.

  2. 02

    Model

    Roofline analysis and a scalability model: where is the real bottleneck?

  3. 03

    Parallelise

    The right strategy for each bottleneck: vectorisation, threads, MPI, GPU or asynchronous tasks.

  4. 04

    Validate

    Numerical correctness and reproducibility, guarded by regression tests.

  5. 05

    Scale

    Strong and weak scaling measurements on the target environment, from workstation to cluster.

  6. 06

    Hand over

    Documentation, benchmarks in CI and training, so your team holds on to the gains.

Is your software computing too slowly?

Tell us what runs, on what, and how long it takes. We will show you where the gains are.

Book a call