As a scientist: The amount I struggled with getting distributed training on a HPC cluster to work vs. how easy it was with Lightning was eye opening. Almost no code change and finally I can run across 20 nodes with 4 V-100 each :). Plus the automatic SLURM checkpoints and restarts <3.
Comments
As a scientist: The amount I struggled with getting distributed training on a HPC cluster to work vs. how easy it was with Lightning was eye opening. Almost no code change and finally I can run across 20 nodes with 4 V-100 each :). Plus the automatic SLURM checkpoints and restarts <3.