Comment on Llama 405B 506 tokens/second on an H200Comments−EgoIncarnate1ynot "an H200", "In the table above, tensor parallelism is compared to pipeline parallelism with each across eight GPUs"−FanaHOVA1yTitle on HN is wrong. The article says GPUs and it's referring to one of their 8xH200 boxes.
Comments
not "an H200", "In the table above, tensor parallelism is compared to pipeline parallelism with each across eight GPUs"
Title on HN is wrong. The article says GPUs and it's referring to one of their 8xH200 boxes.