I agree that this is the case for GPT-3.5, at least subjectively. However, with GPT-4 Turbo, it seems that performance has improved. If they got it to be faster using quantization, then they must have also found a way to offset any resulting performance losses.
Comments
I agree that this is the case for GPT-3.5, at least subjectively. However, with GPT-4 Turbo, it seems that performance has improved. If they got it to be faster using quantization, then they must have also found a way to offset any resulting performance losses.