Comment on LLMLingua: Compressing Prompts for Faster InferencingparentComments−sroussey2yWhat would happen if instead of the long prompt, you just sent the mean of the embeddings of the prompt tokens?
Comments
What would happen if instead of the long prompt, you just sent the mean of the embeddings of the prompt tokens?