I agree with some of your points. Since the author is a non-native English speaker, there might be some grammatical issues in their English expressions. However, this is also constrained by the dataset; it's challenging to obtain sentences that are completely identical in both grammar and semantics. The author's main concern seems to be that when there are subtle semantic differences in inputs, the model shouldn't catastrophically fail. We can see examples like "Going ahead in an even pace," where previous models might even interpret it as moving backward. Or "A human utilizes his right arm to help himself to stand up," where the action of standing up might not even be present in other examples, posing serious problems. However, the author employs a similar approach to adversarial learning, enabling the model to learn expressions of actions that are similar to the original semantic sentences, which is already a significant improvement. We lack real motion data to learn expressions like "Going ahead in an even pace." The author also points out that there's a trade-off between stability and accuracy.
Comments
I agree with some of your points. Since the author is a non-native English speaker, there might be some grammatical issues in their English expressions. However, this is also constrained by the dataset; it's challenging to obtain sentences that are completely identical in both grammar and semantics. The author's main concern seems to be that when there are subtle semantic differences in inputs, the model shouldn't catastrophically fail. We can see examples like "Going ahead in an even pace," where previous models might even interpret it as moving backward. Or "A human utilizes his right arm to help himself to stand up," where the action of standing up might not even be present in other examples, posing serious problems. However, the author employs a similar approach to adversarial learning, enabling the model to learn expressions of actions that are similar to the original semantic sentences, which is already a significant improvement. We lack real motion data to learn expressions like "Going ahead in an even pace." The author also points out that there's a trade-off between stability and accuracy.