Comment on Ask HN: DeepSeek V3's AI Code Review Performance – A Reality Check with DataComments−homarp1yEvaluation: Results were assessed by Claude 3.5 Sonnet V2 for consistencywhy not doing human assessment on top... to ensure the assessment by Claude is correct?conducted a detailed benchmarki suggest you post a sample for other to try to reproduce
Comments
why not doing human assessment on top... to ensure the assessment by Claude is correct?
i suggest you post a sample for other to try to reproduce