The spread isn't as large as I thought from the headline. It's closer to a toss-up than I would have expected.
The length of the response is a huge factor:
Students tended to prefer longer responses. The selected answer was 37% longer on average than the alternatives. The longest response won 47.7% of decisive writing comparisons. The shortest still won 25.0%.
So the score is partially a proxy for longest responses.
Makes me wonder how much the reviewers actually read the text. Were lazy evaluators picking the text that looked the longest or most structured without reading it all?
Yeah it's hard to get good data on what's best. We did read the a good sample amount of essays. Lots of arguments about what made them better, but its uncontroversial that longer ones were more preferred.
Interestingly we trained a simple LORA layer on top of Inkling and it turns out a model can smell which model wrote a response 70% of the time. I wonder if the smell of gemini is just more preferred by students.
Matches my experience. When I need to think or reason about something, between Clause and ChatGPT, the latter is my go to. With more to read, there’s just more to work with. For the same reason it’s bonkers for tech issues. How many times did I get a ChatGPT reply a mile long, do the first thing, and then read more after it didn’t work. Only then realize chatgpt didn’t offer the best answer first.
Way back when I was in high school, the most important part of essay writing was figuring out how to expand the topic I cared nothing about to the required X pages. I'm unsurprised that college students would default to the same thing.
Makes me wonder how much the reviewers actually read the text.
How many of the college students voting have years of ai use under their belt, where ai voice itself has shaped and influenced their stylistic choices of language preference?
Comments
The spread isn't as large as I thought from the headline. It's closer to a toss-up than I would have expected.
The length of the response is a huge factor:
So the score is partially a proxy for longest responses.
Makes me wonder how much the reviewers actually read the text. Were lazy evaluators picking the text that looked the longest or most structured without reading it all?
Yeah it's hard to get good data on what's best. We did read the a good sample amount of essays. Lots of arguments about what made them better, but its uncontroversial that longer ones were more preferred.
Interestingly we trained a simple LORA layer on top of Inkling and it turns out a model can smell which model wrote a response 70% of the time. I wonder if the smell of gemini is just more preferred by students.
Matches my experience. When I need to think or reason about something, between Clause and ChatGPT, the latter is my go to. With more to read, there’s just more to work with. For the same reason it’s bonkers for tech issues. How many times did I get a ChatGPT reply a mile long, do the first thing, and then read more after it didn’t work. Only then realize chatgpt didn’t offer the best answer first.
Way back when I was in high school, the most important part of essay writing was figuring out how to expand the topic I cared nothing about to the required X pages. I'm unsurprised that college students would default to the same thing.
I think this is one of the most awkward parts of essay writing that AI does help with .. which is filling in the words to fill the word count.
How many of the college students voting have years of ai use under their belt, where ai voice itself has shaped and influenced their stylistic choices of language preference?
Hmm, seems like which model teachers prefer would be the more important headline…
You know.. might try to see what we can do here