Skip to content

Comment on Models Don't Go Rogue

Comments

Not sure I should trust an article written by an LLM to make a solid judgment about what other models did or didn't do.

It doesn't read as AI generated text to me. Pangram also suggests it's human-written, for what it's worth. That's not to say that it's correct, just human-written. If the model used in the HF hack did indeed have all of the safeguards manually removed, that would change my perception of the situation, at least.

AI detectors do not work. There are passages in this that have some odd structures that don't feel human to me. I'm sure a human edited this and refined it with some prompting, it's not just rough output from an AI. But a lot of the text feels like it was edited via prompting rather than actual editing or writing.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.