Skip to content

Comment on Anthropic walks back policy that could have 'sabotaged' researchers using Claudeparent

Comments

I guess I don't understand why it's shady. It seems more like a poorly executed decision to enforce a publicly stated policy (it's been against Anthropic's ToS to use their models on frontier ML research for a while now). After all, people found out about this through their published system card.

It is definitely a bad idea to do this without notifying the user, because users who are incorrectly affected will have no way of providing feedback or getting support. And it is also anticompetitive, but if you truly believe that AI is not a normal technology, it is rational.

It's shady because they were going to silently poison your outputs.

It's actually worse than it sounds initially, because Fable isn't actually omniscient when it comes to safety classification. Many people (myself included) had refusals or fallback to Opus 4.8 for seemingly compliant/innocuous requests.

Wouldn't you be pissed off if they decided to sabotage your project despite having done nothing wrong?

The trouble is the silence, not Anthropic setting guardrails. Claude saying "I'm sorry, I can't assist further because it looks like you're [XYZ]" is fine.

We all know the false positive rates for classifiers on Fable. Imagine being a ML researcher working on any kind of ML/AI project that isn't against their ToS, and having your codebase poisoned and sabotaged silently.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.