Skip to content

Comment on Show HN: I made a library for LLM prompt injection/exploit/jailbreak detectionparent

Comments

Yeah, I don't like this at all. If I'm going to evaluate a prompt injection protection strategy I need to be able to see how it works.

Otherwise I'm left evaluating it through wasting my time playing whac-a-mole with it, which won't give me the confidence I need because I can't be sure an attacker won't guess a strategy that I didn't think of myself.

This doesn't even include details of the evals they are using! It's impossible to evaluate whether what they've built is effective or not.

I'm also not keen on running a compiled .so file released by a group with no information on even who the authors are.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.