Skip to content

Comment on We hid backdoors in ~40MB binaries and asked AI + Ghidra to find them

Comments

I'm not an expert but about false positives: why not make the agent attempt to use the backdoor and verify that it is actually a backdoor? Maybe give it access to tools and so on.

So many models refuse to do that due to alignment and safety concerns. So cross-model comparison doesn't make sense. We do, however, require proof (such as providing a location in binary) that is hard to game. So the model not only has to say there is a backdoor, but also point out the location.

Your approach, however, makes a lot of sense if you are ready to have your own custom or fine-tuned model.

Surprising that they still allow to catch the back doors but not use them.

A bad actor already has most of the work done.

Sounds like the pitch writes itself, "you'd better spend a lot of token money with us before the bad guys do it to you..."

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.