Skip to content

Comment on Show HN: Hacker News, without AI

Comments

Here is an approach to filtering LLM written content: https://hnslop.nilsherzig.com/

Instead of filtering keywords this ranks the content using Pangram. But it's also pretty effective at filtering content about AI. Turns out that a lot of LLM tooling projects have LLM written project descriptions/blog posts.

It's amazing to see how many of the "AI hurts my brain" posts are apparently written by an LLM.

nvm sorry, just read the "how this works" comment from OP. https://news.ycombinator.com/item?id=49660301

OPs solution does more than keyword filtering. But its still about filtering out "content about ai" not filtering out "ai written content".

Great point. I think the key is differentiating between "prompt generated" content vs "AI as a Topic" content and then being able to toggle them at will.

I am not averse to prompt-generated content, but it should be marked and quoted as such, just like any other reference. If it's not, it's like the author is double-cheating.

In fact there are no shortage of comments here on HN, especially open-ended questions that could've been plugged into a bot for an answer. Avoiding LLMs altogether is foolhardy. It's a matter of balance, and being able to write a good prompt is a skill - no mental atrophy there. Content derived from non-English sources using English-based LLMs is also an interesting area. [1]

It's possible to obtain AIaaT from https://aibriefs.news/ and then prune it elsewhere if need be. But the best aggregators like that may be machine-generated with a lot of human moderation. I wonder if that is 100% machine generated.

Likewise, HN-AI could well be a decent site.

[1] How much content can be derived from non-English sources using LLMs? https://share.gemini.google/ObjMQHCj8JWj"

Nice! I was considering doing this but the Pangram API is very expensive. Then I considered training my own model and I fortunately stopped at the edge of that rabbit hole.

Yea very expensive (hence me siphoning another page haha).

Im trying to get something cheaper to work, Pangram has some nice docs on how to build something like their service https://github.com/pangramlabs/EditLens, they even have training data online.

The harm a missclassification carries is much lower for a hn post than a masters thesis, so we might be fine with a worse model.

Yeah agree, Pangram puts out interesting material. I recently came across their v4 technical report and they shared a lot more than I would have expected them to.

But see, here I am back at the edge of the rabbit hole, and you're trying to pull me in. I refuse!

If you do end up training a model, send me an email and maybe I can tie it into the site.

Will do, thanks for hcker.news btw, has been my main mobile client for couple of months. Of all things i like the changelog most, its nice to have a quick view about new stuff, especially if the website changes somewhat faster than usual.

Thanks! I always forget to update the changelog for a bunch of stuff. I'm glad you have eyes on it, I'll be more mindful to add to it.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.