Skip to content

Comment on Indexing iCloud Photos with AI Using LLaVA and Pgvectorparent

Comments

I'm using CLIP here generically to refer to families/models generating captions by leveraging CLIP as the encoder - of which there are plenty on "The Hub".

Have you actually done the approach I think you're suggesting for anything more complex than "this is a yellow cat"? Not trying to be snarky, genuinely curious. I've done a few of these projects and this approach never comes close to meeting user expectations in the real world.

Do you an example of a query that should fail by using the CLIP embeddings directly but works with the method describe in the article?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.