I assumed that a specific flavour of LLM was needed, an “embedding model” to generate the vectors. Is this announcement that pinecone is adding their own?
Normally you take your content and run it through an embedding model, inserting the resulting vectors into the vector DB. On a query, for instance, you run the query through the embedding model and query the vector database for the most similar hits to the resulting embedding vector. Similarly reranking is when get get the broad hits from the embedding similarity search and/or BM25, and then a reranker uses the looked up source material to rank the results more finely.
This is building it into the vector DB such that you send it the content and it is "built in".
Seems silly. It's like bundling a stove with cookware. But cookware fit specific niches and have different life cycles. I get that it might cater to some "drop in solution" targets, but seems of no value for most engineered, long-term solutions.
Is this announcement that pinecone is adding their own?
TLDR: they trained their own embeddings model and rely on Cohere for ranking. Pinecone (the database) uses this model automatically to generate and store embeddings.
I assumed that a specific flavour of LLM was needed, an “embedding model” to generate the vectors.
You're mostly right, with one caveat: embeddings models aren't really LLMs in that they're not very large: they just map semantic meaning to numerical space.
This is the golden question. As far as I know, there is no appropriate benchmarking/eval data about this. I think the real value is the first-class integration between their model and their service.
I think the general rule is "the smarter the model the better the embedding" but I can't cite a paper right now. So in theory GPT 4 would give better embeddings (if extracted from the middle layers) but that would be overkill.
Comments
Can someone please explain how this works?
I assumed that a specific flavour of LLM was needed, an “embedding model” to generate the vectors. Is this announcement that pinecone is adding their own?
Is it better or worse than the models here: https://ollama.com/search?c=embedding For example?
Normally you take your content and run it through an embedding model, inserting the resulting vectors into the vector DB. On a query, for instance, you run the query through the embedding model and query the vector database for the most similar hits to the resulting embedding vector. Similarly reranking is when get get the broad hits from the embedding similarity search and/or BM25, and then a reranker uses the looked up source material to rank the results more finely.
This is building it into the vector DB such that you send it the content and it is "built in".
Seems silly. It's like bundling a stove with cookware. But cookware fit specific niches and have different life cycles. I get that it might cater to some "drop in solution" targets, but seems of no value for most engineered, long-term solutions.
There's more technical detail here: https://www.pinecone.io/blog/integrated-inference/
TLDR: they trained their own embeddings model and rely on Cohere for ranking. Pinecone (the database) uses this model automatically to generate and store embeddings.
You're mostly right, with one caveat: embeddings models aren't really LLMs in that they're not very large: they just map semantic meaning to numerical space.
This is the golden question. As far as I know, there is no appropriate benchmarking/eval data about this. I think the real value is the first-class integration between their model and their service.
I think the general rule is "the smarter the model the better the embedding" but I can't cite a paper right now. So in theory GPT 4 would give better embeddings (if extracted from the middle layers) but that would be overkill.