Skip to content

Comment on What is full-text search (and why you need it now)

Comments

I wonder what they're using? I'd guess at Lucene / Solr.

No, we need to make that more clear on the site. It's our own search engine, written to overcome some of the limitations of Lucene. In particular, update of relevance formulas on the fly so you can experiment without reindexing.

What advantages does your search system have over better search systems than Lucene/Solr, such as Sphinx?

Sphinx does not store data, so you have to get it from your database to display results, it doesn't have query auto-suggest, etc.

But most importantly, you don't have to deal with installing, configuring, managing and scaling a sphinx setup. IndexTank is cloud-based, you don't have to deal with servers, RT indexes, hadoop, and so on. We worry about all that.

Thanks for the reply.

ahh.. i wondered when that would come along.

I had a thought that relevance should be computed like a "stream database", but with an index rather than a data store behind it. Maybe something like Streambase - but written in Clojure - on top of the lucene index.

Take it one step further, where the entire index is expressed as s-expressions (i'm a bit out of depth here) - you can basically write a mapreduce job on the index.

However, I'm not so sure that computing relevance can be parallelized - i.e can it be broken down as opposed to be computed as a global problem ?

That's quite a bold move. Lucene after all has the best part of a decade's worth of enhancements and optimisations.

what do you mean by "relevance formula"? You can implement your own Similarity measurement in lucene and change it at query time, so I'm guessing you mean something else?

You can define arbitrary mathematical functions based on document variables stored in memory, and that you can change in real time. You can have as many of these as you want.

http://indextank.com/documentation/function-definition

(Feedback on this documentation is more than welcome, it's a work in progress!)

please pardon me but I still confused: why can't you do this with the standard lucene? It kind of seems like what solr does with function queries[0]. Is the difference in what the variables are?

http://wiki.apache.org/solr/FunctionQuery

The main difference is that you can update the document variables without fragmenting the index. In Lucene you need to reindex the document (i.e. delete it and write a new one). If you update your relevance variables an average of 10 times per document it can get really fragmented with Lucene.

ah, got it. The point is the document fields are updated in place instead of delete+reinsert. Seems interesting.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.