Posts tonen met het label Solr. Alle posts tonen
Posts tonen met het label Solr. Alle posts tonen

woensdag 6 april 2011

Lucid Imagination zet een nieuwe standaard met LucidWorks 1.7

Lucid Imagination, het eerste bedrijf met een commerciële support formule voor het leveren van het open source search product Solr, heeft een nieuwe versie uitgebracht van haar Enterprise class search oplossing: LucidWorks for Enterprises.

Deze enterprise search oplossing begint een concurrent te worden van de commerciële oplossingen. Wat Solr met name mistte was een goede beheerschil én connectoren om met enterprise sources, zoals Sharepoint, te communiceren.

"The reign of costly proprietary search vendors is coming to an end. The initial release of LucidWorks Enterprise was downloaded by nearly 1,000 companies in the first few months, and this combination of an open source base, unprecedented flexibility in the underlying architecture, and enterprise-ready features from Lucid Imagination helps companies with huge appetites for data slash their cost of development and speed innovation, while delivering the best possible user experience." -- Eric Gries, CEO of Lucid Imagination

Lucid Imagination voegt steeds meer functies toe die al jaren in commerciële oplossingen aanwezig zijn.
De Enterprise versie is zeker niet gratis. Voor een Support overeenkomst betaal je $64.000,00 per jaar voor gebruik in een produktie-omgeving.
Deze prijs kan je vergelijken met de prijs van een Google Search Appliance (2 jaar voor 2 tot 3 mln. documenten) of Exalead Cloudview. Die producten kennen echter veel betere oplossingen voor connectoren.

Grote voordeel is wel dat je geen extra licenties hoeft aan te schaffen als het aantal te indexeren documenten groeit.

Bron: http://www.sys-con.com/node/1782391

woensdag 21 april 2010

Apache Lucene gives birth on triples... and...

The Lucid Imagination Blog has today posted a blog about the exciting new thing in Lucene, namely "Triplets". The only thing that I mis in this blogpost is any information about "Apache Lucene gives birth to triplets!".

The blog is nothing more than a requiem to the Apache / Lucene project. How fantastic is the ongoing development of Lucene and Lucene and of course... Solr...

Come on LucidImagineers... you can do more than that... give us an insight in the meaning of this new exciting feature.

Source: http://www.lucidimagination.com/blog/2010/04/21/news-flash-apache-lucene-gives-birth-to-triplets/

maandag 11 januari 2010

LucidImagination releases LucidWorks for Solr 1.4

Finally a couple of months after the release of the Solr 1.4 distribution by Apache, LucidImagination – a company that deliveres commercial level support for the open source Solr engine – releases their certified version of the Solr 1.4 Enterprise Search Server.

This distribution contains a comprehensive manual with lots of information on how to setup and use the search server.

The documentation is outstanding in quality and completeness. But that’s not all.

LucidImagination has succeeded in wrapping the Solr software in a user-friendly installer that makes it possible to get the search environment up and running in no time.

The thing that is missing at this moment is a “ready to run” indexing solution for a filesystem with binary documents like PDF’s and Office-type documents.
When they succeed in packaging that type of solution, combined with a ready to run HTTP crawling environment, LucidImagination has the potential of competing with the “big” and expensive enterprise search providers like Autonomy, Microsoft / FAST and Exalead.

 

Open source search is starting to pick up on the big pie of enterprise search solutions, but lacks the “click and run” possibility that other software vendors offer.

The complexity of getting a Solr solution up and running with real life data sources is what’s holding back the large adoption. It is still too much a toolbox.

woensdag 21 oktober 2009

Solr 1.4 Enterprise Search Server Book review

Today I came across a blog posting about the Book "Solr 1.4 Enterprise Search Server ". As I mentioned earlier I am in the process of reading, dicing and slicing the book.
I will post updates on my journey through the examples and findings, but I thought it would be nice to see a short reviews of the book in de mean time:

http://happygiraffe.net/blog/2009/10/20/book-review-solr-1-4-enterprise-search-server/

This review says more about the structure and topics in the book as were I will give more information on experiences while using the book.

dinsdag 13 oktober 2009

Experiences with Solr 1.4 Enterprise Search Server (Part 0)

This weekend I started exploring the Book "Solr 1.4 Enterprise Search Server".

Although Solr 1.4 is not actually released at the moment of writing, the latest nightly build does provide nearly all functionality and is very stable.

In some of the coming posts on this blog I will share my experiences with the reading of the book and working through the examples. I will also make comparisons with some of the other search engines that I use in my work, like Autonomy, Exalead and the Google Search Appliance.

The EBook is accompanied by example code, download-able from the PACKT website. That example code contains the version that was used while writing the book so every example in the book should work.
The example code contains a fully filled and configured Solr / Lucene instance. This instance consists of some hundred thousands of records pulled from the Musicbrainz database. To use this data you must have a development / test environment with enough diskspace and RAM.

One negative remark about this dataset is that is it mostly "structured database data": a lot of fields with small amounts of data.

Enterprise Search environments that I stumble upon mostly hold lots of unstructured information from documents from filesystems, DMS and CMS systems.
Of course database offloading is a hot topic in BI / enterprise search land, but most information that has to be searched and found come from unstructured documents.

Maybe the fact that database data was chosen says something about the field of operation of Solr / Lucene in the real world.

It would be nice if I could use a more representative data set I could work with. This would make the examples more usefull.

dinsdag 15 september 2009

Reasons for choosing Solr on all for good twisted

Today I read a blogpost on the Lucid Imagination blog about the fact that a Google employee chose Solr as the search engine for the site.
The blogpost cited a part of the testimonial on the allforgood site on with I have to disagree with the reason for choosing Solr

The problem they had was:

One of the top concerns we’ve been hearing from nonprofit organizations who list
volunteer opportunities on All for Good is that their opportunities aren’t
updated on the site as frequently as they need. This happens because All for
Good doesn’t directly receive volunteer opportunities from nonprofits – we crawl
feeds from partners like VolunteerMatch and Idealist just like Google web search
crawls web pages. Crawlers don’t immediately update, they take time to find new
information.

The solution stated:

Today, we’re rolling out improvements to All for Good that will help solve this
problem and improve search quality for users. The biggest change, which you
won’t see directly, is that our search engine is now powered by SOLR, an
incredible open source project that will allow us to provide higher quality and
more up-to-date opportunities. Nonprofits should start seeing their
opportunities indexed faster, and users should see more relevant and complete
results.


Now... why do I disagree with the way the choice for Solr is argumented?

It is the fact that the use of Solr solves their problem of "latency". Remember the biggest problem was that the indexed information was not up to date.
Solr doesn't solve that. Solr is just a service around Lucene. Solr doesn't take care of the crawling part of the problem.

Us experts on search applications and information access solutions know that it is the combination of crawling frequency, the accessibily of the source that has to be indexed (RSS, Web, document repositories, databases etc.) the preprocessing of those diverse formats that determine the speed of the indexing and thereby, search process.

In this case probably Nutch will take care of the crawling part, so the frequency with which updates are processed rely on the speed of that part of the solution. Not the fact that Solr is used...

donderdag 3 september 2009

Solr next new thing for Europian Parliaments

According to some representatives of European Parliaments (UK, Europe, Belgium) , Solr is the next new thing in the field of search solutions for their websites and intranet search.

Today I was at a workshop focussed on the topic of search. The workshop was organized by the dutch Parliament and there were representatives of 8 parliaments of other European countries (Belgium, UK, Denmark, Norway, Sweden, Finland, Austria, Israel).

What stroke me there was the real interest in Solr as a solution for there search needs. I know that Solr is really upcoming but still has a "techy" ring about it... so I thought.

It seems that Solr is on its way to becoming a standard in government land. Faster than I had thought it would be.

Driven by the need for Open Source software and lower license costs, many organisations are willing to try out Solr. They take the absence of support from the vendor (there is no) for granted and are investing in own knowledge or that of implementation partners.

One reason for this "gamble" was very sensible in my opinion:
The field of search is really evolving. Not only from a technical point of view but also from a vendor and marketing angle.
We choose Solr because it is cheap, it is still being developed very fast, and our need for functionality is simple now. At the time we want more functionality, the product probably will have it. It therefore grows with our needs...
When after 3 years it seems that we have made the wrong decision we still haven't thrown away much money. We can than always switch.
I must agree with the choice. Why pay EUR 100.000,00 for a product that promises everything while you are using part of it.
Of course Solr is free but the implementation will cost you money. But even if you spend 50.000,00 on customising Solr it it still half as cheap.

Seems like the way to go for me...