Tag: search

  • WinFS as a GUI to Search

    Microsoft blogger Dare Obasanjo digs at Google’s Desktop Search and points to data visualization as the key to the success of Microsoft’s delayed WinFS. While Google and other “command line” search tools are useful for finding that one file that lays somewhere on the internet or on your hard drive, there is clearly a need for an easier way to browse through results visually. If it’s a picture, display the thumbnail – I don’t care if it’s a jpeg, gif, or Photoshop file.

    The promise of WinFS is that it aims to turn every application [including file navigation applications like Windows explorer] into the equivalent of Outlook and iTunes when it comes to data visualization and navigation by baking such functionality into the file system APIs and data model. Trying to reduce that to “full text search plus indexing” is missing the forest for the trees. Sure that may get you part of the way but in the end it’s like driving a car with your feet. There is a better way and it is much closer than most people think.

  • Google Desktop Search

    In a move that took everyone by surprise, Google announced a new downloadable product that installs on your hard drive, indexes your email, Word, Excel, Powerpoint, and AIM chat logs and adds them to the Google Search results window. The expected move was that Google would launch their own, Google-centric browser but they have once again side-stepped popular wisdom and done something that new and fantastic.

    You’ll do a double-take the first time you run a search after installing Google Desktop Search. Up on top of your results, right under the paid search ads, you see links to personal email and files that contain hits on your query. Instead of bringing the web to your desktop, by putting hits on your desktop files into the Google UI it now looks (and feels) like Google has put your desktop onto the web.

    Rael Dornfest explains what’s going on behind the scenes:

    What’s actually going on is that the local Google Desktop server is intercepting any Google web searches, passing them on to Google.com in your stead, and running the same search against your computer’s local index. It’s then intercepting the Web search results as they come back from Google, pasting in local finds, and presenting it to you in your browser as a cohesive whole.

    John Battelle caught up with Marissa Mayer, Google’s director of consumer web products, and found out that the app is only 400k and runs on only 8 MB of RAM. She also says that the relevance algorithm obviously doesn’t use PageRank but does use 150 other proprietary variables (bolding, font size, etc) to determine relevance.

    Danny Sullivan writes in depth about this new tool going on to say that the Google page that you see when you launch Desktop Search is not actually on the web but is being served up by the web server that comes with the app. This is apparent when you see the address of the URL [http://127.0.0.1:4664/&s=400994545] which is a local address.

    Another benefit is the caching so that you can now quickly peek into the contents of a file without having to wait for Excel to fire up. If there are multiple copies in cache, there’s version history which can save you if you’ve overwritten a file using the same name.

    It’s still in beta so I’ll forgive the fact that it only runs on Windows and indexes only AIM chat and Internet Explorer caches but other than that, this is a most impressive product that redefines its category.

  • Snap.com

    At Web 2.0 this week and saw Idealab’s Bill Gross announce the release of snap.com – a new search engine designed to address comments that Google’s interface is looks like a command line.

    Based on the premise that search engines need to give control back to the user. You can change the focus of your results (with an integration of X1) in real-time and sort results on a number of parameters which change based upon your search query. Snap also believes at its core that it’s important to be transparent so it provides an interface to an area where you can see not only stats related to the top keywords and their related terms but also an area where you can view Snap’s revenues as well as performance reports on click throughs on keywords that you can specify.

    A number of examples showed how this is a different kind of fish.

    Search on “jaguar” then refine my typing in “os”

    Search on “camera” and you get a totally different set of results based on the assumption that you’re looking to purchase a camera. In the results, new widgets to sort on things like price, resolution. Refine the search by typing in “cannon” and see the list narrow itself to just the cameras made by Cannon.

    Search on “walmart” and you’ll see that the result set changes again. Bill explained that this was because they analyzed ISP traffic to determine popular clickstreams and saw that most people searching on “walmart” were either looking for the website, the company (news & quotes), the job page, or directions to the nearest store. The left hand column has these links ranked for easy access and other elements on this page provide shortcuts as well.

    Search on “cars” and you get yet another view which is like a dashboard to the most popular links yet in a different format. This time the screen is divided into four main categories into which most behavior following the search results fell. Here you see further drill down on to Buying, Research, Loans, and Insurance. The same concept applies for that other big ticket item, “real estate” but this time with yet another set of drill down choices.

    These innovations are all very cool and welcome but as I write about it, I realize that variations of each of these exist in some shape or form on Google and Yahoo. Snap’s presentation is so much more graphically rich than what is out there today giving a greater opportunity for advertisers to integrate their messages with the results in a much more meaningful way.

  • From Vivisimo comes, Clutsy the Search Engine

    From Vivisimo comes, Clutsy the search engine which has elicited guffaws over it’s name, coined to evoke a clustered set of search results but instead brings to mind a buck-toothed, all-thumbed version of the more refined Mr. Jeeves. The jury is still out for me on this as I’m so well-trained that I usually am able to throw enough contextual keywords into a query to get to what I’m looking for but Clutsy’s Clustering Engine (sound’s like a book I might read to my 5 year old) does shine when put to the “polish” test. Are you talking about the people, the sausage, the language, or something shiny? Clutsy will parse it out for you and give you several paths to try.

  • More ways to get there. A9 from Amazon

    Amazon’s A9 came out of beta with much fanfare and continues to get rave reviews because of a new feature which keeps track of your search history. This can be quite useful for those that are trying to retrace their steps to get at a vital piece of information. Bookmarks are single points of reference like an address. But the human brain doesn’t always work this way – sometimes it’s easier to remember how you got somewhere. A9’s Search History takes this geographical metaphor to the virtual world of clickstreams.

    I notice that Ask Jeeves has a similar feature on it’s My Jeeves which is currently in Beta.

    The other feature that I enjoy using is the image search which throws up images related to your search query right alongside and in context with the web hits. This is especially useful when doing searches on individuals. For those of you into vanity searches, it’s always a kick to scroll through all the people out there that share your name.

  • Big Blue Masala

    IBM will release a new corporate search engine, the “DB2 Information Integrator” (code-named Masala) tomorrow reports CNet and eWeek.

    The information integrator is able to do this because it can search rapidly across multiple databases, including relational and non-relational databases and structured and unstructured data such as text files, word documents, Adobe Acrobat files, video or audio files, according to Jones.

    “To gather this information up today, they might have to use multiple searches,” Jones said. DB2 Information Integrator can replace all of these searches with a single search that gathers all of the types of information to answer a single question, he said.

    Sounded like a pretty tall order to me. I’ve heard of connectors that can search the closed-caption text of a video but audio has no such meta-data. A scan of the IBM website for Masala doesn’t help either. I then happened upon this IBM Research site that shows early attempts to automatically categorize images on MPEG-7 video files. Once categorized, you can then query the meta-data attached to each image.

    Some further work is needed to iron out the kinks. In the example, looking closer you can see that both Janis Joplin and Peter Jennings were tagged as “animals”

  • Monitoring Blogs

    The San Jose Mercury News did a piece this week on companies turning to new tools to track consumer opinions on blogs. More and more people are beginning to realize that the right blogs, if monitored correctly, can serve as an early warning mechanism for the PR flacks everywhere. With their finger on the pulse of “the next big story,” the more popular blogs can amplify little known facts and points of view to the point where they can get picked up by the popular media and broadcast to the world at large.

    So how does a company keep track of the sentiment of what’s being said in the blogsphere about their product and brand? One of the more interesting tools highlighted in the article is Blabble. Founded by Rochester, NY based web designer, Matt Rice. The concept is called “thought parsing” using natural language processing to aggregate opinions expressed about a set of user-defined keywords to get at overall sentiment.

    Existing software products aggregate listings from blogs, but require the user seeking a view of overall trends or opinions as represented in blogs to read through all the blog listings to make that determination manually.

    Rice says Blabble goes a step farther by incorporating natural language processing that parses blog listings returned in a search into parts of speech so as to extract from them words, phrases and constructions that indicate opinion. “50,000 people may write about a topic, but you don’t have time to read 50,000 listings,” says Rice. “And I probably don’t care about one individual opinion; it’s the aggregate that I care about.”

    Internet Retailer.com

    UPDATE : as of January 2006, the Blabble service will no longer parse the blogosphere. According to the site, “we don’t know what we’re going to do with the technology.”

  • LexisNexis Total Search

    Lexis, the legal research division of Reed Elsevier, announced enhancements to a product called Total Search. Included in the enhancements is the ability to hook into a firm’s document management system and bring back an aggregated set of search results.

    LexisNexis Total Search also identifies, correlates, and links case citations appearing within internal work product and LexisNexis search results. These citations are noted, and access to the internal work product is provided, through a “correlation” icon appearing next to a particular case or code citation within an internal document result, LexisNexis full-text document or in a LexisNexis cite list.

    Not clear from the literature how easy it is to integrate Total Search as either an outbound feed to downstream search engines or as a platform to receive and integrate inputs from other systems such as email and newsfeeds.

  • The Independent on Search

    Charles Arthur, who writes for the UK paper, The Independent puts the Search conundrum in plain English,

    Yet it’s strange that it’s a lot easier to find something on the web than on my desk; and easier to find something on the web than on my computer. You would think one would store the things that are most important in the closest, most accessible locations; after all, you don’t leave your wallet and car keys at the bottom of an unlit stairwell locked in a safe while keeping your entire wardrobe within arm’s length.

    Instead, Google, halfway around the world, is the default for a lot of the horsework, looking up phone numbers, checking facts, finding references…. But why do we let computers make our lives hard for us? Partly because we”ve let them remain stupid. The graphical user interface, with its metaphors of “files and folder and desktop”, has remained unchanged since 1983 when Apple introduced it with the Lisa. . .

    I believe it was Bill Gates who once admonished his developers that it was ridiculous that it took longer to find a file on a local hard disk through Windows than it did to find a document among billions on the open web via Google. This became the rallying cause behind including a SQL based file system for Longhorn.