Friday, November 7, 2008
Best Bets now available with Affinity Search
The Best Bets feature lets you guarantee that specific pages will appear at the top of the search results when your visitors search for particular keywords. For example, when someone enters the term "emergency", you might want a page of emergency numbers to appear first in the search results. Until now, Best Bets were only available with Enterprise Search, but now it has been added to Affinity Search as well. For details on using Best Bets, please see the Search Guide section Search Options.
Thursday, October 30, 2008
Controlling URL length in search output
The URLs for pages matching a search query are usually shown in search output along with the page title and page snippets. Sometimes those URLs are very long, so, by default, the search engine removes characters from the middle of the URL to keep the total length less than 80 characters. A new option to the search engine lets you set that length to some other value. For details, and an example, see the Search Guide section on Output Format.
Friday, August 29, 2008
PDFs that just contain scanned images
Many websites use the PDF format to store documents that have been scanned. PDFs containing scanned documents consist of a series of bitmap images--they don't contain any text and so are not searchable. Up until now, to keep the Blossom spider from downloading these image files, you needed to put them explicitly into the exclude list for an index.
No longer! The indexing process now recognizes PDFs that just contain images, removes them from the index, and instructs the spider not to download the files again. As a result, the page count for a search index will no longer include image-only PDFs.
No longer! The indexing process now recognizes PDFs that just contain images, removes them from the index, and instructs the spider not to download the files again. As a result, the page count for a search index will no longer include image-only PDFs.
Wednesday, June 11, 2008
Searching blogs
Blogs create some challenges for search engines because they often create many paths to the same content and the "last modified" dates are often not reliable.
Multiple paths causes a spider to download the same content multiple times. Also, if the search engine isn't careful, search results might contain the same text multiple times. With Blossom search, the multiple-path problem can be solved with judicious use of include and exclude patterns.
The "last modified" date is used as part of search-engine ranking algorithms as well as to implement sorting by date. Blog systems generate their content dynamically, so the "last modified" date is often reported as today regardless of when the content was actually created. This, of course, makes "sort by date" useless. With Blossom search, the date problem can be solved using a custom indexing filter. See www.blossom.com/search_blog.html for more details.
If you are indexing a blog and want help overcoming these problems, let us know by sending email to support@blossom.com.
Multiple paths causes a spider to download the same content multiple times. Also, if the search engine isn't careful, search results might contain the same text multiple times. With Blossom search, the multiple-path problem can be solved with judicious use of include and exclude patterns.
The "last modified" date is used as part of search-engine ranking algorithms as well as to implement sorting by date. Blog systems generate their content dynamically, so the "last modified" date is often reported as today regardless of when the content was actually created. This, of course, makes "sort by date" useless. With Blossom search, the date problem can be solved using a custom indexing filter. See www.blossom.com/search_blog.html for more details.
If you are indexing a blog and want help overcoming these problems, let us know by sending email to support@blossom.com.
Friday, May 2, 2008
Best Bets for Enterprise Search
Normally, the search engine ranks results based on the quality and number of matches between the words a query and the words in each document. Using titles, descriptions, and meta-tag keywords you can influence the rank for a document with particular queries. Also, using explicit page weights you can push a document up or down for all queries. (For details, see the library paper on Page Ordering.) However, without using the Best Bets feature, you cannot guarantee that a page will be listed first.
With Best Bets you can specify explicitly what pages should be listed at the top of the search results for particular queries. Each Best Bet is described by a set of word patterns, a URL, and a page description. If all of the patterns match a query, then the URL and description are output at the top of the search results. For more details, see the Search Guide section on Best Bets.
Thursday, April 17, 2008
Explicit setting of document dates
The Blossom engine allows results to be sorted by relevance or by date. Usually the date for a document is taken from the file system, but sometimes that date is not accurate, such as when the document is generated dynamically or when the files have been copied. For HTML files, the document date can be set using the HTTP-EQUIV meta tag (see the Search Guide for details.) Recently we have added the same capability to PDF files--the technique is described in the Search Guide section mentioned above.
Monday, March 24, 2008
Dollar and time entities now recognized
The search engine has been upgraded to recognize two more entities, dollar amounts and times of day. Previously the engine recognized postal addresses, email addresses, zip codes, and phone numbers.
An entity stands for a class of items. When an entity is recognized, the search engine can match a query to any member of the class instead of just a keyword. For example, the query "license cost" matches documents that have the word "license" near a dollar amount (in addition to documents containing the word "license" near the word "cost". Similarly, a query "when does the store open" will match documents containing the words "store" and "open" near a time of day.
An entity stands for a class of items. When an entity is recognized, the search engine can match a query to any member of the class instead of just a keyword. For example, the query "license cost" matches documents that have the word "license" near a dollar amount (in addition to documents containing the word "license" near the word "cost". Similarly, a query "when does the store open" will match documents containing the words "store" and "open" near a time of day.
Subscribe to:
Posts (Atom)