Whole words only

Tuesday, February 18, 2020

New spider fully deployed

We have completed the transition to the new spider. It has been tested on every index. If a problem was found with your index, your technical contact would have received an email from Blossom Support advising a fix to the problem. In most cases, we were able to implement the fix and your technical contact was just asked to confirm the change.

The new spider has improved handling of dynamic websites and offers some new control over spidering. Among the changes:

  • The inclusion/exclusion lists are now more powerful. Read about the full capabilities in the Search Guide.
  • Of special note for includes lists is the new $ prefix. It tells the spider to only follow links given in an index file, for example a sitemap file. This is especially useful for richly interconnected sites like blogs.
  • The exclude list now allows multiple wildcards (the * character) and an end of URL mark (the $ character). The ? character is not special, following the syntax of robots.txt files.
  • Redirections (HTTP codes 301-308) now adhere to include/exclude specifications. Redirections are cached between spidering runs, speeding updates.
  • Canonical links are used when possible.
  • Cookies are always saved a resent. This improves the experience on session-oriented sites.
  • From the search configuration page (at https://blossom.com) you can control the speed of spidering.
  • Chunking of document content into logical units (e.g., sentences) has been improved. This is reflected in improved snippets shown search results.
If you see any problems due to the new spider, please let us know by emailing Blossom Support.

Tuesday, January 21, 2020

Major revision of Blossom spider now being deployed

If you look over the issues discussed in this blog, you'll see that many have arisen due to website content becoming more dynamic. Static web pages are becoming rarer, making the job of spidering more difficult. As a result, we have begun testing a significant rewrite of the Blossom spider.  In addition to handling dynamic pages better, the new spider will offer more flexibility in how sites are traversed. This post will be updated as testing progresses.

 If you monitor your web logs, you may notice extra activity from Blossom as we run the new spider alongside the old. You can pick out visits from Blossom by looking at the User_Agent HTTP header. For the production Blossom spider, the agent is Mozilla/5.0 (Blossom); for the new spider it is Mozilla/5.0 (Blossom/Beta).

We have begun rolling out the new spider to handle the regular update of indexes. In some instances, the new spider may require changes to the configuration of an index. (We will notify your technical contact via email if we make changes for you.) Here are some of the changes we've seen that can impact the contents of an index:
  • Stricter handling of redirection URLs. When a request is redirected, either by an HTTP header (e.g. 301 or 302 status code) or by an HTML meta-tag refresh, the redirection URL must satisfy the include/exclude specification for the index.
  • Stricter adherence to the HTTP status code and content type as reported by the webserver. Documents will only be added to the index if they are delivered with a status code of 200. HTML pages must either have a content type of text/html or begin with an identifying tag such as or .
  • Documents limited to 100MB by default. Likely this will only impact PDF files, and usually just PDFs with lots of images.
  • Reading of sitemap.xml and robots.txt are the default.
  • Scanning of URLs in javascript strings has been improved. 

Tuesday, October 29, 2019

New indexer live

A new spidering and indexing engine for Blossom Search is now live with significant changes to improve site coverage and search results. Over time, websites have become much more dynamic with much of the HTML generated at the time of delivery. Dynamic sites present two significant problems:
  1. Links to some content may only be generated by client-side programs.
  2. Multiple links may generate the same content.
These are not, of course, new concerns, but the increasing complexity of websites makes the spidering and indexing tasks more difficult. The new changes address both of these issues. As a result, you may see your search indexes grow, or perhaps shrink!

The index will grow if your site uses sitemaps. The Blossom spider will now routinely look for sitemap.xml in the root directory for a website. If the file exists, it will use the sitemap to help guide the traversal of a website. If you wish to prevent that behavior, log into the Search Configuration page for your index at Blossom.com, follow the "Spidering, Indexing, and Reporting" link and uncheck the box "Read Sitemaps".

The index will shrink if your site contains or generates very similar pages accessible by different URLs. Duplicate recognition has been improved to pick up pages with nearly identical content regardless of the tag structure. Duplicate pages always have the same title, so using unique titles for unique content will prevent the indexer from ever judging very similar pages to be duplicates.

Thursday, May 24, 2018

Use sitemaps to guide Blossom spider

As websites become more dynamic, some links may be generated programmatically rather than specified directly in HTML. While the Blossom spider does search Javascript for URLs embedded in strings, it does not execute Javascript. As a result, URLs generated by string operations can be overlooked.

The spider was recently enhanced to read sitemaps as a way to guide its traversal of a site. By specifying a sitemap in the "include" list for an index, the spider will visit each URL in the sitemap.
 A sitemap is an XML file that lists the URLs on a website. (See https://www.sitemaps.org/ for details.)

For Blossom, the list doesn't have to contain all the URLs on a site; it only needs to include those URLs generated dynamically. Other URLs can be picked up in the standard way by including the site's home page. For example, this include list can be used find all URLs on mysite.com:
https://www.mysite.com
!https://www.mysite.com/sitemap.xml
Notice two things. First that it's okay if there is overlap between the sitemap and other seed URLs in the include list; the spider will remove any duplicates. Also notice that the sitemap line starts with "!". This tells the spider to scan the file sitemap.xml for URLs, but not to include the text in sitemap.xml in the search index.

Thursday, December 28, 2017

Enhanced treatment of PDF files

We have upgraded the PDF text extraction engine to handle more character encodings. You should see better retention of punctuation and better sentence construction. (Identifying sentences in PDF files can be challenging because successive lines in a paragraph may not be adjacent in the PDF data.)

Depending on how PDF files are generated, they may not have a title. Titles are important to the search engine as the text is considered highly descriptive of the document. Also, the title is displayed in the search results presented to your visitors. In the new extractor, if there is no title we use a heuristic that chooses the first non-common line of text in the document as the title. Non-common text is text that doesn't appear frequently elsewhere on the website. Common text is usually boiler plate, such as the name of an entity.

For both PDF and HTML files, we recommend that each document have a descriptive title to help the search engine select the document when relevant and to help your visitors understand what the document contains.

Thursday, December 29, 2016

Secure Search

Blossom Search can now be integrated with secure websites that use the HTTPS protocol. To use secure search, put https instead of http and ssearch instead of search in the search URL of your search form.

For the results page to be labeled as secure by a browser, you will need to make sure that all hard-coded links in your search results Template File (or Head/Tail files) use HTTPS. Also, if you specify a BASE URL in your Template File, make sure it also uses HTTPS. You can find more information about search Template Files in the Search Guide.

The search results generated by the search engine contain links back to your website based on the search query. Those links match the documents discovered by the search spider during indexing. They are controlled by the Include List associated with the search index. Use the HTTPS protocol for all URLs in the Include List to make the links in the search results secure. (The Include List is set on the Search Configuration page at Blossom.com in the Search Engine Settings section.) If you do change your Include List to use HTTPS, you should also trigger a reindexing to flush documents retrieved via HTTP.

If you need help, send email to Blossom Support.

Tuesday, April 19, 2016

Using Best Bets to Guide Search Results

Each week you receive a report detailing the search queries of your visitors. Have you taken a look at the most popular queries historically? The queries listed here are good candidates to add to a Best Bets list.

Best Bets allow you to hand-pick pages on your website to appear at the top of the search results for specific queries. Here is how it works:
  1. You provide a few must-match terms.
  2. You also provide a URL and a short description.
  3. When a search includes all of the must-match terms, the URL and description are listed at the top of the search results.
To specify Best Bets, log in to Blossom.com and follow the link Best Bets in the Search Index Settings section:

  • The Title will be the heading above the Best Bets. The default is "Best Bets".
  • To add a Best Bet, select Add a New Item.
  • If a query contains all of the Terms specified, then the Best Bet will be shown. Terms may include the wild card characters "?", to match any one character, and "*" to match zero or more characters.