Whole words only

Wednesday, January 27, 2021

Robot searches removed from search report

You may have noticed that the number of searches reported in your weekly search report has gone down, perhaps dramatically. We have recently changed the reporting system to ignore searches from known robots.

Robots have long dominated overall web traffic. As a result of their increased sophistication, we've seen an increase in the number of searches carried out by robots. These robot searches skew the search data, making it harder to see just what your (human) visitors are searching for. On your search report the data no longer includes searches from known robots (and you will see the number of robot search that were ignored).

Wednesday, October 28, 2020

New CSS classes to control search engine output

The appearance of search engine output can be controlled in a variety of ways. The simplest is to use parameters in the search URL, for example to control whether the size or date of a matched page is shown. By using a style sheet you can get control over the fonts, colors, and spacing of the text in the search output. Each component of the output is surrounded by CSS classes. Documentation for using CSS is in the Search Guide.

The next version of the search engine contains a few new classes to control the search output:

Blossom_DocBlock: A div that surrounds all of the output for each document matched by a query. Use this to control block-level behavior such as shading, hovering, and selecting.

Blossom_DocType: If a document is not HTML, then an indicator of the document type is added to the title. This class controls the style of the indicator.

Blossom_MoreButtons: Links for the "next" and "previous" pages of search results are displayed when the "more" option is used. This style controls the format of the link text.

Blossom_SearchForm: A div containing the search-again forms.

You can access the new version of the search engine by using "nquery" in place of "query" in your search URL. Use nquery just for testing, as it changes regularly as we test new search engine features. The changes will migrate to the production search engine, "query", in early November.

Update. These features went live on 11/9/2020. 

Tuesday, September 8, 2020

New domain names

Blossom Software has used the domain blossom.com since its founding in 1999. During that time we have often received offers to purchase the name. Recently, the company received an offer that could not be refused.

We have taken this opportunity to distribute functionality across a few domains. This improves security and makes the search engine more resistant to attack. Blossom now uses these domain names:

  • BlossomSoftware.net for the Blossom Software website. Go here to access your Search Configuration page.
  • searchBlossom.com for the Blossom search engine. Use this on your website to invoke the search engine. This domain works for both http and https access.
  • blossoft.com for email. For example, use support@blossoft.com to email Blossom Support.
The most important of these changes is the use of searchBlossom.com, as it impacts your search form. search.blossom.com and ssearch.blossom.com will continue to work for a while, but they will eventually be dropped.

Wednesday, August 5, 2020

First Notice! New URL to access search engine

We have unified access to the search engine for http and https access. Previously, if you used "http://" to call search, you used the server "search.blossom.com". For "https://" you used "ssearch.blossom.com". Using the two servers prevented mixed-content warnings from browsers.

The latest update of the search server can handle both protocols, so two servers are no longer required. Access the new server at searchBlossom.com. Please change your search engine access URLs as follows:

Old: http://search.blossom.com
New: http://searchblossom.com

Old: https://ssearch.blossom.com
New: https://searchblossom.com

This change also affects access to search engine support for popup results:

Old: https://ssearch.blossom.com/search_modal.js
New: https://searchBlossom.com/search_modal.js

Old: https://ssearch.blossom.com/search_modal.css
New: https://searchBlossom.com/search_modal.css

The old URLs will continue to work for several months, but they are scheduled to be retired.

Tuesday, March 31, 2020

Popup search results

Perhaps the trickiest part of integrating Blossom Search with a website is making the search results page blend in with the style of the site. Blossom has many options for formatting the results page, including head/tail files, template files, style sheets, and XML. Blossom is now testing its second-generation search-results overlay or popup window.

The primary benefit of putting the search results in a popup window is that its layout is independent of the page from which the search is performed and thus there is no need to insert the results into a page template. The popup is created in Javascript and styled in CSS. You can find all of the code needed on the Blossom website. Instructions for using the code are in the Search Guide section titled Using a Modal Popup.

Tuesday, February 18, 2020

New spider fully deployed

We have completed the transition to the new spider. It has been tested on every index. If a problem was found with your index, your technical contact would have received an email from Blossom Support advising a fix to the problem. In most cases, we were able to implement the fix and your technical contact was just asked to confirm the change.

The new spider has improved handling of dynamic websites and offers some new control over spidering. Among the changes:

  • The inclusion/exclusion lists are now more powerful. Read about the full capabilities in the Search Guide.
  • Of special note for includes lists is the new $ prefix. It tells the spider to only follow links given in an index file, for example a sitemap file. This is especially useful for richly interconnected sites like blogs.
  • The exclude list now allows multiple wildcards (the * character) and an end of URL mark (the $ character). The ? character is not special, following the syntax of robots.txt files.
  • Redirections (HTTP codes 301-308) now adhere to include/exclude specifications. Redirections are cached between spidering runs, speeding updates.
  • Canonical links are used when possible.
  • Cookies are always saved a resent. This improves the experience on session-oriented sites.
  • From the search configuration page (at https://blossom.com) you can control the speed of spidering.
  • Chunking of document content into logical units (e.g., sentences) has been improved. This is reflected in improved snippets shown search results.
If you see any problems due to the new spider, please let us know by emailing Blossom Support.

Tuesday, January 21, 2020

Major revision of Blossom spider now being deployed

If you look over the issues discussed in this blog, you'll see that many have arisen due to website content becoming more dynamic. Static web pages are becoming rarer, making the job of spidering more difficult. As a result, we have begun testing a significant rewrite of the Blossom spider.  In addition to handling dynamic pages better, the new spider will offer more flexibility in how sites are traversed. This post will be updated as testing progresses.

 If you monitor your web logs, you may notice extra activity from Blossom as we run the new spider alongside the old. You can pick out visits from Blossom by looking at the User_Agent HTTP header. For the production Blossom spider, the agent is Mozilla/5.0 (Blossom); for the new spider it is Mozilla/5.0 (Blossom/Beta).

We have begun rolling out the new spider to handle the regular update of indexes. In some instances, the new spider may require changes to the configuration of an index. (We will notify your technical contact via email if we make changes for you.) Here are some of the changes we've seen that can impact the contents of an index:
  • Stricter handling of redirection URLs. When a request is redirected, either by an HTTP header (e.g. 301 or 302 status code) or by an HTML meta-tag refresh, the redirection URL must satisfy the include/exclude specification for the index.
  • Stricter adherence to the HTTP status code and content type as reported by the webserver. Documents will only be added to the index if they are delivered with a status code of 200. HTML pages must either have a content type of text/html or begin with an identifying tag such as or .
  • Documents limited to 100MB by default. Likely this will only impact PDF files, and usually just PDFs with lots of images.
  • Reading of sitemap.xml and robots.txt are the default.
  • Scanning of URLs in javascript strings has been improved.