Whole words only

Wednesday, August 5, 2020

First Notice! New URL to access search engine

We have unified access to the search engine for http and https access. Previously, if you used "http://" to call search, you used the server "search.blossom.com". For "https://" you used "ssearch.blossom.com". Using the two servers prevented mixed-content warnings from browsers.

The latest update of the search server can handle both protocols, so two servers are no longer required. Access the new server at searchBlossom.com. Please change your search engine access URLs as follows:

Old: http://search.blossom.com
New: http://searchblossom.com

Old: https://ssearch.blossom.com
New: https://searchblossom.com

This change also affects access to search engine support for popup results:

Old: https://ssearch.blossom.com/search_modal.js
New: https://searchBlossom.com/search_modal.js

Old: https://ssearch.blossom.com/search_modal.css
New: https://searchBlossom.com/search_modal.css

The old URLs will continue to work for several months, but they are scheduled to be retired.

Tuesday, March 31, 2020

Popup search results

Perhaps the trickiest part of integrating Blossom Search with a website is making the search results page blend in with the style of the site. Blossom has many options for formatting the results page, including head/tail files, template files, style sheets, and XML. Blossom is now testing its second-generation search-results overlay or popup window.

The primary benefit of putting the search results in a popup window is that its layout is independent of the page from which the search is performed and thus there is no need to insert the results into a page template. The popup is created in Javascript and styled in CSS. You can find all of the code needed on the Blossom website. Instructions for using the code are in the Search Guide section titled Using a Modal Popup.

Tuesday, February 18, 2020

New spider fully deployed

We have completed the transition to the new spider. It has been tested on every index. If a problem was found with your index, your technical contact would have received an email from Blossom Support advising a fix to the problem. In most cases, we were able to implement the fix and your technical contact was just asked to confirm the change.

The new spider has improved handling of dynamic websites and offers some new control over spidering. Among the changes:

  • The inclusion/exclusion lists are now more powerful. Read about the full capabilities in the Search Guide.
  • Of special note for includes lists is the new $ prefix. It tells the spider to only follow links given in an index file, for example a sitemap file. This is especially useful for richly interconnected sites like blogs.
  • The exclude list now allows multiple wildcards (the * character) and an end of URL mark (the $ character). The ? character is not special, following the syntax of robots.txt files.
  • Redirections (HTTP codes 301-308) now adhere to include/exclude specifications. Redirections are cached between spidering runs, speeding updates.
  • Canonical links are used when possible.
  • Cookies are always saved a resent. This improves the experience on session-oriented sites.
  • From the search configuration page (at https://blossom.com) you can control the speed of spidering.
  • Chunking of document content into logical units (e.g., sentences) has been improved. This is reflected in improved snippets shown search results.
If you see any problems due to the new spider, please let us know by emailing Blossom Support.

Tuesday, January 21, 2020

Major revision of Blossom spider now being deployed

If you look over the issues discussed in this blog, you'll see that many have arisen due to website content becoming more dynamic. Static web pages are becoming rarer, making the job of spidering more difficult. As a result, we have begun testing a significant rewrite of the Blossom spider.  In addition to handling dynamic pages better, the new spider will offer more flexibility in how sites are traversed. This post will be updated as testing progresses.

 If you monitor your web logs, you may notice extra activity from Blossom as we run the new spider alongside the old. You can pick out visits from Blossom by looking at the User_Agent HTTP header. For the production Blossom spider, the agent is Mozilla/5.0 (Blossom); for the new spider it is Mozilla/5.0 (Blossom/Beta).

We have begun rolling out the new spider to handle the regular update of indexes. In some instances, the new spider may require changes to the configuration of an index. (We will notify your technical contact via email if we make changes for you.) Here are some of the changes we've seen that can impact the contents of an index:
  • Stricter handling of redirection URLs. When a request is redirected, either by an HTTP header (e.g. 301 or 302 status code) or by an HTML meta-tag refresh, the redirection URL must satisfy the include/exclude specification for the index.
  • Stricter adherence to the HTTP status code and content type as reported by the webserver. Documents will only be added to the index if they are delivered with a status code of 200. HTML pages must either have a content type of text/html or begin with an identifying tag such as or .
  • Documents limited to 100MB by default. Likely this will only impact PDF files, and usually just PDFs with lots of images.
  • Reading of sitemap.xml and robots.txt are the default.
  • Scanning of URLs in javascript strings has been improved. 

Tuesday, October 29, 2019

New indexer live

A new spidering and indexing engine for Blossom Search is now live with significant changes to improve site coverage and search results. Over time, websites have become much more dynamic with much of the HTML generated at the time of delivery. Dynamic sites present two significant problems:
  1. Links to some content may only be generated by client-side programs.
  2. Multiple links may generate the same content.
These are not, of course, new concerns, but the increasing complexity of websites makes the spidering and indexing tasks more difficult. The new changes address both of these issues. As a result, you may see your search indexes grow, or perhaps shrink!

The index will grow if your site uses sitemaps. The Blossom spider will now routinely look for sitemap.xml in the root directory for a website. If the file exists, it will use the sitemap to help guide the traversal of a website. If you wish to prevent that behavior, log into the Search Configuration page for your index at Blossom.com, follow the "Spidering, Indexing, and Reporting" link and uncheck the box "Read Sitemaps".

The index will shrink if your site contains or generates very similar pages accessible by different URLs. Duplicate recognition has been improved to pick up pages with nearly identical content regardless of the tag structure. Duplicate pages always have the same title, so using unique titles for unique content will prevent the indexer from ever judging very similar pages to be duplicates.

Thursday, May 24, 2018

Use sitemaps to guide Blossom spider

As websites become more dynamic, some links may be generated programmatically rather than specified directly in HTML. While the Blossom spider does search Javascript for URLs embedded in strings, it does not execute Javascript. As a result, URLs generated by string operations can be overlooked.

The spider was recently enhanced to read sitemaps as a way to guide its traversal of a site. By specifying a sitemap in the "include" list for an index, the spider will visit each URL in the sitemap.
 A sitemap is an XML file that lists the URLs on a website. (See https://www.sitemaps.org/ for details.)

For Blossom, the list doesn't have to contain all the URLs on a site; it only needs to include those URLs generated dynamically. Other URLs can be picked up in the standard way by including the site's home page. For example, this include list can be used find all URLs on mysite.com:
https://www.mysite.com
!https://www.mysite.com/sitemap.xml
Notice two things. First that it's okay if there is overlap between the sitemap and other seed URLs in the include list; the spider will remove any duplicates. Also notice that the sitemap line starts with "!". This tells the spider to scan the file sitemap.xml for URLs, but not to include the text in sitemap.xml in the search index.

Thursday, December 28, 2017

Enhanced treatment of PDF files

We have upgraded the PDF text extraction engine to handle more character encodings. You should see better retention of punctuation and better sentence construction. (Identifying sentences in PDF files can be challenging because successive lines in a paragraph may not be adjacent in the PDF data.)

Depending on how PDF files are generated, they may not have a title. Titles are important to the search engine as the text is considered highly descriptive of the document. Also, the title is displayed in the search results presented to your visitors. In the new extractor, if there is no title we use a heuristic that chooses the first non-common line of text in the document as the title. Non-common text is text that doesn't appear frequently elsewhere on the website. Common text is usually boiler plate, such as the name of an entity.

For both PDF and HTML files, we recommend that each document have a descriptive title to help the search engine select the document when relevant and to help your visitors understand what the document contains.