Receive the weekly sampler of posts and "Resource of the Week".
Subscribe »

Enter your
email address:

My Account »


Bookmark and Share

Testimonial?
If you find ResourceShelf useful, please supply a testimonial »








Home > ResourceBlog > Article

« All ResourceBlog Articles

 

Bookmark and Share   Feed

Thursday, 15th November 2007

Researching Robots.txt Files: Study shows Google favored over other search engines by webmasters; Say Hello to BotSeer

When Dr. C Lee Giles writes, we read and learn.

Robots.txt Time: Study shows Google favored over other search engines by webmasters

Web site policy makers who use robots.txt files as gatekeepers to specify what is open and what is off limits to Web crawlers have a bias that favors Google over other search engines, say Penn State researchers whose study of more than 7,500 Web sites revealed Google’s advantage.

That finding was surprising, said C. Lee Giles, the David Reese Professor of Information Sciences and Technology who led the research team which developed a new search engine—BotSeer—for the study.

“We expected that robots.txt files would treat all search engines equally or maybe disfavor certain obnoxious bots, so we were surprised to discover a strong correlation between the robots favored and the search engines’ market share,” said Giles of Penn State’s College of Information Sciences and Technology (IST). While the study doesn’t include explanations for why Web policy makers have opted to favor Google, the researchers know the choice was made consciously. Not using a robots.txt file gives all robots equal access to a Web site.

As an example, some U.S. government sites favor Google’s bot—Googlebot – followed by Yahoo and MSN, according to the researchers.

“Robots.txt files are written by Web policy makers and administrators who have to intentionally specify Google as the favored search engine,” Giles said.

This news release is based on findings from the paper: Determining Bias to Search Engines from Robots.txt
7 pages; PDF.
Presented at: IEEE/WIC/ACM International Conference on Web Intelligence

See Also: Giles and Students Create BotSeer
BotSeer is a search engine for robots.txt. Its goal is to provide information about and access to robots.txt files throughout the web by crawling and indexing web robots.txt files and related documents. In addition, statistics about favored robots, comments and robot behavior is analyzed and presented.

See Also: Earlier This Year, Another Search Legend, Matt Koll, Now at Revolution Health, Pointed Out What Dr. Giles Writes About on the DailyMed Site (Available Only to Googlebot for Crawling)

See Also: Will SEO Become The Law For Federal Agencies?


Category:

Views: 1190




blog comments powered by Disqus

« All ResourceBlog Articles

 

Read about the FreePint FamilyFreePint Family

A family of resources to help information workers be more effective, raise the value of information in their organisations and contribute to success. Read more »


FeedLatest Family Articles:


Click to view the article Quilting big data threads
Thursday, 24th May 2012

Recently I have found myself cooing over visualisation maps (and heat maps) of health and well being resources. The content rich data is overlayed with mapping technologies, and some interesting themes and patterns are emerging.


Click to view the article The fallacy of information overload
Wednesday, 23rd May 2012

A lot of the talk around social media in the last year has been around information overload. Social media has provided us with new and exciting ways to create content. But it has also meant learning new ways to manage and engage with social media tools. Are we teetering on the edge of an information overload precipice?


Click to view the article Information overload: fact, fantasy or filter failure?
Wednesday, 23rd May 2012

Information overload is a figment of your imagination. Or a failure of your filter. Or a symptom of your technological submissiveness. Depends on who you ask.


Click to view the article Newsdesk: tracking millions of pieces of information a day
Tuesday, 22nd May 2012

What if you had to sort through 3.5 million articles and social media posts a day and try to pull out the most relevant items for your organisation? What if you then had to cobble it all together into something readable for your top groups and executives in your organisation?


Click to view the article Alacra Compliance adds managerial oversight
Tuesday, 22nd May 2012

Alacra Compliance saves time by aggregating information from both free and fee-based sources and enabling users to conduct an accurate federated search across these sources (coined “simultaneous search” by Alacra).


All Family Articles »
Family Articles by Category »


Tell us what you're working on,
and we'll talk to you about how FreePint can help »


FreePint Family Testimonials

"Fabulous resource to learn of unique tools and insights. Very useful." Manager, Futures and Forecasting, Virginia, USA

More testimonials »






Subscribe

Subscribe to the ResourceShelf Newsletter and receive the weekly sampler of posts and Resource of the Week.

Find out more »

ResourceShelf sponsored by:

Article Categories

All Article Categories »

Archive

All Archives »