My spider, as with all legit spiders, looks for a file called "robots.txt" off the main directory of any website it attempts to "crawl".

If a webmaster dosen't want a search engine to index specific directories, they simply add those directories to their "robots.txt" file, and the spider will not "crawl" them.

In short, if a webmaster dosen't want a search engine to index their content, all they have to do is tell it so in "robots.txt"

I don't know if having a "robots.txt" file is some kind of secret or what.. but very few webpages i've crawled so far have one.

Stovebolt.com has one, and thats why the forum content from stovebolt.com is not included in OldTruckStuff.com's indexing.

I'll go down your trucks list tonight before I go to bed and add any of the pages I don't already have indexed.. thanks!


an idea is only stupid if you think about it rationally.