The way that Miller starts out in the book is explaining how information is stored on the web, which is important to understand in order to achieve great SEO. Miller uses a great analogy comparing the web to a bunch of filing cabinets. “You can walk from file cabinet to file cabinet (surf from website to website), rifle through the file folders in any particular file drawer (browse through the pages on any site), and scan the papers within any individual folder (read the contents of any individual web page)” (Miller). After all the searching through the filing cabinets you probably didn’t find what you were looking for. This is because the filing cabinets (websites) and the content inside (web pages) are created and organized by humans. The problem with this is that humans: A. are not perfect, B. rarely think perfectly logically, and C. do not think alike (Miller). There is an estimated 150 million websites with close to 1 trillion individual web pages says Miller. Each website was created by a different human which means that each website has its own unique organization. Since there are over 150 million websites with their own organization, this means that the internet is nothing more than an unorganized filing cabinet. How search engines work
The best thing about search engines are that they are not powered by human hands like filing cabinets are. Search engines use special software that send out spiders through every page on the internet and stores the findings on a massive computer. The spiders either save the whole webpage or only certain words or phrases, depending on what search you use (because different search engines work differently) (Miller). Then when a person does a search the search engine will pull up the most relevant information from its own database. Each search engine is different because they all have their own database in which they search from. This is why if you search the same keyword in different search engines you may get different results.
How search engines index the web
When a term or keyword is searched through a search engine, for the users stand point, it is very simple.
1) Type in search term
2) Press search button
3) View results
From the search engines side of the search it is much more complicated. The process would look something like this:
1. When the user clicks the Search button, his query is transmitted over the Internet to the site’s web server.
2. The search site’s web server sends the query to the company’s array of index servers. These computers hold a searchable index to the site’s database of web pages.
3. The query is matched to listings in the site’s search index—that is, the index servers determine which actual web pages contain words that match the query.
4. The search site now passes the query to its document servers, which store all the assembled web listings (documents) in the database.
5. The document servers assemble the results page for the query by pasting together snippets of the appropriate stored documents.
6. The document servers send the assembled results page back to the main web server.
7. The search site’s web server sends the results page across the Internet to the user’s web browser, where he views those results. (Miller)
The amazing part is that this whole process, no matter what side you are on (user, or search engine), takes under half of a second to complete! (Miller) The reason why the search engine can produce the results so quickly is because it only accesses it’s own index to find the relevant information. It would take much longer if the search engine had to search the entire internet every time someone did a search.
The spider software for the search engine does periodic searches through its own database to see if there are any website updates, as well as, searches the links on each site to find new content to index. This process is typically done every few weeks (Miller). Google has it’s own spider software called GoogleBot and it searches some sites more often than others. News sites might be searched hourly by the GoogleBot whereas a reference site might be searched every few weeks. (Miller)
Once the spiders have compiled all of their gathered information in the index then the search engine can determine search rank of the sites. Search Rank is what every website owner wants to be on top of. The better the search rank that your site has the higher it will appear on the results page of the search engine (Miller). Even though the process in which the search engines figure out your search rank is very confidential (for competitive reasons), according to Miller each one functions the same way. Each search engine uses its own algorithm to rank each webpage in its index. Since each search engine uses a different algorithm and a different database, your site can be ranked high on one and low on another. Google is perhaps the most secure with their search rank but from what is out there it is known that Google uses three components to determine ranking:
• Text analysis. Google looks not only for matching words on a web page, but also for how those words are used. That means examining font size, usage, proximity, and more than a hundred other factors to help determine relevance. Google also analyzes the content of neighboring pages on the same website to ensure that the selected page is the best match.
• Links and link text. Google then looks at the links (and the text for those links) on the web page, making sure that they link to pages that are relevant to the searcher’s query.
• PageRank. Finally, Google relies on its own proprietary PageRank technology to give an objective measurement of web page importance and popularity. PageRank determines a page’s importance by counting the number of other pages that link to that page. The more pages that link to a page, the higher that page’s PageRank—and the higher it will appear in the search results. The PageRank is a numerical ranking from 0 to 10, expressed as PR0, PR1, PR2, and so forth—the higher, the better.
(Miller)
“It’s important to note that Google’s determination of a page’s rank is completely automated. There is no human subjectivity involved, and no person or company can pay to increase the ranking of its listings. It’s all about the math.” (Miller) Unlike Adwords and CPC, companies cannot buy their page rank. I think this is good because it still gives other sites that rely on organic traffic a chance to have a high page rank.
Even though the search engines are really smart in indexing websites and pages, there are still things that the search engines have trouble finding.
Dynamic Pages – “A dynamic website is one that changes or customizes itself frequently and automatically, based on certain criteria.” Often the URL of the site will change which makes it hard for the spiders to index the site.
Images and Media – Pictures, Videos, Podcasts, Music, Flash, these are all things that the search engine spiders have trouble indexing. The reason that the spiders have trouble indexing this type of media is because it isn’t text. The spiders look for text when searching, so all of your pictures, videos, podcasts, ect can get lost in the search. (Miller)


No comments:
Post a Comment