Showing posts with label computers. Show all posts
Showing posts with label computers. Show all posts

The Invisible Web

I have been working on the net for last several years. I always thought that when I needed to find anything on the net I could use Google, Altavista, Hotbot, Yahoo, Excite, Snap, About, AOL, MSN, Looksmart, Web crawler, Ask Jeeves, Northern Light and covered almost the entire gamut of the Internet for my searches. It was one day when going through a job advertisement on Web Researcher, in qualifications column, I saw ...."knowledge of Invisible web required....". Now what was that..Invisible Web...? I had never heard of it before. Research on the topic produced startling information which I share with you.

What exactly is Invisible web: A vast majority of information available on the Internet does not reside directly on the World Wide Web. It is found in hidden databases that cannot be seen or searched by common Internet search engines. There are many many such databases This vast Ocean of databases is commonly referred to as the "The Invisible Web". It is called Invisible as it is not searched through commonly used search engines. It is also called "Hidden Web" as this part of the Internet hides inside the Data Base. Some People call it the "Deeper Web" because most of the search Engines only search (spider/crawl) the surface of the Internet.

Why is this part of the Internet Hidden/Invisible: To understand this we must understand how search engines work. Search engines visit (crawl, spider) the web servers registering the addresses of Web pages they discover. When they come across a database, they read and index the home page of a database. They do not go inside that database to find out what is stored therein. Thus search engines don't distinguish between the index page of a huge data base and a simple Web page. Although information is there in the Database it is invisible to the search engine because of its technology. Another reason is that the HTML pages we normally post on the Internet are fixed or static pages. Pages of the databases are more active and offer content that is put together from many parts of the database. These are called dynamic pages. Commonly used Search Engines cannot index dynamic pages and so the databases remain hidden.What is the extent of the Invisible web: The Google Search Engine mentions indexing 28.0547 billion pages. This at the moment is considered to be one of the biggest indexing of Web pages. Bright Planet a company which has brought out a white paper on the subject estimates this to be 800 billion to 950 billion web pages. This translates to up to 500 times of the visible web. The Invisible Web contains 7,500 terabytes of information, compared to 19 terabytes of information in the commonly searched Web. There are an estimated 1,00,000 Invisible web sites and growing.

What is the significance of this concept to Web Masters: As web masters especially those working with the data bases it is very essential to know that a conventional search engine will not index their site/ web page. More likely it will only index the Index page of the entire data base once. There is, as yet, no solution to this problem of data base not being found through a conventional search engine.

Why is Invisible Web important: Most of the Invisible Web contains information that can be publicly accessed and available for free. This information is extremely relevant for domain names, market research,and is commercially relevant information.

Gateways to Invisible web: Now the question arises how to search the Invisible web. One way is through gateways, which are the interfaces giving direct access to the thousands of searchable databases on the Web that the major search engines ignore. Some of these are Alpha Search, The Big Hub, WebData.com, Digital Librarian etc.

How to search the Invisible Web: Besides gateways there are specific search engines which search the Invisible web. Most significant amongst these are Lexibot, Infomine, InvisibleWeb

contributed by:Kainaat creations

Boolean Operators in Internet Searching

We all search the Internet frequently for specific information.Supposing you are looking for the topic: Cinematic techniques used by James Joyce in his novels. The first and the obvious way to search for the topic is to type the word James Joyce in the search box of the search engine. This would yield a large number of pages of sites related to James Joyce as also all James and all the Joyce pages. Opening all these pages and browsing them for specific topic of cinematic techniques is a very tedious job. Are there any short cuts ?

Well yes there are short cuts and these are called Boolean operators based on the Boolean logic (after George Boole,18th cent.Mathematician). Boolean operators are used to combine search concepts in a more precise way than is possible with word searching.

The few important ones are discussed below:

AND:
The AND operator narrows a search to include only those web pages that contain both keywords. The search syntax for our above query would be; james AND joyce AND cinematic AND techniques. AND operators are a great way of limiting the numbers of search results because they link two subjects to create a new query subject for the engine.Thus the engine will only search for web pages on the Internet that will include all the four subjects.

OR:
The OR operator broadens searches to include web pages/ web sites that contain any keywords of the search. The search query in our case could be; joyce OR joycean OR joyces. OR operators can be useful when searching for alternative spellings,or searching for synonyms e.g.in our case; techniques OR concepts OR methods

NOT:
The NOT operator narrows a search to exclude certain keywords. Say searching for james joyce cinematic techniques we don't want the search engine to open web sites relating to James Joyce Foundation, James Joyce Society etc. The query will then be; james AND joyce NOT foundation

NEAR:
NEAR operators work similarly to AND operators, retrieving only those web pages that contain both keywords, but NEAR operators further limit the search results by requiring that the keywords be within ten words of each other. This is especially important in our case of a proper name search because when we look for james and joyce the engine will also open pages of james dean, james hill, james watt, james smith etc. and similarly of richard joyce and william joyce etc. Use of NEAR operator will open only pages of james joyce. The query will thus be james NEAR joyce.

IMPORTANT:
It is very important to note the in all the discussion above the operators have been mentioned in uppercase and the searched words in lowercase. This is essential, because if operators are mentioned in lower case the search engines will ignore them as these are the most commonly occurring words in any web page.

The major search engines all permit Boolean searching, but they vary in the syntax they require. Learning the fine points of a particular search engine will greatly improve the precision of your queries. This fine tuning to the major engines like Google, Altavista, Excite, Northern Light, Hotbot, Lycos, Yahoo, Web crawler, Infoseek etc. will be discussed in another article.

Search strategy to Search that elusive Web page

We all have had times on the internet when we don't find the thing that we are looking for on the net.It is not only time consuming but also very frustrating. With the web growing every minute the number of pages is increasing to incredible levels. From 54,000 odd pages in 1994 to about 1.347 billion pages in 2001 to 8.058 billion pages in 2005 the growth has been phenomenal. How often we wished that someone could have classified things for us like our office secretary and kept them in small filing cabinet, subject wise so that without wasting time we could reach for it straightway.

Directories: Search engines crawl the net and store information about every site they visit. They don't categorise or differentiate. Humans all along have been making the effort to simplify and classify the information available on the Internet. This classifications have been called directories.

How are directories created: The Directory researchers work constantly to examine new web pages one-at-a-time. They classify all the pages they find to categories. These start with one main category and then keep going down to sub categories and sub-sub categories .e.g. The Home page of An American Museum of Contemporary Art may thus be classified as :

ARTS & ARCHITECTURE > ARTS > CONTEMPORARY ART > AMERICAN > MUSEUMS

The people working on the directory projects browse the new pages, make value judgements concerning the appropriateness of each page for listing, find a suitable category for listing, and exclude inappropriate content. This whole process is very much like what a librarian does when he selects, acquires, catalogs, and weeds books.

What is an open search: The most common way of searching the Internet is by using keywords. Type the keywords for the topic you are looking for in the search box of the search engine and see the relevant search results. e.g. searching for biography of Robert Ludlum you may type biography, Robert Ludlum. This type of search is called open search. (here we are not discussing the results of this search)

What is a directory search: In directory search we go form one main category to a category to a sub-category to a sub-sub category of the search directory, till we reach the desired level of the topic we searched e.g. searching for the same biography the path would be;

Arts > Literature > Authors > biography > (alphabet) L > Ludlum > topic

When to use open search:
1. When you want to quickly reach to relevant web pages and you have the keywords

2.When searching for all the freshly published content.

3.When you want all the pages on the topic without excluding anything.

When to use Directory Search:
1.When you dont have the keywords.

2.When keywords are not yielding the results.

3.When the open search is giving too many results.

4.When seeking for quality article/ web site on a specific subject.

5.When you want to exclude all junky pages from the search.

What are the top of the line directory projects: Compiling a directory is a time consuming job. Not every web site or Search Engine owner starts a simultaneous project of creating a directory. Looksmart was one of the early pioneers in this field and it subsequently tied with Altavista. Now they together claim to have the Largest directory on the Internet with 350 million indexed web pages. AOL.com is another big player which has a Human compiled directory. DMOZ is also in the field and attempting to create one of the largest open directories on the net. Wikipedia is the latest entrant and going very rapidly everyday.

We can only say in the end that search engines are essential but Directories are also important.Because the search engines will search and search and create chaos, it is only the humans who will bring sanity to this data.

Significance of Meta tags and Key words

How do search engines relate Description, Keywords ,Title and Content in the Body of web page:

Search engines are continuously 'spidering' crawling the web. When they visit a web site they note the URL (universal resource locator) of the web page, Title of the web page, Description Meta tags, Keywords Meta tags, as also Keywords in the text or the body portion of the web page When we submit a search query to the engine The job of a search engine is to find the content which most suitably fits the search query, the engine links up these words with what it has stored in its memory. Depending upon the occurrence of these words in the web pages already spidered by it the results are displayed. It is very essential to understand that different search engines respond differently to the Meta tags. Some use only the Title tag, others use description tag and few others the keywords. So all three tags are important and also all three must have similar content.

Pitfalls of using unrelated Keywords in Meta tags :
It is a common belief that if a Web master has the right keywords his page would figure in the top 10 of a search engine. Meta tags are important as they do help in search engine listing but they are like a double edged sword.Most search engine look at the Keywords, Title and Description of the web page and then they try and relate these to the Actual Content of the page. It is vital that the Title, the Description and the Keywords are in unison with each other as also with what is appearing on the web page. Having a too long description (it wont fit in the description window of the search engine), Keywords should not be repeated. If there are to many Keywords and unrelated to the content on the page the search engine will classify it as spam and put the page out of its listings.

And one final word-please, please don't steal keywords from a popular web site, it is Illegal..

Significance of Meta tags and Key words

What are Meta tags:
Meta tags are a code used in the web page programming (both HTML & JAVA). They are not visible when a web page is being viewed by a browser. However the search engines use the Meta tags to categorise and classify the content of the web page. Web masters use Meta tags for better placement of their pages on the search engines.

What are commonly used Meta tags:
1. Keywords: This Meta Taget of keywords (with a comma after each keyword)

2 Description: Meta Tag gives description of web page's highlights .

3. Title: Defines the title of the web page. (some authors don't consider this as a Meta tag.)

Other less significant metatags:

4. Author: This tag defines who wrote the web page.

5. Generator: This tag defines the program used to create the web page.

6. Robot: This allows a page to be indexed or not to be indexed by search engine..


Where should be the Meta tags located on the Web page:
Any HTML document (web page) consists of a Header portion and a Body section. The information located in the Header is used by search engines. So meta tags should be located in the and the portion.This has t be included for all pages. For framed pages include META TAGS on all individual framed pages and NOT merely on frame set page.

How to use Title, Description and Keywords in the Meta tags:
The Title tag must accurately define the content on the web page. The description tag must in 1-2 short sentences give the gist of your page. The Keyword tag in about 8-15 'key'words should highlight the most significant aspects of the page. If we were to define the Meta tags for this article they would be;



Topics

Society (19) Family (18) health (8) humour (6) computers (5) Art (4) writing (4) Religion (3) Music (1) Science (1)