Afrikaans
Akan
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Corsican
Croatian
Czech
Danish
Dutch
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hmong
Hungarian
Icelandic
Igbo
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Maltese
Maori
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
We humans,
have a keen eye for
visual treat
and love a good eye candy.
That being said, statisticians
were having a hard time
with getting people to listen
to their very important
and relevent data.
Frankly speaking,
even though tables
are easier to consume,
it still does not taste good.
Does it?
So what
to do?
A lot of data
that we deal with
in the real life is comparative.
As in comparing 2 things.
It can be about
which student is tallest
or which candy is cheapest
and so on.
This is where
graphical representation
comes to our rescue.
Graphical representation of data
allows us to understand
the data much more easily
and intuitively than a table.
Our aim here,
is to throw some light
on 3 major types
of graphical representation of data.
That is the bar graph,
the histogram and
the frequency polygon.
You already know
a few things about graphs
from your earlier classes.
Let's build on that.
Let's represent table of heights
in the form
of a bar graph.
Let's bring up the table
of ungrouped data here.
To draw the chart,
I'll start off by drawing
a flat horizontal line
called the X axis.
Where you represent
different height values
of all the students.
Now, as I
go through this data,
I start to add a dot
above the X axis.
So first, I put a dot
at 195.
The next is at 175.
The next at 170
and I keep continuing.
The fifth student
is at 185.
And so is the sixth student.
So I add a dot above that.
This can continue
till I exhaust the complete data.
So if you observe,
the Y axis represents
the number of students
or the frequency.
Because this is called
a bar graph and not
a dot graph,
instead of you using dots
like you just did,
you can start to draw
rectangular bars and extend it
till your corresponding data point.
For example, there is just
one student with the height of
130 cms. So I extend the bar
till it reaches the level of
1 on the Y axis.
The next data value is
135. Has a frequency of 2.
So I extend the bar
till it reaches a value of
2 on your Y axis.
Moving on, we have 155
which appears 7 times.
So the bar for this value
extends till it reaches 7
on the Y axis.
162 goes upto 6.
168 goes up to 4.
Finally 195,
with the frequency of 1.
Remember that the thickness
of all these bars
that you see
is actually of your choice.
But for the sake of clarity,
you tend to maintain
all of them
as the same thickness.
From this you can easily tell
the number of students
that have the same height
by looking at the
height of each of these bars.
This becomes all the more necessary,
when we ask questions regarding
a particular height
and also how many students
have the same height.
We apply the same logic
to grouped data as well.
In this, we have
heights of 60 students
grouped into classes of 10 each.
So we drop the axis again.
This time we have
the classes on the X axis
and corresponding frequencies
on the Y axis.
The class of 130-140
has a frequency of 9.
So the bar extends from
X axis, to reach
a level of 9
on the Y axis.
Similarly, for the rest
of the frequencies.
Making data visual,
makes it leads better
to understand it. Doesn't it?
A similar representation can happen
using a histogram as well.
Let's dive in.
A histogram is just like
this bar graph.
But I will have to make
a few changes here.
Just like a bar graph
we represent the height
on the horizontal axis
but using a suitable scale.
Scale here, becomes very important
because the area
that this bar covers
in a histogram
is very very important.
We can choose the scale
as 1cm equivalent to 10cms.
So each class occupies
a width of 1cm
on this graph.
Also since the first class interval
is not starting from zero
but a fixed non-zero value
we show it on a graph
by marking a Kink
like you see here.
As this has a break
on the axis. Next.
Unlike the bar graph
there are no gaps
in between the rectangles
of the graph.
So, I will have to
knock off all of these gaps
and will have to keep
only the lower class limits
on the graph.
Technically, it is one solid figure.
What you see now
is called a Histogram.
One important thing
that you will need to
keep in mind about a histogram,
is the area of the graph
plays a very crucial role.
In fact, the area of the bar
is directly proportional
to the frequency of that data.
Also the sum of areas
of all the bars
is equal to the
total frequency of all the classes
in the table.
Till now we have dealt with
classes of equal sizes.
What if I have
different class sizes
on the same histogram?
That is, what if
I have to put all students
with heights less than 150
in one bracket
and anyone with heights
more than 170 in another bracket.
Absolutely arbitrary.
So that means
class interval 130-140,
140-150 get clubbed
into one class interval
of 130-150 along with their
respective frequencies.
Likewise class intervals of 170-180,
180-190, 190-200
all get clubbed
in a class interval of 170-200.
And in between, we have 150-160
and 160-170.
That remain as is.
So, that means, now
you have a new table
with classes of different widths.
Tthe first class width is 20.
That is 150 minus 130.
Followed by 2 class widths
of 10 each.
That's 160 minus 150.
And the last one
with a width of 30,
which is 200 minus 170
that you see.
If I were to draw
a bar graph here,
this is how it would look.
For a histogram
on the other hand,
I mentioned that the
areas of the bars are crucial
for accurate representation.
We need to pay attention
to the width and height
of these bars here.
So, the width of
all the 3 classes are
different. Remember
I told you that the
area of a histogram
has to be proportional
to the frequency.
So how do we do this?
We need to bring
all the frequencies in line
with the minimum class width.
The minimum class width here is
10.
The length of the rectangles
are to be modified
to proportionate this class size.
For instance,
when the class size is 20,
as is the first case.
The length of the rectangle
will be 16 times 10
divided by 20,
which is going to be
equivalent to 8.
This is simple cross multiplication.
This way the total frequency
will be 16 in this range.
The next 2 groups
the class widths are the same
as the minimum class width.
Hence you don't need to
change anything.
The last one however,
goes through the same treatment
as the first one.
In this instance,
the class size is 30
and the frequency is 8.
So when the class size
becomes 10,
the length of this rectangle
will be 8 times 10
divided by 30.
That is 2.666
This histogram can now be said,
to be proportional
to the students
per 10 cm interval.
Even though a bar graph
and a histogram look alike,
you might have noticed already
that there are a few differences.
In fact if I bring them together
unless you are a statistician,
chances are,
that you will get confused.
This exercise that
we will do now,
will help you sort out
this confusion.
If I ask you to
collect data about language preferences
of the students and
add it to our
original table.
Now I will be able to
draw a bar graph
out of it.
Now let's try to make
a histogram out of the
language data
that we have collected.
Is that even possible?
Hmm. No it is not.
Infact the data that you collect
can be split into qualitative
and quantitative data.
If you're looking at
colors of the car
on the road,
then the color of the car
which is a data
which is of the qualitative kind
because this describes the
quality of that particular data.
Or if I ask you
the flavor of ice cream
that you like,
that again is a qualitative data.
On the other hand,
data such as heights,
weights, roll numbers, etc.
are data that are
represented by numbers.
Here height is 160 cms tall.
160 is a quantitative data
since it refers to
numerical data.
From the examples
that we have solved before,
you can see
that we can represent both
qualitative and quantitative data
on the bar graph.
Where as we can represent
only quantitative data
on a histogram.
So the next time,
you need to make a graph
be sure to analyze
what kind of data
you are trying to represent.
Now let's start making
a difference table out here.
And let's start populating
the differences as we go about.
Let's bring back the
graph of heights
from the bar graph section.
Now we see that
on the X axis,
each data point is represented
individually. For example,
a student's height of 130 cms
is represented individually
as 130 cms on the X axis.
This kind of data representation
individually is called Discrete data.
And we also
know that we can
construct a bar a graph
using grouped data as well.
Grouped data here, refers to
when a data point is represented
not individually but as a
continuous range of values.
In case of discrete data,
we can have gaps
in between the values
of data points
on the X axis.
The data that you collect
can again be classified
in one more type.
As continuous and discrete data.
When you're talking about
discrete data, there can be
gaps in the data
that you collect.
For example 130 cms
and 135 cms as heights of students
has a gap of
5 in between them.
And when you're talking about
continuous data,
there cannot be these gaps
that you see here.
So, when it comes to a
bar graph, you can represent
both continous and discrete data.
But in a histogram
you can represent only
continuous data.
This is another reason why
bars of a bar graph
are separated by a gap,
since they are discrete values.
Whereas in a histogram
all the bars are clubbed together.
We also cannot
reorder this data
in case of a histogram
due to continuity of the data.
Let's add these 2 points also
into our comparison chart.
Using continuous data means
that the classes have to be
ordered on the graph
as the appeared to us.
On the other hand
having discrete data
in the bar graph
allows you to arrange the variables
in anyway you want to.
When I'm drawing a bar graph
the order in which
I show the elements
on the X axis
is not a problem at all.
I can first show 130-140.
Then show 150-160.
And then I can have 140-150.
But when it comes to a
histogram, I cannot
reorder the data.
This is obvious because
we are dealing with
continuous variable.
This is one more
for the comparison chart.
As you've already seen,
the spaces in between the bars
are not present in the histogram.
It essentially looks like
one big block.
Also the width of the bars
need not be the same
when it comes to a histogram.
Also remember,
that the area of the bar
plays a huge role
in a histogram and hence,
we need to maintain uniformity
of class width through out.
But in the case
of a bar graph,
the width of the bars
are immaterial
to the interpretation
of the bar graph.
There is yet another visual way
of representing quantitative data
and its frequency.
It's called the frequency polygon.
Let's consider the histogram
that we initially constructed
with equal class intervals.
Let me mark this point,
which is the midpoint
of the class interval
of 130-140.
I will call this point
as the class mark.
So class mark is a
mathematical way of saying
mid-point of class interval
which we obtained
by adding the upper
and lower limits of a class
and dividing it by 2.
If we consider
the class interval of 150-160,
its class mark is
150+160/2
which is going to be 155.
Next I will highlight
the class marks
for all other class intervals
as well.
For a frequency polygon,
all I have to do
is to connect
all of these dots.
Or connect all of these
class marks. Well.
I said frequency polygon.
But what is a polygon?
A polygon is a
multi-sided shape.
But before all
it is a closed shape.
So how do we get that?
We add a class interval
before the first one
in the data
and do the same
in the other end
of the histogram as well.
Since, the first class interval
is 130-140 we add another
with 120-130.
This class interval will ofcourse
have a frequency of zero,
since it is not
represented in the table.
We just have to
mark the class mark
for this group. That is
130+120/2
which is going to be 125.
And then we are done.
We can now connect the line
to the X axis.
Doing the same
on the other end,
we add 200-210
to the frequency polygon graph.
Marking the class mark as 205
and closing the figure
at the both ends
gives us the frequency polygon.
Instead of drawing the entire
bar of a histogram,
you just mark the frequency levels
with the Y axis
at the class mark.
Just like a histogram,
frequency polygon's total area
is directly proportional
to the total frequency
of the table.
For the sake of convenience,
let's bring back the histogram
that we drew
in the previous sections.
Let's take the graph
where the frequency polygon
is drawn over the histogram.
So if I join the class marks
you can see
that the chunks of area
are being leftout of calculation.
There are also a few
empty areas
inside the frequency polygon.
To prove to you
that the area of the
frequency polygon
and that of the histogram
are the same,
I will cut the part
which is outside the line.
Flip it all over and see
that it fits exactly
into the empty area here.
The same can be done
for all the bars.
So eventually, we see that
all the triangles
ejected by the line we drew
are included within this
frequency polygon. And hence,
we can visually say
that the total area
of the frequency polygon
is equal to the
total area of the histogram
made by the same data.
Also the area
is proportional to the frequency.
So till now,
we have learnt about
raw data and how
unless it has context,
it is useless. Raw data
can also be made useful
by processing it.
We process data
by means of statistics.
Using methods such as
creating a frequency distribution table.
Frequency is the
number of times
a particular data
appears in a data set.
When the number of heights
are considered individually
we call it ungrouped data set.
It was too much data
to deal with.
So we then
clubbed the heights to create
a grouped data and a
grouped frequency distribution table.
We then decided
that numbers are all together
too boring,
and came up with
graphical methods of representing data.
This includes bar graphs,
histograms and frequency polygons.
Bar graphs is an excellent
comparative tool
and is used mostly
in non numerical context.
Such as comparing 2 items.
The bars in a bar graph,
typically are of the same width
but bare no relevance
to the area that they occupy.
In contrast to it,
in a histogram
the dimensions of the bars
are very crucial.
The area of the bar
is directly proportional
to its frequency.
Consequently, so
the width of the class intervals
must be taken into account
whenever you're attempting
to answer relevent questions.
Whenever the width of the
class intervals is non-uniform,
use the minimum class interval
as a standard
and use cross multiplication
to get an accurate representation
of data on the graph.
When it comes to frequency polygons,
the only thing that
you need to do differently
from a histogram,
is to mark the
class mark on the graph.
Class mark is the midpoint
of all the class intervals.
Instead of an entire bar,
you only make one mark.
Then you connect
all of these dots
and get a line.
To make frequency polygon
out of this,
you need to
close the figure.
To do this, add a class
before the first
and after the last classes
with the same width
as the width of the first
and the last classes respectively.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.