Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome back to the course on deep learning.
2
All right today we're talking about the activation function.
3
Let's get straight into it.
4
So this is where we left off previously we talked about the structure of one neuron.
5
So there it is in the middle we know that it has some inputs values coming in it's got some weights
6
then it adds up the way to calculate the way that some of those inputs and then apply the activation
7
function in step 3.
8
It passes on the signal to the next year and then that's what we're talking about today we're talking
9
about the value that is going to be passed over.
10
So we're talking about the activation function that's being applied.
11
So what options do we have for the activation function.
12
Well we're going to look at four different types of activation functions that you can choose from.
13
Of course there are more different types of activation function but these are the predominate ones that
14
you'll be hearing about and that we'll be using in this course.
15
So here is the threshold function.
16
This is what it looks like.
17
So on the x axis you have the weighted some of inputs on the y axis.
18
You have just you know the values from 0 to 1 and basically the threshold functions are very simple
19
type a function where if the value is less than zero then the free.
20
Thanks ssion passes on zero.
21
If the value is more than zero or equal to zero then threshold function pusses on a 1.
22
So it's basically kind of like yes no type of function.
23
Very very straightforward.
24
Very kind of like rigid type of function either yes or no.
25
No other options.
26
So there you go.
27
That's how it works.
28
Very simple function.
29
Let's move on to something a bit more complex.
30
Now this sigmoid function very interesting formula that we have here you'll see just now there is one
31
divide by one plus each.
32
The power of minus X whereas in this case of course X is the value of the sums of the way that sums.
33
And so yeah.
34
So this is what the sigmoid looks like.
35
It's a function which is used in the logistic regression.
36
If you recall from the machine learning course.
37
So what is good about this function is that it is smooth.
38
Unlike the virtual function.
39
This one doesn't have those kinks in its curve and therefore it's just nice and smooth gradual progression.
40
So anything below 0 is just like drops off above zero.
41
It acts approximates towards one and this sigmoid function is very useful in the final Lehren the output
42
layer.
43
Especially when you're trying to predict probabilities.
44
And we'll see that throughout the course.
45
And then we've got the rectifier function rectifier function even though it has a kink is one of the
46
most popular functions for artificial neural networks so it goes all the way to zero it is zero.
47
And then from there it's gradually progresses as the input value increases as well and we'll see that
48
throughout the course we'll see that in other intuition tutorials and we also see that how we use this
49
function in the practical side of the course and I will comment on this a bit more in a few slides from
50
now.
51
So just remember the direct fire function is one of the most used functions in artificial neural networks.
52
And finally we've got one more function that you will probably hear about.
53
It's the hyperbolic tangent function.
54
It's very similar to the sigmoid function but here the hyperbolic tangent function goes below zero so
55
the values go from 0 to 1 or approximately 2 1 and go from zero to minus 1 on the other side.
56
And that can be useful in some applications.
57
So we're not going to go into too much depth on each one of these functions I just wanted to acquaint
58
you with them so that you know what they look like and what they're called.
59
If you'd like to get some additional reading then check out this paper by a 75 year lot.
60
Have you a lot called Deep sparse rectifies neural networks 2000 paper.
61
And there you will find out exactly why the rectifier function is such a valuable function why it's
62
so popularly used.
63
But nevertheless for now we don't really need to know all of those things.
64
For now we're just going to start applying them which you start using them more and more and more.
65
And so when you feel comfortable with the practical side of things then you can go and refer to this
66
paper and then you will be able to soak in that knowledge much quicker and it will make much more sense.
67
But just keep this in mind that when you're ready when you feel that you're ready then you can go and
68
research paper and get some valuable knowledge from them.
69
So just to quickly recap we have the threshold activation function which goes like this the sigmoid
70
activation function which looks like this.
71
We have the rectifier function and we have the hyperbolic tangent function and now to finish off this
72
tutorial Let's quickly do a few exercise so just do two quick exercises to help that knowledge sink
73
in.
74
So first one is we've got an example here of a neural network of just one neuron and that right away
75
the output layer.
76
And the question is assuming that your dependent variable is binary So it's either 0 or 1 which threshold
77
function would you use.
78
So out of the ones that we've discussed we have a threshold function the sigmoid function the rectifier
79
function and we've got the hyperbolic tangent function in it's in their roll forms which ones would
80
you be able to use for a binary variable.
81
OK.
82
So the answers here are there's two options that we can approach this with.
83
So one is the threshold activation function because we know that it's between 0 and 1 and it gives us
84
0 Anderson umbrellas and then otherwise it gives you once it only can give you two values.
85
It fits perfectly fits this requirement perfectly and therefore you could you say y equals the threshold
86
function of your sway to some and that's it.
87
And in the second case which you could use is the sigmoid activation function.
88
It is actually also between 0 and 1 just what we need.
89
But at the same time you want is just one right so you is not exactly what we need but in this case
90
which you could use it as is the probability of Y being yes or no.
91
So we want Y to be 0 1 but instead we'll say that the sigmoid function Simoun activation function tells
92
us whether it would tell us of the probability of Y being equal to 1.
93
So basically the closer you get to the top the more likely it is that this is indeed a one or a yes
94
rather than a no.
95
And yeah so that's very similar to the logistic regression approach.
96
And those are just two examples.
97
If you have a binary variable.
98
Now let's have a look at another practical application.
99
Let's have a look at how all this would play out if we had in your all natural like this.
100
So in the first layer we have some inputs.
101
They are sent off to our first hidden layer and then an activation function is applied.
102
And usually what you would apply here and what you will see throughout the Scorsese will apply a rectifier
103
activation function so it would look something like that.
104
We apply the rectifier activation function and then from there the signals would be passed on to the
105
output layer where the sigmoid activation function would be applied and that would be our final output.
106
And that could predict a probability for instance so this combination is going to be quite common where
107
in the hidden layers we apply the rectifier function and then output there we apply the sigmoid function.
108
So there we go.
109
Hope you enjoyed this tutorial now you are quite well versed in four different types of activation functions
110
and you will get some hands on practical experience with them throughout this course will be using them
111
all over the place so you'll get to know them quite intimately and you should be quite comfortable with
112
them.
113
But for now this is the knowledge that you need to progress and understand what he's going to be happening
114
further down in this course.
115
And on that note I look forward to seeing you next time.
116
Until then enjoy learning.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.