Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome back to the course on a I I in the previous part we talked about the deep learning
2
Killary intuition we started there.
3
And in fact we actually got all the way to this part and where we talked about learning and now we're
4
going to move on to the actual acting part.
5
So there's there's two parts to distinct parts that we have to remember.
6
So that's the learning part but now he actually he's done all of this.
7
That's beautiful.
8
Now he actually has to take an action he has to decide what is he going to do is going to do action
9
one two three or four.
10
And so how does he do that.
11
Well the way he does it is now given those same values so the values don't change after we've we have
12
these values of compare them with Calcott the last two by arrogated era we've updated the weights but
13
the values don't change in that whole process.
14
To have got the cube values there.
15
They're fixed.
16
We know what they are.
17
All this happens though.
18
Networks updated and out using those same values that we had.
19
What we're going to do is we're going to parse them through a soft max function.
20
And again soft Max as described.
21
I think an annex 2 and we'll talk a bit more about soft max.
22
Further down in or we'll talk about this action selection policy further down in the rest of this section.
23
So just in a few tutorials.
24
But for now we're just going to say we're passing it through a soft next function.
25
Basically what it does is it allows it helps select the best one it selects the best action possible.
26
And there's a small caveat to that.
27
It's not just the best one possible.
28
We'll talk about that in the action selection policy tutorial.
29
But for now let's just say it selects the best action from here it says OK so Q1 you know the likelihood.
30
Basically we know that q values predicted the Q value so it can look at them and say OK so the highest
31
Q value of these just as we did in the simple Q learning algorithm.
32
Ill just look at all these for say the highest values this one I'm going to select that action we're
33
going to take those.
34
And that's pretty much it.
35
That's how he chooses which action take takes takes action and then all of this process happens again.
36
For for the next stage the agent ends up in in our case and the next square of the maze.
37
But generally speaking in the next state.
38
So there we go.
39
That's how we feed in a reinforcement learning problem into a neural network through a vector describing
40
the state that we're in.
41
And once we fit it.
42
There's two parts of the process that happen Part one is the learning.
43
So remember that part where we compare each of the cube values with the target and then we back propagate
44
the loss through the network to update the weights so that our network is learning as we go through
45
this maze or through this environment.
46
And also the second part is of course we have to act we have to select an action and that is where we
47
pass the values through a soft max function and or basically an action selection policy which we'll
48
talk about further down.
49
And then we simply select the action that we want to take and we perform that action and then this whole
50
process starts again.
51
And then maybe the agent gets then maybe the agent doesn't pausa the game.
52
In any case the game ends.
53
And then once again the whole process repeats the agent plays the whole game again and then that stops
54
so basically that's that's another airpark every time the agent you know every time the game ends with
55
a favor beyond fairie that's the end of an airport.
56
And then he starts again and then he starts again and then he starts again.
57
And so on.
58
So that happens and this process happens for every single time the agent is in you in a new state so
59
the state is encoded here so that's important not just for every single game that he plays but for every
60
single state.
61
So he's in a state that goes through his process dates and so on and happens every single time.
62
And so the learning happens and the acting happens as well.
63
So that is deep learning in the intuition behind deep learning.
64
We've got lots more to cover off and then of course practical and in the meantime if you'd like to get
65
some additional information on keep learning.
66
We've got a recommended reading so we've already spoken about Arthur Giuliani's series of blog posts.
67
If you look at simple informal learning Lifton's flow part 4 you will find the part that's relevant
68
to what we discussed today.
69
Note that here he talks about convolutions we are not covering revolutions in this section we're going
70
to be talking about them in the next section of the course.
71
So the difference here is that it's just kind of skip the conclusions part for now and we'll talk about
72
them in the next part of the course but the difference is in evolutions.
73
You're like looking the agent is looking at the image and and therefore he has to process an image an
74
additional complication for now where we're slowly gradually building up to that.
75
For now we're encoding our environment through you look here we're encoding our environment or maybe
76
like look at this one probably in coding our environment as a or in to state the agent is in as a vector.
77
So in our case was very simple vector of values.
78
Sometimes people even in that in that simple may sometimes or as you'll see from this blog post.
79
Sometimes people prefer the one hot and coded version of that state.
80
So basically where every single box of the maze has a.
81
So you have like a vector of for a null case would be 12 values three by four.
82
So it isn't like either either 1 or 0 depending on which elements and which box you're in.
83
In the environment.
84
So in whichever way you decide to code your environment and the state of your environment that's how
85
in coding It's basically a vector.
86
The key here is that it's not a convolution So it's not like an image and there's no convolution volt
87
So this part will come later.
88
For us it starts over here and that just simplifies the process for us to gradually understand better.
89
And of course don't forget that this post is rude and tends to flow and we're using pi torche in our
90
tutorials.
91
So hopefully you enjoy this.
92
A quick intro into a deep convolutional deep not yet deep book learning.
93
And on that note I look forward to seeing you next.
94
And until then enjoy artificial intelligence.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.