Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome back to the course on artificial intelligence.
2
I hope you're enjoying the course so far.
3
And today we're talking about action the selection policies.
4
All right let's get straight into it.
5
Previously we talked about adding a neural network to our simple learning and so far we are getting
6
quite into deep learning.
7
We've talked about the learning part quite a bit including adding some elements to it.
8
And today we're talking about this part we're talking about the acting.
9
So let's have a look.
10
So here we've got what we discussed about acting that once you input the values the parameters are the
11
vector describing the state agent is clearly in that environment then that is after all the learning
12
is done or even before the learning is done.
13
Basically we get all the q values so we're not interested in the learning right now we insist on acting
14
so once we have these key values how do we understand which one we need to use.
15
Well if you think about it.
16
Q values are simply predictions for the cube.
17
So as we did in the simple learning algorithm what did we do we just selected the one with the best
18
of the highest value.
19
Once we have the one with the highest IQ value we just take that action because it just brings us the
20
highest value and that we know that Duval's calculator's immediate reward that we expect to receive
21
Plus the DK factor times the value of the next date.
22
And it's a recursive calculation so why not why wouldn't you take the best value and that's kind of
23
the end of it.
24
But as you can see here it's not as simple here we're using a soft max function and this is where we're
25
going to talk about actual selection policies.
26
So here in reality we don't have to have just a software function.
27
We can have different action selection policies for example we've got Epsilon greedy Epsilon's soft
28
and we've got the soft Macs and those are kind of like the most commonly used action selection policies
29
of course there are others.
30
For instance the most basic one is a very simple action sociables it just select the best.
31
The one with the highest Q value.
32
But why doesn't that action pulse fly and why do we have different types of action pulse action selection
33
policies.
34
Well it all boils down to exploration versus exploitation.
35
And that is the core of reinforcement learning because we already talked about this a little bit that
36
your agent when it's operating in an environment it might predict certain queue values which might be
37
good and it might turn out great it might turn out that those are available and will be forced to explore.
38
So if we for instance in this case predict that Q2 is the best one and then it takes Q To takes action
39
to and it.
40
So from here to Section 2 and then it gets it gets a very negative reward.
41
Then the environment is forcing the agent to go and explode because now he's going to learn that oh
42
actually I thought Q2 was going to be very good but it turned out very bad.
43
So the results are not very bad.
44
So the networks can update itself so next time he's in the state he's going to probably eat my soul
45
just get to it.
46
You know like if it is a very very favorable so you might think that that's like you know you might
47
need a couple of times a couple of penalties or punishments in order to learn it is about action.
48
But maybe he'll already soon learn that I'm going to take a different action and take the wrist action
49
because now it has the best value.
50
So sometimes the environment forces the agent to take different to explore different actions but sometimes
51
the agent might get it find itself stuck in a local maximum it might find that it followed through through
52
its initial exploration and found that oh this is a pretty cool action like I'm going to go right here.
53
And that d'esprit collection.
54
But the problem is that it thinks is the best action simply because it hasn't explored is explored going
55
up his nose or going left is explore going right but it hasn't explored going down from that specific
56
state that it's in and now that it's kind of like biased towards this action and think thinks a good
57
action is going to keep taking it is going to keep getting.
58
He's going to keep taking is actually going to keep getting a good reward.
59
But what if this action would have been even better if this action would have been so much better that
60
if it knew about this action it would actually switch to this action but because it got stuck in a local
61
maximum is getting these good rewards is just going to be reinforced.
62
This is going to keep reinforcing itself that or the violence going to reinforce it that this is a good
63
action to take keep doing that.
64
But really the reality is that there's this other action that hasn't found yet or hasn't even explored.
65
That would have been much better.
66
So what we want to do is we want to come up with an actual selection policy that allows our agent not
67
to get stuck in a local maximum.
68
Yes it's important to you know keep doing the good actions that's the exploitation part.
69
We won't exploit what we've found.
70
But at the same time we still want to explore we never want to stop exploring as like in life you never
71
want to stop learning you stop learning you die.
72
That's things like that that when you're not growing you're dying or something got so you want to keep
73
learning and your agent wants to keep learning.
74
And that's where these action selection policies come in.
75
So we've got three you listed here so the first one is Epsilon greedy it's a very simple one it sounds
76
pretty complex in the sense that like it's got a cool name and usually things with surgical names.
77
It's actually not.
78
So basically what it does is it will select the one with the best Q value and epsilon like Epsilon you
79
might hear other places it's just like a selection policy.
80
So in this case we're using it to slick so our out of Al-Q values are by sales like the one with the
81
highest Q value all the time except for Epsilon percent of the time.
82
So for instance if you set epsilon to 10 percent then you're going to or 0.1 than 10 percent of the
83
time that the action is going to be selected at random.
84
So 90 percent of the time you're still going to be selecting the best action based on the highest value.
85
But 10 percent of the time is going to be selecting a random action.
86
Uniform it is going to be absolutely randomly taking an action or if you said epsilon to zero point
87
five for 0.05 that means that 95 percent of the time the agent is going to be taking the action with
88
the highest value.
89
But 5 percent of the time it's still going to be selecting and random action.
90
So it's going to be going out there and exploring.
91
So Epsilon's soft is very similar to the way that does kind of like why it's called FCL greedy because
92
then you're greedily selecting the action the good action except for that little episode.
93
Some of the time.
94
So the lower the EPS deal they'll lower the Lepp Epsilon the more greasily you're selecting that kind
95
of the action that is the optimal action and the less you're leaving the less chances you leaving for
96
exploration Epsilon's soft is the opposite.
97
So basically you're selecting at random you're selecting one minus Epsilon cent of the time.
98
So if you epsilons like 0.1 to 10 percent then only 10 percent of the time you're taking this action.
99
And 90 percent of the time you're selecting a random action.
100
So very very simple just inverted algorithms and a soft Max is kind of like the next step from or it's
101
it's a more advanced version I would say over epsilon of epsilon greedy algorithm although they both
102
have merit and they both have a place.
103
We're going to be using self-finance in our coding in our practical sort of thing.
104
So that's what we're going to talk in a bit more detail about soft max.
105
So let's have a look.
106
So let's move on to your next hopefully.
107
It's pretty clear about Ebsen agrees it's a pretty straightforward algorithm.
108
Select this one.
109
Most of the time except for sometimes go and explore.
110
And now we also see why it's important to do that exploration so that we don't end up in local maximums
111
in our in our optimization process so now we're going to talk a bit more about soft Macs.
112
There's a tutorial on soft marks at the end of the course.
113
I think it's an annex number two where we talk about the concept of Maxim's because you refresh a little
114
bit here so there we're talking about neural networks and by the way we're all going to be covering
115
convolutional.
116
We're not covering evolution neural networks in this section.
117
Of course in this section we're still using a vector.
118
But in the next section of the course when we're we're creating an AI to play Doom we are going to be
119
using convolutional neural network so it could be beneficial for you to look at in relational neural
120
networks and then take a self max function or you can learn a bit more about soft Max.
121
After you take the convolutional neural networks and of course later on.
122
But here's a quick refresher So here we've got our convolutional neural network which decides whether
123
it's a dog or cat.
124
So here we've got the voting process between these neurons and this one says that it's a it's got the
125
features you know the fluffy ears What's the pointed pointed face type of thing and the kind of the
126
features are the types of eyes the eye with eyes look all these features that belong to a dog.
127
So it's a 95 percent chance that it's a dog and the 5 percent chance that it's a cat.
128
But the question is how did we get in that Tauriel we're talking about how do we get these values to
129
add up to one.
130
Well whatever convolutional all our whole neural networks are the convolutional neural network plus
131
the fully connected Lares whatever it's bad out whatever the values that we apply to soft max function
132
are here.
133
This is where we introduced the formula for the soft next function.
134
Is what it looks like.
135
And then we got these values.
136
And so basically that's a quick refresher.
137
This is the formula for the soft Max.
138
It's what it does is it takes however many outputs you have doesn't matter.
139
It will take them and it will squash them all into values between 0 and 1 regardless of how big they
140
are just by it's for me you can see that there's a total sum at the bottom so these devices are going
141
to be zero and in.
142
And also the all these values are going to add up to one always.
143
And so that's that's very beneficial for us because when we're using the soft max function what happens
144
is we get these values we select this best view value.
145
But in reality what happens is these values that we get there are actual numbers right.
146
So this is some kind of numbers.
147
They don't have to all add up to one and don't have to be between 0 and 1.
148
Just some numbers.
149
But when we apply soft Max we don't just select the best one we actually get numbers like that so we
150
get our numbers in the range between 0 and 1 and that are also that also add up to 1.
151
And so what other thing do we know that adds up to one.
152
Well probabilities we know that probabilities always have to add up to 1 so that is why we can say here
153
we've got q values but here all of a sudden we've got soft or we've got probabilities.
154
So we can say that the likelihood of this being the best action is 90 percent.
155
This lesbian section 5 percent 2 percent 3 percent because we know the higher your value the better
156
the action.
157
So if we squash them to 0 to 1 then these become possibilities and we can deal with them as such.
158
And therefore now is when the action is selected and that's how we come up with Q2.
159
But if you look at it closely this isn't a strict 100 percent and these are not Saroo 0 percent.
160
So this is a 5 percent to 3 percent.
161
So the most natural way to apply the soft Max in order to preserve exploration in the algorithm is to
162
use these exact probabilities as how often we're going to be taking that action.
163
So these probabilities actually present the distribution of these actions that we're taking so basically
164
soft Max makes it very easy for us to come up with a way to combine exploitation and exploration.
165
So the best the best action will always have the high probability because it has highest Q value and
166
therefore here we're going to be just going to use these as our distribution or we're going to say okay
167
we're going to be taking Q2 90 percent of the time but 5 percent of the time we still get to be taking
168
Q1 and 2 percent of the time we get to 3 and 3 percent of the time we're going to be taking Q4.
169
And the beauty here is also that as these values update as and as the agent goes through the network
170
more and more and more it becomes more familiar with with the environment and therefore these updates
171
so this value for instance might become like it might might ascertain that this value is actually less
172
or this actually is higher and so these probabilities will also change as an agent goes through.
173
So even though here we've got Choo-Choo.
174
Nobody is to say that sometimes 5 percent of the time to be more precise we'll be selecting Q1 as the
175
action to take and sometimes or action one will be taking action one.
176
Sometimes will be taking action through a two action three two percent of the time and action for will
177
be taking about 3 percent.
178
So every action has a chance to play in this process as long as we have enough iterations an agent goes
179
through lots and lots of times through these states that they're in.
180
And that's that's how this that's how any kind of deep learning algorithm works that you want to do
181
this many many times so that you learn from experience and therefore as you can see here it's a very
182
natural transition to.
183
We're not just randomly like an Epson angry algorithm and not just randomly selecting the actions we're
184
selecting them based on their soft max values which makes it makes it like has some logic behind it
185
not just not just that random 10 percent of the time we're selecting a random action but there's some
186
logic behind how we're doing it and based on the key values that we've explored.
187
And so that's the action selection policy that we're going to be using in this course.
188
You're welcome to definitely check out Ebsen greedy action section Polsce if you like but we're going
189
to be predominately using the soft Max action section policy and I've got an interesting reading for
190
you.
191
So this is called adaptive Epsilon greedy exploration in reinforcement learning based on value differences
192
it's the 2010 article.
193
And it's interesting because Mike Michel I'm not sure how to pronounce Michelle and Miquel toxic introduces
194
a different type of Algren's and adjusted Epsilon greedy algorithm and called the VDB VDB algorithm
195
or epsilon greedy VDB algorithm you can see here.
196
And he actually compares compares to the Ebsen greedy and soft Max and it's an absolute greedy algorithm
197
which basically the main idea behind it is to adjust the value of epsilon depending on the state the
198
agent is in.
199
So if if the agent is very certain about the state in then Epsilon should be smaller so they should
200
be less exploration if the agent is answered Epson's should be higher should be more exploration.
201
So it is a 2010 article.
202
I'm not sure if it's if this new proposed algorithm is widely used or is as being accepted in the community
203
or or if artificial Times has kind of a way from this this suggestion.
204
But nevertheless it will definitely help you reinforce your knowledge about action selection policies
205
which we discussed the Epsom Ingredion the soft Naxal help you ill give you an opportunity to compel
206
Subha site and also see in which direction people actually think when they want to improve artificial
207
intelligence so if you're ever planning on creating really interesting algorithms that are pushing the
208
edge of Elche artificial intelligence and pushing the envelope in this space then this could be a good
209
way for you to see in which direction people think sometimes when they're trying to improve the norms
210
of artificial intelligence or the norms that existed back then in 2010.
211
So there we go.
212
Hopefully you enjoyed today's tutorial about the action selection policies and we learned about abseil
213
greedy Epson salt and the soft Macs and now you're even more prepared for the practical side of things.
214
And on that note I look forward see your next step.
215
And until then enjoy AI.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.