Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranî)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome back to the course on artificial intelligence.
2
I hope you're excited about today's tutorial because we are taking our very first step into the world
3
the I.
4
And today we're talking about reinforcement learning.
5
It's a very important story because it will underpin everything else is going to be happen in this course.
6
So let's get started here.
7
We've got a little maze and this maze is our representation of an environment and that's what we're
8
going to be dealing with in this course.
9
We're going to be dealing with certain environments in which our artificial intelligence is going to
10
be performing it's going to be taking actions it's going to be looking to beat these in my going she'll
11
be looking to win in these environments.
12
And here we've got an agent.
13
The agent is our artificial intelligence.
14
That's the person or that's the mind that's going to be navigating these environments and learning from
15
the feedback that their minds are going to be giving it in order to perform certain actions.
16
And so the way it works is the agent perform certain actions in this environment.
17
And as a result the state in which it is in will change so it might be further or closer or more to
18
the left more to the right.
19
It might have sort of the other parameters that describe it state and those parameters.
20
So the state is going to change because of the action takes and it will also get rewards based on the
21
action.
22
So every time it takes an action the state will change and it'll get reward.
23
Now bear in mind sometimes it might happen that it won't change the state the action won't change a
24
stay or there won't be a reward for taking that action.
25
In that sense it was.
26
But nevertheless the agent's going to keep doing that was going to be taking actions cheating the state
27
getting rewards changing action taking actions changing the state and getting rewards.
28
And by doing that process it's going to be learning about what was going to be exploring the environment
29
understanding what actions lead to good rewards and favorable states and what actions the two rewards
30
an unfavorable state.
31
And this is a very simplistic representational very global problem.
32
So if you think about it environments actually don't have to be just mazes.
33
It's not just about getting out of a maze or finding a treasure in a maze.
34
An environment can be pretty much anything in life.
35
So imagine you waking up in the morning and cooking an omelette.
36
So in order to make that omelet you need to go through certain steps you need to get the salt get the
37
eggs get the frying pans which to fire on and so on and it does sound like a routine mundane thing.
38
But it's become routine because you've done it so many times.
39
But in reality it's an environment where you're performing certain actions you're taking that you putting
40
the fire on you putting a frying pan on the fire you're putting all the eggs into the frying pan and
41
you put some salt on the eggs and you're turning over and so on.
42
So as you can see they are CRN actions actions which are taking in certain states and those actions
43
lead to certain other states and sometimes reward.
44
So for instance when you put the fire on and you wait wait wait wait wait you take an action of wait
45
wait wait wait too long and then you put the eggs in into the frying pan.
46
The rewards are going to be very negative.
47
It's all going to burn.
48
On the other hand if you do all the all the correct actions in the correct time so it's also very important
49
to understand that actions should be taken at the correct points in time.
50
So for instance putting the salt in the frying pan before you put the eggs in might not be the best
51
idea.
52
You might want to take that action of putting the salt into the frying pan after the eggs are in there
53
so that in a different state.
54
So it's important to remember that.
55
And at the same time so if you take all the correct actions in the correct order in the correct states
56
your final reward could be that you get an omelet which you can eat.
57
And so that's a very basic activity in your life but if you think about it it is actually an environment
58
and you are the agent going through this environment and perform a task you don't really need to learn
59
anything because you already know it pretty well.
60
But at same time you could learn maybe you could learn how to make a better omelet or especially if
61
it's your first omelet that you're making you're probably going to screw it up.
62
But you will learn from that because you will understand what actions lead towards states and routes
63
and anything else in life.
64
For instance even trading on the stock market and you know buying and selling and getting certain feedback
65
from the market in the sense of return positive or negative returns.
66
That's also an environment that's you participating in that environment as an aged.
67
Driving a car is also an environment where you can turn the steering wheel you can accelerate you can
68
break and so on and you're getting feedback from the environment and you know one of those feedbacks
69
is the policeman giving you a speeding fine if you're going above the acceptable or allowed speed limit
70
on that highway.
71
And therefore from there you learn that that's not something that should be done because it leads to
72
a negative reward.
73
So rewards don't have to be just at the very end of the process.
74
They can be throughout the journey throughout the process.
75
So those are a couple of examples.
76
And in terms of a I the simplest way to think of reinforcement learning is like training a dog when
77
you train the dog you to give it certain commands and if it obeys those commands then you give it a
78
reach you give it like a biscuit or something if it doesn't Abeles Kamaz you tell it that it's a bad
79
dog or you just don't give it a treat.
80
And through that process it learns what certain commands or what it needs to do what action it needs
81
to take in certain states and the states are the commands that you're giving it.
82
And based on that it will get some certain rewards of course in the world of AI.
83
It's not that complex.
84
You don't have to give the treats.
85
You don't have to have like a bag of biscuits with you every time you just give it a plus one or a minus
86
one so it's a huge advantage that in the world of AI we've created these AIs ourselves.
87
So the rewards that we're giving them if you think wow this is really cool rewards are giving them they
88
don't actually exist they're just a plus or minus one or plus a one or a zero or something.
89
So it's all nonexistence all imaginary stuff.
90
But at the same time it leads to great results as we can create these amazing things these amazing artificial
91
intelligence as by this amazing artificial intelligence by just providing rewards we don't really exist.
92
Plus and minus one doesn't cost anything but same time release results.
93
So very similar to real world.
94
And you know for example Dokes But here the rewards are digital and just numbers.
95
And with that in mind we can talk about about robot dogs I love this example so this is just around
96
in pictures not necessarily that exact robot dog you know that is trained through reinforcement learning
97
some of the robot dogs especially the older ones you'd have an algorithm in there.
98
And this is actually a good example of the difference between preprogramed agents and reinforcement
99
learning agent so you could have a robot dog which is preprogrammed to how to walk it will say.
100
So in the in the algorithm behind the dog in the software will say OK so in order to walk you need to
101
move your left leg forward left front leg forward then your back right leg forward then your front right
102
leg forward then your back left leg forward and repeat that action and you know that's that's the definition
103
of walking is a function inside this dog.
104
And then it might have you know how to sit how to stand and things like that.
105
Whereas in a robot dog that is trained through reinforcement learning what happens is you don't preprogram
106
it.
107
This is the key concept to everything here that you don't have any algorithm inside that is hard coded
108
into the dog.
109
Instead you have what we'll be discussing in the future.
110
You have this reinforcement learning algorithm which is told that OK so the goal is from to get from
111
where you are now not knowing anything to that to the end of the room for example.
112
And here are the certain actions you can take.
113
You can move your right foot you can move your left foot you can move your right back foot you are left
114
back foot so here all the degrees of freedom you can do you can move it like this you can move like
115
that so like a list of actions you can take and your rewards are every time you take a step forward
116
you get a plus one every time you fall over.
117
You get a minus one and that's all there is to it.
118
And then they just leave the dog and let it figure it out on its own.
119
So the dog tries to stand up it falls then it realizes that OK I shouldn't do that action that led to
120
me falling because every time I fall I get a minus one which is not good for me then.
121
So does the other action that helped him stand up and then it figures are just experiments experiments
122
experiments tri's things randomly and then figures out that it can make a step forward by moving its
123
right front foot and he gets a plus one and realize oh I should do more of that.
124
OK cool so it now learns that it should do more of this and less of that.
125
And through this learning process it quickly very quickly understands how it can walk.
126
And those those dogs that figured out on their own can actually sometimes walk better than dogs that
127
are preprogramed because really preprogrammed things we look at the real life dogs and or you know we
128
use our own imagination how to do it whereas a reinforcement learning dog can optimize things on its
129
own.
130
And because in AI sometimes it can get even better results.
131
And that's how they can train these robot.
132
The same robot dogs to play soccer.
133
You can train a normal dog to play soccer because you know simply the whole approach is different.
134
And it's not something that you know probably a normal dog has been trained to do or has ever done in
135
its process of its evolution.
136
Whereas a reinforcement learning robot dogs can very easily understand how to play soccer as long as
137
you tell them what the rewards are what the goals are what the possible actions they can take.
138
So that is how reinforcement learning works.
139
In general there's a quick overview of reinforcement learning.
140
I hope that got you very excited about was going to come next because it's a completely different world
141
compared to preprogram solutions a hard program hardcoded solutions where you have the if else conditions.
142
This is very different.
143
And we're going to be talking more about that.
144
In the meantime we've got some additional reading for you so if you'd like to have some supporting materials
145
Here's a great article which you can look and look into.
146
It's called simple reinforcement learning with tensor flow.
147
It's got ten parts.
148
The link is here and you'll find the full clickable link on.
149
In the course of resources by Arthur Giuliani's 2016 article and you can follow along this course and
150
also get additional information from that article.
151
But bear in mind that that article is tends to flow where as in this course we are using pi torche so
152
different implementation but implantations but at the same time you might pick up a few things here
153
and there that might supplement your learning that we're going to be doing in this course.
154
So great articles follow you in if you're considering following it for sure.
155
Still just in case.
156
Check out that that first part and see if you like it see if you'd like to read it a bit more.
157
And then we've got specific to this tutorial a border enforcement learning there's a paper by Richard
158
Sutton which is called reinforcement learning.
159
One introduction is the 1998 papers are quite old but at the same time you can learn a bit about reinforcement
160
learning some of the examples like that omlet example and other examples of where reinforcement learning
161
can be applied and just a general overview of reinforcement learning.
162
If you are looking for some additional reading and on that note we're going to wrap up this tutorial.
163
Can't wait to see you next time.
164
And until then enjoy AI.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.