All language subtitles for 007 What is reinforcement learning-subtitle-en

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranî)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

1

Hello and welcome back to the course on artificial intelligence.

2

I hope you're excited about today's tutorial because we are taking our very first step into the world

3

the I.

4

And today we're talking about reinforcement learning.

5

It's a very important story because it will underpin everything else is going to be happen in this course.

6

So let's get started here.

7

We've got a little maze and this maze is our representation of an environment and that's what we're

8

going to be dealing with in this course.

9

We're going to be dealing with certain environments in which our artificial intelligence is going to

10

be performing it's going to be taking actions it's going to be looking to beat these in my going she'll

11

be looking to win in these environments.

12

And here we've got an agent.

13

The agent is our artificial intelligence.

14

That's the person or that's the mind that's going to be navigating these environments and learning from

15

the feedback that their minds are going to be giving it in order to perform certain actions.

16

And so the way it works is the agent perform certain actions in this environment.

17

And as a result the state in which it is in will change so it might be further or closer or more to

18

the left more to the right.

19

It might have sort of the other parameters that describe it state and those parameters.

20

So the state is going to change because of the action takes and it will also get rewards based on the

21

action.

22

So every time it takes an action the state will change and it'll get reward.

23

Now bear in mind sometimes it might happen that it won't change the state the action won't change a

24

stay or there won't be a reward for taking that action.

25

In that sense it was.

26

But nevertheless the agent's going to keep doing that was going to be taking actions cheating the state

27

getting rewards changing action taking actions changing the state and getting rewards.

28

And by doing that process it's going to be learning about what was going to be exploring the environment

29

understanding what actions lead to good rewards and favorable states and what actions the two rewards

30

an unfavorable state.

31

And this is a very simplistic representational very global problem.

32

So if you think about it environments actually don't have to be just mazes.

33

It's not just about getting out of a maze or finding a treasure in a maze.

34

An environment can be pretty much anything in life.

35

So imagine you waking up in the morning and cooking an omelette.

36

So in order to make that omelet you need to go through certain steps you need to get the salt get the

37

eggs get the frying pans which to fire on and so on and it does sound like a routine mundane thing.

38

But it's become routine because you've done it so many times.

39

But in reality it's an environment where you're performing certain actions you're taking that you putting

40

the fire on you putting a frying pan on the fire you're putting all the eggs into the frying pan and

41

you put some salt on the eggs and you're turning over and so on.

42

So as you can see they are CRN actions actions which are taking in certain states and those actions

43

lead to certain other states and sometimes reward.

44

So for instance when you put the fire on and you wait wait wait wait wait you take an action of wait

45

wait wait wait too long and then you put the eggs in into the frying pan.

46

The rewards are going to be very negative.

47

It's all going to burn.

48

On the other hand if you do all the all the correct actions in the correct time so it's also very important

49

to understand that actions should be taken at the correct points in time.

50

So for instance putting the salt in the frying pan before you put the eggs in might not be the best

51

idea.

52

You might want to take that action of putting the salt into the frying pan after the eggs are in there

53

so that in a different state.

54

So it's important to remember that.

55

And at the same time so if you take all the correct actions in the correct order in the correct states

56

your final reward could be that you get an omelet which you can eat.

57

And so that's a very basic activity in your life but if you think about it it is actually an environment

58

and you are the agent going through this environment and perform a task you don't really need to learn

59

anything because you already know it pretty well.

60

But at same time you could learn maybe you could learn how to make a better omelet or especially if

61

it's your first omelet that you're making you're probably going to screw it up.

62

But you will learn from that because you will understand what actions lead towards states and routes

63

and anything else in life.

64

For instance even trading on the stock market and you know buying and selling and getting certain feedback

65

from the market in the sense of return positive or negative returns.

66

That's also an environment that's you participating in that environment as an aged.

67

Driving a car is also an environment where you can turn the steering wheel you can accelerate you can

68

break and so on and you're getting feedback from the environment and you know one of those feedbacks

69

is the policeman giving you a speeding fine if you're going above the acceptable or allowed speed limit

70

on that highway.

71

And therefore from there you learn that that's not something that should be done because it leads to

72

a negative reward.

73

So rewards don't have to be just at the very end of the process.

74

They can be throughout the journey throughout the process.

75

So those are a couple of examples.

76

And in terms of a I the simplest way to think of reinforcement learning is like training a dog when

77

you train the dog you to give it certain commands and if it obeys those commands then you give it a

78

reach you give it like a biscuit or something if it doesn't Abeles Kamaz you tell it that it's a bad

79

dog or you just don't give it a treat.

80

And through that process it learns what certain commands or what it needs to do what action it needs

81

to take in certain states and the states are the commands that you're giving it.

82

And based on that it will get some certain rewards of course in the world of AI.

83

It's not that complex.

84

You don't have to give the treats.

85

You don't have to have like a bag of biscuits with you every time you just give it a plus one or a minus

86

one so it's a huge advantage that in the world of AI we've created these AIs ourselves.

87

So the rewards that we're giving them if you think wow this is really cool rewards are giving them they

88

don't actually exist they're just a plus or minus one or plus a one or a zero or something.

89

So it's all nonexistence all imaginary stuff.

90

But at the same time it leads to great results as we can create these amazing things these amazing artificial

91

intelligence as by this amazing artificial intelligence by just providing rewards we don't really exist.

92

Plus and minus one doesn't cost anything but same time release results.

93

So very similar to real world.

94

And you know for example Dokes But here the rewards are digital and just numbers.

95

And with that in mind we can talk about about robot dogs I love this example so this is just around

96

in pictures not necessarily that exact robot dog you know that is trained through reinforcement learning

97

some of the robot dogs especially the older ones you'd have an algorithm in there.

98

And this is actually a good example of the difference between preprogramed agents and reinforcement

99

learning agent so you could have a robot dog which is preprogrammed to how to walk it will say.

100

So in the in the algorithm behind the dog in the software will say OK so in order to walk you need to

101

move your left leg forward left front leg forward then your back right leg forward then your front right

102

leg forward then your back left leg forward and repeat that action and you know that's that's the definition

103

of walking is a function inside this dog.

104

And then it might have you know how to sit how to stand and things like that.

105

Whereas in a robot dog that is trained through reinforcement learning what happens is you don't preprogram

106

it.

107

This is the key concept to everything here that you don't have any algorithm inside that is hard coded

108

into the dog.

109

Instead you have what we'll be discussing in the future.

110

You have this reinforcement learning algorithm which is told that OK so the goal is from to get from

111

where you are now not knowing anything to that to the end of the room for example.

112

And here are the certain actions you can take.

113

You can move your right foot you can move your left foot you can move your right back foot you are left

114

back foot so here all the degrees of freedom you can do you can move it like this you can move like

115

that so like a list of actions you can take and your rewards are every time you take a step forward

116

you get a plus one every time you fall over.

117

You get a minus one and that's all there is to it.

118

And then they just leave the dog and let it figure it out on its own.

119

So the dog tries to stand up it falls then it realizes that OK I shouldn't do that action that led to

120

me falling because every time I fall I get a minus one which is not good for me then.

121

So does the other action that helped him stand up and then it figures are just experiments experiments

122

experiments tri's things randomly and then figures out that it can make a step forward by moving its

123

right front foot and he gets a plus one and realize oh I should do more of that.

124

OK cool so it now learns that it should do more of this and less of that.

125

And through this learning process it quickly very quickly understands how it can walk.

126

And those those dogs that figured out on their own can actually sometimes walk better than dogs that

127

are preprogramed because really preprogrammed things we look at the real life dogs and or you know we

128

use our own imagination how to do it whereas a reinforcement learning dog can optimize things on its

129

own.

130

And because in AI sometimes it can get even better results.

131

And that's how they can train these robot.

132

The same robot dogs to play soccer.

133

You can train a normal dog to play soccer because you know simply the whole approach is different.

134

And it's not something that you know probably a normal dog has been trained to do or has ever done in

135

its process of its evolution.

136

Whereas a reinforcement learning robot dogs can very easily understand how to play soccer as long as

137

you tell them what the rewards are what the goals are what the possible actions they can take.

138

So that is how reinforcement learning works.

139

In general there's a quick overview of reinforcement learning.

140

I hope that got you very excited about was going to come next because it's a completely different world

141

compared to preprogram solutions a hard program hardcoded solutions where you have the if else conditions.

142

This is very different.

143

And we're going to be talking more about that.

144

In the meantime we've got some additional reading for you so if you'd like to have some supporting materials

145

Here's a great article which you can look and look into.

146

It's called simple reinforcement learning with tensor flow.

147

It's got ten parts.

148

The link is here and you'll find the full clickable link on.

149

In the course of resources by Arthur Giuliani's 2016 article and you can follow along this course and

150

also get additional information from that article.

151

But bear in mind that that article is tends to flow where as in this course we are using pi torche so

152

different implementation but implantations but at the same time you might pick up a few things here

153

and there that might supplement your learning that we're going to be doing in this course.

154

So great articles follow you in if you're considering following it for sure.

155

Still just in case.

156

Check out that that first part and see if you like it see if you'd like to read it a bit more.

157

And then we've got specific to this tutorial a border enforcement learning there's a paper by Richard

158

Sutton which is called reinforcement learning.

159

One introduction is the 1998 papers are quite old but at the same time you can learn a bit about reinforcement

160

learning some of the examples like that omlet example and other examples of where reinforcement learning

161

can be applied and just a general overview of reinforcement learning.

162

If you are looking for some additional reading and on that note we're going to wrap up this tutorial.

163

Can't wait to see you next time.

164

And until then enjoy AI.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.