All language subtitles for 019 Deep Q-Learning Intuition - Acting-subtitle-en

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranî)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
uz Uzbek
vi Vietnamese
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

1

Hello and welcome back to the course on a I I in the previous part we talked about the deep learning

2

Killary intuition we started there.

3

And in fact we actually got all the way to this part and where we talked about learning and now we're

4

going to move on to the actual acting part.

5

So there's there's two parts to distinct parts that we have to remember.

6

So that's the learning part but now he actually he's done all of this.

7

That's beautiful.

8

Now he actually has to take an action he has to decide what is he going to do is going to do action

9

one two three or four.

10

And so how does he do that.

11

Well the way he does it is now given those same values so the values don't change after we've we have

12

these values of compare them with Calcott the last two by arrogated era we've updated the weights but

13

the values don't change in that whole process.

14

To have got the cube values there.

15

They're fixed.

16

We know what they are.

17

All this happens though.

18

Networks updated and out using those same values that we had.

19

What we're going to do is we're going to parse them through a soft max function.

20

And again soft Max as described.

21

I think an annex 2 and we'll talk a bit more about soft max.

22

Further down in or we'll talk about this action selection policy further down in the rest of this section.

23

So just in a few tutorials.

24

But for now we're just going to say we're passing it through a soft next function.

25

Basically what it does is it allows it helps select the best one it selects the best action possible.

26

And there's a small caveat to that.

27

It's not just the best one possible.

28

We'll talk about that in the action selection policy tutorial.

29

But for now let's just say it selects the best action from here it says OK so Q1 you know the likelihood.

30

Basically we know that q values predicted the Q value so it can look at them and say OK so the highest

31

Q value of these just as we did in the simple Q learning algorithm.

32

Ill just look at all these for say the highest values this one I'm going to select that action we're

33

going to take those.

34

And that's pretty much it.

35

That's how he chooses which action take takes takes action and then all of this process happens again.

36

For for the next stage the agent ends up in in our case and the next square of the maze.

37

But generally speaking in the next state.

38

So there we go.

39

That's how we feed in a reinforcement learning problem into a neural network through a vector describing

40

the state that we're in.

41

And once we fit it.

42

There's two parts of the process that happen Part one is the learning.

43

So remember that part where we compare each of the cube values with the target and then we back propagate

44

the loss through the network to update the weights so that our network is learning as we go through

45

this maze or through this environment.

46

And also the second part is of course we have to act we have to select an action and that is where we

47

pass the values through a soft max function and or basically an action selection policy which we'll

48

talk about further down.

49

And then we simply select the action that we want to take and we perform that action and then this whole

50

process starts again.

51

And then maybe the agent gets then maybe the agent doesn't pausa the game.

52

In any case the game ends.

53

And then once again the whole process repeats the agent plays the whole game again and then that stops

54

so basically that's that's another airpark every time the agent you know every time the game ends with

55

a favor beyond fairie that's the end of an airport.

56

And then he starts again and then he starts again and then he starts again.

57

And so on.

58

So that happens and this process happens for every single time the agent is in you in a new state so

59

the state is encoded here so that's important not just for every single game that he plays but for every

60

single state.

61

So he's in a state that goes through his process dates and so on and happens every single time.

62

And so the learning happens and the acting happens as well.

63

So that is deep learning in the intuition behind deep learning.

64

We've got lots more to cover off and then of course practical and in the meantime if you'd like to get

65

some additional information on keep learning.

66

We've got a recommended reading so we've already spoken about Arthur Giuliani's series of blog posts.

67

If you look at simple informal learning Lifton's flow part 4 you will find the part that's relevant

68

to what we discussed today.

69

Note that here he talks about convolutions we are not covering revolutions in this section we're going

70

to be talking about them in the next section of the course.

71

So the difference here is that it's just kind of skip the conclusions part for now and we'll talk about

72

them in the next part of the course but the difference is in evolutions.

73

You're like looking the agent is looking at the image and and therefore he has to process an image an

74

additional complication for now where we're slowly gradually building up to that.

75

For now we're encoding our environment through you look here we're encoding our environment or maybe

76

like look at this one probably in coding our environment as a or in to state the agent is in as a vector.

77

So in our case was very simple vector of values.

78

Sometimes people even in that in that simple may sometimes or as you'll see from this blog post.

79

Sometimes people prefer the one hot and coded version of that state.

80

So basically where every single box of the maze has a.

81

So you have like a vector of for a null case would be 12 values three by four.

82

So it isn't like either either 1 or 0 depending on which elements and which box you're in.

83

In the environment.

84

So in whichever way you decide to code your environment and the state of your environment that's how

85

in coding It's basically a vector.

86

The key here is that it's not a convolution So it's not like an image and there's no convolution volt

87

So this part will come later.

88

For us it starts over here and that just simplifies the process for us to gradually understand better.

89

And of course don't forget that this post is rude and tends to flow and we're using pi torche in our

90

tutorials.

91

So hopefully you enjoy this.

92

A quick intro into a deep convolutional deep not yet deep book learning.

93

And on that note I look forward to seeing you next.

94

And until then enjoy artificial intelligence.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.