All language subtitles for 029 Object Detection - Step 3-en

af Afrikaans
ak Akan
sq Albanian
am Amharic
ar Arabic
hy Armenian
az Azerbaijani
eu Basque
be Belarusian
bem Bemba
bn Bengali
bh Bihari
bs Bosnian
br Breton
bg Bulgarian
km Cambodian
ca Catalan
ceb Cebuano
chr Cherokee
ny Chichewa
zh-CN Chinese (Simplified)
zh-TW Chinese (Traditional)
co Corsican
hr Croatian
cs Czech
da Danish
nl Dutch
en English
eo Esperanto
et Estonian
ee Ewe
fo Faroese
tl Filipino
fi Finnish
fr French
fy Frisian
gaa Ga
gl Galician
ka Georgian
de German
el Greek
gn Guarani
gu Gujarati
ht Haitian Creole
ha Hausa
haw Hawaiian
iw Hebrew
hi Hindi
hmn Hmong
hu Hungarian
is Icelandic
ig Igbo
id Indonesian
ia Interlingua
ga Irish
it Italian
ja Japanese
jw Javanese
kn Kannada
kk Kazakh
rw Kinyarwanda
rn Kirundi
kg Kongo
ko Korean
kri Krio (Sierra Leone)
ku Kurdish
ckb Kurdish (Soranรฎ)
ky Kyrgyz
lo Laothian
la Latin
lv Latvian
ln Lingala
lt Lithuanian
loz Lozi
lg Luganda
ach Luo
lb Luxembourgish
mk Macedonian
mg Malagasy
ms Malay
ml Malayalam
mt Maltese
mi Maori
mr Marathi
mfe Mauritian Creole
mo Moldavian
mn Mongolian
my Myanmar (Burmese)
sr-ME Montenegrin
ne Nepali
pcm Nigerian Pidgin
nso Northern Sotho
no Norwegian
nn Norwegian (Nynorsk)
oc Occitan
or Oriya
om Oromo
ps Pashto
fa Persian
pl Polish
pt-BR Portuguese (Brazil)
pt Portuguese (Portugal)
pa Punjabi
qu Quechua
ro Romanian
rm Romansh
nyn Runyakitara
ru Russian
sm Samoan
gd Scots Gaelic
sr Serbian
sh Serbo-Croatian
st Sesotho
tn Setswana
crs Seychellois Creole
sn Shona
sd Sindhi
si Sinhalese
sk Slovak
sl Slovenian
so Somali
es Spanish
es-419 Spanish (Latin American)
su Sundanese
sw Swahili
sv Swedish
tg Tajik
ta Tamil
tt Tatar
te Telugu
th Thai
ti Tigrinya
to Tonga
lua Tshiluba
tum Tumbuka
tr Turkish
tk Turkmen
tw Twi
ug Uighur
uk Ukrainian
ur Urdu
uz Uzbek
cy Welsh
wo Wolof
xh Xhosa
yi Yiddish
yo Yoruba
zu Zulu

Original subtitles

1

Hello and welcome to this new tutorial today in this to all we are going to define the detect function

2

that will do the detections exactly like what we did for open city only this time it is not going to

3

be based on open city.

4

It's going to be based on deep learning because we're going to do that detection through the SSD model

5

single shot multi-book detection.

6

So excited to start because this time we're really taking things at the next level.

7

All right so let's do this.

8

The first thing very important to understand is that exactly like before we are going to do a frame

9

by frame detection that is that the detect function that we're about to implement will work on single

10

images it will not do the detection on the video directly it will do the detection on each single image

11

of the video and then using some tricks with actually image IO we will manage to extract all the frames

12

of the video apply the detect function on the frames and then reassemble the whole thing to make the

13

video with the rectangles detecting the dog and the humans.

14

And Carol Carol is actually going to be detected on just one frame you'll see.

15

But anyway that's frame by frame detection.

16

That's the first thing important to understand.

17

And now let's start to implement that function.

18

So as usual we start with that to define a new function that we need to give a name to these functions

19

we are going to call it detect.

20

Just like that.

21

And now we need to specify the arguments that this function is going to take.

22

So this function detector is going to take three arguments.

23

The first one is the image on which the detect function is going to be applied to detect the object.

24

That's our first argument and we're going to call it Freyne.

25

Now as opposed to before with open C we only need to input as arguments the frame and not the great

26

image.

27

We only need to put the original image here the frame the original frame in color and not the grayscale

28

because remember we had to do this before because Open City only works on great images.

29

But here we're doing something totally different and therefore we don't have to take the black and white

30

version of the frame.

31

All right so that's our first argument then the second argument will be net which will be the SS The

32

neural network the single shots multi-book detection neural network and then the third and last argument

33

is transform because there is going to be some transformations applied to the image but not to put them

34

in black and white just to make sure that the images are compatible with the new one that work you know

35

the images will be the input of the neural network and therefore they have to have a certain format

36

and this transform argument that is the final argument of this detect function will transform the images

37

so that they have the right format to get into the new network.

38

All right.

39

And then it's important to understand what this function will do.

40

So as you might have guessed it will do the detections on the images the single images one by one.

41

But what exactly is this function going to return.

42

Well it will simply return this same frame but with the rectangle detecting the objects the dog and

43

the humans and not only will there be direct Englebert Also there will be the label on the rectangle

44

So we'll see some rectangles with the label humans because there are several persons on the video and

45

one rectangle with the label Doug.

46

So there will be everything that will be totally clear and amazing detection on the video.

47

I can't wait to show you this.

48

All right.

49

So there are three arguments and now we're ready to go inside the function to define what we wanted

50

to do.

51

All right so the first thing we have to do is to get the height and the weight of the image and the

52

frame the frame that is the argument here on which the function is applied.

53

So we're going to introduce two new variables.

54

Hide and with right and to get these height and width of the frame we're going to do that very efficiently.

55

We're going to get it from our frame obviously because this is some information specific to the frame.

56

And then this frame has some attributes.

57

One of them is shape and shape is actually an attribute that returns a vector of three elements.

58

The first one is the height of the frame.

59

The second one therefore of index One is the width of the frame and is there one of index 2 is the number

60

of channels.

61

So the number of channels means that if you have a black and white image you will have one channel and

62

if you have a color image you will have three channels for red blue and green.

63

But we just want the height and width.

64

So we're going to get the first index which is zero corresponding to the height and the second index

65

one corresponding to the width.

66

But the correct way to do this is actually to type a colon here and then to because that means we're

67

taking the range from zero to two but with two excluded.

68

So we're just taking zero and 1 and therefore we're taking the hide and do it.

69

All right.

70

So that's the first thing we had to do.

71

And now we're going to do several transformations to go from the original image which is our friend

72

right now to a torch variable that will be accepted into the as is the neural network.

73

So there is a series of transformations to do before getting to this torche viable.

74

The first one is to apply the transform transformation to make sure that the image has the right format.

75

That is the right dimensions and the right color values.

76

That's the first transformation we need to make.

77

Once we have done this transformation then we will need to convert this transform frame from a number

78

array because it will still be an entire array from a number of array to a torch tensor.

79

That's not the same.

80

That sensor is a more advanced matrix A more advanced array and therefore that's exactly what is our

81

second transformation.

82

We convert the transform frame from an umpire array to a torch tensor then that's not all we will need

83

to do a third transformation which will be to add a fake dimension to the torch sensor and that fig

84

damage and will respond to the batch.

85

And then finally the fourth and final transformation to do before it is ready to go into the new one

86

that work will be to convert it into a torch variable.

87

Remember this variable closets we imported here is a class that converts a torch sensor into a torch

88

variable that contains both the tensor and a gradient.

89

And this torch variable will then be an element of the dynamic graph which will allow us later to do

90

some very fast and efficient computation of the gradients during backward propagation.

91

So there we go we have four transformations to make.

92

We'll make them in the next tutorial starting with the first one.

93

And then finally we'll be able to feed the neural network with this image.

94

So let's match this.

95

And until then enjoy computer vision.

Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.