Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome to this new tutorial today in this to all we are going to define the detect function
2
that will do the detections exactly like what we did for open city only this time it is not going to
3
be based on open city.
4
It's going to be based on deep learning because we're going to do that detection through the SSD model
5
single shot multi-book detection.
6
So excited to start because this time we're really taking things at the next level.
7
All right so let's do this.
8
The first thing very important to understand is that exactly like before we are going to do a frame
9
by frame detection that is that the detect function that we're about to implement will work on single
10
images it will not do the detection on the video directly it will do the detection on each single image
11
of the video and then using some tricks with actually image IO we will manage to extract all the frames
12
of the video apply the detect function on the frames and then reassemble the whole thing to make the
13
video with the rectangles detecting the dog and the humans.
14
And Carol Carol is actually going to be detected on just one frame you'll see.
15
But anyway that's frame by frame detection.
16
That's the first thing important to understand.
17
And now let's start to implement that function.
18
So as usual we start with that to define a new function that we need to give a name to these functions
19
we are going to call it detect.
20
Just like that.
21
And now we need to specify the arguments that this function is going to take.
22
So this function detector is going to take three arguments.
23
The first one is the image on which the detect function is going to be applied to detect the object.
24
That's our first argument and we're going to call it Freyne.
25
Now as opposed to before with open C we only need to input as arguments the frame and not the great
26
image.
27
We only need to put the original image here the frame the original frame in color and not the grayscale
28
because remember we had to do this before because Open City only works on great images.
29
But here we're doing something totally different and therefore we don't have to take the black and white
30
version of the frame.
31
All right so that's our first argument then the second argument will be net which will be the SS The
32
neural network the single shots multi-book detection neural network and then the third and last argument
33
is transform because there is going to be some transformations applied to the image but not to put them
34
in black and white just to make sure that the images are compatible with the new one that work you know
35
the images will be the input of the neural network and therefore they have to have a certain format
36
and this transform argument that is the final argument of this detect function will transform the images
37
so that they have the right format to get into the new network.
38
All right.
39
And then it's important to understand what this function will do.
40
So as you might have guessed it will do the detections on the images the single images one by one.
41
But what exactly is this function going to return.
42
Well it will simply return this same frame but with the rectangle detecting the objects the dog and
43
the humans and not only will there be direct Englebert Also there will be the label on the rectangle
44
So we'll see some rectangles with the label humans because there are several persons on the video and
45
one rectangle with the label Doug.
46
So there will be everything that will be totally clear and amazing detection on the video.
47
I can't wait to show you this.
48
All right.
49
So there are three arguments and now we're ready to go inside the function to define what we wanted
50
to do.
51
All right so the first thing we have to do is to get the height and the weight of the image and the
52
frame the frame that is the argument here on which the function is applied.
53
So we're going to introduce two new variables.
54
Hide and with right and to get these height and width of the frame we're going to do that very efficiently.
55
We're going to get it from our frame obviously because this is some information specific to the frame.
56
And then this frame has some attributes.
57
One of them is shape and shape is actually an attribute that returns a vector of three elements.
58
The first one is the height of the frame.
59
The second one therefore of index One is the width of the frame and is there one of index 2 is the number
60
of channels.
61
So the number of channels means that if you have a black and white image you will have one channel and
62
if you have a color image you will have three channels for red blue and green.
63
But we just want the height and width.
64
So we're going to get the first index which is zero corresponding to the height and the second index
65
one corresponding to the width.
66
But the correct way to do this is actually to type a colon here and then to because that means we're
67
taking the range from zero to two but with two excluded.
68
So we're just taking zero and 1 and therefore we're taking the hide and do it.
69
All right.
70
So that's the first thing we had to do.
71
And now we're going to do several transformations to go from the original image which is our friend
72
right now to a torch variable that will be accepted into the as is the neural network.
73
So there is a series of transformations to do before getting to this torche viable.
74
The first one is to apply the transform transformation to make sure that the image has the right format.
75
That is the right dimensions and the right color values.
76
That's the first transformation we need to make.
77
Once we have done this transformation then we will need to convert this transform frame from a number
78
array because it will still be an entire array from a number of array to a torch tensor.
79
That's not the same.
80
That sensor is a more advanced matrix A more advanced array and therefore that's exactly what is our
81
second transformation.
82
We convert the transform frame from an umpire array to a torch tensor then that's not all we will need
83
to do a third transformation which will be to add a fake dimension to the torch sensor and that fig
84
damage and will respond to the batch.
85
And then finally the fourth and final transformation to do before it is ready to go into the new one
86
that work will be to convert it into a torch variable.
87
Remember this variable closets we imported here is a class that converts a torch sensor into a torch
88
variable that contains both the tensor and a gradient.
89
And this torch variable will then be an element of the dynamic graph which will allow us later to do
90
some very fast and efficient computation of the gradients during backward propagation.
91
So there we go we have four transformations to make.
92
We'll make them in the next tutorial starting with the first one.
93
And then finally we'll be able to feed the neural network with this image.
94
So let's match this.
95
And until then enjoy computer vision.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.