Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romanian
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
1
Hello and welcome to the last step of this module to object detection before we start the homework.
2
But here we go with the last step and this last step is going to be about the training of the SSD.
3
So we're not going to implement the whole code that we'll try and the SSD simply because it's a huge
4
code.
5
And besides you're going to see that you're going to need specific requirements to be able to run that
6
training.
7
So this tutorial is going to be about explaining how you could train the SSD.
8
So I'm going to show you the file how to execute it.
9
And mostly I'm going to show you the data set on which the SSD was trained.
10
So in case you want to train the SSD with some other object you would understand the approach of how
11
to do it.
12
But remember you're going to see that you're going to need some pretty advanced system.
13
So as you can see right now I'm on a web page and this is the web page that contains the data sets on
14
which the SS was trained.
15
This dataset is not image net.
16
If you were hoping for that but the Pascal visual object classes dataset.
17
So if you remember in the code we saw Virk the OSI you know for classes to have the mapping between
18
the classes and integers.
19
Well Varg stands for visual object classes and this is already a huge dataset not as huge as image net
20
but a huge one that you will see contains lots and lots of images.
21
So you can find this page at this address but don't worry the next it will be an article which you will
22
find the homework folder that will contain all the tools to train the SSD Plus the address of this website
23
so that you can have a look at this better.
24
All right so let's go through this web page so as you can see the Pascals Virk project provides standardized
25
image data sets for object recognition.
26
So exactly what we're doing and besides provides a common set of tools for accessing the data set and
27
annotations.
28
So it's not only a data set it's also an advanced structure of data set that helps a lot for the training.
29
And also you can see that there were some challenges but now finished.
30
But anyway it is definitely useful data set to train the SSD.
31
And so now if we have a closer look at what these days that are well we need to scroll down to that.
32
What challenges from 2005 to 2012 and here are the data sets you can see there's several of them from
33
2005 to 2012 and the ones we have the ones on which the SSD was trans and I'm going to show you the
34
code and the data sets themselves and how you can execute the code to train the SSD on these two data
35
set.
36
These two data sets are the Vark 2007 and the Vok 2012.
37
So let's have a closer look at where they are right now and we need to scroll down even more.
38
Actually at the bottom of the page there we go.
39
So here you have this table that's the most important part of this page.
40
That's where you can see how the data sets are structured and made and what they contain exactly.
41
So let's have a look at the first one.
42
Remember we have the 2007 data set and 2012 sort of the 2007 data set contains 20 classes.
43
Remember I told you that the pre-trained the model we have can detect between 30 to 40 object.
44
Well among these objects we have some persons Well the S's can detect any person in a video or in some
45
images then several animals some birds cats cows dogs horses sheeps some vehicles airplanes bicycles
46
boats buses cars you can do some computer vision for some driving car if you have a self-driving car
47
business I would highly recommend to integrate as is the model for the computer vision inside yourself
48
in cars are also motorbikes and trains of course.
49
And then we have some other objects like some bottles chairs dining table potted plants sofas and TV
50
monitors.
51
All right.
52
And I think we have even more but there you go.
53
You can see you have plenty of objects to detect.
54
So really really useful.
55
And now if we go to the 2012 dataset of 2012 we can see we also have 20 classes and the data contains
56
eleven thousand five hundred and thirty images containing 27 450 animated objects so we're going to
57
have a look at these objects specifically because as I told you I've prepared for you a folder to training
58
as is the folder that contains not only the code but also these two data sets.
59
We're going to have a clear look at images.
60
So we're going to look at some of them.
61
And so what I'm going to do right now is walk you through this folder that I've prepared.
62
So let's open Anaconda up.
63
There we go.
64
We need to connect the virtual platform of course because the training is made with by torch again and
65
some other tools that are preinstalled individual them.
66
So there we go.
67
Let's not forget to connect to the virtual platform applications on virtual platform.
68
There we go.
69
And now we launch spider so spiders launching and I'm going to show you this training as is the four
70
that I've prepared for you.
71
I put it in the module to object detection color which I'm going to go to right now so I'll explore
72
it on my desktop that's the computer vision it is that folder and if we go to module 2 you'll find a
73
new folder.
74
This folder training as is d that contains all the tools and data sets for you to train the SSD.
75
But I have to specify something very important now.
76
Remember at the beginning of this oil I told you that you need an advanced system to train the SSD while
77
this advance system is simply to have kuda on your machine.
78
Kuda is like an accelerator that is connected to every geographic cards and that allows to speed up
79
considerably the training computations.
80
Therefore the training implementations that were made are only compatible with kuda.
81
I'm going to tell you a fun fact now if there was an implementation of the training implemented without
82
kuda I can tell you for sure the training would take several months or even years.
83
This is not a joke.
84
The training of the SSD is first of all made with thousands and thousands of images.
85
But if you train that without kuda it would definitely take several months.
86
You can not trained the SSD on your CPQ you have to train it with a GP you and the best you can find
87
is kuda.
88
So just to set things clear this implementation is only kuda compatible and therefore you will only
89
be able to execute the code if you have an NDA Geographic card with kuda and labels on your machine.
90
But even if that's the case.
91
While it would take some time and you would not have your SSD trained in the next hour.
92
But I think it's still interesting for you to see how you can run the training and therefore that's
93
exactly what I'm going to show you right now.
94
But first let me walk you through this training SSD folder starting with the data so you can see we
95
have this data folder that contains several files.
96
Here you can have the scripts which are shellscript and which allow to download the datasets Vogue 2007
97
and Vok 2012.
98
So to do that you simply need to open a terminal and then type S H space Virk 2007 dot S H then execute
99
and this will download word 2007 and then you can type another command as a space walk 2012 that s h
100
and then press enter to the queue to download.
101
Also 2012 but I've already done that for you.
102
I've downloaded this dataset and here they are there in this folder Vark dev kit and you have Virk 2007
103
and Vark 2012.
104
So let's have a look at Vocht 2007.
105
First you have several subfolder but the one I want to show you is this one.
106
GPL images which contains lots and lots of images.
107
As you can see let's open randomly.
108
Some of them so let's see this one for example number 8.
109
It's a chair.
110
All right so that's the ground truth for the chair.
111
Remember Kyrios intuition lectures need to start with the ground truth to ground truth is the truth
112
that there is in fact a chair right here on this image.
113
So that's the ground truth that will be compared to the prediction of the SS doing the training just
114
before a possible error is back propagated into the neural network.
115
All right so let's close this let's open another one.
116
Number 14.
117
It's a car a yellow cab probably in New York I would say maybe.
118
Or some other cities in the US.
119
But there we go.
120
That's to detect a car so it's a ground truth for a car number 31 then a train.
121
There we go.
122
You can even detect train.
123
And let's open a few more.
124
So that's a pigeon to detect birds.
125
Ground Truth for the bird.
126
And now 53 a cat cute little cat.
127
And one last one hundred.
128
Well let's go further.
129
That's open the number one thousand five hundred eighty six.
130
And that's a beautiful horse jumping.
131
And speaking of horses well that's going to be the homework for this section.
132
It's going to be it's going to be a fun homework for you to relax and enjoy playing some object detection
133
and some really really nice videos.
134
We're going to do some action on beautiful horses running on a field.
135
So there we go.
136
You have the object.
137
And now let's have a look at the other data set because the training is made on the two data sets at
138
the same time.
139
So have to scroll back up.
140
There we go.
141
Varg 2017 and now let's have a look at 2012.
142
Same subfolder is.
143
But this time with different images.
144
All right.
145
Here they are.
146
So the 2012 we actually have images of several 2007 2008 and probably 2009 there we go and up to 2012.
147
So let's open any of them for example this one from 2010.
148
Beautiful dog with a somewhat suspicious look.
149
But I love dogs.
150
I hope you liked the dog bouncing on the field as we detected during this much too.
151
I really like this Doug.
152
So let's have let's open another one let's open one image from 2011.
153
There we go.
154
This one.
155
Well two nice persons probably a mother with her daughter.
156
So very nice.
157
Let's open one from 2012 now and that will be our last.
158
So there we go the less ground truth we have.
159
Oh another kid and we can see several persons.
160
This one is actually interesting.
161
There are several ground truth in this image.
162
So maybe let's open the last one to look for a last object.
163
So that's another person and another person so probably all the images from 2012 are persons.
164
Now here we have a horse and a person so that's basically combining several different classes for the
165
day to be able to recognize several objects in an image.
166
All right so we're going to stop here and now now that we're done with the dataset I'm going to show
167
you the code and I'm going to show you how to run the code.
168
First let me just scroll back up to find my way back to the folder.
169
There we go.
170
Scripts done data done and going back to the main training as the folder.
171
All right.
172
So in this folder you also have besides the data you also have the layers which was the same folder
173
as before when we made the detection.
174
So that's because in order to train the SSD we need some of the functions and modules of the SSD.
175
Then we have the SSD implementation that contains all the architecture of the SSD Anchorage to have
176
a look at this.
177
And we have our trained file which I'm going to open right now because data from this file that we're
178
going to execute the training of the SSD on these two datasets Word 2007 and book 2012.
179
All right.
180
And then we have you Teal's which contains some tools for the training like augmentations so augmentations
181
are some ways to increase the amount of images so that we can have even more material to train our SD
182
on even if we have already lots and lots of images.
183
You can actually check that in our deep learning course we have some tutorials explaining in more details
184
how Image augmentations work and then finally we have this for that that contains some weights.
185
So make sure to understand that these are not some weights of some pre-trained model because actually
186
right now we're doing some training but these are just some initialized weights.
187
You imagine that they need to be initialized in some way a way that should be compatible with the future
188
of data the way weights during the training.
189
So these are the weights and now we're going to go through this train file and mostly we're going to
190
execute the code to show you how the training is done.
191
It's actually going to be very easy we simply need to select the whole code and then press command or
192
control press and to to execute.
193
But before we do that I would just like to show you some of the arguments here.
194
So basically these are some parcels containing all the arguments that you can change if you want to
195
change the way the training will be done.
196
So for example you can change the argument related to the learning rate here the default learning rate
197
is 0.01 but you can change it to 0.05 for example if you want to try a different training then you can
198
also change the parameter for the momentum.
199
The way to the gamma parameter.
200
Well these are all the parameters that you can change the hyper parameters to do some type of parameter
201
tuning of the training.
202
But again remember that training is taking lots and lots of time so it's mostly for the research and
203
developers working on the state of the art the models.
204
That's not for us to do right now anyway.
205
But then you have some other parameters I mentioned.
206
Image net which is another big huge dataset on which you can train your DeBerry morals and computer
207
vision models and therefore if you want to train your SSD on a different data set then Virk 2007 invoked
208
2012.
209
Well here is the parameter related to that which contains the path to the dataset data if get that exactly
210
the path data as this father and then there it is this folder containing Vok 2007 invoke 2012.
211
All right so let me go back.
212
So basically that's how it works.
213
You know you have several hyper parameters you can change them to experiment different trainings and
214
the rest is some implementation of other functions like.
215
Exactly and wait and it function that initializers the way it's the proper way.
216
So we're going to see that in more details in module 3 because in modules we will implement from scratch
217
the training of the Ganns.
218
Generally our research on that works because this will be possible to do without kuda and without thousands
219
of lines of code.
220
So you have several functions and mostly we have.
221
There you go the train function that does the training of the SSD on all the images that are contained
222
in the Vok folders.
223
So there we go now to finish the tutorial.
224
I'm going to show you how you can execute this file so very simply you just select the whole code again.
225
Make sure that you have an empty Geographic card and kuda enabled and that's not the case you will get
226
an error and be relieved.
227
Even if we had a non a version of this training.
228
Well it would take months or years.
229
And therefore this would be completely useless.
230
And that's exactly the reason why the developers have not implemented and none could have version of
231
the training.
232
That is because it would be completely useless.
233
So I'm going to you now.
234
There we go.
235
The train is on its way and it's basically going to take several hours so we're just going to stop here.
236
It was just to show you how the training works how the data sets were structured and to show you a little
237
behind the scene how the SS is trained.
238
And now we hope you understand why we had to use a pre-trained model to do our detection.
239
All right so now we're going to move on to the homework.
240
It's going to be a fun and exciting homework about detecting some beautiful horses running on a field.
241
I can't wait to show you this.
242
And until then enjoy computer vision.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.