Afrikaans
Akan
Albanian
Amharic
Arabic
Armenian
Azerbaijani
Basque
Belarusian
Bemba
Bengali
Bihari
Bosnian
Breton
Bulgarian
Cambodian
Catalan
Cebuano
Cherokee
Chichewa
Chinese (Simplified)
Chinese (Traditional)
Corsican
Croatian
Czech
Danish
Dutch
English
Esperanto
Estonian
Ewe
Faroese
Filipino
Finnish
French
Frisian
Ga
Galician
Georgian
German
Greek
Guarani
Gujarati
Haitian Creole
Hausa
Hawaiian
Hebrew
Hindi
Hmong
Hungarian
Icelandic
Igbo
Indonesian
Interlingua
Irish
Italian
Japanese
Javanese
Kannada
Kazakh
Kinyarwanda
Kirundi
Kongo
Korean
Krio (Sierra Leone)
Kurdish
Kurdish (Soranรฎ)
Kyrgyz
Laothian
Latin
Latvian
Lingala
Lithuanian
Lozi
Luganda
Luo
Luxembourgish
Macedonian
Malagasy
Malay
Malayalam
Maltese
Maori
Marathi
Mauritian Creole
Moldavian
Mongolian
Myanmar (Burmese)
Montenegrin
Nepali
Nigerian Pidgin
Northern Sotho
Norwegian
Norwegian (Nynorsk)
Occitan
Oriya
Oromo
Pashto
Persian
Polish
Portuguese (Brazil)
Portuguese (Portugal)
Punjabi
Quechua
Romansh
Runyakitara
Russian
Samoan
Scots Gaelic
Serbian
Serbo-Croatian
Sesotho
Setswana
Seychellois Creole
Shona
Sindhi
Sinhalese
Slovak
Slovenian
Somali
Spanish
Spanish (Latin American)
Sundanese
Swahili
Swedish
Tajik
Tamil
Tatar
Telugu
Thai
Tigrinya
Tonga
Tshiluba
Tumbuka
Turkish
Turkmen
Twi
Uighur
Ukrainian
Urdu
Uzbek
Vietnamese
Welsh
Wolof
Xhosa
Yiddish
Yoruba
Zulu
Welcome to simply learns YouTube
channel. Artificial intelligence and
machine learning are transforming the
way business operate, make decisions and
innovate. From personalized
recommendation on streaming platforms
and intelligent chat bots to
self-driving vehicles and advanced
healthcare systems, AI and machine
learning are powering some of the most
impactful technologies of our time. And
AI and machine learning engineering
combines programming, mathematics,
statistics, data science, machine
learning expertise to create intelligent
applications that deliver business
value. In this complete AI and machine
learning engineering course, you will
learn everything from the fundamentals
of AI and machine learning to advanced
concept used in the modern intelligent
systems. We will start with the
programming and mathematical foundation
then gradually move into data analysis,
machine learning algorithms, deep
learning and generative AI. Throughout
the course, you will gain hands-on
experience with industry standard tools
and frameworks such as Python, NumPy,
Pandas, Kikit Learn, TensorFlow,
PyTorch, and others. You'll also learn
how to collect and prepare data, build
predictive models, train neural
networks, evaluate model performance,
and deploy machine learning solution in
real world environments. By the end of
this course, you'll have a strong
understanding of complete AI and machine
learning life cycle and practical skills
required to pursue a career as a AI and
machine learning engineer. Having said
that, let's take a look at today's
agenda. We'll start off with module one,
which is introduction to artificial
intelligence and machine learning.
Module two is Python programming for AI
and ML. Module three is mathematics,
statistics, and probability for machine
learning. Module four is data
collection, cleaning, and
pre-processing. Module five is
exploratory data analysis and data
visualization. Module six is machine
learning fundamentals. Module seven is
supervised learning algorithms. Module
eight is unsupervised learning
algorithms. Module 9 is model evaluation
and feature engineering. Module 10 is
deep learning and neural networks.
Module 11 is natural language
processing. Module 12 is computer vision
fundamentals. Module 13 is generative
AI, LLMS and AI agents. Module 14 is
envelopes, model deployment and AI
engineering workflows. Module 15 is real
world AI projects. Module 16 is
interview question and answers. Hope I
made myself clear with that agenda.
That's it. If these are the type of
videos you would like to watch, then hit
that subscribe button with the bell icon
to get notified whenever we host. Also,
just so that you know, if you want to
upskill yourself, master generative AI
and land your dream job or even grow in
your career, then you must explore
Simply Learn's cohort of various
generative AI training and professional
certification programs. Simply learn
offers a variety of masters
certification and post-graduate programs
in collaboration with some of the
world's leading universities. Through
our courses, you will gain knowledge
along with work ready expertise in
skills like Python, Aentic AI, AI
automation systems, LLMs, and over a
dozen others. And that's not all. You
will also get the opportunity to work on
multiple projects led by industry
experts working on top tier
service-based and product companies.
After completing these courses,
thousands of learners have transition
into an AI and machine learning role as
a fresher or moved on to your higher
paying job and profile. If you're
passionate about making your career in
this field, then make sure to check out
the link in the pinned comments and in
the description box to find an AI and
machine learning program that fits your
experience and areas of interest. So
let's get started with our AI and
machine learning engineer full course
with a small quiz. What is NLP? Is it
network layer processing, natural
language processing, neural learning
platform, or is it native language
programming? Please let us know your
answers in the comment section below.
Now over to our training experts.
>> Once upon a time in the quiet town of
Newite, there lived a curious teenager
named Arya. She wasn't like most kids in
her school. While others were busy with
sports or music, Arya was fascinated by
machines, especially the idea of making
machines think like humans. Her
curiosity began one evening when she
asked her grandfather, who used to be a
computer engineer, "Can machines ever
think?" Her grandfather smiled and said,
"That's what artificial intelligence is
all about." Arya's eyes lit up.
Artificial intelligence? What's that? So
he began to tell her a story, not a
fairy tale, but a real story about the
science and ideas behind machines that
learn, decide, and sometimes even
surprise their creators. Artificial
intelligence, or AI, is the science of
making machines that can do things that
normally require human intelligence.
This includes tasks like recognizing
faces, understanding speech, making
decisions, and even playing games. But
AI isn't magic. It's built through
programming, mathematics, and data. Arya
imagined a robot that could talk like a
human and help with homework. Her
grandfather nodded. That's one kind of
AI, but there are many types. He
explained that AI isn't just about
robots. In fact, most AI systems are
just computer programs running inside
machines we already use, like phones,
laptops, or even refrigerators. Her
grandfather told her that AI comes in
two main types, narrow AI and general
AI. Narrow AI is the kind we see today.
It's designed to do one specific task.
For example, the AI in a smartphone that
unlocks the screen by recognizing your
face is only good at that one job. It
can't cook or write a story. General AI,
on the other hand, would be as smart as
a human, able to learn anything and do
many tasks. But this type of AI doesn't
exist yet. It's more of a dream for now.
Arya asked, "How do these machines
learn? That's where machine learning
comes in," her grandfather replied.
Machine learning is a type of AI that
learns from data instead of being told
what to do step by step. "Imagine
teaching a dog to sit. You show it how,
give it treats, and repeat. Over time,
the dog learns. Machine learning works
the same way. You feed it data and it
finds patterns. For example, if you want
a computer to recognize pictures of
cats, you show it thousands of cat
pictures. It starts to see what cats
usually look like. Furry whiskers,
pointy ears. Over time, it learns to
tell a cat apart from a dog or a chair.
The program that does this learning is
called a model. A model is like a brain
built by the computer using the data it
was given. The more data it gets, the
better it learns. But how does the
computer know what a cat is? Arya asked.
Her grandfather said, "That's thanks to
something called a neural network. It's
a method used in machine learning that's
inspired by how our brains work. A
neural network is made up of layers of
tiny parts called neurons. These are not
real brain cells, but math functions.
Each neuron takes in numbers, does some
math, and passes the result to the next
layer of neurons. Imagine passing a note
through a group of friends, and each one
adds or changes a word before giving it
to the next. By the end, the note may
have transformed in a useful way. That's
what a neural network does to data. It
turns it into something meaningful, like
recognizing a cat in a picture. The more
layers a network has, the more complex
patterns it can understand. When a
network has many layers, it's called
deep learning. To get a neural network
to work, it needs to be trained.
Training is the process where the model
is shown lots of examples so it can
learn. Training involves giving the
model data and letting it guess
something like whether a picture has a
cat. At first, it guessed badly, but
then it compares its guess to the
correct answer. If it's wrong, it
adjusts itself using a method called
back propagation. Back propagation is
like checking your math homework. If the
answer is wrong, you go back, find where
you messed up, and fix it. In AI, this
helps the model improve step by step.
This cycle of guessing, checking, and
adjusting is repeated many times. The
model slowly gets better at the task.
Can AI make mistakes? Arya asked. Oh
yes, her grandfather said AI is smart in
some ways but not perfect. AI only
learns from the data we give it. If the
data is bad, the AI will be bad. This is
called bias. For example, if a face
recognition system is trained mostly on
photos of light-kinned people, it might
not work well on darkerkinned people.
Also, AI doesn't really understand the
world. It only sees patterns in numbers.
It doesn't know what a cat feels like or
why we love them. That's why AI can
sometimes be fooled by simple tricks
like weird images that a human would
never mistake for a cat. AI is
everywhere, her grandfather explained.
It helps recommend videos on YouTube,
powers voice assistants like Siri or
Alexa, drives some cars, and even helps
doctors find diseases and scans. But not
all AI is harmless. It can be used for
spying, spreading fake news, or making
decisions that affect people's lives,
like who gets a loan or a job? That's
why it's important for people to
understand how AI works so they can ask
good questions and build it responsibly.
Arya asked, "Will AI take over the
world?" Her grandfather laughed. Not
like in the movies, but it will change
the world. The future of AI depends on
how people choose to use it. It can help
solve big problems like climate change
or disease. But it also needs rules and
careful thinking. Just like fire or
electricity, AI is a tool, a powerful
one. If used wisely, it can do great
good. Arya sat back, her mind buzzing.
She had started the day wondering if
machines could think. Now she knows that
while they don't think like humans, they
can do amazing things through learning
data, and clever programming. She smiled
and said, "Maybe I'll build an AI
someday." Her grandfather smiled, too.
Just remember, it's not about making a
machine smart. It's about making it
useful and fair for everyone. And from
that day on, Arya started her journey
not just to understand AI, but to shape
it with care, creativity, and curiosity.
>> We know humans learn from their past
experiences, and machines follow
instructions given by humans.
But what if humans can train the
machines to learn from their past data
and do what humans can do and much
faster? Well, that's called machine
learning. But it's a lot more than just
learning. It's also about understanding
and reasoning. So today we will learn
about the basics of machine learning. So
that's Paul. He loves listening to new
songs.
He either likes them or dislikes them.
Paul decides this on the basis of the
song's tempo, genre, intensity, and the
gender of voice. For simplicity, let's
just use tempo and intensity for now.
So, here tempo is on the x-axis, ranging
from relaxed to fast, whereas intensity
is on the y-axis, ranging from light to
soaring. We see that Paul likes the song
with fast tempo and soaring intensity
while he dislikes the song with relaxed
tempo and light intensity. So now we
know Paul's choices. Let's say Paul
listens to a new song. Let's name it as
song A. Song A has fast tempo and a
soaring intensity. So it lies somewhere
here. Looking at the data, can you guess
whether Paul will like the song or not?
Correct. So Paul likes this song. By
looking at Paul's past choices, we were
able to classify the unknown song very
easily, right? Let's say now Paul
listens to a new song. Let's label it as
song B. So song B lies somewhere here
with medium tempo and medium intensity.
Neither relaxed nor fast, neither light
nor soaring. Now, can you guess whether
Paul likes it or not? Not able to guess
whether Paul will like it or dislike it.
Are the choices unclear? Correct. We
could easily classify song A. But when
the choice became complicated as in the
case of song B. Yes. And that's where
machine learning comes in. Let's see
how. In the same example for song B, if
we draw a circle around the song B, we
see that there are four votes for like
whereas one vote for dislike. If we go
for the majority votes, we can say that
Paul will definitely like the song.
That's all. This was a basic machine
learning algorithm also. It's called K
nearest neighbors. So this is just a
small example in one of the many machine
learning algorithms quite easy right
believe me it is but what happens when
the choices become complicated as in the
case of song B that's when machine
learning comes in it learns the data
builds the prediction model and when the
new data point comes in it can easily
predict for it more the data better the
model higher will be the accuracy there
are many ways in which the machine
learns it could be either supervised
learning unsupervised learning or
reinforcement learning. Let's first
quickly understand supervised learning.
Suppose your friend gives you 1 million
coins of three different currencies. Say
1 rupee, 1 and 1 dirham. Each coin has
different weights. For example, a coin
of 1 rupee weighs 3 g. 1 euro weighs 7 g
and 1 dirham weighs 4 g. Your model will
predict the currency of the coin. Here
your weight becomes the feature of coins
while currency becomes their label. When
you feed this data to the machine
learning model, it learns which feature
is associated with which label. For
example, it will learn that if a coin is
of 3 g, it will be a 1 rupee coin. Let's
give a new coin to the machine. On the
basis of the weight of the new coin,
your model will predict the currency.
Hence, supervised learning uses labeled
data to train the model. Here, the
machine knew the features of the object
and also the labels associated with
those features. On this note, let's move
to unsupervised learning and see the
difference. Suppose you have cricket
data set of various players with their
respective scores and the wickets taken.
When we feed this data set to the
machine, the machine identifies the
pattern of player performance. So, it
plots this data with the respective
wickets on the x-axis while runs on the
y-axis. While looking at the data,
you'll clearly see that there are two
clusters. The one cluster are the
players who scored high runs and took
less wickets while the other cluster is
of the players who scored less runs but
took many wickets. So here we interpret
these two clusters as batsmen and
bowlers. The important point to note
here is that there were no labels of
batsmen and bowlers. Hence the learning
with unlabeled data is unsupervised
learning. So we saw supervised learning
where the data was labeled and the
unsupervised learning where the data was
unlabeled. And then there is
reinforcement learning which is a
reward-based learning or we can say that
it works on the principle of feedback.
Here let's say you provide the system
with an image of a dog and ask it to
identify it. The system identifies it as
a cat. So you give a negative feedback
to the machine saying that it's a dog's
image. The machine will learn from the
feedback and finally if it comes across
any other image of a dog, it'll be able
to classify it correctly. That is
reinforcement learning. To generalize
machine learning model, let's see a
flowchart. Input is given to a machine
learning model which then gives the
output according to the algorithm
applied. If it's right, we take the
output as our final result. Else we
provide feedback to the training model
and ask it to predict until it learns. I
hope you've understood supervised and
unsupervised learning. So let's have a
quick quiz. You have to determine
whether the given scenarios uses
supervised or unsupervised learning.
Simple, right? Scenario one. Facebook
recognizes your friend in a picture from
an album of tagged photographs.
Scenario two, Netflix recommends new
movies based on someone's past movie
choices.
Scenario three, analyzing bank data for
suspicious transactions and flagging the
fraud transactions. Think wisely and
comment below your answers. Moving on,
don't you sometimes wonder how is
machine learning possible in today's
era? Well, that's because today we have
humongous data available. Everybody's
online either making a transaction or
just surfing the internet and that's
generating a huge amount of data every
minute and that data my friend is the
key to analysis. Also, the memory
handling capabilities of computers have
largely increased which helps them to
process such huge amount of data at hand
without any delay. And yes, computers
now have great computational powers. So
there are a lot of applications of
machine learning out there. To name a
few, machine learning is used in
healthcare where diagnostics are
predicted for doctor's review. The
sentiment analysis that the tech giants
are doing on social media is another
interesting application of machine
learning. Fraud detection in the finance
sector and also to predict customer
churn in the e-commerce sector. While
booking a cab, you must have encountered
search pricing often where it says the
fair of your trip has been updated.
Continue booking. Yes, please. I'm
getting late for office. Well, that's an
interesting machine learning model which
is used by global taxi giant Uber and
others where they have differential
pricing in real time based on demand,
the number of cars available, bad
weather, rush hour, etc. So they use the
search pricing model to ensure that
those who need a cab can get one. Also,
it uses predictive modeling to predict
where the demand will be high with a
goal that drivers can take care of the
demand and search pricing can be
minimized. Great. Hey Siri, can you
remind me to book a cab at 6 p.m. today?
>> Okay, I'll remind you.
>> Thanks.
>> No problem.
>> Artificial intelligence, machine
learning, and deep learning represent
the evolution of computer science
towards creating intelligent systems. AI
is the broader concept striving to build
machines capable of humanlike
intelligence. ML is a subset of AI
emphasizing algorithms that learn from
data to make predictions or decisions.
DL in turn is a specialized branch of ML
that employs deep neural networks to
model complex patterns. Imagine an AI
powered voice assistant like Apple Siri.
It utilizes ML to understand and respond
to user queries, learning from
interactions over time. Deep learning
comes into play when Siri recognizes
speech patterns or interprets natural
language using neural networks to
process intricate features. The better
it becomes at understanding diverse
accents or refining responses
exemplifying the continuous learning
inherent in these technologies. AI seeks
to emulate human intelligence. ML
harness data for learning and DL employs
deep neural networks for intricate task.
The integration of these technologies
manifest in everyday applications,
transforming how we interact with and
benefit from intelligent systems. This
technology enables voice interaction,
allowing the device to play music, set
alarms, present audio books, and provide
up-to-date information on topics like
news, weather, sports, and traffic
reports, etc. Let's move forward and see
what is machine learning. Machine
learning is a subset of artificial
intelligence that focuses on developing
algorithms and models capable of
learning and making predictions or
decisions without being explicitly
programmed. ML systems leverage data to
recognize patterns, adapt and improve
their performance over time. There are
several types of machine learning.
Number one, supervised learning. The
algorithm is trained on a label data set
where each input is associated with a
corresponding output. Number two comes
as unsupervised learning. Unsupervised
learning deals with unlabelled data to
find inherent patterns or structures
within the information. And then comes
the reinforcement learning. This type
involves training agents to make
sequences of decisions by interacting
with an environment. And then comes
semi-supervised learning.
Semi-supervised learning combines
supervised and unsupervised learning
elements typically using a small amount
of labelled data and a larger pool of
unlabelled data. Let us move forward and
see what deep learning is. Deep
learning, a branch of machine learning,
focuses on algorithms inspired by the
human brain structure and functionality.
It excels in processing vast amounts of
both structured and unstructured data.
At the heart of deep learning are
artificial neural networks, empowering
machines to make decisions. The key
distinction between deep learning and
machine learning lies in data
presentation. Machine learning
algorithms typically demand structured
data while deep learning networks
operate through multiple layers of
artificial neural networks allowing them
to handle diverse data formats. So let's
start with the difference between
artificial intelligence, machine
learning and deep learning. And this
we'll show in a table form. So starting
with the definition.
So definition of artificial
intelligence. So broad field of machine
learning or creating machines with
intelligent behavior is artificial
intelligence. And when we talk about
machine learning, it's the subset of AI
focusing on algorithms learning from
data. And then comes the deep learning
that is specialized subset of ML using
deep neural networks. And now we'll see
the difference with the learning
approach between all these three. So in
learning approach artificial
intelligence can include rulebased
systems, expert system and more. And in
machine learning, it learns from data
patterns without explicit programming.
And then comes the deep learning where
it learns hierarchical representation
using neural networks. And if we talk
about scope, it encompasses various
techniques beyond learning from data.
And in machine learning, it primarily
focus on learning patterns from data.
And then the deep learning, it
specifically utilizes deep neural
networks for complex task. And now we'll
move to the next difference. And we'll
start with an example. So in artificial
intelligence, the example is autonomous
vehicles, chatboards or expert systems.
And for machine learning, it's spam
filters, recommendation systems, image
recognition. And in deep learning it is
image and speech recognition natural
language processing. And now we'll see
the difference for the data
requirements. So it depends on the
specific application and problem solving
approach. And in machine learning it
requires labeled or unlabelled data or
training. And for the deep learning it
relies on large amounts of labelled data
for training deep networks. And now for
the complexity artificial intelligence
addresses a wide range of task including
those beyond ML. And in machine
learning, it deals with moderate to
complex task depending on algorithms.
And for the deep learning, it is well
suited for intricate task often
requiring substantial computational
resources. And now see the flexibility.
So for the artificial intelligence, it
can be rule- based, evolving and
adaptive. And for the machine learning,
the flexibility adapts to patterns and
the changes in data. And for the deep
learning, it adapts to hierarchical
representations and diverse data types.
And then comes the training process. So
in artificial intelligence, training
process varies based on specific AI
techniques used. And in machine
learning, training involves feeding data
and adjusting model parameters. And in
deep learning, training involves
optimizing neural weights and
structures. And now we'll talk about the
applications between all these three
terms that is a IML and deep learning.
So for artificial intelligence the
applications are robotics, natural
language processing, game playing and
for machine learning it's predictive
analytics, fraud detection and
healthcare diagnostic and for the deep
learning that is image recognition,
speech synthesis and language
translation.
>> Now you guys must be thinking why should
I consider a career in AI? Well AI is
not just a passing trend. It's a seismic
shift that is reshaping our world and
creating new venues for innovation and
discovery. Now by embracing a career in
AI, you become a part of dynamic field
that thrives on solving complex problem,
pushing boundaries and making a profound
impact on society. The demand for AI
professionals is skyrocketing across the
industries from healthcare, finance,
entertainment, transportation.
Organizations are actively seeking
talented individuals who can harness the
power of AI and drive their business
forward. But what skills does it take to
become an AI engineer? How can you
embark on this thrilling journey? We
have the answer to all your questions.
Some steps are crucial to master the
field of AI and become an AI engineer.
Let's go through them real quick. So the
first step is to establish a strong
foundation in mathematics and
programming. Start by gaining a solid
understanding of critical mathematical
concept such as linear algebra, calculus
and probability theory. Additionally, it
is crucial to become proficient in
programming languages like Python which
is commonly used in AI and develop
coding skills. Next, you need to pursue
a degree in relevant field. Earn
bachelor's or master's degree in
computer science, data science, AI or a
related discipline to acquire a
comprehensive understanding of AI
principle and techniques and after that
you need to acquire knowledge in machine
learning and deep learning. Familiarize
yourself with ML algorithms, neural
network and deep learning frameworks
like for example TensorFlow, PyTorch to
train and optimize models using real
world data sets and afterward engage in
practical projects. Gain hands-on
experience and demonstrate your skills
by working on AI projects. Building a
portfolio of projects that showcase your
ability to solve AI problems can make a
strong impression on potential
employers. After that, collaborate and
network. This is really important.
Engage with AR communities, attend
conferences, and participate in online
forums to connect with professionals in
this field. Collaborating with others
can enhance your learning experience and
open up new opportunities.
Seek internships or entrylevel positions
where you can gain practical experience
through AI internships or entry-level
roles in industry or research
institution. Now this will provide
valuable exposure and help you further
develop your skills. After that
continuously learn and adapt. In the
fast-paced world of AR, it is very
important to stay updated on new
developments, explore specialized areas,
and embrace emerging technologies and
tools. Continual learning and
adaptability are essential for pursuing
a successful career as an AI engineer.
Now that you're familiar with the steps
involved in the journey of an AI
engineer, let's discuss the essential
skills you need to know to become an AI
engineer. So, here's a breakdown of the
skills needed. First one is having
strong programming abilities. This
typically refers to expertise in one or
more programming languages commonly used
in data science and machine learning
such as Python or R language. Now,
proficiency in programming allows you to
write efficient and scalable code for
data analysis, modeling and algorithm
implementation.
Next, you need knowledge of machine
learning algorithms. This involves
understanding and familiarity with wide
range of machine learning algorithms
including both supervised and
unsupervised techniques. You should be
able to select and apply appropriate
algorithms for specific problems as well
as evaluate and optimize their
performance. Next skill is proficiency
in statistics and mathematics. Sound
knowledge of statistics and mathematics
is fundamental for data analysis and
machine learning. You should be
comfortable with statistical concepts,
hypothesis testing, regression analysis,
probability theory, linear algebra and
calculus.
Now after that you have acquired a good
amount of knowledge of these skill set,
we'll move on to our next skill which is
having familiarity with deep learning
frameworks. Now deep learning has gained
significant popularity in recent years
and familiarity with deep learning
frameworks like TensorFlow, PyTorch or
Keras is valuable. Now these frameworks
provide tools and libraries for
building, training and deploying deep
neural networks for tasks such as image
recognition, natural language processing
and time series analysis. Next, you need
experience with big data technologies.
Dealing with large scale data sets
requires knowledge of big data
technologies such as Apache, Hadoop,
Spark or distributed computing
frameworks. Understanding how to
process, store and analyze data
efficiently in distributed environments
is very essential. Now after you have
gotten experience with big data
technologies, now it's the time to move
on to our next skill which is having
excellent problem solving and analytical
skills. Now these skills will enable you
to break down complex problems, identify
key factors and develop efficient
solution.
Now you should be able to adapt at
critical thinking, troubleshooting and
debugging to handle real world
challenges in data science and machine
learning. So guys, remember to stay
updated with the latest advancements in
the field and continue learning to stay
at the forefront of data science and
machine learning. So that's all we had
for you in this AI engineer road map. Do
you know how AI has become so fast? It's
now replacing entire teams in some
industries. Yes, it's true. Over 50% of
companies are already using AI to
automate jobs. AI tools are writing
emails, creating content, and even
giving job interviews. And while some
people are worried AI will take the job,
I let you in on a secret. AI is also
creating tons of highpaying roles. The
catch, you need the right skills to get
it. And that starts with learning with
the right programming language. Now,
I've tested a whole bunch of them.
Python, C++, R, Java, you name it. And
in this video, I'm breaking down the top
five programming languages for AI that
you need to know if you want to build a
career, land real jobs, and actually
stay relevant in the age of AI. We will
cover what each language is best at, how
to start learning, what kinds of AI jobs
they lead to, and yes, how much you can
earn with each one. All right, first up,
we've got Python. And honestly, this one
is the most valuable player of the AI
development. Just like the star player
in a sports team, Python is the go-to
language that everyone relies on when it
comes to building AI system. So, why is
Python the AI king? Let me break it
down. Simplicity and readability. Now,
Python is super easy to learn. It's
almost like writing in plain English.
You don't have to worry about
complicated code. If you're just
starting out in programming, that is
definitely the language you are going to
feel most comfortable with. It's got
this userfriendly vibe that makes it
simple even for people new to coding.
Second of all, it has got endless
libraries. Now, Python is packed with
tools. We call it libraries like
TensorFlow, PyTorch and Scikitlearn.
Think of these library as pre-made
toolkits that make AI development way
easier. They save you a lot of time
because instead of building everything
from scratch, you can use these
libraries to quickly train your models
and run algorithms. It's like having a
shortcut to building AI system. It has
also got rapid prototyping. If you need
to test your ideas quickly, Python is
perfect for that. You can build a model,
a simple version of your AI system in no
time. So whether you're working on
machine learning models or neural
networks, fancy word for AI system that
learn like the brain, Python let you
prototype or build a quick model fast.
So I know you must be wondering now what
kind of AI jobs can Python land me? It's
a great question. With Python, you could
land jobs like data scientist, machine
learning engineer, or an AI researcher.
Now these jobs typically pay between
around six lakh to 15 lakh peranom.
That's the salary range. But the best
part is as you gain more experience and
expertise that number will go way
higher. So how do you start learning
Python? You don't have to break the bank
to learn Python. You can get started
with free platforms and YouTube
channels. Simply learn even offers a
free comprehensive course in Python and
I'll leave the link for you to check it
out. And the best part is Python has got
huge community. So if you ever feel
stuck, there's always someone out there
who's ready to help you. Next, we'll
talk about C++. Now C++ isn't as
beginner friendly as Python, but it's a
beast when it comes to performance heavy
applications. If you're working on
realtime AI like self-driving cars or
high frequency trading algorithms, then
C++ is where you want to be. But why did
we choose C++ for AI? First of all,
because of its speed and efficiency.
Now, C++ is all about its speed. It's
the language you want when you're
working with large data sets or AI
applications that need to be super fast.
Second of all, it has got lowlevel
memory management. Now, C++ gives you
full control over memory, which is
essential when you're building AI system
that require extensive computation and
realtime performance. But isn't C++ more
complex than Python? Definitely, yes.
But if you're diving into AI
applications that require high
performance, think computer vision or
robotics, C++ is unmatched. It's a bit
trickier to learn, but if you want to
build realtime AI systems, it's worth
the effort. Roles like AI software
developer or computer vision engineer
are your goto with C++. The salary range
for these roles is around 8 lakh to 20
lakh peranom depending on the project's
complexity and your experience. Third on
a list is Java. This one's a workhorse
in the world of AI. And if you're aiming
to work on enterprise level AI projects,
then Java is definitely a language you
want to know. Now, it's not the first
choice for small scale AI projects. But
when it comes to big scalable systems,
Java is untouchable. So why Java for AI?
Because of its scalability. Now, you can
think Java as a beast when it comes to
handling large scale applications. If
you're working on AI system that need to
process huge data sets or manage complex
computations, then Java can handle it
all without breaking a sweat. It's
designed to scale which makes it perfect
for enterprise level AI projects where
big data is involved. Platform
independence. One of the best things
about Java is its right ones run
anywhere feature. It doesn't matter
which platform you're using, whether
it's Windows, Mac, Linux, Java can run
all of it without issue. This is a huge
win when you're building AI systems that
need to operate across multiple
platforms. Mature libraries. Java has
been around for decades and because of
that, it's packed with reliable
libraries for AI. Libraries like Qua,
H2O make implementing machine learning
models or building AI system a lot
smoother. These libraries we tried and
tested so you know you're working with
solid tools. Let's talk about what jobs
can you actually land with Java. Now
with Java you're looking at some big
roles in the AI world. Think of AI
solution architect or AI backend
developer. These positions are not just
highly respected but it also comes with
a solid salary range typically between 7
lakh to 18 lakh peranom. And with
experience, well, let's just say that
number can easily climb higher. Now, you
must be thinking, how do I get started
with Java? Now, if you're already
familiar with object- oriented
programming, learning Java will be a
breeze. And if you're new to it, don't
worry. You can start with some great
resources like a YouTube channel or
LinkedIn Learning. There are plenty of
courses that will teach you how to use
Java for AI from the ground up. Now,
let's talk about R. This one's for all
data science enthusiasts out there. If
you're diving into statistical AI and
the language built specifically for
handling massive data and performing
complex statistical analysis, R is your
goto. So why R for AI? Because of its
statistical power. R is packed with
tools for statistical modeling. So if
you're working on AI projects that need
data analysis before you even start
applying machine learning, R makes it a
breeze. It's got everything you need for
analyzing trends, finding patterns and
building strong predictive models. It
has also got a feature of its data
exploration and visualization. One of
the R's strength is its data exploration
and visualization capabilities. You can
easily plot, chart and analyze your data
to uncover insights. This makes art
perfect for the datadriven side of AI
development where understanding your
data is just as important as building
the models. But wait, can I still work
in AI if I learn R or is it just for
data analysis? Absolutely. R is
fantastic for AI projects that rely on
statistical methods and data analysis.
It's actually the language of choice for
roles like AI data analyst or
quantitative analyst where you'll be
building predictive models or analyzing
data trends to make decisions. Now these
roles are in high demand and the salary
range typically falls between 6 lakh to
12 lakh peranom but with experience you
can definitely push those numbers
higher. Now to get started with art, you
can find tons of free resources on a
plat new kid on the block that's growing
fast in the AI space. It's relatively
young compared to Python or C++. But
trust me, it's making a huge impact. And
here's why. Now, Julia was created back
in 2012 by a group of researchers who
wanted a programming language that could
handle the complex calculations required
for scientific computing and they nailed
it. But why did Julia grew so fast?
Well, it's been picking up speed because
it combines the performance of C++ with
the readability of Python. You get the
speed and efficiency that C++ is known
for, but with Python's clean and easy to
write code, it's like the best of both
worlds. Let's talk about why did we
choose Julia for AI? Because of its
speed and simplicity. Now, Julia's speed
is one of the biggest advantages. It's
designed for high performance computing.
So if you need to run complex AI models
or process tons of data, Julia will do
it in a fraction of the time it would
take in other languages. And the syntax,
it's also super easy to read and write.
So you're not sacrificing convenience
for performance. It has also got the
feature of scientific computing. Now
Julia is optimized for AI task like deep
learning and numerical analysis. You can
think AI applications in robotics, data
science, and scientific research. Now,
if you're working on projects that
require heavy computations or advanced
AI models, Julia's is your go-to. So,
why isn't everyone using Julia yet? It's
still growing, but Julia community
expanding rapidly, and more libraries
and frameworks are being developed every
day. It's quickly becoming a top choice
for high performance AI, and it
continues to evolve. And of course, I
expect to be even more popular. So, is
Julia better than Python or C++? Now the
answer is it depends. Now if you're
building scientific AI applications that
require high performance, Julia is a
fantastic option. It's still growing but
the community expands. Julia will
quickly become more powerful in the AI
space. Julia is perfect for roles like
AI developer in the scientific or
numerical computing space. Salaries can
range from 7 lakh to 15 lakh peranom
especially if you're working with
advanced AI. So guys there you have it
the five best programming languages for
AI. So whether you're interested in
machine learning, realtime AI or data
science, there's language for you. Each
of these will help you land AI job you
want and give you the tools you need to
build powerful AI system. Which one are
you going to start with? Drop your
thoughts in the comment section below
and let's talk about it. And if you
found this video helpful, hit that like,
share, and subscribe button to get more
AI tips and career advice by simply
learn. get started with the onboarding
and interface including the subscription
plan. As you can see here, it is
offering us three major plans. Now, now
there is a free version of manuals you
can use on a day-to-day basis. But make
sure to know that everyday credits are
six rupees.
>> But make sure everyday credits are only
300 to limited. But only 300 credits
will be assigned to you on a everyday
basis. Now, the first plan is $20 per
month, which gives you 300 fresh credits
every day, 4,000 credits per month,
in-depth research for everyday task,
professional website for standard
outboard, insightful slides for regular
content, task scaring, and wide
research, early access beta features,
and 20 concurrent task, 20 schedule
tasks. Now again if you are working in
an organization which where you can auto
you have to auto too many stuffs you can
upgrade to a $40 or $200 plan. Now $20
is for a person single usage because
it's only 300 fresh credits per day.
It's total of 4,000 per month as well.
So when it so when it comes to $40 plan
you can consider sharing it with two to
three people. Again it's 300 credits but
8,000 credits per month. all the other
things plus plus you'll get an addition
of 4,000 more credits to work on. Now
when it comes to 200 you'll get a
firstly you'll get free cloud computing
where you don't have to worry about the
storage and stuff and here it is 40,000
credits per month an organization which
uses automation tools a lot more can use
this now we are going to start by
understanding the manusi interface and
the first thing we need to lock out at
the hub which is basically your main
dashboard. Now before we start giving
task to manus AI it is very important to
understand how credits works because for
many users credits can be a little
confusing in the beginning. When you
open the dashboard you will notice that
man's AI shows two different credit
counters. The first one is the daily
refresh credits. These are the credits
that refresh every day. For example you
may see around 300 credits per day. The
important thing to remember is that
these are use them or lose them credits.
That means they reset every 24 hours and
if you don't use them, they do not carry
forward in the next day. So these are
the daily credits which are good for
regular task, quick experiments, small
research work, testing prompt or even
trying out different features inside
Manusa. The second credit counter is
your monthly pool. This is your main
credit balance for the month. For
example, if you're on a standard plan,
you may need something like 4,000
monthly credits. These credits are more
useful for larger and more complex task.
So, if you ask manus AI to do something
longunning like researching a topic
deeply, creating a report, browsing
multiple sources, analyzing information,
or even completing a multi-step
workflow, then this monthly pool gives
you the main runway to complete those
bigger tasks. So just remember this
simple difference. Daily credits are for
everyday use and reset every 24 hours.
Monthly credits are your larger credit
pool for bigger tasks throughout this
month. Now the next important thing in
the dashboard is the active task window.
This is where manusci shows the tasks
that are currently running and this is
one of the most powerful parts of the
platform.
Unlike a normal chatbot where you can
ask one question wait for one answer,
Manos AI can work on multiple task at
the same time. For example, on this
subscription you can run up to 20
concurrent task at once. That means
manos can work on multiple request in
parallel. Maybe one task is researching
a topic, another is preparing a
document, another is analyzing a website
and another is organizing the
information. For the free users, the
limit is usually lower around five
concurrent tasks. But the important
thing is not just the number of tasks.
The important thing is that these tasks
are asynchronous and cloud-based. This
means once you start a task, Manus AI
continuously working in cloud. You do
not have to keep watching the screen the
entire time. You can start a task, close
the browser, disconnect from the
internet, and even come back later. and
Manus AI can still continue to work on
that task in the background. This is
what makes it feel less like a normal AI
chat port and more like an AI worker.
You're not just asking a question and
waiting for the reply. You're assigning
work, letting the agent process it, and
then checking the results once the task
is completed. So before using Minus AI
for real workflows, always understand
these three things. Your daily credits
reset every day. Your monthly credits
support bigger and longer tasks and your
concurrent task window shows how many
jobs Manus AI is currently handling for
you. Once you understand this dashboard,
it becomes much easier to manage your
credits, plan your task properly, and
use Manus AI more efficiently. Now that
we have understood the dashboard and
credits, let's move on to the next
important part of Manus AI interface,
which is the goal, input, and task
planning area. Now this is where you
actually start working with manus. In a
normal chatbot we usually give a small
instructions one by one. But in Manus AI
the idea is slightly different. Here you
give a highle goal and manus plans the
steps needed to complete that goal. So
in the main input box let's type a
simple goal such as research the top AI
tools for content creation and create a
comparison report. So let's start.
research the top AI tools for content
creation and create comparison report.
Now here we have assigned a proper goal
to Manus AI. Now notice what happens
after we enter this prompt. Manos does
not directly jump into the final answer.
First it create a task plan. This is
where you will see a to-do list or
step-by-step structure showing how Manus
is planning to complete the task. For
example, manus may break the goal into
steps like understanding the topic,
searching the AI content, creation
tools, collecting useful information and
comparing those tools and finally
preparing the report. So here you can
see the steps. In simple words, manus is
taking one big goal and breaking it into
smaller actions. This view is very
important because it gives us a chance
to review the plan before the agent
starts doing heavy work. Before manos
begins browsing, opening pages,
analyzing sources and consuming more
credits, we can quickly check whether
the plan looks correct. For example, in
this case, we should check is manners
searching for the right type of tooth.
Is it planning to compare them properly?
Is it going to create a final report as
we asked? If the plan looks correct, we
can continue. But if the plan looks
incomplete or slightly wrong, we can
stop and adjust the prompt before moving
forward. This helps us avoid wasting
time and credits. So the key point here
is simple. In manus AI, we don't need to
write every step manually. We can give
one clear goal and manus will create a
plan for completing it. But before
allowing the task to continue, always
review the documentation decomposition.
But before allowing the task to
continue, always review the
decomposition view. This helps you
understand how the agent is thinking and
whether it is moving in the right
direction. So in this example, our goal
was to research the top AI tools for
content creation and create a comparison
report. And Manus turns that single bowl
into the structured task plan that can
review before execution. This is what
makes Manus air different from a regular
chatbot. It does not just answer
immediately. It plans the work first,
shows the direction and then start
completing the task. Now that Manus has
understood our goal, the created task
plan, the next step is execution. It's
already executing. This is where Manus
AI actually starts working on a task.
You can think of this part as a hand in
the platform. The goal input is where
Manus understands what we want. The
planning view is where it decide how to
do it and the execution view is where it
actually performs the work. Once we
approve or continue with the task,
manage begins completing the steps one
by one. The interesting part is that we
can watch this happen in real time. On
one side, you will usually see the
progress list or task steps. This shows
that manus has completed what is
currently doing and what is still
remaining. The next is that you can see
manus actually taking action. For
example, if the task is repeat, for
example, if the task requires research,
you can see the agent opening websites
and browsing pages. If the task requires
collecting information, it may take
screenshots, extract details, or even
organize the data. If the task needs a
structured output, manuals may update a
spreadsheet, write content, run the
code, or even build an interactive
artifact. So instead of only showing the
final result, manos shows the workflow
while it's happening. This is useful
because we can understand how the agent
is working not just what answer it gives
to the end. Now another important thing
is to understand here is the sandbox
environment. Manos does not directly
operate your local computer. It works
inside a cloud and is created for the
task. Inside this sandbox, miners can
browse websites, collect information,
test the ideas, run code, fill forms and
build outputs without affecting your
personal system. For example, if we ask
miners to research AI tools and prepare
a vision report, it can browse different
website, collect the required details,
organize them and then create the final
report inside this workspace. And for
more advanced task, the sandbox can also
help maners create things like websites,
slide decks, spreadsheets, dashboards,
and other interactive files. This is one
of the major reasons maners feels
different from the normal chatbot. A
regular chatbot mostly gives a text
responses. But maners can actually
perform actions inside a controlled
environment. So while the task is
running, we should keep an eye on two
things. First the progress list to
understand which step manus is working
on. Second is realtime action view to
see what the agent is actually doing.
This makes the whole process more
transparent. You're not blindly waiting
for the final output. You can see the
agent browsing, checking information,
organizing the data and building results
step by step. So in simple terms, the
execution view shows manus in action.
The sidebyside workflow helps us track
the task in real time and the sandbox
environment gives Manus a safe cloud
workspace where it can browse, run code,
collect data and create useful outputs.
This is the part where Manus moves from
planning the work to actually doing the
work. So we'll get back to this task
once it is completed. Let's start with a
new task. Now that we have seen how
Manus works inside the browser, let's
look at how can Manus be on normal web
interface. Now here you can even connect
a different apps such as Gmail, browser,
meta and you can add other connectors as
well if you're planning to automate any
kind of workflow. Now here when you come
to the desktop side you will have a
mobile app as well as the desktop app as
well. Now if you come to settings you
may find an option called integration.
This is where you can actually connect
manos with platforms like slack,
telegram or even line. So as you can see
here there are connectors. This is
useful because it allows you to interact
with manus through a messaging apps you
already use. For example, instead of
opening the browser every single time,
you can just delegate a task, check the
progress or monitor updates from a
messaging app. So if you're working with
a team, Slack can be useful. If you want
quick mobile access, WhatsApp, Telegram
or Lion can make it easier to stay
connected with the agent. Part two manus
AI. Let's continue. The main benefit is
remote control. You can start monitoring
task even when there is no sitting in
front of the main system. Now the next
advanced feature is the desktop my
computer feature. You can download the
computer version here in the desktop
app. This is available when you have
Manus desktop app installed in your Mac
or a PC. Here Manus can request access
for your local machine for specific
action. For example, it may need you to
read a local file, open a folder or run
a terminal command. But the important
thing is to notice that manus does not
get a fully access automatically. There
are permissions grades. There is a
permission gate when manus wants to
perform an action on your computer. You
will see prompts like allow once or
allow always. From a safety point of
view, allow once means you are giving
permission only for that specific
action. Always allow means you are
allowing that type of action more
regularly depending on the setup. So
while showing this, this is especially
useful when you want manos to work with
files on systems, run scripts or even
help with local development tasks. Now
the third advanced area is the web app
builder. This is where manage becomes
even more powerful. In the web app
builder interface, you can see manage
generating a live interactive web
application. This is not just writing a
text or giving code snippets. It can
actually build pages, connect the
databases, structure the app and prepare
it while working with the project. For
example, if we ask manus to create a
simple landing page or a small web app,
it can generate a layout and add
interactive sections, connect the
required backend logic, and even support
things like database setup and SEO
optimization. The best part here is that
you can watch the agent work step by
step. You can see it creating files,
updating the design, testing the pages
and even improving the final output. So
this part is useful for users who want
to build something practical like
websites, dashboard, internal tool,
product page or even prototype without
manually writing everything line of code
from scratch. To summarize this section,
so now let's test the same logic. Now
let's ask minus AI to create a web
landing page for a skincare brand. So
create a brand. Now to summarize this
section, the browser is the main place
where you use manus AI which is this.
The messaging integration help you
delegate and monitor task remotely. The
desktop app gives you manus control
access to your local machine with
permission prompts. And the web app
builder helps manus create live
interactive web project. So this is what
takes manus from being just a web- based
AI agent to something that can connect
with your communication tool, your
computer and your real project works.
Now as you can see there are approaches
here. This is the code for the entire
web page. Let it generate. I'll show you
the output since this is just running in
the first step. There are more three
steps involved in this. So we'll get
back to this once this is done. Now we
are going to see where Manus AI becomes
really powerful which is deep research
and data. The main idea here is very
simple. Manus AI is not just a chatboard
that gives one quick answer. It can work
more like an autonomous research worker.
That means you can give it a goal and it
can plan the task, browse multiple
sources, collect the information, cross
the check details and organize the
findings and finally create a proper
output. So instead of manually opening
20 tabs and copying the nodes, checking
the resources and building the report
yourself can handle a larger part of
that workflow for you. Let's start with
a we'll just use a practical prompting
as of now. Now for this demo, you can
just type in research the top CRM tools
for small business and create a
comparison report with pricing, key
features, pros, cons, best use cases and
source link. So I've given the exact
same prompting. Now once we enter this
goal, manus first creates a plan. This
is important because the task is not
just asking for a simple answer. We are
asking manage to research multiple CRM
tools, compare them and prepare a
structured report. Once the task starts,
notice how manners does not depend only
on one search result. It begins visiting
different websites and checking product
pages, pricing pages, review platform,
blogs, and others available sources.
This is what we call multi-source
research. For example, if MinusAI is
researching CRM tools, it may check
official websites for pricing, review
platforms for user feedback and
comparison articles for feature level
difference. The important thing here is
that Manos is not just collecting random
information. It is trying to cross
interface the details. So if one website
mentions a price, Manos can compare it
with the official pricing page. If one
source mentions a feature, it can check
whether the same feature is also listed
on the product website. This helps
improve the quality of the research.
Now, while the agent is working, keep
your attention on realtime interaction
view. On one side, you can see the task
progress. On the other side, you see
minus browsing websites, opening pages,
taking screenshots, reading the
information, and updating its findings.
This makes the process more transparent.
You're not blind. You're not blindly
waiting for final answer. So you are
actually seeing how the agent is
collecting and organizing the
information. Another important thing is
to notice how manage handles small
problems during the search. Sometimes a
page may not open. Sometimes a link may
be broken. Sometimes a website may be
JavaScript heavy and difficult to read.
In manual workflow we would have stopped
and find another source assets. But
maners has planning layer that can
create recovery steps. So if one source
does not work, it can try another
source. search again or adjust the path
without needing constant human help.
This is why manners is useful for
research heavy tasks. At the end, the
output should not just be a paragraph
summary. A good result should be a
structured artifact like a comparison
table or a full research report. For the
CRM example, the final output can
include tools, names, pricing, key
features, pros, cons, best use cases,
and source links, which we'll check back
in a few minutes. If you can move beyond
one short answers and prefer a full
research workflow across multiple
sources, manusi is the tool. Now let's
move on to which is wide search. This is
more advanced credit intensive feature.
So we'll get back to all the three in a
minute. We'll get back to all the three
outputs and I explain what was the exact
steps required. So coming back to wide
research. In normal research, the agent
may explore sources step by step. But in
wide research, the idea is very
different. Wide research is designed for
scaling. Instead of checking a few
sources one after the other, it can
explore many sources in parallel. Manus
described wide research as using
parallel multi-agent orchestration where
many agents can work across large
research space at the same time. So this
is not meant for basic questions like
what is CRM or even give me five tools.
This feature is better for high impact
research tasks like market analysis,
competitive research, industry reports,
investment research, product research,
or even strategy planning. For example,
we can use a large version of the same
CRM topic. Run wide research on the CRM
software market for smaller businesses.
Compare major players, pricing, trends,
AI features, and even customer
sentiment, market positions, or even a
growth opportunities. This kind of
prompt is much broader. Here we're not
only asking for a tool comparison. We
are asking miners to understand the
market from a different angles. It may
explore companies, websites, review
sites, market reports, competitors,
pages, product documentation, user
discussions, and other public sources.
Now, before starting wide research,
always explain the credit part clearly.
This type of task can consume a lot more
credits than a normal research would.
Since wide research explores a large
number of sources and runs a much
heavier workflow, it can cost
significantly more credits. So we should
use it for important research work, not
for a simple Q&A. This is important for
learners. Think of it like hiring a full
research team for one task. You would
not only use them for a small
definition. You would use it when the
output has real business value. So the
main takeaway here is use normal
research for focused task. Use why
research when you need a large scale
high depth analysis across many sources.
Now next we'll move on to it is useful
because it shows how manuals can move
from raw data to a finished business
report. Here for example let's say let's
upload a CSV file. So here I've taken a
random data set from Kaggle and I've
uploaded it. It says loan data set. Now
let's give it a prompt saying analyze
this loan data and create a report
showing revenue trends, top performing
products etc. So let's just say analyze
this loan data and create a report
showing the trends. So mind you I have
already cleaned this data and executed
using AI which is in collab but still it
took me like proper an hour to create
it. So let's just leave it. Now as you
can see this is where maners becomes
different from normal AI tools. It does
not only look into the file and guess
the answer. It can work inside a cloud
sandbox. Inside this sandbox, miners can
write and execute code such as Python to
process the data. So if the file
contains thousands of rows, miners can
calculate totals, averages, trends,
category performance, product
performance, regional performance, and
other useful metrices. While this is
happening, show the executional view.
You may see man is reading the file,
writing the code, running analysis,
checking the output and generating
charts. This is manus is not producing
text. It is actually performing mini
data and this is workflow. After
processing the data, Manus can also
create visualization. For example, it
can generate charts showing loan
prediction data, which category will
take more loan, etc. Then the final
step, it can synthesize everything into
a business report. Now, what does a good
report include? what the data shows,
which products are performing well,
which areas need attention, which trends
are visible and what actions the
business should take next. So from one
uploading of file and one prompt, Manus
can complete an end to end workflow. It
can pass the data, run the code, create
charts, interpret the results and write
a final report. This is why Minus is
very useful for business users,
analytics, marketers, sales teams,
founders, and students learning data
analysis. Now let's see how Manis AI can
work on autonomous research worker. It
can browse multiple sources, analyze the
data, run code and prepare structured
reports. Now in this module we will see
manus AI as a creator. This is where
manus move from just giving answers to
creating finished functional artifacts.
So instead of only asking manus to
explain something, we can ask to build
something. It can create web apps, slide
text, posters, infographic, visual
content and also complete project assets
from a single natural language prompt.
So let's get started. So here let's give
manners a single prompt. Now let's ask
it to create a landing page for AI
productivity tools for students with
sections for features, pricing,
testimonials, FAQs, and call in action.
Now can you notice what happens here? We
are not giving miners a full design
document. We are not writing code. We
are not explaining very section step by
step. We're only giving it an idea.
Manus takes this idea, understands the
goal, creates a plan, decides the page
structure, writes the content, designs
the layout, and starts building the
page. This is important. Manus is not
just giving us text response. It is
creating a clickable portfolio. So, as
you can see, it already started creating
the This can be very useful for
developers, product managers, startup
founders, marketers, and business teams.
If someone has an idea and want to
quickly see how it might look at a
website, manners can help create the
first version very quickly. Instead of
spending hours preparing a wireframe or
explaining the idea to a designer or a
developer, we can just use maners to
create a rough working version. Then we
can share it with the team, client or
stakeholder for feedback. Depending on
the tunnels can also help with more
advanced parts like databases, payment
flow, SEO friendly structure and
deployment related steps. But for
beginners, the main thing is to
understand this manual can move from an
idea to a functional prototype. Also,
this type of task is more resource
inensive than a simple chat response.
Building a web page may consume hundreds
of credits. Sometimes around 500 to,000
or even more depending on the
complexity. So before running a web app
task, always check the estimated credit
usage. This is because minus is not only
writing text, it is planning, coding,
testing, building and sometimes handling
deployment steps as well. Now let's move
on to the next part. Now, now let's ask
manus to create a slide deck. Now let's
ask the manus AI to create a text on the
future of AI agents for business teams.
Now once we give this prompt, Manus
starts planning the slide tech. It does
not randomly create slide. It first
creates a proper structure. For example,
it may begin with an introduction, then
explain what AI agents are, why
businesses are using them, their
benefits, use cases, challenges, and
finally a conclusion. This is what makes
the output useful. It's not just a set
of separate slides. It's a structured
visual story for research heavy topics.
Miners can also browse the web, collect
useful information and include cited
points. This makes it useful for
business presentation, research decks,
pitch decks, training models, and
internal reports. While the task is
running, look at the interaction view.
You can see manners creating the
outline, preparing slide content,
improving the design, building the final
deck and once the deck is ready, you can
usually download in a businessfriendly
format like Pex. So we can still open it
in PowerPoint and make final manual
changes. Now this is very important
because manus gives us a strong first
version but we can still fine-tune it
the slides based on our brand audience
or even presentation style. Now let's
move on. Let's come back to this later.
Let's see what are the outputs for all
the prompting that we have given. So
firstly I have asked it to create a
landing page for a skincare brand. Now
as you can see there is a skincare brand
where you can also edit these. So the
name is given the benefits products
purifying tensor what is the cost in
dollars. You can edit the landing page.
This usually used to take days for an
UIUX designer to design the entire page.
is just done with a small prompt. So you
have products, reviews, shop now and if
you come here ready to transform your
skin, the shops, new arrivals, colle
collection, support, etc. This doesn't
look like it's just done from a
prompting. Now if you come to the second
one, let's we had asked to compare the
top CRM tools. Let's see what's the
answer for that. So here the prompt was
to research the top CRM tools for small
businesses and create a comparison
report with pricing, key features, pros,
cons, best use cases and course link. So
as you can see let's open this report.
So here we have a summary where small
business CRM section is the best
approach as a trade-off among these easy
tools. Now is it comparing all the
things that we have given? The first one
it's HubSpot sales hub starting price
key features pros cons best use cases
and source link all the things are
present usually if you use a person they
used to browse through every single
website they could find and create such
kind of report now it's done in just a
small prompt that I've given now let's
move on to the next one which is loan
data analysis this is the most useful
tool for data analyst because we spend
hours. They spend hours cleaning the
data, visualizing trends, what graphs is
suitable for what kind of data,
normalizing the data and so many other
steps. Now, if you can just upload a
file and ask it to create all the
reports and all the things necessary to
take a business decision, this will be
the most useful tool for data analysts.
Now, let's see the answer for this. As
you can see the graphs are there. Let me
just open. You can give a prompt where
which kind of graph you want, what
against what graph you want etc. You can
see the credit history, marital status,
property area, education,
self-employment and dependency all
against approval rate. Now if you come
here there is a summary as well which is
a report. Now the summary is that the
report analyzes 614 do applications
using the uploaded loan data set
covering applicants demographic income
co-licant income etc. The data set shows
422 applications were approved
presenting an overall 68% while 192
applications were rejected representing
a rejection rate of 31%. Now as you can
see we have approval and rejected rate
and the data overview. What are the
data?
Now here this is a very small data set
and I took almost an day to work with
this data set and create modeling etc.
This is done within a few minutes and
this is amazing because it takes a lot
of time cleaning the data set knowing
the data how to understand the data.
This sorts out all the problem. Now
coming to the next one content creation.
So here I had asked manusi to research
the top AI tools for content creation
and create a company report. So here as
you can see there is a report that is
given. Let's preview it. So here the
heading is there explore tools download
the report again this is treating as
like a website that has all the
information. So here you can see tool
distribution by category text generation
tool is like one etc. Pricing tier
distribution 63.2 to AI tools directly.
The first thing is chat GPT which is
probably mostly consumed I think. Next
is Jasper AI. Then we have Canva AI and
next Grammarly Ptory Morph AI Descript
Midjourney
Surfer SEO Gemini Claude Copy.ai etc. So
here you can see the price also what is
best suited for there is a free version
also. So it's given free version as well
rating. This is amazing for content
creation because usually we don't get
pictures which give the exact direction
or exact ratio of the exact numbers that
we found online. It's either we have to
create from scratch. So this is amazing
for content like you have ratings, you
have pricings, you have to compare them,
select the top tools to compare. Let's
compare chat chibity and Jasper sorry
Jasper and chat chibity and also Canva
AI all three are equally used. Now let's
deselect them and copy.ai. Now as you
can see copy.ai is a little bit less on
ratings. Oh my god this is too good to
be a tool. This is literally AI to work.
Let's check out the next one which is
landing page for AI productivity. Now
again this also will be a landing page.
So it's basically like a website. Now as
you can see we have the heading college
study flow features pricing testimonials
FAQs study flow AI is your personal AI
tutor study planner productive companion
get instead explanation organize your
listings there's a free trial watch a
demo powerful features for the success
everything you need to excel in your
studies all in one place etc. And you
have the pricing as well. This looks
like a legit platform website that has
no flaws. There is literally a review
rating also frequently asked questions
which is common in most of the websites.
And then you have the down at 2024 study
flow AI all rights reserved. Next let's
see if the slide deck is ready. Now
let's play the PPT. It's about the
future of AI agents for business. So as
you can see first is the heading
footages. The next one is core ship with
this AI agents change the unit of work.
And then you have what and all things
are changing. Why now the agent stack is
maturing. AI agents are not just smarter
chat bots. What are the difference
between chatbot co-pilot AI agent agent
portfolio? The new team model in human
agent collaboration business teams will
adopt agents by functions scale agents
required enterprise architecture
governance adoption. It's a legit PPT to
explain each and every single step of AI
agents future. Manus AI is literally
describing how AI is put to work not
just give a text response.
>> All right, so we're ready to start. uh
we are going to start with this you know
first course which is going to study the
basics of Python. Python will be our
primary focus for the entire program. Um
we will use co-pilot. So there's there
will be co-pilot material um later on in
the program but like in this first
course we're going to be focused on
Python and and for most um things we
will be using Python. Um even when we
use co-pilot it will produce Python code
everything we do will be in Python. I
think one of the things is by you know
by the end of the program if anything
else you guys will be in a much better
position with Python. You'll be better
Python coders by the end by the end of
the program. If you don't learn anything
else you'll get better at Python. I
promise. Uh because that's you know all
of our examples all of our demos
everything we do will be in Python. So
you'll you'll get better at it. uh for
sure and we'll have a lot of practice to
do that. Okay. So this first lesson is
all about an introduction to what Python
is. So if you're completely unfamiliar
with it, totally fine. We will uh get
you up to speed and talk about the
fundamentals and how to set everything
up on your own computer and talk about
the various ways to um utilize Python.
that some of it will involve a setup you
can do on your own computer. Some of it
will involve some cloud resources um so
that you don't need to set anything up
on your computer if you don't want to.
Um we'll have options there which will
be nice. So I will show us those and
walk us through those. But this first
lesson all about the basics uh and
getting set up. So, um what's
interesting is like at the beginning of
every lesson, we usually have this uh
kind of um engagement or discussion. Uh
but you know, we've I kind of already
asked you guys about this of uh uh if
you're familiar with programming, if
you're familiar with Python. Um but one
thing I want you to think about a little
bit is that um especially as we go along
and learn about what Python is is why is
Python the
chosen language for AI? So why is it the
one that everyone uses uh to do AI? And
I think what you're going to learn is
that it has a really amazing ecosystem
that has been around for a long time
that um supports AI in particular. So,
Python is the go-to for anything AI,
data science, machine learning, anything
in that sort. Uh, because it's been used
for so long for that and it has such a
uh community and ecosystem around it.
That's something we're going to learn.
It's also really easy to learn and use,
which makes it nice to to be uh kind of
an introduction to the field. It doesn't
take a lot to get started in it.
because it's so easy to work with. Um, I
can tell you as someone who's gone
through that experience, like I studied
mathematics in college and in graduate
school and studied like probability and
statistics, but I was able to teach
myself Python primarily and use that to
get into kind of data science and
machine learning in the industry.
So, and I think that's a common story is
people and I've seen that from many
learners coming from uh different
backgrounds. Uh they've been able to
pick up Python pretty easily because
it's a very easy language to understand
and and syntax of it and there's so many
tools within it that make it really easy
to work with.
So, um I promise it won't be as uh
daunting as it may seem even if you're
coming at it from zero experience. Uh, I
think you'll find this is the perfect
way to get into programming and get into
data science and and AI and machine
learning because it's so easy to pick up
and learn and it has such a nice rich
community ecosystem.
So, just wanted to mention that.
Okay. So, some of our objectives for
this first lesson will be to talk about
programming languages in general and um
programming in general. So maybe you
know more generic than Python just you
know what are what do general programs
look like? What are some of the building
blocks of programs that are important?
What are some of those uh key principles
of programming that we will want to
follow as well? Even if we're doing
Python for AI purposes.
Um so just talk about programming in
general and then kind of zoom in on
Python as we go along. One of the things
we'll be interested in doing is just
getting you guys set up. So talk about
how we can configure Python for you to
use on your own machine. Um but also
have some options that don't require
installing anything on your own machine.
Uh which is nice. Um and then as I said,
we'll kind of zoom in on Python, talk
about its benefits, uh some of the nice
features. I've kind of already mentioned
it. Really big community around it, easy
to learn. We'll just talk about those
more in detail. Talk about um why it's
so popular in the AI world. Um,
and then we'll get into some very
fundamental things specific to Python.
So once we talk about the background,
get you guys set up, we'll go into uh
some of the syntax basics, things like
identifiers, things like indentation,
comments, um, some of the basics of the
code that are going to be important for
you to kind of get started with. Um and
then talk about some of the basic data
types that Python offers to manipulate
and work with data which of course is
important um when you know as we go
forward and and do anything with data
which of course with AI we will be
interested in doing um but that's these
are the objectives of just the this
first lesson. As we go forward we're
going to learn about many other basic
topics within Python. So things like how
to write functions, how to build
objects, how to manipulate our flow of
the program with like things like if
else statements, things like loops.
We'll learn all about that in kind of
the next lessons after this one. But
this is all the content for this lesson.
I anticipate today
we will get through all of this today
and then get into the second lesson
which will um get into those kind of if
else and loops. So we'll we'll get we'll
I'm sure by today we'll get into those.
All right. Any questions on kind of what
we're going to learn in this first
lesson? So mainly trying to get you guys
set up, give you some background on
Python and then towards the end of the
lesson um get into some basics of the
syntax is kind of the goals I would say.
Okay. Okay. So when we talk about
programming um what do we mean by
programming in general? It's really uh
synonymous with instruction. So
programming really means giving or
writing instructions for a computer to
perform tasks. Um so these instructions
we write down in what we call code. But
those those are just telling the
computer what to do. And of course the
computer's not going to do anything
unless we write down these instructions.
So these instructions can do really
powerful things. They can power, you
know, whole applications, things that we
use every day like Microsoft Word,
PowerPoint, Excel, those kind of things.
Um they can automate tasks. They can um
power websites. Um they can do AI,
right? So we can have um things like
chat GBT and Alexa and Siri, etc., etc.
Um these are all powered by instructions
telling the computer what to do.
One of the things that we will get
better at as we go along is figuring out
how to write these instructions in
Python. Python is going to be the
language we write those instructions in
um and and they will be executed by a
Python um program. But we should think
of programming in general as just
instructing the computer what to do just
at a high level. Right?
So when we talk about these
instructions, they have two ways of
being executed by the the computer. Um
and roughly these break down into what
we call interpreted languages and
compiled languages. So that the code
that we write which is um representing
the instructions that we write can be
executed um in one of these two ways.
Let me start with the left. So the
interpreted languages.
This means that the computer is
literally executing the the instructions
line by line by line when we run the
program. So there is no
translation of anything. It's just
literally taking our instructions and
running it line by line, instruction by
instruction essentially. Um, now the
advantage to doing this is that it's uh
easier to debug because the instructions
are going to be executed one by one. So
it can hit an error pretty quick. If
there's a mistake in one instruction,
nothing else will run. Um, however, it's
also slower because we're going to take
it one instruction at a time. Um, and so
the the uh this way of running programs
tends to be slower, but it's also easier
to work with, which is why we're so
interested in Python. It's in this
bucket of what we call interpreted
languages. So a lot of scripting
languages find themselves in this bucket
of being executed one line at a time. No
translation needed by the machine. It
just reads our instructions and executes
it. The thing that does the execution is
called an interpreter.
Um, and Python has an interpreter that
we will get you guys set up with on your
own machine that can execute Python
code. So you need an interpreter. The
interpreter just executes your
instructions line by line by line. Um,
so some examples would be like Python.
That's what we're going to study in this
um entire program. But there's other
languages like JavaScript, Ruby,
um Pearl, many others that are uh
interpreted. They require an
interpreter, but they execute line by
line by line and there's no intermediate
translation of anything. Um it's kind of
executed as is. Now, contrast this with
compiled languages, which are uh kind of
a different piece. they these these
instructions have to be translated into
something the machine can understand in
order to execute. So there is an
intermediate step of what we call
compiling the code um into uh basically
a translated version of your
instructions so that the machine can
execute it. Now there's a trade-off
there. Doing that can make it more
difficult to develop and it can take
longer to debug because you have to go
through this translation step every
single time through the compiler.
But when you run the code because it's
already been translated into this
machine format, it's a lot faster. Um,
so some examples of languages like this
are C, C++, Java,
um, Go,
but uh, we won't really be working with
those. We'll just be sticking with
Python. But if you have experience with
those languages, those you're probably
familiar with this, you have to compile
the program first before you can execute
it. But we are going to be in this
interpreted world. If you know and it's
okay like if none of this makes sense,
that's okay. Just understand that um
generally interpreted languages are
going to be more user friendly because
they're they're easier to execute. They
don't require as many moving parts as
what a compiled language would require.
which is nice for us, right? Nice for
Python. That's what we're going to be
interested in working with. Uh kind of
um yeah, they're kind of rel So, so the
question is are JavaScript and Java
related? Kind of. Um, JavaScript is kind
of like the um the the
scripting version of um some of the same
concepts we see in Java, but Java is the
compiled um it it requires a a special
kind of what's called a Java runtime,
which is a a compiler to translate the
Java code into um machine code that the
Java runtime will execute. JavaScript is
not like that at all. It can actually be
ran in a web browser which is um
JavaScript usually powers a lot of like
front-end websites are usually powered
by JavaScript and Java usually powers
more like backend
um applications like actual software
programs are usually would be coded in
Java. JavaScript is going to be used
more for like building a website. But,
you know, I'm not an expert on that
really, but that's kind of my
understanding of it. And if anyone is an
expert on those differences, feel free
to let us know in the chat. But, uh,
that's my that's my basic summary of
that. Okay. So, we have interpreted
languages. That's where Python falls
under. So, it just um summarizing that,
it's going to be easier to work with
those, which is great for us. That's
another reason why Python's so easy.
It's interpreted, meaning that
everything executes. We don't need to
worry about compiling things, which is
nice. Um, but also in terms of
programming, there's also uh categories
of how the instructions are written that
you can bucket different languages into.
So for example um some language are are
more um procedural in nature meaning
that you write out all the instructions
exactly kind of line by line by line.
You don't really organize things at all
in your instructions.
Um so some examples would be like C and
Pascal
are more like that. Um then on the
opposite end of the spectrum is kind of
object-oriented
in which case you uh build your code and
organize it around the idea of
everything being an object. And so some
uh Python actually falls into this
category where um uh most things in
Python are objects and you manipulate
objects and objects have data to them.
They have things they can do and
interact with other objects. Um, so
think of it just as a way we will
organize our instructions.
Python allows us to organize it around
the concept of an object. We'll learn
about what that means as we go along,
but just realizing that some programming
languages break down along these um kind
of buckets here. Um, Python is also a
scripted language, meaning you can write
out your code in a individual script and
you can e that you can have an
interpreter that executes that script.
Um, so you don't need to organize all
your code inside of an object. So for
that reason, Python super flexible.
That's another reason why it's so nice
to use. It actually falls into both of
these buckets on the right, which is
very convenient. We can have basically
this means we can have a lot of
organization or very little organization
depending on how we want to set it up.
Yeah, Roberto. So even though there are
different types so Java is compiled and
Python is interpreted
um they are both object-oriented meaning
so think of the this slide as telling
you how the instructions are organized.
So how they are executed is different.
So, Java requires a a compiler to
execute things. Python requires an
interpreter.
This is more about how the instructions
are organized. So, Java and Python both
allow you to organize your code into
objects.
Um, but what's nice about Python is it
also falls under the bucket of
scripting, meaning that it allows you to
organize things into scripts, which is
less organization than it would be in
into objects. We're actually going to
learn about objects later on in a future
lesson, like how to build objects and
what they mean.
So yeah, even though they're different,
they're both object-oriented, which just
means that you can organize your code
into objects. Python allows that. So
does Java. So does C++.
Uh many many languages allow for um
organizing your your code into objects.
So we're going to learn about that.
It's It's not that one's better. They're
just um I I would put them at different
So, let me draw this. I would put them
at different spectrum, different ends of
the spectrum on organization.
So, scripting
is very loose. Basically, you it's more
like a an individual um uh set of
instructions to do one task. you can
just have and you can have many
individual scripts to do many small
tasks. Um, and then on the other end of
the spectrum, think about it as like
you've organized your cabinet into many
folders and many like uh you know many
pieces of organization that are we would
call objects. Um so objectoriented
programming OOP is kind of on the other
end of the spectrum when it comes to
like level level
of organization.
Does that make sense? So scripting very
loose. It usually scripting is is um
reserved for like one task and it's um
you're just writing out your
instructions to accomplish that one
task.
um which is helpful for like automation
of things because you're going you're
usually automating like a single task.
Um so it's very loose. It's not very
organized into nothing is organized
necessarily into objects. Um very loose
organization. Object-oriented is much
more structure to it and things being
put into objects um in order to
manipulate and work with objects
throughout the program. Yeah, it's not
that one's better. I think it's more
just use case dependent. Um there are
times where it actually will benefit us
from using objects. Um and I think the
thing to pay attention to on this slide
is that look at where Python falls into.
It actually falls into both. Meaning
that we can have things very loose and
easy to work with because scripting
usually will be faster and easier to
just write something to to accomplish
one task. But we have the flexibility to
organize our code into objects if we
want to. which will be better for
bigger tasks that require more
organization
like training a neural network or
building an LLM.
Those bigger tasks would benefit from
organization.
And then uh finally on this slide um
there are languages that are built on
the concept of um their their entire way
of writing instructions is more in a
functional way meaning everything is
based on operating uh functions and
variables. Um and so there are some
languages like that has and scholar are
very popular ones. Um but that is can be
very difficult to learn. It's it can be
difficult but very nice in some ways
because uh it can be very natural to
think of um you manipulate like giving
instructions to computer in a functional
way. Think about it as like applying a
function to a variable.
Um that makes sense but writing your all
of your instructions in that way can be
kind of difficult to learn. So for that
reason I think these languages are more
difficult to learn but they can be very
powerful. Um and they find themselves
very useful in like operating on big
data.
Um so if you ever heard of like Spark um
Spark operates with uh Scola for
instance um but uh we won't really focus
on functional. It's kind of its own
paradigm.
Um but uh again like Python is where our
focus will be. It allows us to be really
organized, loosely organized. Nice
flexibility there.
So, so far
based on these two slides, I'm showing
you that Python is interpreted, which is
easier and faster to work with. Um, not
faster to run, but faster to get up and
running because you don't need to
compile things. That's nice from our
perspective.
And it's also has very good flexibility
when it comes to organizing our
instructions, organizing our code can be
very loose in scripts, could be very
structured in in objects.
Okay. Okay. So generally no matter how
uh no matter what language it is um when
you process those instructions generally
things are going to be organized
even if it's in a script or if it's
object-oriented
um you're generally going to have the
very beginning of the program um kind of
setting up the input then the middle of
it really processing that and doing
something with that. So that's usually
like the bulk of the logic is in the
processing phase and then generally
you're producing some output. So that
could be like a model prediction, that
could be um a a graph that you've built
from your code um whatever that output
is. But generally it flows this way.
This is this is makes sense, right? Of
course there's input, you're
manipulating that input in some way and
then you're producing some output. I
think that all makes sense. That's a
very logical way to flow.
Um
now that's not to say that within this
processing step there may not be
um iteration like of course there may
may be times where we need to as part of
the processing kind of iterate and do
multiple passes of processing. Um so the
processing could be a lot. We could be
doing a lot. We could be doing a little.
Just depends on what we're actually
doing. So, if we're reading in some data
as the input um and then we're just
doing some simple um slicing and dicing
of it, that's some easy processing and
maybe producing a graph or producing a
metric, something of that sort, that's
pretty easy to do. But if we're training
a neural network or training a model,
the processing step can take a while and
it may, you know, be very iterative in
nature. So it just depends on what we're
doing and those instructions.
But no matter what, most of our programs
will flow in this way kind of input
processing output. It makes sense. It's
very logical.
So what are some principles that we
should abide by when we're writing our
code? So this this would really be for
any language, but of course for Python
that we are interested in. Um so
something we're going to be interested
in doing is um basically avoiding
repetition where we can. So instead of
having copy paste everywhere, we will
generally favor organizing our code to
some degree. Meaning we will utilize
functions where it makes sense and
objects where it makes sense to organize
things. And also instead of um having
very repetitive code, we will favor
using uh loop structures that can
iterate over um things many times
instead of us us having to write all
those out one by one by one. So we're
going to learn about these tools that we
have at our disposal, but they will help
us organize our code, avoid repetition
all over the place. One of the things we
want to avoid is having the same code
repeated all over the place. If if we
find ourselves doing that, we should
really put that code into a function or
maybe into an object so that we can
reuse it. So, we're really going to
favor like reusability of things,
re recycle, reuse, you know. So, we're
going to learn how to do that, how to
build functions, how to build objects.
But that's something we're going to
favor uh when we're when we're
programming. It's something you should
be on the lookout for. If you find
yourself writing the same code over and
over just in different spots, um that's
probably a clue you should organize that
into a function so you can just call
that function wherever you need to
rather than copying all that code. Okay,
so we're going to avoid repetition.
Now, the the reason we're going to do
that is to uh you know keep everything
simple. We want to make sure things are
clean, simple, understandable.
Um, we don't want to h we don't want to
have overly complex things that are very
difficult to follow. So, one of the
things that is going to be really nice
about Python is it lends itself very
well to being simple because it's going
to be so easy to actually read and
understand um, you know, understand
what's going on. But one of the things
that falls in line with this is like um
for instance naming things
appropriately. So instead of just
calling everything in our code like X Y
and Z if somebody comes along and reads
oh I see your code has an X Y and Z that
may not make sense. You know we would
want to be more thoughtful with the
names of our variables and names of our
function. So instead of XYZ, maybe we
would use something like name or place
or you know something appropriate to
identify this is what this is. So think
about that when you're writing your code
is try to make it understandable. Name
things that somebody else reading it
would understand what it is if they see
that name. So that's that's a mistake I
see a lot of people make when they first
start. It's okay like when you're first
getting started and practicing to name
things like X, Y, and Z. I think that's
fine. Or like ABC.
Um,
but does that make sense? Like if
somebody else was reading it, they see
XYZ in the program, that may not make
sense, you know? So, but if it has a
good name to it, you could say, oh, like
I see this is somebody's name that this
variable is referring to or this is um a
particular object that this is referring
to. Um, it's not just kind of an
abstract X or Y or Z.
Yeah, no spaghetti. Yeah, that's that's
what uh that's what a lot of people
refer to that as. Uh just sloppy,
unorganized, um hard to understand code.
One of the things that's great about
Python is it's naturally very
understandable. So like I don't think we
will have that issue as much as if we
had other languages, but it's still
possible.
So these are things we'll learn as we go
along. I'm just trying to get it into
your mind a little early here. Name
things appropriately is main one of the
main pieces of advice I can give here.
Um
so the next tip is to organize things.
This goes along with avoiding
repetition. So organize
um let's put things into functions.
Let's put things into objects where it
makes sense. If we know we're going to
reuse that um let's put it into a
function. And so we're going to learn
about how to do that. But generally this
is good practice if you find yourself
writing um uh code to do something and
it turns out to be um
it turns out to be uh something you know
you're going to reuse or it turns out to
be more than a handful of lines of code.
Generally you want to organize that into
a function so that uh it's clear
this is what this code is doing. This is
what it's responsible for. it's obvious
um you know that it's organized into
into uh that unit of work essentially.
So we are going to practice this. This
is something we're going to get good at
I think as we go along because we're
going to favor organization where it
makes sense.
Okay.
So readability. One of the things is
using meaningful names. I kind of
already mentioned that. The other thing
is using good comments. So, we're going
to learn probably today how to make
comments in our Python code, which is
going to be helpful to orient yourself
or another reader of it to, hey, this is
what this function does. This is what
this line of code is doing. Um, I can't
tell you how many times, you know,
people write code and then it they
themselves come back to it a week later
and have no idea what it's doing. That
happens all the time. It's even happened
to me. So, uh, comments are your friend
in that regard. and that um they don't
really cost you anything to put comments
in there um to say to to kind of
highlight this is what this piece of
code is doing and you can make a note to
yourself right within the code. That's
what comments are. They're basically
notes to yourself. Um
so we're going to learn about that
today. How to write comments and and
what that looks like in the code. The
other thing is indentation. you know,
Python supports uh indent like you have
to indent. So, that's not really going
to be an issue. Some languages don't
really support that, especially the
compiled ones. They don't enforce
strictly indentation. They enforce other
things like braces and and semicolons
and such, but um our our Python code
will be properly indented uh by
necessity because otherwise it won't
work. So, um, that's something we're
going to learn about too today is how we
indent things and why that matters.
We'll talk about that.
Um,
I see a question from Sherry. Is Python
a program that can be programmed with
simple language? Yes, it's very easy to
uh it it's Python is a very natural
language to program in because um yeah
it's very simple uh simple languages
used all over the place. I think it's
going to be really easy to learn. I
think it'll be really easy to pick up.
At least that's my hope and I think it
from my experience it is. As I said I
was someone who did that and I've worked
with many learners who've done the same.
So yes, I think it'll be pretty easy to
pick up, very simple.
Um, and then the other thing is we can
do uh we can find our errors very
quickly. Now, because this is an
interpreted language, we can run things
one line at a time and we we will
quickly hit errors
uh early on in our code if if we have
them. So this will be nice and Python
provides really good um error messages
um to say hey like this is what's wrong
with your code you should fix it this
way um essentially like giving you a
clue into what needs to be fixed. Um so
so this is something uh that we will
practice with as we go along is kind of
um finding errors and what to do with
them. Um, but because it's interpreted,
we will run across those very quickly.
Unlike with compiled language, which is
harder to debug because you basically
have to compile everything, hope that it
compiles. If it does, then you have to
run things. Um, it just takes longer to
get through that debugging phase. But
with the with Python, it's very quick.
You get a very quick feedback loop on if
your code's working or not, which is
nice. A lot of votes for C. I agree. AC
is the correct answer here. So the
interpreter is the thing that will
execute the code line by line. So it
doesn't do everything at once. It
actually goes line by line, which is why
you can stumble onto your errors quickly
because if you're going line by line
um and you have an error on this first
line, you're never going to reach these
other lines, right? You're it's just
going to show you this is where your
error is. it's on line 101 or whatever
it is and you know it's going to show
you where the error is. So it's going to
go one at a time and execute those. Um
it's not going to convert the code into
machine language. That's what a compiled
language would do, not an interpreted
one. Um and uh they do require an
interpreter. So D is just completely
wrong. It's the opposite of that. It
does require it. So the interpreter is
the thing that is executing the uh code
line by line. So what is Python in
particular? So it is a as we've already
seen an interpreted language meaning
that it requires an interpreter to
execute it. It's going to be executed
line by line by that interpreter. Um it
has capability to be object-oriented. It
also has capability to be scripted.
um which is just in relation to how it's
organized. One of the really nice things
is it is what we call dynamically typed
or what you would say dynamic semantics.
We will see what this means but
basically it means that we don't have to
declare what every piece of uh what
every variable or every piece of data is
inside of Python. We can let the
interpreter interpret that which is
nice. It makes things really easy to
work with. We don't need to say okay
this is an integer. This is a
floatingoint number. This is an array.
This is you know with a lot of program
especially compiled languages
programming languages you have to do
that because you have to tell the
compiler this is what this piece of data
is. But with an interpreter the
interpreter can as the name suggests
interpret that. It doesn't need to know
what everything is in terms of its data
type, which is which makes it really
easy to code. On the cons of that, it
can make it more prone to error because
you're not really enforcing types. So,
there is somewhat of a trade-off there.
But um for our purposes the dynamic
semantics make make it so that um the
interpreter can dynamically understand
what data is um based on how it's being
used which is great um for us like it
makes it just quicker to get up and
running and started and and working with
data. We don't need to declare what its
type is which is um static semantics.
Um now Python itself amazing programming
language that's used across many
different applications um such as data
science, automation, machine learning,
AI. It's also used in to build software
even um not sure if you guys know this
but there's um some really famous
software that's written in Python. Um,
one of the most famous is Instagram at
Meta is completely coded in Python,
which is it's over like 20,000 lines of
Python code, which is pretty amazing.
But um so of course it's been really um
heavily used in AI and machine learning
and such but it's also as a programming
language been used for other things like
more pure software applications which is
what makes Python really nice is it's so
simple so easy to learn. Um so for that
reason uh it is going to be great for us
to get started with especially if you're
coming in with basically no programming
experience. The other thing about Python
is it has uh as I said earlier like a
really big ecosystem uh meaning that
there's many different packages and
modules within those package packages
that do things already. So we don't
what's great about Python is we won't
need to reinvent the wheel on so many
different things like if we need to
build a plot if we need to train a model
and and use a specific type of model
that likely already exists in a package
somewhere. And what's great is they're
almost always open source meaning we
don't have to pay for anything. You can
just use it out of the box which is
fantastic. So there's within Python
there's so many ways to do things
especially in the AI and machine
learning world that we'll just borrow
those and use them in our own code um
which helps uh you know with um getting
up and running very quickly. We don't
need to reinvent things. We can just use
things that already exist um which is
fantastic. So that ecosystem really
benefits machine learning AI. Um because
they they already exist. We don't need
to spend our time rewriting all those
things. Um and so that's something we're
going to learn as we go along is like
how to install those, how to import
those, how to use those in our own code,
those those packages that already do
something for us. So we don't need to
come up with it on our own. we just need
to use it properly. Okay, so there's a
little bit of history. Python was first
invented in the late 1980s by a guy
named Guido Van Rossom in Amsterdam. Um,
where it gets its name is after the old
comedy series, you guys might be
familiar with it, the Monty Python
Flying Circus Show. Um, and so that's
where it's got its name. um you know it
was first created then but has since
taken on a really big role in the
especially you know I keep saying in the
AI community so much so that it has its
own software foundation that kind of is
responsible for maintaining it they meet
regularly they come up with improvements
um they come up with new versions of
Python
uh for example Python 3.14 just released
in October which is a major release. Uh
they hadn't had one in a while and that
one is uh 3.14. So it's kind of known as
Python.
Um which was a big milestone. Um but you
know they have uh they've had many
different versions over the years. It's
been maintained and developed by this
software foundation. Um and people are
actively working on it at many large
companies. So for instance, Meta has a
big group that is working on um Python
improvements. Microsoft as well, um
Google, all of those guys have groups
kind of working to improve Python
because they all use it. And so what
they typically do is work on it, open
source it, and then the community gets
to use those tools, those packages,
those tools, those improvements. Um so
it's it's actively um utilized across
many big companies actively uh
maintained by them or contributed to by
them. So that's that's really great. Um
you know Python was originally derived
from other language um other languages
uh as kind of a trying to find like a
mixture of some of the best of all
worlds. But its main like driving force
in why Python came to existence from
these other languages is it just its
ease of use. People really wanted
something like super easy to get up and
running and something really natural.
Um and so we will as we start learning
the syntax of it I think you guys will
understand why it's so easy. But um
that's that's what led to the
inspiration is just people wanted
something easier to work with not as not
as uh strenuous to kind of get up and
running.
What open source license is it? Um,
that's a good question. I think it's the
MIT license, but I could be wrong on
that.
You could look it up. If you go to
python.org.
Yeah, if you go to python.org, I think
it might talk more about what the uh
license structure is there. I want to
say it's MIT open license, but
I've I'm really not 100% sure on that.
Okay. So, what are some of the benefits
of working with Python? And these are
things you will experience as we go
along, but just wanted to call them out.
Um, the flexibility of it. As I said, it
can be really organized into
object-oriented or it can be loosely
organized into scripts. So, that
flexibility alone is really awesome. um
which has allowed it to power many
different things like um APIs, web
pages, full-blown applications like
Instagram, um chat, GPTs, like actual uh
AI, LLMs.
Um you know, it has so much flexibility
there to power so many different
applications.
Um probably the biggest benefit,
especially to us, is its ease of use.
Um,
uh,
oh, thank you. Some Tim just posted it.
It's the the GNU,
uh, public license. Yes.
Oh, never mind. It's a Python software.
It has its own. Okay, perfect. Thanks
for sharing that. Thanks for sharing
that. Yeah, I wasn't completely sure
which which license it was, but it is
open source. Um, and people do make
their own kind of derivations of Python.
But as I was saying, one of the benefits
of Python is how easy it is to learn. I
keep emphasizing that because it's true.
Once we get into it, you will see this.
I promise it'll be easy to learn, easy
to pick up. Um, and it's designed in
that way. Designed to be very minimalist
as a language, which is great.
um it has a lot of things that come with
it and it's kind of built into Python, a
lot of capability. So we call that the
standard library. It's just the things
built into Python. It has a lot of
capability out of the box. Um you know,
not only that, but it has a large
community that's developed so many
different packages that do things for
us, especially in the AI world. So
that's another great thing, kind of a
robust community developing these
packages that help us get things done.
Um, readability. So because the code is
so simple, it's also easy to read. So
you can usually read other Python code
and quickly understand what it's doing
which you know makes for easy um easy
understanding of other people's code
easy understanding of code in the
community and kind of almost like it's
selfdocumenting because it's so easy to
read. So that that simplicity that ease
of use lends itself well to being really
readable. You can usually just take a
look at the code, easily read it,
understand what it's doing, which is
great, like great for you guys learning,
great for taking a look at the demos and
examples that we will do. They're very
readable.
Okay. So why has Python really dominated
AI? So this is a valid question like
even so it's used for many different
things. It's a programming language. So
it can build application and I've given
you the example of Instagram and there's
many others um that are built off of
Python code. Why is it so useful for AI
in particular?
mainly
uh some of the reasons we've already
talked about mainly how easy it is to
use lends itself well for AI because um
that has allowed people to kind of
quickly get up and running and test out
their algorithms, test out their models
just really quickly with Python. That's
great. The other things listed on here
are certainly big reasons as well. So
for example, it has so many community
libraries, those those packages that um
have AI models and AI tools that we can
reuse that people have built these up
over years and years and years. Um so
it's to our benefit to reuse those and
not have to reinvent everything and we
can get quickly up and running with
those which would be great.
The other thing is Python, it lends
itself very very well to working with
data in general. Very easy to work with
data, very easy to load it in from
external sources, query it, work with
it, visualize it. Python is so adept at
that. Um, so that's what makes it really
nice at doing machine learning and AI
because so much of it is manipulating
data. So, um, for that reason alone,
Python is so popular in the AI community
just because of its ability to work with
data. It's so easy. This is something
we're going to really focus in on like
in our next course when we talk about
data science.
But, um,
just the ability and the power of it to
work with data makes lends itself well
to AI uh, capabilities.
Um, the other thing is I mentioned the
rapid prototyping. You can quickly build
a model in Python because the code is so
easy. So, and there's so many libraries
already can quickly prototype. Um,
it has obviously a big community around
it that's building out these packages,
writing documentation, maintaining it
from an open source level. So, that's
another reason it's very popular. Um,
Python's also used with other
technologies. So, it does have
capability to integrate with other
languages. So for instance, Python can
one of the most popular integrations is
Python can work with C and C++. So
sometimes that's necessary to integrate
with those to do certain things. Um so
Python has been extended to work with
other languages. So sometimes there's
other uh necessary support from other
like things in other languages that are
necessary to power something in AI. um
for example working with GPUs
and doing things in deep learning. Um
there's been a lot of integration with
uh working with um C tools. Now will we
do that? No, it's already been done for
us and some of these packages. But um
the pure ability of Python to do that is
really powerful and it gets taken for
granted honestly because you don't see
that it's underneath the hood and it's
abstracted away from you when you work
with those Python packages. But there
was a lot of work that went into it to
integrate it with other kind of other
programming languages.
Okay.
So as an example like I mentioned the
Instagram one. So Netflix for instance,
all of their recommendation is powered
by Python. So when you open up Netflix
or really any streaming service for that
matter, they're going to use Python to
deliver those recommendations and
produce those personalized
recommendations. Um Spotify as well for
like music. Um nearly all recommendation
algorithms are written in Python.
And in this program, we are actually
going to learn about recommendation
systems. So that'll be pretty fun way
down the road when we get into machine
learning. We'll talk about how do we
build a recommendation engine,
but um they're all done through Python
for for example. So really cool uh use
cases there.
So one of the things I wanted to address
is how AI itself is changing coding. So
you guys may be aware of this, but
obviously there's been a huge um kind of
explosion in generative AI tools that
can help write documents and write
emails and write text and all these
things. One of the things they can do is
write code. So um one of the big areas
where AI is changing coding is it's an
its ability to generate code for us. And
so um throughout this program like we
won't shy away from that necessarily
and I encourage you guys to use AI tools
as you see fit to help your own
understanding and help your own
productivity. Um
you know we still will go through the
fundamentals so you can understand it
but the AI tools can definitely be a
supplement to help. Um it's just that I
think you guys will understand it better
going through the examples that we do we
do together and so that when AI
generates code you will be able to
understand it and also be able to debug
it right because it's not always going
to be perfect. So that's always the
catch with AI is that you know it
doesn't always produce perfect answers.
Um but the at the very least we will be
able to you know debug things and
understand things better so that uh we
can catch those errors.
Um so obviously like AI is also besides
flat out generating it it's also
suggesting what should be there. So, uh,
some of the code editors really do a
good job at that, suggesting things, um,
picking up on what you should produce
next. That's going to be, um, very
interesting as we get into, uh, some of
the platforms that you guys will work
with to write your Python code. They
will have that ability. Um, so, uh, the
other thing is like there's some cloud
tools that, um, don't require writing
much code at all and they can just do
things. So, in other words, you can
power them by prompts. You're not really
writing code. You're just writing
natural language and then they do
something. Um, they generate the code in
the background and they execute
something. Um we will learn about those
things uh later on in the program
especially because we we will cover
generative AI in the future
um towards the end of our program. So if
you're wondering like are we going to
cover LLMs? Are we going to cover how
these things get generated? Yes. It just
will be um later on in the program.
Okay. A lot of votes for B.
Yeah, pretty unanimous on B. I think I
agree with it. Yeah, B is definitely the
right answer. So, all of the
recommendation systems which we will
learn how to build ourselves later on
are written in Python and um they uh are
machine learning models that make the
recommendations and that machine
learning is driven by data um and all of
that data is manipulated in Python
um and used to train uh models that do
the recommendations. That's all
happening in Python.
So, we're going to talk about getting
you guys set up on your own machine and
talking about the different development
environments we can use to actually work
with Python code. Um, before we go into
that, any questions about anything we
covered so far?
Everything's good so far. Yep. And you
know again if you have experience in
Python I recognize that it is going to
be a little slow in beginning. Um it's
mostly to get us really oriented to some
background around Python and get us set
up and then we will be doing you know uh
getting into the syntax and all that uh
coming up shortly. So we will actually
be learning Python specifics coming up
soon. But you know we're going to um get
everything set up first.
All right. So, let's continue then.
Thank you guys for that.
So, um it turns out that there are many
tools in the community for developing
Python code. And so, um you might hear
this word ID. It is short for integrated
development environment. This is a piece
of software that helps you write and
test Python code. So, and there's many
out there. There's a bunch on this list.
We are going to focus on a few options.
There's even more than what's on this
list, but we're going to focus on a few
options. These IDs are designed to
really help you write Python. They
provide many tools in the background
that make your life easier when you're
working with Python. So, for example,
they can provide syntax highlighting.
They can tell you when you have a syntax
error. Um, almost like a spell check for
Python.
Um, they can help you run Python code
right within the window. Um, they can
help you organize your projects. Uh,
they can do a lot of different things.
Um, and so there's many tools out there
that can do it, and it's really a
personal preference which one you use,
but in this program, we're really going
to focus on a few of them to to showcase
those options because they're very
popular options. Um, and then, uh, allow
you guys the flexibility to choose which
option makes the most sense for you. So,
generally, that's going to be mostly a a
preference.
um mostly a preference as to which one
you're the most comfortable with, but I
want to give you guys the option to uh
explore
the various options that are available.
Uh Roberto, is there one that stands out
as an industry standard? Um there's a
couple that you see like honestly the
two of them that we will study uh in
this coming up in the next few slides
are the industry standard which are
going to be VS code Microsoft VS code
and then Jupyter notebooks. So these two
are going to be uh ones that we will
study in particular and use throughout.
Um
so so yes we will those will be industry
standards. PyCharm's also very popular.
Um so I don't want to rule out PyCharm.
I know a lot of people who use it. So um
I would encourage you to explore PyCharm
as well if you want to but we are not
going to do that uh in in these slides
but um I would check it out and see if
you like it. Um it's another very I'm
putting a an asterisk next to it because
I think it's one of the more popular
uh yes uh yeah we're going to do
descriptions.
um requirements uh I'll try my best to
give those but honestly the requirements
will be given when you install them. Um
so the other thing I want to say is we
will have a couple options that don't
require you to install anything. So I'm
going to showcase those as well. So
there's a couple options that are um we
won't have to install anything because
they're going to be cloud-based.
Okay, I'll show you those.
Okay. So, but yeah, VS Code, I think VS
Code and Jupyter notebooks are are
probably the industry standard most
popular uh idees.
Okay. So, what we would recommend in
this program and the ones that we will
use the most uh throughout are going to
be these three. Visual Studio Code, also
known as VS Code, Jupyter Notebooks, and
Google Coll Collab, which is Google's
hosted
um Google's hosted version of notebooks
essentially. Um so
I will showcase each one of these and
give you some examples of how to set it
up and examples of how to work with it.
Um, and that's what we'll do over the
course of the next few slides and the
next uh bit of time is I'm going to go
through each one of these and kind of
show you what you would need to do to
get it set up. Um, now that being said,
excuse me, these two are ones that you
will install.
These two you would install locally on
your on your own machine.
And this one is um uh cloud hosted
by Google and it's free. Um all of these
are free but uh the first two VS code
and Jupyter notebook you would install
on your own machine. Collab you would
just access through your web browser. It
is hosted by Google. So that's an
advantage. You don't really need to
install anything. And for that reason um
sometimes we will favor Collab. Uh and
for other reasons too. Collab has some
really nice features if you've never
used it. Um, but notebooks, um, Jupyter
Notebook and Collab are very similar.
They're very similar. Collab just has
its own spin-off on on the notebook, um,
type of file that Jupyter Notebooks work
with. And it's um, like I said, kind of
cloud hosted. So, I'm going to I'm going
to walk us through each one of these and
explain to you what they do, what they
look like, and then we will um I'll set
up each one of them uh kind of in a live
demo so you guys can see. Um but uh we
throughout the program, it will really
be up to you which one of these you want
to use. There's no hard requirement to
use any one of them. It's really going
to be your preference which one of these
tools you want to use to work with
Python. Whatever one you feel
comfortable working with, that's the one
you should use.
All three of these are very popular in
the industry. So, you're not missing out
by using one versus the other. Um,
they're all very popular. Even Collab, I
know it wasn't on the screen, but it is
widely used in in the community and the
industry.
Uh, no system recommendations for
training LLMs. Um, no, because we don't
we won't really focus on that until the
end. When we get to when we get into
generative AI, we'll talk about that.
When we get into generative AI, we'll
talk about that.
So, yeah, we're not we're not focusing
on LM in the beginning. That's that's an
advanced topic for us.
What is my personal preference? Um, I
like Visual Studio Code. Um, personally
I that's what I use for my day-to-day
work is uh Visual Studio Code. I like
Visual Studio Code and I like Collab a
lot. Um, so you know, we'll talk about
this, but one of the advantages to
Collab is that it has free access to
GPUs, which is huge for doing things
like uh neural nets. Um, so we will lean
on collab quite a bit later on
uh later on when we um actually get to
deep learning and neural nets. We'll
because collab has free access to GPUs.
I'll show us that. It's it's really
nice.
And when you do anything with neural
nets, it usually benefits you to have a
GPU access. Um
so that'll be nice. But I usually do
most Python coding inside of VS Code. It
supports Python pretty pretty well.
What is more commonly used in the
industry? Um,
the two most popular are Visual Studio
Code and and Notebooks. Jupiter
notebooks.
They're both like you can't go wrong
with either one.
Those
two are really popular. Jupyter
notebooks and Visual Studio Code are
really popular. There's there's both of
those you would be okay with. Either
one.
Let me start with Visual Stu Studio
Code. So, um now Visual Studio Code
is a more general code editor. So, it's
actually you can edit lots of different
languages inside of VS Code. Um, so you
could do Java, you could do C, you could
do Scala, you can do Go, you can do all
kinds of languages are supported inside
of Visual Studio Code. So it's a really
fantastic product for programming in
general, not just Python. Um, it has
built-in terminal support. It has
co-pilot integrated into it, which is
nice for AI, like generative AI
assistance working with your code, which
is nice. Um, of course it supports
Python, which is what we are interested
in. Um, it has it has Python tools. I
will show us which ones we should
install as part of VS Code so that we
can work with Python files and
notebooks. Um, so it's it's a really
great code editor in general, which is
why I like using it. Um, but in
particular, it's pretty good at working
with Python. it it supports Python
pretty uh deeply. Um so and for that
reason VS code is really really popular.
But just keep in mind you can actually
use it for many different types of code
that uh that people write uh JavaScript
um Java as I said like many languages
are supported inside of Visual Studio
Code. So it's a more general code
editor. It happens to be really great at
working with Python though.
All right. So, I'm going to show us a
demo on setting up VS Code. Now, we are
going to do this for each one of these.
For Jupiter and for Collab, I'm going to
I'm going to do similar demos. So, um
don't worry, we'll get to those, but I
want to start with VS Code to show you
kind of how to get that set up and what
it looks like. Um, so where you can find
this demo
is inside of the demos that I mentioned
earlier in the reference material. So
I'm going to I'm going to jump over to
that. Let me show you guys.
So I'm back in the LMS. You guys will
want to download the demos. I think
somebody linked it earlier in case this
didn't show up for you, but we're going
to be inside of the demos and we're
going to do demo one for lesson one. We
do lesson one, demo one, which is going
to be the VS Code demo.
So, the main steps that we're going to
do is just going to be to point you to
where to install Visual Studio Code. So,
it is a it is an application is a free
application you can install on your
machine. Um, so,
uh, you will want to follow this link
that is within the demo file, this
code.vvisualstudio.com/d
download and download it for your
particular platform. So, if you're on
Windows, obviously, choose the Windows.
If you're on a Mac, um, choose Mac and
make sure that you choose the right, one
of the precautions is to choose the
right Mac platform. So, if you have like
an M1, 2, M3, M4 Mac, choose the Apple
Silicon
um button. If you're on an older Mac, um
then you'll want to use the Intel chip
one. Um
uh if you're on if you happen to be on
Linux, which I don't probably most of
you are not, but if you are, um you want
to download the right uh distribution uh
version.
But, uh follow this link first. So
that's the first step. Very easy step.
Just go to that site, pick your right
platform and uh go ahead and download
the installer. And mostly we will be
walking through the steps in the
installer. And then um I will show us
what it looks like once it's installed
and then show you a couple additional
steps that are actually not mentioned in
this file that I think are worth doing
to get you set up.
Uh yes, we will be doing Jupiter next.
Yes, we we'll we're going to be covering
VS Code, Jupiter, and Collab. I'm going
to show us examples of all of those.
Okay, let me ask you guys. Were you guys
able to get to the download page and
start that download and installation of
VS Code?
able to do that.
Any issues with that?
Okay. Yeah, it's just like any
yet I love I love the optimism
yet.
Uh already having both of them
installed. Okay. Yeah. No, if you
already have it installed, I mean,
great. I'll show so if you if you
already have VS Code installed, great.
You can sit tight. I will show you um a
couple of extensions that you'll want to
add for Python support
if you have it installed already. I'll
show us how you can use it with Python
in particular.
Okay.
If you already have it installed,
perfect. Looks like you have it
launched.
Still working on it. Okay. So, these
these instructions um uh show an example
of someone that would be on a Microsoft
platform um walking through the
installation.
Uh if you're on a Windows, you probably
want to create a desktop icon. You
definitely want to add it to your path.
And this just shows what's being
installed. So this is all the install
wizard on Windows. Nothing that exciting
there. So this if you follow all these
steps, you will have it installed. I
hope you have enough disc space. Uh I
don't think it's too big.
I don't think it's too too massive. I
forget how much space it takes. I don't
think it's that much.
I don't think it's that much. But um
yeah, hopefully you have enough.
So if if you don't
uh if you do not have enough disc space
um don't worry because we're going to do
collab which doesn't require you
installing anything. So you can always
use that option. All right. So if if for
some re let me just say that too just
even if if it's not a dispace issue if
you have an in any installation issues
no worries because we will work with
collab and Google that is going to be
cloud hosted that you don't need to
install anything you just need a Google
account
okay a free Google account
um so no worries at all if you cannot
get any of these things installed the
which are going to be Jupiter and
uh Jupiter and VS Code.
Where do we go? I haven't said yet. It
just I'm just making sure it's installed
for folks.
I'm going to I'm going to go over to it
in a second, but did we generally get it
installed and do you have it open? So,
if you once you get it installed,
uh once you get it installed, then open
it.
Yeah, you need to get it installed. Uh,
it should be this first. It should be
this link here.
Follow this link to get it installed.
Oops, I pasted the wrong link.
Let me find I'll copy and paste the
link. But yeah, take take a moment to
get it open. Once you have it open, just
sit tight
if you want to.
What does it say?
Yeah, feel. So, for you guys seeing the
co-pilot features, um,
click click use AI features. I think
that's okay. Yes. Um, you'll you'll
likely want copilot. Yes.
Click click okay on that.
That's the link, by the way, for the
download
in case uh we needed to get to it.
Okay.
So, I'm going to go over to VS Code
and show you what it looks like on uh my
end.
Okay. So, you should have something that
looks roughly like this. I don't have
anything open. I don't have any files
open. Uh just kind of have a blank
screen here. Um, but if you I would
recommend uh using the AI features if
you can. Um, I think that'll come in
handy later on.
Um, are we
comfortable uh moving forward? I want to
show us the extensions that support
Python. So, right now when you first
when you first install this, it does not
work with Python out of the box. We have
to install a couple extensions inside of
here to get it to work with Python. I'm
going to show us how to do that.
Don't worry about tuning any settings.
No, don't worry about doing any of that
at this stage. Don't really need to tune
anything. We just need to get Python
support.
Okay.
So, you guys with me on this main page?
You can use your corporate. Sure. Sure.
Yeah, you can you if you have it. If you
have co-pilot and want to use your
corporate, you can use that. That's
fine.
But you guys are with me on the main
page because I'm about to show us uh I'm
about to show us the extensions we need
to install to work with Python.
Okay, really important because this
isn't this is not in the documentation.
Um, no, no need to reinstall. Um, you
can I'll show you how to add that
through the extensions. No need to
reinstall.
You can add it as an extension. Yeah.
Okay.
So, let me ask you guys on the left,
do you see
this
little box icon that if you hover over
it says extensions?
Do you see that? you. There may be other
things here too, but at least that one
with the extensions.
Okay, so we do see that one. Okay,
so what we want to do,
no, I wouldn't I wouldn't uninstall.
That's okay because we're actually gonna
install Anaconda to get Jupiter. I
wouldn't un I wouldn't I would cancel
that if you can because you're going to
want that for Jupiter as well. I
wouldn't uninstall Anaconda.
I wouldn't uninstall. But I mean, if
it's already going if it's already doing
it, that's okay. We'll just reinstall it
later. All right. So, back to the
extensions. So, let's click on the
extensions.
Okay. So, do we see something like this
that has a search bar for extensions?
Do we see the search bar for the
extensions?
Okay. What do you think? We're going to
search for
Python.
Python.
We're going to search for Python. Yeah.
So, you are going to want to install the
official Python extension from
Microsoft. It is this one that has the
blue check mark next to Python.
Uh, so there now there are other ones
here,
but we just want the one that says
Python
from Microsoft. Do we see that extension
when you type in Python? Do we see that
one?
So just so it should just say Python. It
should be Microsoft.
Uh it's really popular. It has a lot of
downloads. Over 192 million downloads as
an extension.
It's from Microsoft.
Okay. Click on that.
Click on that.
And then you should see an install
button. It I already have it installed.
So it says uninstalled. Right here there
should be an install button. Install the
Python extension.
So out of 192 million installs,
really popular extension.
Are you guys able to install it?
You want to install that? It should be
pretty quick.
It should be pretty quick. It's not that
big of an extension.
So, what this does is
just the Python Sherry. It's just a
Python one. If you go into the
extensions and then search for Python,
it is just the one. It's just this one
that says Python and it's from
Microsoft.
Python blue check mark Microsoft.
You want that one.
And then you want to click on that one
and then hit the install.
Um, Roberto, is that for a co-pilot?
Is that for a co-pilot? I
maybe try closing it and reopening it.
Try closing VS Code, reopening and
retrying the install.
Um, no, we're not opening any folders
right now. We're not opening it. We're
just installing the extension.
That's all. We're just installing the
extension.
We're not opening any project folders.
just installing the extension.
Were we were we able to install that?
I know there's a lot by Microsoft, but
there should just be one that that says
Python.
there. So see how the name like this
name is this name here is eyesore. This
name is Python debugger. This one is
pilance.
Just the one that says Python.
Just that one.
That's the one we want. Only that one
right now.
Okay. Perfect. Perfect.
Okay, great.
Okay, so one more extension for you
guys. So once you install that one, I
have one more for you that you want to
install.
Are we ready for that one? One more we
want to install.
Okay, we're ready for the next one. So,
the next one you want to install
is the Jupiter extension,
which is the Jupiter.
It's this one. It's the very first one
here on my screen. So, it's it says
Jupiter
and it's from Microsoft.
Okay, we want to install that one.
Jupiter and it's from Microsoft. want to
install that one.
So, this one has 98 million uh installs.
You want to install this one.
Did you guys find that one? So, you want
to type in Jupy
Ter and it should be the Jupiter
extension here
that is uh from Microsoft.
So you want to install that one.
Great.
Now what does this one do? This
extension will allow you to work with
Jupiter notebooks inside of VS Code if
you want to.
So you Jupiter notebook has its own
standalone program which we will look at
next.
But you can open you can have those
files, those Jupyter notebook files be
compatible with VS Code and open them
and edit them and run them inside of VS
Code if you want to. So this extension
gives you the flexibility to work with
notebooks inside of VS Code. So you
never have to leave VS Code if you want
to work with notebooks. Um,
so this is a good extension if you
really want to work with notebooks and
stay inside of VS Code.
Yes. Uh when you Yeah. When you install
install an extension, it might it might
install a couple other dependency
extensions. Yes. But that's okay. Those
are required. That's okay.
That's that's that's okay.
All right. How do we feel? Good. Uh did
we get those installed?
Did we do were we able to get those
installed?
Okay, here is how we will test that it
all worked. So, we're going to do
something really simple.
Here's how we will test that it worked.
Let me go out of here and back to our
files.
So, out of the extensions, I just went
to the top button where it's the little
file um icon and um I am going to
um
go up to the very very top where it um
so you guys see on your VS Code window
where it says file, edit, selection,
view. I'm just going to create um
I'm just going to create uh a new
new file.
So, do you guys see that where where you
say file edit selection view? Click on
file and then click on new file.
You should see what I see on this screen
right here.
If you see
if you see Python and Jupyter notebook
then you know those are installed
correctly.
Do you guys see these options text
Python and Jupyter notebook?
Great. So what that means is we we can
now create those kind of files in the
future. We can create notebooks. you can
create Python files and VS Code will be
able to work with those.
If you don't see Python, that means your
Python extension didn't install yet or
you didn't install it. So, you want to
go back to you want to go back to your
extensions and make sure you installed
Python.
So go go go to this button over here,
the extensions,
type in Python,
and then make sure you install this
Python extension.
Okay. So, you're going to install the
Python extension and you're going to
install the Jupiter extension,
which is this, and install both of
those.
Make sure those are installed. If
they're installed and you still didn't
see that when you went to file um new
file,
if you don't see those, then um try
exiting VS Code and relaunching it.
Okay? Try exiting VS Code and reopening
it and seeing if you can make a new
file.
Okay? But it should be under uh at the
top file and then new file
and then you should see those options
Python and Jupiter.
Once you have those extension installed,
you may need to close out of VS Code and
reopen it to see that.
Okay, perfect. after you relaunched.
Okay, perfect. Yeah, you may need to
relaunch so that it can show the it can
show the extensions.
Yeah,
perfect.
Okay, perfect. So, that's set up for you
guys. So, um Perfect. It's set up for
you guys. Uh we will work with it in the
future, but just wanted to make sure it
was installed and set up. Once we start
working with Python, um I will show you
guys how to how to work with it. Um but
glad that's set up for now.
Uh what issue are you having uh Romero?
Is it not showing? It's not showing
Python or Jupiter for you when you do
file new file.
It's not showing those.
You may need to exit VS Code and reopen
it.
You uh Sil, yeah, you can you can make
one. We're not going to do anything with
it right now.
It's not going to you're not going to do
anything with it right now, but um
it's make sure you're searching for it
with a Y. It's J U P Y T E R.
You have to search. You have to So when
you go when you click on the extension,
search for JUP
Y. It should be the first thing that
shows up with Jupy Ter.
It's this Jupiter one from Microsoft.
I kernel I'll So the let me show us let
me show us that later. The kernel you
have to um you have to have a Python
interpreter.
So you may need to install a Python
interpreter to to be able to run the
kernel. So, I need to show us that. Um,
but I I don't want to get into that
right now.
Save what to
Oh, wherever you want. Wherever you want
on your own machine. It's up to you. It
doesn't really matter. Just wherever you
want.
It doesn't matter. It's up to you.
All right. So, what I want to do is uh I
want to take a break. Um because now,
you know, I said after two hours, we'll
take a longer break. Um so, we will now
we'll take a 10-minute break. Now, um if
you're still having any issues, um we
can try to get you set up at the end of
class. Um but we are going to set up.
So, coming up after our break, we're
going to take a 10-minute break. Coming
up after that, we'll we'll go and
install Jupyter Notebook. And then after
that, we will look at Collab. So, you're
going to have multiple options to run
Python. Not So, if this wasn't working
for you, that's okay. We'll try a
different route.
Okay? I will try a different route. Um I
I know Collab will work for you because
that is hosted by Google and really easy
to get working with. So at the worst
case scenario, Collab will work for you.
I know it. Um but we'll try to get
Jupyter Notebooks installed for you as
well. But if you're having issues with
VS Code, let me know at the end of
class. We'll try to get you set up,
okay?
You're still having issues with it.
But um what we're going to do right now
is take take a 10-minute break.
So let's try to be back um in about uh
10 minutes. Let's call it an even um
let's call it an even
uh what will we be covering? Um
installing the other installing the
other um Python setups. So Jupyter
notebook and working with collab. And
then we will get into the basics of
Python's the syntax. So we're going to
talk about indentation, identifiers, um
maybe if we have time, basic variable
types, data types. Yep. So we'll get
into Python.
We will get into Python today.
All right. So let's jump over to
uh Jupiter notebooks. So um what's so
special about Jupiter?
Well, it turns out that uh Jupiter is a
platform for running what are called
notebook files. So obviously we just
installed the Jupiter extension in VS
Code which will allow us to run
notebooks in VS code but Jupiter has its
own notebook platform and that's what
you will install in this setup. Um,
notebooks are special. They are um
really great um Python code files that
give us the ability to execute isolated
what are called cells of code. So we can
run one cell at a time and test and
debug the execution of that single cell
without affecting any of the other
cells. So, um, notebooks are great for,
uh, running code live and interactive.
When we do a lot of our demos in this
program, they're all going to be in
notebooks. Um, so that we can kind of
run things one cell at a time. Um,
uh,
no. So without notebooks you either have
to run you run like a Python script like
a Python file um which is a py file and
usually you have to either run that
through a debugger or run the entire
script at once. You don't really get
code isolated into individual cells
which is really nice with notebooks. The
other thing is notebooks are easily
sharable.
So you can share a notebook with
somebody else and they can open it and
see all of your inputs and outputs in
the notebook which is really nice like
all of the outputs get saved into the
notebook. Um which is nice. So and
notebooks uh especially in the Jupiter
platform are going to have all the data
science libraries available to them. So,
uh, if you're people usually love doing
notebooks for working with data, um,
really easy to work with data inside of
notebooks and and build things like
plots. You can display your,
uh, you can display your graphs really
easily inside of the notebook and then
share your notebook so other people can
see your graphs. Um, so notebooks are
really awesome like interactive
environments for running code. Um and we
will favor notebooks uh as our primary
way of running code throughout the
program. Now where you open those
notebooks is up to you. You can open
them in VS Code. You can open them in
the Jupyter notebook platform. Uh
you can open them inside of Collab and
run notebooks in Collab. Uh notebooks
are very very popular.
Why isn't running in notebooks the
default? It's because uh not all code
runs inside of cells. Like applications
are not going to be well suited for
notebooks. Like Instagram is not running
in a notebook. Uh it's more structured
into actual Python files and actual uh
more structured programs are going to be
not in a notebook. Notebook is more for
prototyping and debugging and uh
executing small chunks of code to test
it out. It's not for writing larger
programs like an like a
an LLM application like a chatbot would
generally be in not in a notebook. It'd
be in like a Python file.
Uh cells versus class objects. So cells
are just small uh think of them as small
little environments to execute our code.
Um class objects are actual chunks of
code that define an object. They're
they're different things.
Yeah, different things. We'll we'll
learn about objects. Um and we will
certainly see what cells are as we go
through. I'm going to show you an
example of a cell coming up when we
install Jupiter.
But uh let's talk about let's uh go
through the installation of Jupyter
notebook so you can see what a notebook
looks like. I think that'll be helpful
to orient.
So let's go over to that demo. So this
is going to be demo two
uh demo two inside of um uh lesson one.
So we're going to go over to that.
Everybody has this one. Okay, perfect.
Okay, so you're going to follow this
instruction. Now, what this is going to
do is first
um
No, this has not this is not going to be
anything to do with VS Code. This is
going to be a different platform. This
is going to be Jupiter.
Where is this? This is the
This is the demos.
This is uh demo two inside of that demos
folder that we said uh
to to uh grab all the demos
from your LMS.
Does anybody have that uh demo 2 PDF
they can upload? I I think somebody
uploaded all of them earlier, but if you
have demo two, want to upload it real
quick? I don't have the PDFs.
if somebody wants to share that.
So there so they're different. Um VS So
what I was saying is you can open
notebooks inside of VS Code and the
thing that allows you to open notebooks
in VS Code is the extension.
So yes, if you're going to work with
notebooks in VS Code, you need the
extension installed. But you can use the
standalone Jupiter platform
to work with notebooks. It's up to you.
If you like using VS Code,
um if you like using VS Code, you can do
it that way. If you like uh the Jupiter
platform, you can do it that way. It's
up to you. It's just a preference. I'm
giving you guys options. That's my goal
is to give you options and let you guys
choose what you're most comfortable
with.
Okay. And we're and we're taking time to
do that now in the beginning of the
program, right? Because we're going to
be doing a lot of Python examples coming
up as we start learning Python. So, it's
it's valuable to spend that time now. I
know it can seem a little slow, but I
promise it'll be worth it so that you
guys have options for running your
running your code.
Yes. Thank you guys for uploading those.
appreciate it. Those are the demos you
want to uh follow along with.
Okay, so the first step here is going to
be to install Anaconda. Now, you may be
wondering, what is Anaconda? I thought
we were talking about Jupiter, and
that's a valid question. Anaconda is a
what's called a distribution of Python.
So Anaconda is a program a software a
collection of software programs that
give you a version of Python with a
bunch of packages
uh with a bunch of packages already
installed.
Um and then
uh one of those is the Jupiter package
so that you can run Jupyter notebooks.
And what Jupyter notebooks will be
is a uh basically a web browser
application that will open up a notebook
editor in your web browser. So that's
ultimately what we're going to do, but
we are going to install it via the
Anaconda distribution
uh via the Anaconda distribution of
Python.
So that's where we're going to start is
with the initial download of Anaconda.
Oh, it's no no skipped registration.
Okay, let me let me uh open the link.
I think there I think there's a way to
find it without having to do the
registration.
There's a way to get to it without
having to do that. I'm going to find it
real quick.
Oh, you can't. Okay. So, if you can't
install it, that's okay. We will be able
to work with notebooks in collab and you
can work with notebooks inside of VS
Code. That's fine, too.
Yeah, I'm getting I I'm going through
the registration process so I can um I
can show you that install.
Okay, let me share my screen.
Did you guys get to once you go through
the like setting up your account, do you
get to this page?
Do you get to this page for those of you
going through? Yeah, that looks right
for you, Ashish. That looks right.
Do you guys get to this page though when
you get through your like account setup?
Okay, you got to this page. Okay, so
then choose your correct Windows or Mac
down. You want to be over here on the
left. You want to do Anaconda
distribution.
This is what you want to do. So, choose
the right one. And if you're on an M1,
M2, M3, you're going to do the silicon.
If you're on an older Mac, you're going
to do the 64. And then obviously, if
you're on a Windows, you should be
clicking over here to do Windows. But
you want to do the Anaconda
distribution, not Minion. Okay. So,
click on the installer for Anaconda
distribution.
Okay? And then let that install. Now
while that's installing let me explain
something about the difference between
uh I think it was asked earlier what's
the difference between um Anaconda
uh as the default Python. So Anaconda
as I was saying earlier is a version of
Python that has a bunch of data science
and machine learning packages already
installed for you. So uh it comes with a
bunch of packages that are already
installed. So if you use that Python
um that Python has a bunch of packages
built in with it that you don't need to
go out and install. So, Anaconda is a
very popular version of Python for
people to install that are working in
data science, AI, ML. Very popular
version because it already comes with a
bunch of packages that you would use for
manipulating data for doing machine
learning or doing anything in AI. So,
it's it's a very um popular
distribution. It also comes with
Jupiter, which is why we wanted to use
it because it comes with the notebook
capability out of the box.
Okay. So, I'm going to launch. So, when
this is done installing, you want to
launch the program that gets installed
called the uh Anaconda Navigator.
So, it should install a program on your
machine called the Anaconda Navigator.
Do you guys have that? Did anybody get
through and and you have that program?
The Anaconda Navigator.
You don't need any advanced ones. You
don't need any advanced options.
Just the just the defaults. All the
defaults
should be good.
Still downloading. Okay. I'm going to
show you
I'm going to show you what the navigator
looks like once you once you have it.
That's okay if it takes a little bit of
time to download. That's okay. Um,
basically once you download it, um, you
just have to click a couple more buttons
and then you can access Jupiter.
Okay, let me share my screen and show
you what you like. Once it installs,
this is what it should look like. It's
okay if it's taking a little bit of
time.
You should have something that kind of
looks like this, which is the um
dashboard that has the different
programs available to you to you.
Um
do you guys see something like this? If
you have the navigator,
do you see something like this?
which is the which is the like when you
open the navigator program, you should
see something like this that has a bunch
of different um
you do. Okay.
It's if it's taking a little bit of time
that's okay.
Yes. Na Anaconda Navigator is how you
launch Yes.
Anaconda Navigator is how you launch it.
So yeah, you want to open that. Now the
the whole reason to come here
is so we can launch Jupiter notebooks.
So we can launch Jupyter notebooks. Um
this is the program we ultimately want
to launch. This is going to
uh allow us to open notebook files, edit
them, run code cells. I'll show you what
a notebook looks like in a second. But
but once you have Anaconda installed,
open the navigator and then launch
Jupiter notebook. It's just one extra
step. Launch the Jupiter notebook.
What that should do is launch the the
notebook.
Uh it should launch the web browser of
your like whatever you have as your
default web browser. It should open the
notebook program in every in your web
browser. So if it's Chrome, Firefox,
Edge, whatever your default web browser,
it's going to launch the notebook
program in the browser.
Okay.
So, I'm going to launch it and then I'll
show you what it looks like. Again, if
it's taking you a little bit of time,
that's okay. Whenever it's done,
how did I get to these icons? Just
launch. Do you have the Anaconda
Navigator program?
It should have got It should be
installed.
Open the Anaconda Navigator program.
It should have been installed with the
Anaconda installation.
All right. Was anybody able to get to
this the Jupiter this? So, it should
launch in your browser. Anybody
able to get to that?
Fantastic. Fantastic. I'm glad some of
you guys are able to get to it. And if
it's it's not yet, that's okay.
Remember, when it's done installing,
you're going to go to Anaconda Navigator
and then
uh launch Jupiter Notebook. That's
That's what you're going to do.
That's okay, Roberto. It's okay.
All right. I do want to I want to show
you guys a notebook. I just want to show
you what it looks like. What I'm going
to do is I'm going to
um open a notebook by going to new and
then Python 3 notebook. So you can open
a folder, you can open a terminal, you
can open a text file. I'm going to open
a Python 3 which is a notebook. You so
it's a it's a a notebook powered by
Python.
So I'm going to click on that which will
launch a new notebook in a new tab.
And here I am in the notebook editor. So
now I am in a notebook editor screen. So
if you go when you first launch Jupiter
you you can navigate to notice that that
notebook got created here where I
currently am on my machine. I could
navigate to I could navigate to
documents and then I could um you know
create a new file there or I could make
a new folder here and and do it that
way. Um but uh I am uh just editing this
notebook right here within this um
current folder that I'm in.
Okay. So, do you guys remember when I
said that code gets executed in a in an
isolated cell?
Do you remember that?
Um,
this is what a cell looks like. And you
can make new cells by hitting this plus
button.
So, you hit this plus button over here,
you can make new cells. So, if you hit
plus++,
I'm making a bunch of cells.
Now, what's really cool about cells?
Yeah, Tim just discovered this. What's
really cool about cells is you can
change them to be text or code. So, if
you change it to markdown,
I can write markdown text in here to say
this is my notebook. And then if I run
this, it's going to display as text.
So if I run that cell which uh when I'm
editing it I can click run and it will
render that as text because I changed
the cell type to markdown. Markdown is
just a flavor of text style.
So otherwise we can write some Python
code. Now, what I want you guys to type,
I'll type this in the chat to verify
everything is working is I want you to
type print
hello world.
I want you to type that
inside of a cell
and then
and then hit run.
And it should run that code.
And you should see you should be able to
see uh you should be able to see that
Shift enter. Yep. You can whenever
you're on a cell, you can hit shift
enter. It'll run the cell.
You can That's okay. You can always You
can go back and watch the video. So,
this is being recorded. You can go back
and watch the video. I I know it's a
little frustrating. It's still
installing for you, but go back and
watch the video. And I definitely
encourage once it's installed to go back
and try this, which would be just
launching your Anaconda Navigator
and then launching Jupiter.
If you don't have, by the way, if you
don't have Python 3, um you may need to
uh exit your navigator and reopen it.
Okay, you may need to exit your your
navigator, reopen it so that you can
launch Jupiter again.
Were you guys able to run this in a
cell? For those of you that have Jupiter
running, were you able to run this?
Nice. And it worked for you. Okay,
perfect. Perfect.
So this is what I meant by this is an
isolated cell. So notice that we can run
this
and it doesn't affect
um
Sure. Sure. I hear you. I I hear you.
Update the doc. Uh I can How about I
post it in our um our Slack channel? By
the way, do you guys have access to the
to the Slack channel?
Okay, I can post it there. I can post
the instructions to get there.
Okay, I can post it in our our uh
cohort's uh channel.
I hear you. You don't want to search for
our video. I I hear you.
Uh I don't have the link on hand, but
you can get to it through the LMS.
So if you go to the LMS and go to
uh there should be
um
there should be a link to get to it
within there. It's should be like over
here on the right.
I don't have the link I don't have the
link to it off hand. Yeah,
but there should be a way to get to it
from the LMS.
Yeah, there should be a banner here. I
don't know why I don't have it, but
should be there.
Okay. So, if you're just getting things
installed, how do you get to here? Um,
you open the navigator.
Open the navigator.
Syntax is hello world.
It's just inside of it's just that print
hello world.
Um, open the navigator.
Open the navigator which looks like
this.
Let me share my screen.
Okay. Open the Anaconda Navigator that
got installed.
Then click launch on the Jupyter
notebook program. So you should have
this at least. You may have other ones.
Click on this launch. Uh click on this
launch and then you can launch the uh
Anaconda Navigator.
Okay.
If it if it's a little stuck, that's
okay. We're going to move on. We're
going to go to Collab, which can run
notebooks as well. So, if it seems a
little stuck, that's okay. I will post
in our Slack instructions on how to run
this.
That's okay.
All right. But what I wanted to do,
what I wanted to do before we move on to
collab is I just wanted to show you I
wanted to call out a couple things about
notebooks.
Um
is that uh a couple things about
notebooks. One is that notice that these
cells are very isolated. Whatever I put
here
um does not affect what I had before. So
I can add numbers like that and it can
um compute that and this does not affect
this. So so this is why notebooks are so
great is you can document the notebooks
with with mixing in text and code like
we do here. Um you can run code in its
own cells.
uh you can run code in its own cells and
then you can um have that very isolated.
So I could jump down here and run
something and that doesn't matter that
there's nothing here like it doesn't
need to be in order. I can um you know I
can uh run stuff out of I can overwrite
this
um and run that and it produces the
output. Uh
so you know many things we could do uh
inside of notebooks that are really
fantastic for just quickly prototyping
and running Python code inside of cells.
So it's very nice that way. So notebooks
are notebooks are really nice. You can
also like I could share this file. So
this produces a file on my machine. Uh
if I go back to the navigator
um
Anaconda navigator Jerry Anaconda
navigator
um
can you read value of variables from
another cell? Uh you you have to store
them into variables. So I could I could
call this uh x
and then I could refer to x later.
We'll learn about that. We'll learn
about that with variables.
But yes, you can kind of do that with
variables.
All right.
Um,
what do the numbers after?
Which numbers? these the ones in the
brackets.
Oh th so those are which cells we've uh
executed. So I executed this one first.
So it it's it's number one. And then I
executed um
uh this I think I did this second. So it
it's text. It doesn't really get one of
those. And then I did this one third.
And then I did um I think I did this one
fourth and then it got overwritten with
the fifth. So it just tells you like how
many executions you've done and what is
which number execution that was. Then I
did this one sixth.
It just keeps track of your executions.
Okay. So somebody asked about a kernel.
What is a kernel? So uh the kernel is um
the kernel is basically the interpreter.
So it's the thing that that the kernel
is just a a um a copy of the interpreter
that the notebook is attaching to in
order to run. So the notebook can't run
anything because remember a pi Python
needs
Python needs an interpreter to run its
code. So in notebooks we basically
create like a virtual copy of the
interpreter called a kernel. Um and you
can actually have many kernels based on
your um your base interpreter. So what's
nice is Anaconda
um Anaconda comes with
uh an interpreter for you and then you
create kernels that are virtual copies
of that um that are virtual copies of
that uh interpreter so that you can run
your Python code against it. Remember
you need an interpreter but notebooks
attach to kernels. Kernels are like
virtual interpreters.
Um, and you can have many kernels based
on the original interpreter. So the
kernel is literally just think of it
like the computer that's powering the
notebook. That's all. It's just the
compute engine that's allowing you to
execute your Python code. So every
notebook has an associated kernel.
And what's interesting is if you restart
your kernel, you lose all your data. So
all of these outputs that we have um you
would lose if you restarted your kernel.
So if I restart um now I like since I
restarted this is not going to know what
x is. So if I try to print x again it's
going to say I don't know what x is
because I restarted my kernel. I lost
all that data.
But I can redefine it. And then there it
is. And notice that my iterations
restart
um my iterations restart when I uh
restart my kernel. So if if I go back
and restart the kernel again
and now if I run this, this will be
first. So notice how that restarts to
first. This will be second. This will be
third.
Try restarting. I don't know what that
is.
I don't know why that
Yeah, choose the Anaconda. Either one.
Choose. You want to use Anaconda as your
But what that's saying is what do you
want to use as your interpreter to to
build your kernels. So, choose one of
those. That's fine.
So, yeah, Anaconda requirements, laptop
requirements. Um,
you need a little bit of you need a
little bit of RAM. Uh, you need a little
bit of RAM to run the notebooks because
you need some memory. Um, you don't need
a lot of it though. I'd be surprised if
you didn't meet the requirements. It's
not that much, but you do need some.
I'm not sure the exact. I'd have to look
that up on the Anaconda website.
Oh, it must have been full to start with
or pretty full. I'd be This doesn't take
up that much space, I don't think.
Was it pretty full to begin with?
I would assume. I don't think this takes
up that much space.
Uh, what I want to do then, I want to go
over to collab. Okay. How do we feel
about the notebooks? I maybe if it's
still installing for you, give it a
little time. Open the open the
navigator.
Let's try Coll. I guarantee you Collab
will work for you if you're still having
issues with if you're having issues with
Jupiter.
No worries. Let's just try collab. I
promise that will be a lot easier, be a
million times easier, I think, than than
working with Jupiter. Okay, great. So,
we will continue then.
All right, let me jump over to our
final
um
demo with setting up a collab notebook.
So I'm just going to jump into doing
that on in the interest of time.
Uh
so
would you be taking up additional
sessions too other than uh so we're I'm
going to be the instructor for all of
the courses in this program. So you're
you're stuck with me
for all of those. Does that make sense?
like all of the all of the uh AI
engineer program uh courses.
Yeah. Yeah. Stuck or be excited. It's
going to be one or the other. Probably
not an in between feeling.
Hopefully. Cool. Hopefully. Hopefully
good. Yeah. Like I said, I've taught
this many times. Uh, I think it would be
uh I think it'll be good.
You are stuck. Okay. Well, we're going
to get you unstuck with Collab. I would
not worry about getting Jupiter set up.
If it's not working for you, we're going
to ditch it and we're going to use
something else that works. I promise
it's not a I promise getting Jupiter set
up is not that important relative to
getting at least one of these options
that works.
So, if it's not working over on Jupiter,
I'm not worried in the slightest about
it because there's going to be plenty of
options to run run Python code. In fact,
we're going to do one next which is
going to be with um with Collab. So,
we'll do that. Um so, let me jump into
that. Let me share my screen here.
Um,
learning a lot already. Great. That's
great. Glad to hear that. Thank you.
Okay.
Thank you guys. Appreciate it.
All right. Let me go to the demo. Demo
three.
All right. So, what you want to do
essentially is to go to this website,
um, which I have here. I'm going to, uh,
paste it in the chat. Um, so what you
want to do is go to Google's website for
their Collab platform, uh, which is, so
Collab is a, um, notebook platform that
Google hosts. So, you don't need to
install anything. You just go there in
your favorite web browser, log in with
your Google account. In fact, I don't
even think you need to be necessarily
logged in. You can in order to save your
notebooks to your drive,
but um you go there and you basically
open up a notebook and start working
with it right away. And it's fantastic.
Their notebook environment already has a
lot of packages installed into it for AI
and machine learning. So that's f that's
really great. Um
uh once you get to the page um you
should log in though if you have a
Google account. Only reason I say that
is because it will save your notebooks
to your drive automatically so that you
it will automatically save your
notebooks just like as if you're working
in a Google doc. So that's great. So
that um it saves your work
automatically.
Um, so please, you know, I would
recommend getting a Google account if
you don't have one for free. Logging in
using Collab is completely free.
Um, so it's a fantastic platform. Um,
when you go to that site, uh, assuming
you've logged in, you want to click on
that lower left blue button where it
says new notebook. You can see it in
this screenshot. And I I'll open up one
in a moment on on my screen. But do you
guys see this screen right here that's
in this screenshot that has the new
notebook on the bottom in the lower
left?
No. From that site, what do you see?
Oh, so you're already in a notebook. It
you're already in a notebook. Like it
says, "Welcome to Collab.
Oh, okay. So, it already opened the
notebook for you. Okay, that's that's
fine. That's fine. I'll show you uh I'll
show you what that looks like. That's no
problem. That means you're already
inside of it.
Okay.
Okay. So, then we're pretty much in the
notebook environment and we can start
running code. Let me let me hop over to
Collab and show you guys what it looks
like.
Let me stop sharing that and jump over
to
collab here.
Okay.
So, if you're in the welcome to collab,
um that's fine or you can start a new
notebook. Let me assume that we've
opened up welcome to collab. So, you're
in this screen. What you want to do if
you're in this screen is just go to go
up to file and do new notebook.
Just go to file, new notebook in drive.
Just do that. File new notebook
and this will create a new notebook for
you which will start fresh a blank
notebook.
Okay.
Were you able to do that? If you guys
were folks able to get here to this uh
blank notebook
one way or the other, you clicked the
blue button to hit a new one or you went
up to file and did new notebook.
Yes. Okay.
So, there we are. Without doing all the
Jupiter install steps, we're in a
notebook. Look how easy that was, right?
So, why didn't we just start with this?
Um
so yeah so we're in Google's notebook
platform uh which is a fantastic
platform and uh what's great about this
is you can export these now these are
pyb which is which is uh if you're
curious what that extension means it's
short for interactive python notebook
okay IPIB
so these are the files that you can open
in Jupiter if you have uh if you notice
when you open up Jupyter notebook book
earlier it was a IP YMBB
um inside of VS code when you work with
notebooks they are IP YMBB so IPMBB is a
notebook file and it can be opened in
any one of these three platforms right
the Jupiter from Anaconda the uh VS code
can open IPMB and you can also upload
your own notebooks here if you have them
on your own machine you just go to file
upload notebook and And then it will
open up a box where you can choose which
file on your machine to upload. So you
can upload your own notebooks, which
will be uh great when we get into um
demos. We have demo notebooks for you
guys that we'll work through with our
code. You can upload those into Collab
and work with them directly inside of
here.
So let's try running something. Let's do
the print
hello world.
So, um, you want to type that in and I
can paste it in the chat for you guys
and then you want to you want to hit
either shift enter or this play button
right next to the cell.
Okay, so that might take a moment
because it's booting up your uh your
kernel.
Uh, but then it should run and you
should see the output. Now, this is
going to look very similar to Jupiter,
just slightly different.
We're inside of Collab
and we started a new notebook.
We're just inside of a blank notebook
for now. And we uh are just within this
first cell and I I'm doing hello world.
Were you guys able to run that?
We didn't. But we could we could open a
notebook in VS Code because we installed
the extension. We did that. Remember we
installed the Jupiter extension. So we
can open notebooks in VS Code and we can
run them there. I just didn't show that
to us. Uh we might do that later down
the road, but you do have that
flexibility to run things there if you
want to.
Okay. One thing I want to show you guys
that's really cool. So, um, one thing I
want to show you is if you go up to
runtime
and then go down to change runtime type.
Do you do you guys have that? Change
runtime type. So, if you click on
runtime
and then change runtime type.
Do we have that? Click on that. Click on
change runtime type.
And look at our options. We can choose a
GPU for free.
So we can swap over to a GPU kernel
which is fantastic for training deep
learning models and we can use that GPU
for free. This is one of the reasons
that uh Collab is so amazing is they
give free access to a GPU. So if you
don't have one on your own machine um
you can use the GPUs from Collab for
free.
Yeah, go ahead. I mean, there's no
nothing wrong with it. So, uh, what it's
going to ask you to do is, uh, terminate
your current kernel because you're
connected to a CPU basic kernel. It
wants you to disable that so you can
swap over the GPU. Click okay. That's
okay.
All right. And then we are now um, we
click save. And that will swap us over
and connect us to a GPU. So, if you how
you know that you're connected to a GPU
is if you go over to um
if you go over to this box on the right.
Do you guys see that one where it says
RAM and disk? If we click that,
it will show us our resource resource
usage. And you should see GPU RAM
available of 15 gigs.
So, you have 15 gigabytes of GPU RAM
available that you can use.
So remember, you just click this RAM
um
you should click this RAM uh
uh
sorry this RAM and disk.
Roberto, did you swap over the runtime
to
uh Yeah, it should pop up. Okay, then
you should be able to click on this
Yeah, it might take some time to connect
to one because what Google has to
allocate one to you um and then it has
to like connect it over the cloud. It
can take a minute. Yeah, it can take a
minute. It has to allocate one to you.
So the question is which one is better?
Um,
so for the vast majority of things,
the CPU, the standard CPU runtime, which
is the default, is going to be better
for the vast majority of things. The
only time the GPU is really going to be
beneficial is when we start doing deep
learning and training neural networks,
then the GPU will be really beneficial.
It will speed up the training time by a
significant amount.
I can tell you like I trained a uh
neural network for images for image
recognition. Uh it took me it took two
hours on the CPU and then when I swapped
it over to GPU it took less than a
minute
took less than a minute and it was
taking two hours on the CPU.
So yeah, training neural nets on a GPU
when we get to that is going to be
beneficial. So if you're not using
Collab right now, that's okay, but in
the future when we get into deep
learning, you're likely going to want to
use Collab to swap over to the GPU for
free.
Now, they do rate limit you,
so it's free, but you can max it out in
a day and then they cool you off for 24
hours, which I have I have done, uh,
unfortunately. So, like, if you max out
that RAM and you use it too much in a
24-hour period, they will, uh, not allow
you to connect to a free GPU for another
24 hours.
So, I doubt you'll run into that
situation, but I have before
if you're just if you're just using it
so much.
No. So, unfortunately,
uh, no.
So, unfortunately, no. You cannot, um,
Collab doesn't connect to your local
resources. It only it does the cloud
Google's cloud resources. So no, you
can't use your own through collab. But
yes, you could use your own GPU through
VS Code. I will show us how to do that
later when we get into deep learning.
I will show you that later.
We we won't need to worry about that
now, but later on, yes, that'll be
important.
Uh it doesn't show GPU. Make sure you
swap over the runtime to go to change
runtime type and make sure you pick GPU.
Make sure you go away from
No, you should use Collab. I wouldn't
You don't need to buy anything. You just
use Collab. Just use Collab for sure.
Collab's free. It does everything you're
going to need for the class.
Yeah, I I highly advocate for Collab. It
So, by the way, if you're curious, like
Collab came about because Google wanted
the the machine learning research
community to have access to GPUs for
free to um develop like machine learning
and uh deep learning models. So, uh it's
been around for a while. I remember
using Collab um probably almost 10 years
ago and it used to be it used to be very
lucky if you got a GPU. You used to like
you used to have to click and then hope
that you would get allocated one and
sometimes you wouldn't and I would sit
there and have to refresh and try to
hope that I would get a GPU but now it's
it's like readily available which is
fantastic.
No, you're But you're joining it at a
good time because I'm telling you, the
GPU was very difficult to get. I would
always try to switch over to that and I
would rarely be able to.
So,
pretty good. But like I said, like if
the CPU is perfectly fine for everything
we're going to do, except when we get
into deep learning, you're going to want
to swap that over to GPU. But that's
going to be for we have a while till we
get to deep learning.
We have a lot to learn between now and
then.
Okay.
What do we think? Do we like collab?
We're comfortable with it. Feel free to
use it. Feel free to use VS Code. Feel
free to use Jupiter. Whatever you want
to use, okay? There's options, right? I
hopefully you have options that work for
you. Um they are all used in the
industry. So you're not missing out on
if whatever you use, people use it of
these three people use all of them.
So feel free to use whatever is easiest.
Yeah, I can show that real quick. Yeah,
let me go back over to it.
I'm going to be real quick on it though
because I want to make sure we get over
to our other material.
Okay, let me show you how you can run a
notebook. Let me show you how to run a
notebook. So I'm going to go to file,
new file, and open a Jupyter notebook.
Okay. So it'll open a new. Now notice
notice the extension of it
is
MB. That should be no surprise. That is
the universal kind of interactive Python
notebook file.
Okay. So the biggest thing you have to
do when you open a notebook in VS Code
is you have to
tell it what kernel to connect to. So
have to go to select kernel
and then what you have to select is the
Python environment. And luckily if you
installed Anaconda
you have a built-in Python environment
which is going to be your uh which is
going to be um the
you know which is going to be the uh
Anaconda that you installed.
So I have Anaconda here. Now I have a
lot of other ones but the
uh Anaconda is here. Say it's this one.
Does that so when you
when you uh
Yeah, you have to install you have to
install a Python environment. Yes. So
you want to install Anaconda first and
then you can run your then you can run
your notebooks.
And then you just uh run your code as
usual
and then we can run that.
Yeah, it but like it's working as if you
know the same kind of notebook that we
have inside of Collab, the same kind of
notebook we have in Jupiter.
You you have to have a Python installed
for this to work. So you go to you go to
Python environments
and then choose a Python environment.
You could try to create one. I'm not
sure if that'll work for you. Create
Python environment. You could try that,
too.
Yeah, that's fine. Any anyone will work.
Any Python will work. You just need to
pick a Python. Anyone will work.
Okay. Yeah, if it defaulted to something
that's fine, too.
And then we can generate more cells
and run cells.
But yeah, that's the thing is you're
going to want to install um Anaconda
most likely because you need a Python
version installed on your machine in
order to run this.
Perfect.
Okay.
All right. So, what I want to do is jump
back over to our notes so we can
continue along. Um again feel free to
use whatever
platform works for you. Collab,
doing notebooks in VS Code, doing
Jupiter notebooks, whatever works for
you, please feel free to use that. There
is no wrong way of using it. Whatever is
best suited to you and you're most
comfortable with, please use that
option.
Uh, it's lowercase P. That's why
lowercase P. Capital P is not a function
in Python.
Lowerase.
Yeah, go with Collab. Yeah, if you're if
you're on a machine, you can't install
anything, go with Collab. That's totally
fine. That's why it's there is for the,
you know, convenient kind of cloud
aspect to it.
Um,
feel free to do collab for everything.
That's totally fine.
I will use collab from time to time as
well.
All right, let me uh go back to our
notes then
and pick up from uh syntax. So, I'm
going to go back to Let me share my
screen. Go back to
Can you use Collab on your phone? I've
never tried it. I would be surprised.
Maybe an iPad.
Maybe like a tablet. It could work
pretty well.
Phone. I'm not so sure.
Yeah. Go ahead. Try it and let me know
how it works.
Try it and let me know. I really don't
know. I'm curious now to try that.
All right. So, I'm going back over the
notes. We're going to finish up today uh
what the time we have left to go through
some syntax. So really uh
really getting into um into Python like
the actual code of it so we can get
started on that and start working our
way through it.
Uh the difference so the the difference
is um you will be executing py files
with the within the terminal. So you'll
be running Python files instead of cells
in a notebook. you're just you're
running a you're running a Python script
rather than individual cells.
Okay, so there's a difference there.
And the reason the reason we choose
notebooks is to run individual cells.
It's just easier.
Same syntax,
same syntax. It's just the code is not
isolated into cells.
still Python.
All right.
Um, let's go forward into the syntax,
start learning about it.
All right. So, something we need to
learn about is how do we properly write
Python code? What is the syntax to it?
So, some things we're going to need to
learn about are how to write proper
identifiers, which are names.
Identifiers are just names for
variables. So, we need to know what's
allowed, what's not allowed. We need to
talk about what the indentation means
and why do we need it in Python. I want
to show you guys how to write comments
because that's really important to
leaving notes to yourself or others
about the code and then talk about um
generally how we produce output and how
we can accept input um from a user or
someone interacting with our code. I
want to talk about all these things.
We'll see how how much of what we get
to.
But let's start with the identifier. So
what is an identifier in programming?
This is really for any programming
language. An identifier is just a name
we give to something inside of our code.
So it's a name we give to a variable, a
name we give to a function, a name we
give to an object. Um so any name we
give to something in our code, like when
we set something equal to x, like x
equals 3 + 3. um that thing the x is the
name we are giving or assigning to a
result or a variable or an object. So
anytime we write down a name in our code
of something there are certain rules
that those names have to follow and
these are something we will um pick up
as we go along but I wanted to call them
out here. So um when we name anything in
Python
generally they have to follow these set
of rules meaning they have to be a combo
of lowercase or uppercase letters either
one's allowed it can be even be a
mixture of lowercase and uppercase
numerical digits are allowed in the name
that's okay any digit 0 to nine it's
okay and underscores are okay
underscores are Okay. And there's no uh
minimum or maximum length. So names can
be really long, they can be really
short. Um of course they should be
meaningful. So when we name something,
it should not remember we want to kind
of get away from naming everything X or
Y or A or B um because those names may
not mean much when we look back at the
code. So even though those are valid
names from an identifier perspective, we
want to be really meaningful when we
name something. We name a variable, name
a function, name an object. Um,
here's one catch.
The name cannot start with a number. So
I can't name something uh just the the
number zero or the number one um because
I can't start with that. Now, it can
include that.
So, if I need to include a number in the
name, as long as it doesn't start with
it, that's okay. But names cannot start
with a digit. That's just one rule of
Python. Anything that you're assigning a
name to,
like a variable, function, whatever,
cannot start with a number or else it'll
be invalid.
Okay?
So I'll show us examples of that later.
Yes, they can start with underscores.
Yes, that's okay. It can start with
underscores. Of course, it can be lower,
uppercase. It can start with It cannot
start with a digit. It can have digits.
They just can't be the first character
of the name.
Yes.
Um, now special symbols cannot be used
in the name. So you cannot have a
percentage, dollar sign, exclamation
point, hyphen,
pound symbol, at amperand symbol, at
symbol. None of those can be used in the
name. So those symbols are not
recognized.
So if you try to include those in the
name of something,
that will produce an error. So we don't
want to do that. The other thing we want
to avoid is naming something in the same
name as something that already exists
internal to Python. So those things are
called keywords. So there are certain
keywords that have a meaning in Python.
They are built into the language. We
cannot reuse those. They're basically
reserved. Um so something like class is
reserved because it means something. It
means you're declaring a class. We'll
see. We'll talk about what that means
later. Or something like global cannot
be used because it declares something as
global. Um,
so
you know, we'll learn what some of those
keywords are. There's a list of them
that are in the Python documentation,
but we want to avoid naming things after
builtin
uh uh functions or builtin keywords. Um
so so in fact we've already used one
which is the print function. You know
when we printed out hello world we would
want to avoid naming something print
because print means something. It exists
as a function. We don't want to name
something print
right that would that would produce it
because it would produce confusion. The
interpreter would see that and say oh do
you mean the function print or you
trying to name something print? it
wouldn't know. So, we want to avoid
naming something that already exists
inside of Python like print or like
class global. Um, there's many others.
Okay. Lastly, and this one always throws
people off, is that when we name
something uh that is case sensitive. So,
if you name something lowercase A, that
is a completely different variable or
completely different function than if we
were to name something capital A. These
are different. They're treated
differently. So, Python will think that
those are two different uh objects or
variables or whatever the case is. So be
really careful with case sensitivity.
Python is case sensitive.
Lowercase A will not be treated the same
as capital A. And whenever you're naming
something, so if I have a variable and I
I I set lowercase A equal to five and
then I set um uh capital A equals to 10.
Then if I um print A, that would produce
five.
But if I print capital A, that will
produce 10. It's not the same. So it is
case sensitive.
These would be two different names.
Lowerase A and capital A.
Okay.
So, we're going to do examples with
these, but these are just some rules we
have to abide by in the syntax if we're
naming anything like naming our
variables, naming our functions, naming
our objects as we go along in in the
course, right? We just cannot The main
one that trips people up is we can't
start with a digit and we can't use
words that already exist like print.
Oh, is my video stuck for people?
Was it stuck?
Oh, okay. Always let me know because it
might it might be.
Always let me know because it definitely
could be. So, it's better to know than
not to know.
Okay.
Oh, no worries. Like I said, always
always feel free to to let us know cuz
um it would if it is, then it's good to
call it out. So, no worries about that.
Any questions about these names? Do does
it make sense about like how we name
things matters and there just are
certain rules that we have to follow. Um
we want to avoid these wacky symbols.
Um, you know, we want to avoid naming
things that already exist. We want to
avoid starting with a digit. Otherwise,
it's going to be a pretty standard like
lowercase, uppercase mixture,
maybe occasionally with an underscore
mixed in there. Um, or or digits even.
As long as we don't start with one,
that's okay.
Yeah, Tim, that's a good reference. the
PEP. So PEP
are the set of guidelines that um the
Python Foundation has kind of agreed
upon as um here's what you should use as
your style guide. Here's what here's
what the community believes is the best
style for Python. Those are good to
read.
Yeah, those are those are uh good
references for really like formatting
and styling your your Python code uh to
be in line with kind of what the
community expects.
Okay.
Okay. So, let me give you some examples.
Um so, the ones on the left are going to
be valid. So, we can name something my
class. We can name something var_1
that's okay.
Count
um that's okay. Uh
but if we have
um like on the right if we have
something that starts with a digit that
would be bad. So so this one is no good
because it starts with this number.
That's not good. Um this name has this
wacky at symbol in it. that's no good.
So, this would be a bad name for
something that would produce an error.
Remember, the interpreter is going to
see that and reject it essentially and
say, "You can't name something this.
It's not valid." Um, same thing with
trying to name something global. This is
a keyword that already exists in the
language. The interpreter is going to
see that and get confused. It's not
going to know if you're talking about
the keyword that's built in or you're
trying to come up with your own name.
It's not going to know. So, it's just
going to throw an error.
Um, so again, we want to we want to keep
things simple. We want to use
underscores where it makes sense. We we
don't want to start with numbers. Um, we
can use a mixture of lowercase and
uppercase. That's fine.
Um, this is a good variable name rather
than if I just called something X.
That again, we're trying to avoid that's
something I always see in the beginning.
I think is okay in the beginning, but
it's something we really want to be
conscious of is naming our variables
very meaningfully.
Like count is going to be more
meaningful if we're keeping track of a
count of something. We would rather call
that count than if than if I just called
it X. Because if you read the code,
which do you guys believe me? Like when
you see it, you kind of know exactly
what it means. X or count? What do we
think?
Which one like has more meaning to it
when you see it? You know exactly what
it's keeping track of. X or count?
Yeah, count.
I would agree with that. Count. Yep.
So, it's just an like that's just a
single example of trying to keep track
of things in a meaningful way. That's a
good name to give to a variable. That
would be uh you know keeping track of
something the count of something
rather than if we just called it x or y
or a or b.
All right, I want to talk to you guys
about indentation next. So now we know
we have to name things appropriately and
the interpreter will give us an error if
we don't name things appropriately.
What about indentation?
So indentation
is a way for Python to understand what
code gets executed together.
Okay.
So
and it it also indicates that I am
breaking the flow of the code from one
section to the next. So the indentation
is really important to signify to the
interpreter there is a new section of
code that has to be considered
um before I move on. So um you should
always use indentation
whenever you have a colon like we see a
colon here with if else statements. Now
we haven't learned about if else
statements but we will. But notice how
we have the colon there and the
interpreter would be okay with this.
This would work.
Okay,
which is a simple statement of saying is
five greater than two? Yes. So in the
case that it is, let's run this code.
But we're only able to run it because
the interpreter recognizes it's
indented.
So the indentation is really really
critical as it makes the interpreter
understand what should be next. The
interpreter understands what should be
next after this statement. Um like an if
statement or a loop statement. Um we
will always have indentation. So this
would actually uh throw an error because
there's this is not indented. This is at
the same level and if you have collab
open you could try this for yourself.
Um if you had collab open it you could
try it for yourself is like this would
this would throw an error where it says
I am expecting indentation but you did
not have have any.
Does it matter how many spaces?
Uh you yes you want to use four spaces.
This this should be four spaces here.
Spaces or tabs?
Uh I'm only laughing because it's a
it's a kind of a controversial question
in the community. Some people get really
upset over
one or the other. I'm not one of those
people. I don't really care. they so
most code editors
uh set the tab automatically as four
spaces. So a tab will do the same thing
as if you manually just did four spaces.
It doesn't really matter in that case.
So either way,
yeah, you're so the ide will do that for
you. The IDs will generally do that for
you because they know it should go on
the same line.
Would it work as on the same line? Yes,
in some cases it will, but not all. It
depends on how complex it is. But what
do you think is easier to read
from a readability perspective? Which is
easier
if it's all in one line or is it more
readable and easier to follow if it's
indented?
Yeah, the that's the purpose. So yes, I
I agree. indented makes it easier to
read. So that's another reason Python
really enforces indentation is to make
it easier to read. There's a reason they
do that. It's to make it easier to read.
Okay?
And that's one of the best selling
points of Python is how easy it is to
read and work with. The indentation
really helps. So to summarize this, we
are going to have to use indentation.
Anytime we have
uh a statement with a colon. Anytime we
have a statement with a colon, we're
going to have to have an indentation
immediately follow it. And there are
certain statements that have a colon
like if, else, else if, and any loop,
any looping statement. Now, all of these
things we're going to learn about, we'll
learn about it in our next lesson.
But anytime we have a colon, this is
signaling the interpreter, okay, there
needs to be a block of code following
that, which is, yeah, as you say,
Romero, it's like a a hierarchy. Yes,
that's a great way of thinking about it.
It's saying, okay, I should check this,
then do this if that's true.
that tells the interpreter this is only
going to be executed in the event that
this is actually true. Otherwise, I'm
going to keep going.
Okay.
All right. Any questions about the
indentation? This is something we're
going to learn about more as we go
along. When do we use indentation and
when do we not? We're going to learn
about it when we get into the if else
and the loops which we will study.
But do we do we let me ask you guys
this. Do we understand the idea or the
intent behind indentation?
Do we roughly get that idea? We don't we
don't know yet when to use it. I get
that. But more the intent or the purpose
of using it is to really like section
things off.
Yeah.
Okay. Good. Good. Good. Good. Glad to
hear that. Okay.
Okay.
Let's wrap up today by talking about
comments. So, uh what are comments?
These are like annotations or notes to
yourself that are completely ignored by
the interpreter.
So when your code gets executed, the
comment does not play any role in what
gets executed. The interpreter will
actually just completely ignore it. The
moment it sees the comment, it will just
ignore it and go to the next line.
So the purpose of it is for humans to
leave a note to other humans reading the
code and that is very powerful is to be
able to read those to to leave those
notes and not have it affect the actual
code that's uh that's actually executed.
So there's multiple ways to make
comments inside of Python. The most
basic is to use the pound symbol. So the
remember we cannot use pound symbols to
name anything.
This is why because the pound symbol is
res reserved for making comments. So you
you when you have a pound symbol like
this uh that immediately signals to the
interpreter everything else that follows
that on this line is a comment. Any
other text that follows that on that
line is a comment. And usually your
editor like VS Code, Jupiter, Collab
will color that differently, maybe like
a a like you can see in here, this is a
Jupiter example. You can see it's kind
of a light gray,
greenish gray
um to signal that this is a comment. Um
now you may be asking why would we have
comments? Again, you are going to look
back at code weeks later,
especially in this program. You're going
to look at code in review and be like,
what what was I doing there? If you
leave a comment, you'll remember what
you were doing there, why you had that
line. Um, and not only that, like when
you share your code with others, which
in the real world you would be doing,
collaborating with others, right? Adding
in those comments can be really
beneficial to do. So
you will see me
throughout the program. I'm going to
leave a lot of comments on our demos and
our notebooks that we work on together
in the live sessions. I will leave
comments mainly to call out certain
things like I will say this is a really
important step or we are doing this
because I will leave a lot of comments
and I encourage you guys to do the same
in your own code. Um, remember they're
free. They're they get ignored by the
interpreter. They don't affect anything.
They are notes to yourself. So, use them
accordingly. Um, and you know, there's
actually multiple ways to leave
comments, but this is I'll show us those
as we go along. But this is the uh most
basic is you just you you type in a
pound symbol and then everything else
that follows that is uh is your comment.
Okay.
All right, guys. That's it for today.
Um, what a great first session. Thank
you guys. Thank you guys for being
patient. um going through the setup of
some of those tools. I hope you landed
on one that worked for you. Um you know,
use that one going forward, please. If
it's collab, use that. Jupiter, use
that. VS Code, use that. Whatever you
you uh feel most comfortable with,
please use that. Um we have a lot to
cover still. You know, we're going to um
continue on Wednesday. Uh we were we
will uh continue talking about Python.
We're just getting our feet wet a little
bit on on Python. A lot more to cover.
We're going to get into the the more
nitty-gritty of the code. So, it'll be
really fun. We'll cover if else loops,
um how to control the flow of our
program. Um we will do that. This is
where we left off was writing comments
uh in Python code, which I I did want to
remind us of how to do that. it is going
to be using the uh pound symbol to
initiate a comment. And basically the
Python interpreter will ignore
everything else on that line. Uh it it
treats all of that text as a comment.
And again like comments are free. You
might as well use them to your advantage
to kind of uh leave a note to yourself
of hey this is what this code is doing.
Um so that when you come back and read
it uh you can understand it better. So I
encourage you guys like when we do demos
uh and we will do a lot of demos um
especially today leave comments you know
put comments in there so so you can make
a note to yourself what this code is
doing um so you will get I think it'll
be good to get in the habit of leaving
comments uh to kind of mark up the code
to to kind of remind yourself oh this is
what it was doing when you look back at
uh in the future.
Okay. So we had ended with that.
What I wanted to do was move on into the
next slide. So talk about um basically
how we display output to the screen
which we've already seen an example of
when we did the hello world which is the
print function on the right. So this is
a by the way this is a Python function
and you know it's a function because
of these parentheses. So these
parenthesis signal that this is a
function because it expects some sort of
input to go inside of those parentheses.
And and the input that would go inside
of there is going to be text like some
sort of uh some sort of text that
belongs inside of quotes. and whatever
we put there um will display on the
screen. So that's useful for us to like
display information
um print we would say we are printing
out information to the screen. Um so if
we want to know the value of a variable
or the value of something that we are
doing a calculation with or uh you know
sanity check something in our code we we
can print it out which would be using
the print function and it will display
that value onto the screen. So we will
use the print function quite a bit. Um
you know we haven't learned what
functions are but uh functions in Python
are you know um designed to be uh chunks
of code that execute and do something
and they take arguments and you know it
takes an argument because of the
parenthesis that is um signaling that
there should be some something inside of
this parenthesis here which is going to
be uh text. So whatever text you want to
display or maybe some variable you want
to display um that would go inside of
there. So we'll get the hang of using
the print function as we go along but
just wanted to call out that's the
primary methodology of kind of um
displaying something on the screen if we
want to print function. Um now the
reverse of that is uh asking a user to
uh input some data. So that would be
this input function and um this is
something that uh as you can see an
example below is we can put some text
inside of this parenthesis. So again a
function it has those parenthesis that
signals it's a it's a function. Um,
and we can put some text in there which
would be kind of what displays in to the
user as kind of a prompt like here. Uh,
enter your name and that would display
on the screen and then there would be a
box next to it. I'm going to show us
this. I'm going to actually run this
inside of a notebook in a minute. But
then there would be a box that displays
that that would say um hey you know
enter your name and then you can type
input in uh and and then when you hit
enter it will save that input into into
this variable called name. So um and
remember name this is a valid identifier
because it starts with a lowercase n um
which is fine and it it has all valid
characters. It doesn't have any wacky,
you know, uh, pound symbol or at or
anything crazy. So, it's it's a decent
identifier. Um,
so name name would be okay. And so input
is whenever you want to get whenever you
want to allow the user to input
something like it'll bring up a text box
and they can enter some data. Um, and
that will be saved in this whatever
variable you name this you set equal to
input. Um, and then you can see like as
soon as we put that in, we immediately
dis we can display it. So we we print
hello and then comma name which
references whatever we stored whatever
the user input there. So I'm I'm going
to show us an example of that. Um, but
input is what get is our primary way of
getting input from the from the the user
in a text box so we can use that data in
our program.
Print is our primary way of displaying
data that we already have in our code.
We can print it which will display it.
Um
so we're going to see many examples of
these as along but just wanted to call
out those two. These
functions by the way are built into
Python. So we don't need to create them
ourselves. They already exist. They're
already built into Python. Um nothing
special we need to do to use them. We
can just use them right out of the box.
So again, we'll see this in our in our
code examples that we're going to do in
a minute.
Um, where exactly would an end user be?
So maybe we ask them for some input. Um,
and then we do like so we ask them for
some like their name, their email, their
uh date of birth, those kind of things.
We can ask in the input and then we
maybe we store them in a database or we
do something with it in the Python code.
So um whenever you want to accept input
from a end user that's when you would
use this input.
It just depends on the application
right on the application like what kind
of input data you you want to accept
from the from the user.
Okay. Okay. So, I'm going to show us
this.
Um, before we do that demo,
um, let me ask you guys, which of the
following do do we remember from Monday?
Which of the following identifier names
follows Python's rules
and best practices for readability?
So not only so you should be looking for
the answer choice here that follows the
rules but also is a meaningful name.
A lot of different choices. Okay.
By the way,
which let me ask let me I'll come back
to the answer to the original question,
but let me ask this alternative
question. Which one of these is not
valid? Meaning it would Python would
throw an error how you use it.
Which one of these is not valid in
general?
Cool. Great. It is C. You guys were
right on top of that. Very good. So C is
not valid. It names of things cannot
start with a number. So So that's um C
is completely invalid in general and
that would produce uh an error.
Right. Starts with a number. Exactly.
Which we cannot do. Now starting with an
underscore is okay. that's allowed. So
that's not an issue. And having a number
be second after the underscore is okay
as well. So technically A and B would
follow the rules. So those definitely
follow the rules. Um now are they
readable and meaningful is the question.
I would argue that possibly not. Var 123
is pretty generic. I would argue that
even though it's valid, like Python
would not have any errors with that uh
variable name, it's not very meaningful.
It's too generic. It's almost as if we
just called something X. We just called
something var 23, that's probably not
going to be meaningful to us and we're
not going to understand what that really
represents. If somebody were to come
along and read it and see VAR 123,
that's probably not that great of a
name. It's not telling us exactly what
that represents. So, I would say A is
likely not um
A is likely not uh a good choice and C
we know is invalid. So, really I think
the only two options you could argue are
B and D. I think D is a really good
answer. It um it it follows the rules.
Uh underscores are fine. Everything is
lowercase. That's fine. Um so it's valid
but it also is meaningful as a name
right so final result value um we we
should probably like in our code we
would have context we would know what
that means okay this is our final result
um so that's a that's a pretty good name
for something um you know this one is
okay it's just not that readable 321
customer details DB table it's okay I
don't think It's um it's not the worst.
It it definitely would work, but it's um
kind of a clunky name. I'm sure we could
come up with something better, but it
would work technically. There'd be no
issues with it.
Okay. So, I think D is probably the best
choice, but B is valid, too. I think B
could work for this. D and B, I think,
are okay.
Okay. Good. Good. You guys are right on
top of that. you have a good I think you
have a good feel for what the allowed
names for things are which is good.
Okay.
Um
I'm going to then swap over to this
demo. So you guys should have the uh
demos um and we kind of went through
some of those first few last time to get
you set up on Cola and Jupiter and VS
Code. Um, I am going to be using Collab
for most of these, but feel free to use
whatever you want to use. If you want to
use Jupiter, if you want to use VS Code
and run your notebooks on your own
machine, feel free to use whatever you
want to use. I'm going to be using
Collab just for the simplicity of it.
Um, so, so this demo will walk us
through, um, opening up a new Collab
notebook and then running those input
and print. So, some examples with input
and print. Um, so we'll do that
together. Let me go over to that demo.
So, if you're following along, we are
going to be doing um demo 4. So, it
should be lesson one, demo 4.
Um, do you guys have this? Give you a
moment to to pull that up. Lesson one,
demo 4.
You guys have access to this one. So, we
did we did one, two, and three on
Monday, which were just getting those
environments set up. So, this is demo 4.
Um, which again, I know step one says
open collab. Feel free to open your own
notebook in VS Code or open your own
notebook in Jupiter as well, whatever
you're most comfortable with. Um,
I'm going to be using the collab to to
do this, but feel free to use whatever
works. You're we really just need a
notebook to be able to run this code.
So, however you're running notebooks,
whether that's in VS Code or Jupiter or
Collab, I any of those, either one is uh
perfectly fine.
So, so step one is to open up a
notebook. I'm going to do it in Collab,
which is what this says. Um, and then
you can make a new notebook and then
rename it to my first program. I'm going
to do that in a second. And then, um, so
I'm going to walk through this live with
you, but just showing you some of the
steps we're going to do. Um, the first
thing we're going to do is is just
practice doing the print hello world
again so that we can um, execute a print
statement. So, we'll practice that.
We're going to make a second cell. Um,
which we can do in Collab or VS Code or
Jupiter by hitting the plus button.
There's usually a plus. Uh, you can see
it here. Uh, in multiple places in
Collab, you can do it right below an
existing cell or there's always a plus
code here, which is kind of what you
have in Jupiter. Usually in Jupiter, you
have a plus button. So, you can just hit
hit that plus button, it'll make a new
cell. Um, and so we'll make a new cell
so we can write some more code.
Um,
and then in this new one, we are going
to practice doing some comments.
We're going to practice doing some
comments and then um see how we can do
uh some more print statements. Okay, so
let's do that. Let me jump over to
Collab. Let's walk through these first
few steps together. Um, and then uh
we'll come back to this and finish out
the rest of the steps because we're also
going to do input. So I'm going to show
you how to do these uh input which will
you can see here like the input is going
to create a text box where you can put
input and it will you hit enter it will
save it for you. So input allows you to
get input from the keyboard
and save that into a variable to use for
later.
Okay. So, let's jump over to
um let's jump over to
I'll show us I'll show us in a second.
How do you rename it?
I'll show you. Let me jump over to
collab.
Um
okay.
So, I am over lesson one, demo 4. Yep,
that's the one we're doing.
Okay. So, I am in Collab. I'm going to
start a new notebook.
Start a new notebook in Collab. Uh, so
now I'm here. I'm just on a fresh
notebook. Um, nothing that interesting
going on. Here's how you rename it. just
go up to this box on the left
and almost like a Google doc just just
uh click into that name and then start
typing to erase it. So see how I'm like
hovering over that name and then I'm
clicking on it and then I can start
typing to erase it. So I can we can name
this my
first program
and then hit enter and it will save
that.
Oh yeah. So if you're in VS Code um do
to to rename it do file and then save as
and then you can give a new name to it.
file, save as.
Okay, that's how you can rename it in VS
Code.
All right, let's do let's do the first
step. Um, let's do print. So, we're
going to do print. So, type in print and
we can uh we can do parenthesis.
Um, remember this is a function. So, we
need the we need the parenthesis to
signal that we want to put some text
inside of this print function. And then
you want to do uh you want to do quotes.
You want to do quotes, the double quotes
there, in order to allow us to put in
some text. So, Python will interpret
what's inside of the quotes as text and
it will display that text. So we can do
hello world my first
Python program.
Okay. And then we can run it. So feel
free to put whatever text in here. It
doesn't really matter exactly what it
is, but you put some text in there
between the parenthesis and then hit
run.
And the notebook will take a second to
connect. And then there it is. Right? So
then you see the the text displayed on
the screen.
Try that out. Are you guys able to run
the print
in your Jupyter what whether it's
Collab, whether it's uh Jupyter
notebook, whether it's VS Code. Can you
run the print
install?
Yeah, I installed that um in VS Code.
Yep. Try installing that.
Okay, great. You guys were able to run
that. Very good. Very good. Okay.
No, you don't want to save it as a JSON
file. You want to save it as a pyb just
like this. See how this one is uh IP
YMBB?
That's the format you want. Remember
that is interactive Python notebook.
You want that file IPY MB.
Uh perfect. Yeah, you get you got it to
run.
You don't have any extension? No. If
you're in VS Code, remember from Monday,
you need to install the the Jupiter
extension.
If you're in VS Code, you got to install
the Jupiter extension.
You have to manually so manually save
it.
You can save it as a py. I would do ipy
so you can open it in collab.
Type it yourself.
Type overwrite what's there and type it.
Type in um my notebook whatever the name
is.
Type it out yourself if you can. like
save as and then
type out the full file name yourself.
Now let's practice a comment. Let's
practice a comment. So let's build let's
do a new code cell. So we made a you can
either do it here. If you hover over
your cell, you can hit plus to build a
new code cell or you can hit plus here
to make a new code cell. So let's do
that.
You should be saving.
Don't worry about the type. Just type in
the namey imm. I don't think you need
to.
Or you could just hit what you could do
is you could just hit save and then in
your file explorer you could just rename
it.
So if you just save it will save it to
the default location and then just and
then just rename it.
So maybe try that route. Just just do
save. Just save it. And then it should
it should try to save it as IPymbb.
Okay. Okay. Let's practice. Um, so the
next step in the demo, if you're
following along in the demo document, it
wants us to do, so we did the print. We
want to do um a practice some comments.
>> Okay, perfect. Uh, let's practice some
comments. So, um, remember I told you
that we can do, uh, comments with the
pound sum. So, this is a
comment. So practice writing a comment.
Remember you start a comment with a
pound symbol.
Um it will get ignored
by the interpreter
interpreter. So feel free to type in
whatever text you want. I'm just
reminding us that whatever the comment
is is going to be ignored and we can
have whatever code below that that we
want to have and that comment will get
completely ignored. So let's do another
print. So write a comment,
hit enter. Immediately below that in a
new line, let's do another print.
This code
gets executed.
So we know this print statement is going
to get executed, but this comment is
going to be ignored by the interpreter.
So let's run that.
So this code gets executed. This comment
gets completely ignored,
right? That comment gets completely
ignored, which is great. Try writing a
comment. Are you guys able to write
comments?
So write a comment and then try writing
a print statement right after it.
And and feel free to put whatever text
you want inside the comment. And feel
free to
uh for the comment, is the space after
the pound symbol required? No, it's not.
So, we could test it out. So, I removed
the space. Doesn't matter. It It's just
for readability. I usually like doing
that so that I have some space after it.
And this is a little It's just a little
bit more readable, right? It's not like
mixed together.
It's just for readability.
Great. You guys wrote a comment. Okay.
Perfect. Perfect. We're able to write
comments. Really great. Okay.
Okay.
If you put multiple code lines, do we
need any separator? Like, no, they just
go on new lines. So, do you mean like a
second print statement? Let's We could
try that. Let's do a secondary print
statement. So, we can do print.
Um, this one is on the next line. No
separator
needed.
Do you see that? See how it's on its
own? I did a print right below this
other print. And as long as they're on
their own line, that's okay. They just
need to be on their own lines. They
don't need any separator.
If we run this, then this one gets exe.
Then see how this is now printed out
below it. Right there.
Is there any character limit on the
comments? Uh, no. There's no character
limit. Um,
but
there's no character limit, but a good
practice is to not like you don't want
this to be super long and to to take up
the whole screen, right? Because then
it's not really readable.
So, there's no limit, but you don't want
to have overly
long comments. You want to keep them
kind of concise and short.
So just so you can read them and they're
they don't take up a lot of space.
Not able to add print statement below.
Why? Why is that?
You should be able to should be able to
have a print right below this print.
Shouldn't be anything that make sure you
close this parenthesis. Make sure every
print needs to close the parenthesis
and they all you also need to close the
quotes. So close this quote, close this
quote within within the print
that needs to be done. So you should be
able to run I'll paste this for you guys
in the chat. Should be able to run this
All right. One thing I want to show you
guys is just like the demo says in the
word document, um you can do multi-line
comments. So if you need to do a lot of
comments, all you need to do is triple
quotes. So triple quote,
then um triple quote, and then
everything in between.
That's interesting that it did that.
Yeah. So, we can do a pound symbol,
pound symbol, pound symbol,
pound symbol, and that that should all
work. So, we can do that.
Yeah, I think it's a collab thing, but
normally in in like Jupiter or in
Python, it it will work just fine. But
like in collab, I think they don't like
the triple quotes.
But yeah, do you guys see do you guys
see how I just did it like this with the
pound symbols? That's okay, too.
Everything between these
uh pound symbols is a comment and is
ignored. So now we should be able to run
that. So there we go. Everything gets
ignored there. Does that make sense to
us? The pound symbol comments
does the does the using the pound
symbols. So notice how we use that to do
multiple lines of comments. So we did
one here, we did one here. We can have
as many we can have
um as many
uh comment lines as we want and they
will all get ignored.
What is those? It's supposed to be
multi-line comments, but for some reason
it's not working. Um, it so the the
triple quote is supposed to be like
representing that you can have a whole
block of comments.
I don't know why it's not working in
collab for me.
It's working for you. Okay. Okay. I
don't know why it's not working.
Single quote.
It's still It still displays here, which
I don't get why that's happening.
It's kind of weird to me.
Yeah, I don't get why inconsistency. It
usually It usually works for me. I don't
get that at all.
Still still doesn't work for me. I don't
know why that doesn't
Yeah.
I don't get why that's not really liking
those triple quotes. Oh well. I mean,
not a big deal. We can just do
Okay, we can do we can try single.
Still doesn't work.
Yeah. Oh well, we can do a pound symbol.
That will always work. Pound symbol is
honestly more popular anyway. Most code
that you see in the wild will have pound
symbols wherever they're doing um
wherever they're doing uh comments. So
that it's fine. Just use a paddle for
now.
Uh yeah, that's correct. I don't know
why that's that's correct. Um I don't
know why collab doesn't seem to like
that. It should be ignored
um generally with the triple quotes, but
uh that's okay. I'm not too concerned
about it for now. I guess what you and
when I do comments, you're usually going
to see me using the pound symbol
anyways. It' be very rare that I would
need to do uh quotes.
Yeah, it's weird that collab doesn't
work very consistently. That's okay.
All right. What I want to show us is I
want to move on to the input. So, I want
to I want you guys to see
I want you guys to type in this code
here that will take input from a text
box and save it into a variable called
name. So, the code we're going to do is
going to be like this. It's going to be
name equals input
and then we'll put um please
enter your name.
Okay.
So this, by the way, I'm going to
comment this code here. Um, this code
should
create
a text box for us to put in our name.
Okay, so that's what should happen. So
when we run this, um, it should pop open
a text box right below this. And we can
type in our name and hit enter. And when
we do that, it will store that result in
this variable called name, which we can
use uh wherever we want to in the code.
So if I hit run, there's that text box.
Do you guys see that? There's the text
box. And see how it says, please enter
your name. And so we can type in our
name.
And we hit enter. And there it's stored
in the name. We can even um display name
by doing print and then the name which
will display uh the name that we stored
when we did the input.
So try this one out. Try this code out
for yourself. Try typing input
parenthesis
and then you want to have some text
there. It doesn't matter exactly what it
is, but something like please enter your
name or enter your name.
Try that out. And then it should store
uh you should be able to type in the box
that shows up. Hit enter on your
keyboard. It should save that. And then
you can um print it out. You can print
out that name which will um
display that whatever we typed in
before.
What does it look like, Roberto? What
does it look like? Were
you Were other people able to run this?
Oh, yeah. Thank you, Melanie. Yeah, I
see that. Perfect.
Perfect. That looks good to me.
Uh, you don't need a space. um it just
looks nice, right? It's so that's a good
practice to have the space so that uh
this code is um evenly spaced out and it
looks nicer on the on the screen.
Um name equals input print hello there
uh name
You need Yeah. So, uh, Roberto, you need
a you need a comma after after the
quotes.
After the quotes, you need a comma after
the quotes to signal to Python that
you're putting in you have you have some
text and then an additional input.
So, it need it needs to be more like it
needs to be like this. print. Um, hello
there.
And then you need an extra comma.
See how I have an extra comma after the
quote. You need you need that.
Sorry. Now, Roberto's uh Kiati.
Hope I'm pronouncing that right.
Okay. So, do we feel good about input
and what it does?
Perfect. Do we feel good about input and
what it does? It It brings up a text
box.
It Did you hit enter, Roberto? To like
Were you able to type something in and
hit It's going to run until you hit
enter.
You have to type in the text and then
hit enter into the box.
So, let me rerun this. So, it See how
it's still running? See how this like
it's going to keep running forever until
I type something in
and then when I hit enter it will stop.
What does your code look like?
Okay, that looks right.
Try try stopping it and rerunning it.
Try try hitting the stop button and then
rerun it.
um you should so yeah you should put
that in a different cell. So if you if
you separate your code into individual
cells so you could do like um you could
do name. So we could we could separate
this. So this code is the only code
that's running in this cell.
That doesn't make sense. Something else
is
that doesn't make sense cuz like this
collab tab is only taking up 235
megabytes.
So something is
chewing up your memory that's not really
I I can't imagine. Are you using collab?
You can see like it's not using that
much. Only 240
230ish.
Yeah, I don't think I don't think Collab
is the culprit unless you loaded in some
really massive data or something.
I can't imagine that's the issue.
You did. You loaded in data. That's
I mean Yeah. Then it's going to it's
going to take in memory. Oh, okay. Okay.
Okay.
Uh MJ, what are you on? Are you on
Yeah. Could you screenshot it?
If it's not working for you, could you
try collab? If Could you try collab just
for the sake of like getting it running?
Things should work in Collab pretty
easily.
You're using Collab and nothing's
working. Uh, are you making sure it's a
code cell and not a text cell?
It's not a text cell like this,
which would be like,
this is where it will be blue.
Did you have that? You need to make sure
it's code. Yeah.
And when I run that, it's going to be
it's going to display text. Yeah.
Okay. Great.
Glad that it's working. Great.
Okay.
Um All right. One more. Uh one more
example what I want to show you guys is
how to do how to include the name in a
print statement. So if we do something
like print. So um we can include the
name in a print statement. So if we do
something like print and then we have um
hello there and then we have um this and
then we have welcome to Python.
um this will
uh this will display all of that
together. So notice that we can have as
many um pieces of information that we
want to display kind of one after the
other as long as they're separated by
these commas.
So we have this uh text,
this text because text is stored in that
variable. Um, then this text and then we
print that all out and we can have this
whole collection of text displayed to
the screen. Try that one out.
Oh, they do the same thing. They do the
same thing. So, the comma and the Sorry.
Yeah, I just noticed the demo does a
plus. They do the same thing in Python.
So, we can swap that over to a plus.
Both of them work.
They have the same I shouldn't say they
do the same thing, but they have the
same effect.
They have the same effect.
Actually, there's no You need a little
bit more spacing here. So, the comma
gives you a little bit better uh
spacing.
So, what the Let me break this down.
what the so plus
plus um adds together
uh text and so what we're doing here
technically is adding all our text
together and then displaying it. Um, so
plus as together text and then the comma
um,
uh, prints out multiple pieces of text.
So they they have the same effect, but
yeah, you can use you can use either
one.
Okay.
Um, one thing I wanted to show you guys
too, by the way, in Collab, if you're
working inside of Collab, I want you to
hover over your name variable.
So, if you just take your mouse and
hover over that,
do you guys see what it says here?
Do you see how it says string name and
then it has the value of that, which is
which is my name. So, that's something
cool about Collab is if you hover over
variables, it will tell you what their
type is. Now, we haven't learned about
types, but any text inside of quotes is
a string. It's it's what we would call a
string. We're going to learn about that.
And
um
we it also displays what data we
currently have stored in that variable.
So all you have to do is hover over a
variable um to to see what the value is.
Yeah, that's yeah, that's kind of a
limitation of VS Code. That's true.
It doesn't show you immediately on
hovering.
Don't see the value on hovering. So, um,
click into the cell. You have to click
into the cell and then hover over it.
Click into the cell and then hover over
it. It should it should work. Yeah, you
have to click on the cell or whatever
cell you're on and then uh hover over
that and it should work.
Uh, Mariel asks, "How do we integrate
that Python code to a client application
for a user to enter a value?
um we would likely have a different set
of code to do that. Um there is Python
code that can get a UI and uh we we will
see that um later on in the in the like
way later on towards the end of the
program. We'll see that um we can we can
write Python code to do a UI essentially
to to make like a almost like a web page
for someone to enter some input. We'll
see that uh much later on. So, we're not
going to get to that right now. It's
really complex.
Um what is the purpose of having
multiple cells? It's so that we can run
individual pieces of code within those
cells. It allows us to isolate, right?
Because I can run I can run code inside
of these cells and they don't affect any
other cell. So, it's it's just for like
debugging and isolation, which is nice,
right? I don't need to worry about
running all of it at once. I can run one
cell at a time.
Okay. Any other questions?
Um, can we execute multiple lines
together? Yes, we did that. Here I had
multiple. So, I'll I'll show you again.
I can do um print
um this is one statement
and then I can come down and do uh print
um this is another and then maybe I can
do some math.
So you can have as many lines as you
want
within a cell.
Within a cell, you can have as many
lines of code as you want.
Is there a way to tell it the order the
cells execute? Um, you no, if you if you
go up to um if you go up to run all,
it's going to run them all in order from
top to bottom. Uh, in order to tell
which cells to execute, you can
rearrange them. You can always like So,
I could rearrange these cells, by the
way, by I think there's a way to move it
down.
So, I can move it down. So now I'm
rearranging. So you can move cells. I
think you can even drag and drop them.
So notice how I took the one that's at
the very top and I'm moving it down.
Otherwise, you have to click, right? You
just have to like I can run them in any
order. If I just click like if I click
here, it will run that one first. If I
go back up here, it will run that one
next. So you just click around which
ones you want to run. Does that make
sense?
I can run them in any order as long as I
click on whatever order I want to do it
in.
Okay.
All right.
Perfect. So, that that wraps up that
demo. I hope it was informative. I hope
you saw the the print statement. Um
we're going to see that many times. The
input statement. Um that's pretty cool.
Um and you got to run you got to run
some Python. So, if it's your first time
ever doing programming, congratulations.
You ran some Python code. That is really
exciting. Um, so glad we got to do that.
Um, let's go back to our notes
and then we'll um
let let me share the screen.
Okay.
So the next thing on our agenda is to
cover variables and data types. So I
just said like text is the string data
type but let's learn about all the
different data types that are going to
be available to us inside of Python and
let's talk about variables. Um it's
going to be a good discussion. So um I
think what we'll do is we'll take a
fivem minute break now and we come back
and we can start this uh discussion
about variables and data types. Um, so
let's take uh a fivem minute break.
And so let's try to be back um around
uh 8:30.
Okay.
Try to be back around 8 8:30.
Okay. So, what are variables? These are
um
basically our way of storing data to
make it easier to reference them and
manipulate uh throughout our program.
So, we've actually already used a
variable. We we used one in our uh demo
we just did where we called the input
the name. Uh we stored that input into a
variable called name. And so um
variables just really are a reference to
some data. That's all they are. They
allow us to reference that data
throughout the program. We can store
information into a variable and then
access it throughout our code. Um so on
this screen are some examples of
variables. Now, variables have names,
which is why I said usually we want
those to be meaningful. Like X is a
valid name, but it's not that
interesting of a name. It doesn't give
us that information much information
about what it's what it really means.
So, probably not the best name. Um, but
we have things like uh we can we can
store some text inside of this variable
called name. We can store a number
inside of this um variable called price.
we can store uh a true or a false value
inside of this variable called
is_active.
Um and so variables will show up all
over our code and uh they are basically
our way to reference some values. Now
these things over here are basically
different types of data that we need to
learn about, right? So we need to learn
about what is a 10 versus what is in
something inside of quotes. is it a
string versus something that has
decimals which is a floatingoint number
versus something that is true or false
which is a boolean value. We need to
learn about those data types. But notice
how all of these are being referenced by
a um by a variable that has some name to
it. Okay. So the variable is this guy.
It is our reference to that data. Um and
we will use variables throughout um so
that we can have you know references to
information in our code.
So variables are fundamental um to to
working with Python.
Um now variables can store different
kinds of data. So I just alluded to
that. And so the different types of data
available to us in Python kind of fall
in these two different categories. one
being single values or what are known as
scalar values. So these are things like
integers. So the number 10, the number
1, the number 2,323,
those are all whole number integers. Um
floats, which are anything with a
decimal.
So 32.3,
3.14,
um 1.2, those are all floating point
numbers. Um, booleans only have two
options. They only have true or false.
So, they represent kind of a binary uh
value um which we say is true or false.
And um then we also have um complex
numbers which are which have imaginary
uh parts to them. We won't really be
dealing with complex numbers too much so
I wouldn't worry about them. But in
reality, Python supports working with
the uh complex numbers and doing complex
math. But uh so so complex numbers just
have kind of a real part and an
imaginary part to them. Um wouldn't
worry too much about that. Again, we're
not really going to work with those ever
throughout throughout the program, but
it does exist. Python supports it. So
scalar data, single values, think
numbers, think single numbers like
floats, think integers, um single uh
true or false values. So these kinds of
data can be stored into variables.
On the opposite end of the spectrum are
aggregated types that we are storing
multiple things.
So we're going to learn about all of
those, but um probably the most common
and one that we've already dealt with is
going to be a string. So a string is
technically an aggregated type because
it has multiple characters that form,
you know, an overall uh string, which is
a string is usually you you know it's a
string because it's inside of quotes,
right? It's inside of these double
quotes or single quotes. Um, Python
actually doesn't care about quotes
really in terms of if it's single or
double as long as you're consistent with
it. Like if you if you start with double
quotes, you should end with double
quotes. If you start with single, you
should end with single. Python doesn't
really care either way. Um, so strings
are going to represent um collections of
characters. Um we are going to talk
about sets which are basically like u an
array of unique values. Um so we'll talk
about sets we'll talk about lists which
are a really important structure. It's
basically an array that can hold many
different types of data. Um so we'll
talk about list. We'll talk about
tupils. So you if you see that word
tuple e that is um people some people
pronounce it tuple. I I like to call it
tupole, but um that is going to be very
similar to an array. It's just going to
have slight differences and uh if you
can change it or not. Tupils you
actually cannot change once you create
it. Um versus list you can modify list.
You can add things to it. You can remove
things from it. Tupils you cannot. So
we're going to learn about those
differences as we go along and start
working with these different types of
data.
Um but they are designed to hold
multiple values, right? So you can see
in that example that list has integers,
it has strings, it can it can hold
multiple types which is if you're coming
from other languages is generally not
the case. Um like arrays in Java, arrays
in C, they can only hold one type of
data in the array. They can't hold
multiple.
Um
so uh then finally a dictionary. A
dictionary is if you're coming from
other languages, it's like a map, a
hashmap. Basically it allows you to have
uh keys mapped to values. So it's a
really dictionaries are highly useful
for storing information where we want to
reference like this value maps to this
value. So for instance in this
dictionary the string a maps to one and
then the string b maps to uh you know
two and or whatever it maps to. And this
will allow us to look up values in the
dictionary. So we could look up, hey,
what is the value stored at key A or
what is the value stored at key B? Those
kind of things. Dictionaries will be
incredibly useful. We're going to
explore all of those more as we go along
in the lesson, but um for right now, it
should be making sense that there are
some data types that store multiple
values like array or sorry, lists, um
dictionary, strings, and then there are
some data types that only have a single
value like a single number like a float,
integer, um boolean.
Okay, so more to come on aggregated
data. We're going to work with those,
learn about the differences, learn about
what it looks like in the code to work
with the set, a dictionary, tupil, list,
but those generally hold multiple values
or can hold multiple values whereas um
scalar data is only going to hold one.
Okay.
All right. So uh so as we said earlier
um you know integers, floats, booleans,
they only hold a single value. By the
way, inside of Python, if you ever want
to see what the type of a variable is.
So let's say we know we have a variable
called name. We can always check the
what data type it is by by using the
built-in type function. So we can use
type and then pass in that variable
and this will display what data type it
is. So um if we stored the value 42 in
some variable called int, if we um
displayed if we did type um if we did
type of this it would uh produce int
which would say okay this value is an
integer versus 3.14 that's going to be a
float versus capital t true that's going
to be uh the boolean type bool.
Okay. So, uh we have integers, we have
floats, we have booleans, all of which
we will use throughout and we'll see
where we will use those one versus the
other. We'll learn about that.
Um as I said, complex. So, uh just
showing you here that those exist
obviously. Um like I said complex has uh
a real part and an imaginary part which
you can access separately. So if you
store a value as a complex you can uh
access its real and imaginary parts
separately which you may need to do for
some type of uh calculations.
Um again we won't really work with
complex numbers in in this program. So
not a big deal for us but it is
supported
and you know a lot of um mathematical
packages in Python will use complex
numbers uh if they need to but we won't
really do it in this program. There's
not really a need to for us.
All right. So aggregated data um we have
those strings which we've already seen.
Those are the things inside of quotes.
We have sets which are going to be
collections of data um that are unique
basically only allowing one uh copy of
those elements inside the set. We're
going to learn about that. Um list which
is going to be a collection of items
which we can change, we can add things
to it, we can remove. Um lists are
really awesome uh structure in Python.
Um, what I want you to see right now
though is you can start to see the
syntax differences, right? So, like a
set, um, a a set is where we have, uh,
this brace. Notice that a set is created
with a curly brace versus a list which
is created with a bracket. So, right
away, like when you see a brace, you
should be thinking either a set or
dictionary. Those are the two things
that are created with a curly brace. Um,
and you know it's a dictionary because a
dictionary will have the colon which
will map I'll show you that on the next
screen. But that will map things from
key to value. Um, depending on if you
know left and right of the colon. Um,
but do you guys see that like the syntax
difference of a list? A list has a
bracket set has a curly brace. Um,
that's just one small difference. you
know, we're going to learn like what is
the actual difference between a set and
a list, but that's just one I'm pointing
out right now.
Um, what does mutable mean? So, mutable
uh means that we can change it. It's
it's able to be changed. So, immutable
would be we cannot change it.
Yeah. And and one thing about a list
that's really nice is every list has a
natural ordering to it which is actually
really beneficial. So a list has a
notion of the first item, the second
item, the third item, the fourth. That's
really important for accessing data
within the list. Okay. So lists are
really powerful. Um
yeah. So a a set the reason it shows
it's in a different order is because a
set does not maintain order. A set never
maintains order because um it a set is
not you do not access items by by order.
So that's just something unique to a set
is that it doesn't have a natural order.
So every time you print it out, it will
display in a different order.
Potentially it's random. It's random
order when you when you display it. A
set is just meant to be a general
collection. Think of it like a bucket.
Like here's this bucket of items that I
have.
It's just a collection of items. A list
actually maintains an order, a
consistent order of items.
This is different.
So we'll talk more about that when we
get into those.
Okay.
All right. So I wanted to show you also
the tupole in the dictionary. So a
tupole
is also a collection of items. Now the
tupil is ordered. So it's like a list.
It's ordered but it is immutable.
Meaning you cannot change a tupole. So
once you create a tupil you cannot
change it or else you'll get an error.
Python will tell you hey this is
immutable I can't change this. So if you
try changing being if you try to add
something to the tupil if you try to
modify one of the entries in the tupil
like if I try to if I go in and try to
change this a um to a d
um this would not be allowed. This would
this would throw an error. The
interpreter would say hey you're trying
to change something that cannot be
changed. So tupils are immutable but
they have a benefit beyond a set of
actually being ordered. So there there's
a natural ordering to a tupil where this
is the first item, this is the second
item, this is the third and every time
you display a tupil will be in a
consistent order. But tupils are not
like a list. You can't change it. So
tupils are useful for situations where
you want ordering, but you don't want
anybody to change any of that data
that's in the tupil. It's it's not
changeable.
Mutable meaning it just means changeable
like you can modify it. If if something
is mutable, you can modify it.
Immutable like a tupole is immutable. We
cannot modify it once we create it.
That's what it is.
Okay.
Yeah. Okay. So then finally a
dictionary. Now you by the way um look
at the tupole. See how it's created with
a parenthesis.
So that's different than the curly
brace. That's different than the
bracket. Right? So a tupole you know
it's a tupil because of the parenthesis
and the items are separated by a comma
just how just how they are in a set and
just how they are in a list. Um
so so the the parenthesis gives it away
that it's a tupole. Um now look at the
dictionary and the dictionary is um a
collection of key value pairs. So this
is a key value pair. This is a key value
pair. Um this is a key value pair and on
and on. We can have as many as we want.
And one thing I want you to notice about
this is there is no restriction on the
data types of the keys and the values.
So keys can be integers, keys can be
strings, values can be integers, values
can be strings, values could be floats,
values could even be other dictionaries
or lists. So value like we could have
what's called a nested dictionary where
we actually have something mapping over
to another dictionary.
That's totally possible in Python. So we
can have dictionaries that part of the
values inside of the dictionary actually
have our dictionaries themselves and
that would represent kind of a nested
structure there. So for instance this
name could map to a dictionary with
everybody's name in it. Um or it could
map to a list um you know
could map to a list it can map to
whatever it could map to a tupole. Uh so
you there's really no restriction in
what the keys and values uh are going to
be.
Uh Brent is it more efficient than the
other uh is what more efficient than the
other methods? Just want to clarify your
question so I so I answer it properly.
The tupole versus using an array.
Yeah. Yeah. Yeah. So uh these are all
good questions. So um the tupole
is guaranteed not to be changed. So it
is a little faster when we are looking
up items like when we are referencing
items. It's a little bit faster because
uh we know that it's not going to be
modified ever. So everything is going to
be consistently in the same spot. So
like whatever's first is going to stay
first, whatever's second is going to
stay second and on and on. So tupil is
is nice in that sense. A list can be
changed. So whatever is first may not
guarantee to be first in the future. We
can modify it. We can remove things. We
can add things to the list. So we can
expand. The list is very like dynamic.
The list. So the list is less efficient
because it's way more dynamic. Does that
make sense? Like it can change. You can
add you can keep expanding the list by
adding things to it. You can shrink the
list by removing things from it.
So list is way more dynamic which for a
lot of scenarios is useful,
right? We want to be able to add and
remove and modify things.
Um but a tupole is more rigid in the
sense that once you create it, you
cannot change anything about it.
Yeah. Yeah. So a dictionary is good for
Yeah. Like a phone book would be a good
example of a dictionary because you with
a dictionary you're usually looking up
things. So you so like a diction in a
phone book you have a name that maps to
a phone number.
Um so yes you have a you have that a
dictionary will map a key to a value
just like a name would be mapped to a
phone number. So yeah a phone book makes
a lot of sense.
Um,
a list, a list is like any is like a
normal like like your grocery list. Like
you may add things to it, you may remove
things from it, you may change things on
it. It's very dynamic. Um, a tupole is
kind of like a fixed um set of data
that's ordered in some way. So maybe
like um what you would see on on a on a
letter like you have your name, you have
your address, you have um your zip code,
like you kind of have those and it it
should stay that way in order to mail
the letter kind of thing.
Uh can you convert a tupil to a list?
Yes, you can do vice versa. You can
convert a tupil to a list and you can
convert um you can convert a list to a
tupole. Yes, you can convert between
them.
I'll show us examples of that later.
Okay. So, just to recap there,
tupil is not changeable, but it has an
order. So, it has a natural ordering to
it. Whatever is first is first. Whatever
second is second, third and third. So,
you can access things based on their
position within the tupole. That's
really nice. But you cannot modify
anything about a tupil once you create
it.
Okay. A list has an ordering to it. You
can access things based on their
position. But a list is dynamic. It is
mutable. Meaning you can change it. You
can change values. You can add things to
it. You can remove things from it. Okay?
So very dynamic. That's what a list is.
Um dictionary. It maps keys to values.
No restrictions on what those keys and
values can be.
All right. And then a set. A set is
think of it like a bucket. It just has
things in it. A set has no order to it.
So you cannot access things based on
their order. And every time you uh
display the set, you can get a different
ordering. Um
but a a set only is special in that it
only allows unique items. So if you try
to put multiple copies of a piece of
data, it's only going to keep one of
them. So a a set is like a bucket with
only unique things in it.
Okay. So and sometimes that's really
useful is to know like what are the
unique values? Uh a set would help us
maintain that.
Any questions about those? You know, we
have to we have to work with this and
see this in the code and we will. But
just any questions right now about these
different types of data that we're
talking about.
Okay.
Very good.
All right. Let's talk about assignment.
So, what that means is um Oops. Let's
talk about assignment which means that
we will be um taking a variable name and
assigning data to it. Now we've already
seen this. We already saw it in our demo
where we did input. We did name equals
input.
So the equals symbol is how we assign
values to a variable.
That makes sense, right? It's very like
self-explanatory.
But um what we should think about with a
variable is really the fact that a
variable is is a reference to that data.
Okay. So when we say x= 34, we are
assigning 34 to the name x. So x becomes
a variable which is referencing the data
which is an integer 34. Right?
What's really interesting about that and
this is how you can kind of test your
intuition of the fact that this is a
reference is if we come along and have
another name Y and we set that equal to
X.
This is just saying that we are creating
another reference that is equal to the
reference we already have. Now, why
would we ever do that? Probably we
wouldn't. That's kind of redundant. But
this just proves that they're ultimately
references because when we display x, we
get 34. Of course, that's what we stored
the value 34
uh referenced by x.
And then when we print y, we get the
same number, right? We get 34. And why
does that happen? Because we we
literally declared y equal to x. Meaning
y should reference the same data that x
does. Okay, so as variables they are
equal meaning that um X is being
assigned to Y meaning Y should reference
the same data that X does. So they they
uh contain the same data. Now what's
interesting is if you print out the ID.
So the ID is the internal
um the internal memory address
of of the reference.
Um now it usually we don't care about
that but this is just to prove the point
is that you can see these are the same
address. These are the same. That's by
design because we're saying okay I have
this reference X which is referencing
this data 34. it's stored at this
address. Um, and then when I come along
and say, okay, y equals x, that's just
the same reference. You see how it's the
same exact address,
same reference.
So, just proving that variables are
literally just references to data. They
allow us to reference that data, which
is really, really, you know, nice. So,
we can reuse x throughout the code. Um,
we can reuse name. we can you know
whatever we create we can reuse.
Um if you look over to the right we have
an alternative example which um now
resets y to store a new value. So
instead of saying y equals to x we
actually overwrite y and reassign it to
the integer 78. That's a new piece of
data right 78. So now if you look at
their their uh references, they're
different. These are different. And that
makes sense because now they're pointing
to two different uh pieces of data,
right? X is pointing to 34. Y is
referencing to 78. So of course they're
going to be different uh different
addresses. And this is a bit of a typo.
This should say ID of Y
because we're ref we're talking about Y.
It's a bit of a typo there.
Okay, so hopefully this now this example
is just to reinforce the fact that when
we use the equal sign, we're setting
equal we're setting a variable name
equal to a piece of data, right? And
that is creating a reference to that
piece of data.
That's all we're that's all we're saying
with this. So we are assigning a piece
of data to that reference X or Y or
whatever it is.
Okay.
All right. Let me ask you guys. Um, what
is the default data type of a variable
assigned using the input function? This
is an interesting question. We didn't
actually cover this, so I'm really
curious to see what you guys think about
this.
A lot of votes for for string.
Let's get a few more.
Perfect. Yeah. So, water votes receipt
it is a string. So, that that begs the
question like what happens if we input a
number? Like what happens if we put in a
two? What happens to that? You know that
two will actually be read in as the
string two. So it would be So if we use
the input and we it pulls up that text
box and we put in a number like two
um and we set that equal to the variable
x, whatever we name that name x,
whatever. What that really means is x is
going to be um equal to the the um x is
going to be equal to the
uh string 2. So that's something to be
cautious about with the input is it
always assumes the input data is going
to be a string. So luckily there's a way
to convert between strings and numbers.
So if we wanted to turn this into the
actual number, what we would do is use
the the data type function int, which
would convert uh this would convert it
over to the numerical two. Would
actually convert it from a string to an
integer. We just use int. Or we could
use like if we if somebody put in a
decimal like 2.5
then um we could do a float
of 2.5
and that would convert that over to uh
the the number.
Okay, let me actually show you guys
this. Let me go over to Collab real
quick and show you guys this. I know
it's not in a demo, but I think it'll be
better if I just show you what I mean by
this because this is an important point
with input.
So, let me uh stop sharing there. Let me
go over to Collab for a second so I can
show you literally what this means.
So, go back into the notebook here. So
what I want to show you is that um when
we do input
the default type
is string.
So for instance when I do um
when I do uh uh value equals to input
and let's try um enter your age.
Oops. Enter your age.
And then we uh run this.
So we enter the age. Now this is going
to be read in as a string. So even
though I'm putting a number there, it's
actually going to be read in as a
string. So now
look at what the type of value is.
It's a string. Do we see that? So
string. So this number even though we
put in a number it gets it the the input
function always converts it to a string
no matter what we put there. If we put a
decimal if we put a a large number it's
always going to assume it's a it's a
string. So luckily
um we can convert to an integer
by using int the int function.
So um we can print sorry we can say
value
uh or we can do int value which which
will convert that 32 string because
right now if I were to um just display
value it's a string 32. You can see it
inside of the quotes. But now when I do
this uh and I can run that now it's an
integer. Do we see that now it's
actually a number
which is great. It no longer has those
quotes. It's actually going to be
treated as an actual integer which which
may be useful for calculations or
storing it or whatever whatever we need
to do with it. So that's just one piece
of caution with the input is if you're
working with numerical data it's going
to treat it as a string. We have to
convert it.
Okay
questions on that. Does that make sense
to us? like the input's always going to
accept the input as a string. So if we
want to work with it alternatively
um we should convert it.
Uh you can yeah so like you could
convert um if I did this if I wrapped
this around in the int function that
would automatically
take whatever we put whatever this
returns would automatically be um cast
over to an int. So we could do that. So
let me show you that. So when I run
this, I can put in 32
and it it's like automatically going to
be casted to an integer. So there now
it's an integer. Does that make sense?
Like when I wrap this int around the
input, it's going to automatically
convert
Uh, what did you put in the input box?
So, yes, you'll get an error if you
don't put in a valid integer.
So, let's put in like if I put in my
name,
this is going to be this should be an
error because I don't know how to
convert this string over to a number. It
doesn't make sense to do that, right?
So, this should be an error.
Right? That will be an error because
it's a string.
So why did you get an error? Uh input
enter your age value int value.
Uh did the did the text box show up?
Maybe try separating it into a different
cell.
Try try putting the other two lines in a
different cell. Um, you need the text
box to show up and then you need to
enter something.
Yeah.
Okay.
All right. Does this all make sense? Any
questions about this? About the input
function.
Okay.
Good. Okay, let me go over back to the
notes then.
Okay.
All right. So, we have another demo.
We'll do that now. Uh I was just kind of
doing one, but let's go back over to
this will be demo five. Let's do that.
So, we're going to practice assigning
different values um to variables and
displaying them just so you get in the
habit of being able to create your own
variables and just go through that kind
of one more time. We'll do this one
relatively quickly um and then uh move
on.
So, this will be uh demo five.
So, let me pull that one up for you
guys.
Okay, let me share my screen.
All right. So, this is going to be demo
five. Um,
now again, like feel free to use
whatever platform you've been using.
Collab, Jupyter Notebook. I know this
instruction says set up a Jupyter
notebook. Feel free to use whatever you
want. You can use Collab. Um, whatever's
been working for you to build your build
your notebooks. So obviously this this
looks a little different than collab but
it's because it's the Jupiter. Um so we
create a notebook.
Now what I want you to see
is this takes the approach of everything
we just did. Let me zoom in on this. Uh,
I know that's a little small,
but this is doing everything we just
said we could do where we
um essentially take
So, I just want to zoom in on this. Um,
notice that we
uh take the um input and this will be
saved as a string.
Um,
so this will be saved as a string and
this will be saved into this name. And
for instance, this will be saved as a
string, but we convert it over to an
int, which is exactly the kind of
example I just did, right? Where we take
take an input, we convert it over to to
uh int.
Does somebody have Yeah. Does somebody
have the demos available? like if if
somebody doesn't mind sharing those in
the chat. I again I don't have the PDFs.
They should be from your LMS. They
should be in the reference material.
There should be a demos folder that you
can download. If somebody has those and
doesn't mind sharing them.
They have that folder of them, like a
zip folder of them, that'd be fantastic.
Yeah. Thanks. Thanks. This is This is
the demo we're going through currently.
Perfect. So, for you guys having trouble
navigating the demos, please download
this zip folder.
Download the zip folder that that these
guys are uploading. Thank you so much.
Download the zip folder so you have all
of them.
Please take a moment to do that.
Okay.
Uh,
copy the code and got an error at height
value. Use foot, not meter. I mean, it
shouldn't matter. It, you know, you
should just be the point of that one is
to put in a decimal.
How to create a new file. Um, what
platform are you on? Collab.
I don't know what platform you're on.
Collab. Uh, just go to file, new
notebook.
New notebook in drive, I think is what
it's called.
Do you see that? It should be like it
should be at the top. There should be a
file and then new notebook.
Let me go over to it.
Uh,
this one. You don't see this
file. It's at the top. The top of the
notebook. Do file and then new notebook.
You don't see new notebook.
Uh if you if you don't see that, just go
to a new tab. Just go to a new tab and
go to um Google Collab.
You can always do that. Just go to just
start a new um just go to Google Collab
and then it will let you like launch a
new notebook. So just just do that. Just
do a new tab if it doesn't work.
Okay.
So, by the way, one of those examples
was entering a float. So, it looked kind
of like this. So we had um our our
height is equal to float and then we had
uh input and then we had um enter your
height and then this was um uh some sort
of uh this should be some sort of
decimal value. So let's say it is um I
don't know uh 5.7
whatever that is uh feet it doesn't it's
just some decimal um and then we hit uh
we hit enter that will store the height
as a float so that when we um display
the height uh it will be rendered as a
float appropriately right that's what
that that's what should happen
that's the point of that It just needs
to be some decimal. It should work.
All right, let me go back to the demo
document.
All right, were you guys able to run
some of these? Like, were you able to
run some of the inputs and change them?
So, try these out on your own real
quick. like try doing int and then input
for enter your age. It should convert
that. You should be putting in a number
or else you'll get an error and it
should convert that over.
I by the way I wouldn't worry about this
last one uh because we haven't learned
about the comparison operator yet which
is this equals equals. So we'll learn
about that in a in a little bit in a few
minutes. So, don't worry about that one
too much right now. But at least these
first few should make some sense and we
should be able to do.
Were you guys able to run one of those
and convert over the the float or int
and do the input and convert it?
Did that work for you?
Give it a try.
Let me clear that. Any questions about
that?
Should look something like this.
Good. We're good on that on converting
over the input. Okay, perfect. Sounds
like Sounds like we're able to run that
and uh it was okay.
What are you entering for the feet?
Like, are you literally entering like
quotes?
Yeah, that's not going to work when you
do that because it's going to um there's
a string f, there's a character there.
it's not going to be able to convert
over to.
So if you did if you did 6.4 that would
work.
Any decimal should work. But like the f
is a character. So the the float doesn't
know how to convert over a character,
right? Yeah. So so that's not going to
work. You need to put in a decimal to to
be able to convert over to the number.
Okay.
Very good. Very good. Let's go back over
to our notes so we can continue along.
1.7. Yeah. If you have any if you have
any character, it's not going to work.
It's not going to work. You need to put
in you need to put in a decimal.
All right, let's talk about operators.
So, these are going to be really
important. Um,
let's talk about operators so that we
can uh
uh be able to compare things and work
with things. Um so let's let's talk
about Python operators.
So what are operators? What do we mean
by that? In Python, operators are
special symbols or keywords that perform
operations. So as the name suggests,
it's performing some level of operation.
Um which means that the interpreter
should do some sort of logical
operation, mathematical operation,
relational operation to produce a
result. Um, so usually that means
there's going to be multiple variables
that are going to be used to do some
operation between. So an example of an
operation would be like adding,
subtracting, multiplying. That's an
operation. But we can have logical
operations like taking the um logical
and or logical or of things. We'll see
what that means. But um in Python,
there's many situations where we want to
we want to be able to do operations
between variables. Whether that's simple
mathematical or maybe some type of
relational like testing if a value is in
a list. That's an important operation.
Is 10 in my list? Is five in my list? Um
those are important operations. So we
want to learn about these operators and
they're going to be really important for
us going forward is because these will
be very standard. um things we will use
as we uh go along. So we're going to
spend some time talking about operators.
Um so it turns out in Python um you can
kind of group operators into many
different categories. Um there's going
to be standard arithmetic operators.
Those are your everyday things like
plus, minus, um division,
multiplication. Um, assignment
operators, which we've already seen, is
things like equals, where we're setting
a reference equal to something. We've
already seen that. That's an assignment.
Comparison, which is things like greater
than or less than. Those are important
for comparing values, comparing
variables. Um, logical operators are
going to be something like and and or,
which will um do a logical operation
between two two boolean values. That'll
be important. And then we have a
collection of miscellaneous operators.
Um those will be things like is
something in a collection like is five
in a list? That's an operator. So we'll
talk we're going to talk about all of
these but just pointing out that there's
many different categories of operators
in Python.
Okay, let's first talk about the
arithmetic operators. So these are going
to be your standard everyday um uh
operations between numbers. So if we
have numerical values like integers or
floats, we can do math between them.
That makes sense. Like that should be a
capability of Python and it certainly
is. We can add things, we can subtract
things, we can multiply things, we can
divide things. So um here are all those
operators. We have plus minus the
asterisk is a multiplication. So x
asterisk y will multiply those together.
So if we have two variables, one of them
is 50, one of them is four, we do x
asterisk y, that's going to multiply
them together to get 200. Pretty pretty
straightforward. Um
division is one that we should be
careful of. Of course, like we don't
want to divide by zero. So if you I if
the uh this secondary value that we end
up dividing by is zero, that'll give us
an error. Um the interpreter will say,
"Hey, you're trying to divide by zero."
We can't do that. It'll it'll produce an
error. So that's the only thing we have
to be on the lookout for with division.
Just don't want to divide by zero.
Um
so all these are pretty standard. I
think they all make sense.
Hopefully they do to you. I think
they're all pretty standard. you know,
the kinds of things you'd see on a on a
basic calculator. They all make sense.
They should exist. Now, here's some more
exotic ones. Um, I don't know if you
guys have ever seen the the modulus
operator, also known as modulo. This is
one that returns the remainder of a
division. Okay? So the the percentage
sign is a mathematical operation between
two numbers that returns not the
quotient like not the actual division
result but the remainder. So 50 divided
by four
um you know four goes into 50 um it goes
in there uh uh 12 times evenly but it
has two left over right. So there the
remainder there is two. So the result of
x mod we would read this as x mod y or
modulo y um returns two. So if you're if
you're unfamiliar with the modulo
operation that seems a little bizarre
that you take these two numbers
um oops it seems a little bizarre that
you take these two numbers and you like
do this operation and you get a
remainder result but it's actually a
very powerful operation. Um the reason
being is that sometimes we want to know
what the remainder is more than we want
to know what the quotient is. For
instance, things that are very like
cyclic in nature. Um so maybe we cycle
through a collection and we want to know
like how many times do we cycle through
and then we have something left over
which is the remainder. Um so the modulo
operation is pretty useful. You could
also check like if a number is even or
odd using this. Like so if you modulo by
two and it returns zero, that means it's
even, right? Because that means there's
there's nothing left over when I divide
by two. So modulo is kind of a nice way
to check if a number is even or odd. Um
so modulo is a pretty nice uh operation.
We'll use it from time to time. Uh but
that is the percent operator. So x
percent y will look for that remainder
of the division. Um now there is also a
double slash operator which is the
integer division operator. This is kind
of the reverse of modulo. It takes the
largest integer quotient that that uh we
can do from a division perspective. So
remember I said 50 / 4. We can divide 4
into 50 12 times evenly and we have two
left over. So the integer division will
just return to us an integer always
which will be that quotient.
So this is the quotient
um and this is the uh remainder of 50 /
4. So the integer division returns to
you that whole number like the largest
number of times that that number goes
into the other. So 12 times evenly
obviously there's a remainder there but
um but but yeah so integer division that
one's useful if we want to know like how
many times can I fit a value into
another value a whole number of times
and that happens from from time to time
we may need to know that.
Okay last operation here is exponent. So
the exponent is the asterisk asterisk
operator. Um so that raises a number to
a power. Um so for instance like x star
y or asteris y would mean that we are
doing an operation like 5 to the 4th
power um which is 625.
Okay. So asterisk pretty useful. Like
probably the most common asterisk would
be squaring something which would be x
um star star 2 which would would would
be um x squared. So I mean that's a
pretty common operation there is to
raise something to the second power
maybe the third raising something to the
fourth probably less common but um the
the asterisk asterisk operator is is how
we do exponents in Python.
Okay,
so these are all basic arithmetic
operations we can do between variables
in Python. All right, any questions on
those? Do those kind of make sense to us
from a syntax perspective?
Pretty straightforward, I think.
Hopefully nothing too surprising there.
Um,
do you used to use module all the time
for date date calculation? Yeah. Yeah.
like when you uh find out how many like
days how many weeks or where you are in
the week, you cycle through like uh
modulo 7 or something.
That make sense?
Okay,
very good.
Okay, I want to talk about assignment
operators now. Now we've already seen
this which is the basic equal sign that
is a data assignment operator right so
that means that we are setting a value
equal to a reference so we are storing
data inside of this reference variable a
we use the basic equal sign as our
assignment operator so that equal sign
is called the assignment operator now
what's really awesome is we can combine
this basic assignment operator with our
arithmetic ones to update values
um and modify them uh as kind of a
shortcut to say uh so so for example
like a plus= 5 really represents the
fact that I want to reassign a to the
result of a + 5. So this means take
whatever it is add five to it and
reassign it to the value of a. So this
is the same thing as if we just shortcut
it in and Python will recognize if we do
plus equals 5 it's the same thing. So
and actually we can do that with any of
these arithmetic operators. So if we
want to take a variable multiply it by
two and reassign it to that variable we
can use star equals. So like a asterisk
equals 2 is the same thing as if we were
to reassign a to the value of a * 2.
Does that make sense on the
reassignment portion of that? So plus
equals divide equal modulo equals star
star equals would exponent something and
reset it back to the variable.
um minus equals we'll subtract and
reassign that back to the variable. So
you know x minus equ= 3 we'll subtract
three from x and re and basically update
it right reassign it back to x.
So so that's pretty useful like whenever
we need to do an operation and add like
um you know a a very typical
reassignment is to do like a plus equals
1
That's a very typical reassignment
because what this is the same as is a
equals a + one. So that's like a single
increment of a. We're just updating it
by one.
So plus equals 1. We may see that from
time to time.
A loop coming on. Yeah. Yeah. These are
used in like while loops. Yeah. Like you
do plus equals and you increment it
until you reach a certain condition.
Yeah.
Now, if you're coming from other
languages, if you have programming
programming experience, you're coming
from other languages, Python does not
have an increment operator like plus+. I
wish it did, but it doesn't. So, like I
know in in Java and I think C they have
um you can do like uh a plus+ or
actually reverse you can do plus a but
um that does not exist in Python
unfortunately. You have to do the plus
equals reassignment. So they don't have
an increment operator. You'd have to
you'd have to do just plus equals one to
do the same effect as plus+.
So I I know some people ask about that,
but yeah, doesn't exist unfortunately.
All right. Any questions about
assignment? It's just really the equal
sign and we can tack on the arithmetic
to do some type of basic math and
reassign to the variable.
Hopefully the fact that we're using a
single equals makes sense. Where people
get confused all the time is the
difference between a single equal sign
and multi and two equal signs which
we're going to see. Two equal signs
means something completely different
than a single equal sign. Single equal
sign is an assignment. We are taking
data and storing it in a reference
variable,
right?
But multiple equal signs, we're going to
learn about what that means. That's
actually a comparison.
It's something different.
All right, we'll continue. Thank you
guys. All right, so we're talking about
uh comparison. So, uh we're going to
talk about a few operators that allow us
to compare two values. Now, this is
going to be useful as we go forward
because sometimes we want to know when
is a value bigger than something or less
than something or equal to something,
not equal to something. Those
comparisons are going to be useful.
um and we have a collection of operators
to do that for us. So again, one that I
think a lot of people get confused on is
the um equals comparison operator which
is uh the double equals symbol. So a lot
of people get confused on that. What is
the difference between a single equal
sign and a double? This double equal
sign is checking if two values are
equal.
Um
so for instance we have uh these two
numbers x and y they're both integers
that are 20. We check if x equals equals
to y and that returns true because
uh these two values are the same. They
both equal 20. So when x equals equals y
that is a true statement. So these
that's something to realize is that
these comparisons are things that return
booleans true or false because a number
is going to be bigger than another yes
you know true or false they they are a a
uh comparison that gives us a kind of a
yes or no answer. Um
so the equals equals checks if two
values are the same and then the uh not
equals operator which is uh an
exclamation point with an equals um
checks to see if two values are
different. So they are not equal. So for
instance if we had um uh 45 and 24 we we
uh do x not equals y that would return
true.
Um now if we had these two values as
before and we checked here x not equals
to y um this would be false because they
are equal right so um
not equals to checks if values are
different so that's a simple comparison
are they not equal um so so in this case
that would return true
so these are pretty useful if we want to
compare directly is a value equal to
another we use the equals equals If
they're different, we use the not
equals. And we're going to have
different scenarios where we will use
those.
I also want to call out the basic, you
know, greater than and less than. So the
this first one is the less than
operator. It is uh going to be obviously
returning true when a number is less
than another number. So when we have
things like 20 and uh 30, this x less
than y would return true because 20 is
definitely smaller than 30. So this
returns true. Um
and then greater than checks if a number
is bigger than another. So that
comparison uh x bigger than y in this
case would uh return true as a
comparison. So again these are all
operators that check uh comparison
between two numbers that will be uh
really useful as we go forward and start
to work with data and numbers and we do
comparisons.
uh we will do those all the time later
on.
Now there's also scenarios when when we
want to know is it less than or equal
to. So that operator just tacks on an
equal sign. So less than equals
is the less than or equal to operator.
So for instance 10 less than or equal to
30 that is true um because 10 is
certainly smaller than 30. But um it
would have been true even if x was 30.
That would also be true because 30 uh 30
equals to um 30 would equal to 30. That
would be a true statement.
Um greater than or equal to same same
scenario. We have a greater than and
then we have an equal sign right after
it. This returns true if something is
bigger than or equal to another number.
So here's an interesting one. We do 30
bigger than or equal to 30. That returns
true because 30 equals to 30. That that
makes sense. So less than or equal to
bigger than or equal to we can we can do
with these simple operators.
Uh is greater than greater than similar
to the usage of brackets? No.
Uh so greater than or greater than is
what's called a uh a bit shift operator.
It's a little bit different. Um I I
would I'm going to save any explanation
that just just look that one up is what
I'll say. It it does like a bit um a bit
manipulation which is um a bit of a bit
of a hassle to deal with but we we won't
ever use greater than or we won't ever
use greater than greater than. It's it
does some sort of a shifting operation
like a bit mathematics which we we don't
need to do.
Okay. So those are comparisons. Um let's
look at our logical operators. So now
these ones are going to be really really
interesting and useful when we get into
controlling the flow of our program. Um
so logical operators are used for
combining conditional statements. So
conditional statements are things that
return these are statements that return
um true or false. So they return a
boolean and we can it's it's like we are
combining them together in certain ways.
Okay,
so the and operator, let's look at that
one first, which in Python is the
literal word and. So that's very nice.
It's it's literally the the keyword and.
Um, and what this does is it takes the
result of some boolean comparison and
some other boolean comparison and
returns true if both of them are true.
So and will only return true as a
combination if both individual
statements are true. They both have to
be true. The moment one of them is false
and will return false.
So this is useful for doing a
combination of things where we want
every individual thing to be true. So a=
1. This is a true statement because a
equals 1 and then b= 2 is a true
statement. So both of these would be
true. So therefore when we combine them
with the and this overall combination is
true.
So keep that in mind. These operators
are ones that combine individual logical
statements or conditional statements.
Right?
Okay. Now the one that is less
restrictive than and is the or statement
which um is used when you only want at
least one of the statements to be true.
So if we want to combine these things
and only require at least a minimum of
one to be true, we use the or statement.
So for instance, a= 1 is true because a=
1. So that's true. and then B equals
equals to 2 is false. So this one is
false. But that doesn't matter from the
perspective of or because we have a
minimum of one of these statements being
true. So or when we use the logical
combination of or um we just need either
or to be true. So a= 1 is true. So this
overall returns true.
So or is something that will combine
conditional statements and return uh
return true if at least one of them is
true. If all of them are false or it
would return false because none of them
are none of them would be true.
Okay. So we have and we have or and then
we have not. So not is an interesting
one. Not essentially reverses a boolean.
So if we have a statement that is
inherently true and we put a not in
front of it, it will invert that to be
false. If we if we have something that
is false and we put a not in front of
it, it will return uh true.
One of the interesting examples in
Python and this trips up people all the
time is the fact that Python treats zero
very specially. So the integer zero
is
oops the integer zero is inherently
treated by Python as false.
So Python treats zero as false and then
every other integer as true. Basically
being Python is indicating that it is
something that is not zero. uh anything.
So like um B equals to one would be
treated as true because it's as long as
it's something that's not zero
then Python treats that integer as as a
true boolean essentially. Um so why
that's interesting is if you put a not
in front of this this would actually
return true because not false what is
the opposite of false? It is true right?
So not false it would return true. So we
we will see not from time to time. Uh
not shows up when we want to negate
something. So when we um you know you
know maybe we have a an iteration an
iterative loop and we say while not
finished and we you know then we will
execute a bunch of statements while we
continue to not be finished and then the
moment that that it finishes then it
then the loop would be over. So not is
powerful to kind of invert uh trus to
falses and falses to true.
Um so maybe we want to check if
something is not empty. Meaning that um
if it's empty
uh if it's not empty that would be
false. Not empty um you know maybe it
would return true. So not is something
we will uh see from time to time as a
negation operator logical negation.
Okay.
Any questions on
uh any questions on these operators?
These now these we're going to use these
in the control of the flow of our
program.
One more example for not. Yeah. So a
pretty typical case for not would be
something like this where um
uh maybe we have some code oops I always
forget to swap over to this maybe we
have some code that checks so if we have
a list
if we have a list and it's uh empty
okay and it has nothing in it let's say
it has nothing in it we could we could
have some code that says like if um if
not
uh list
um then we then do something. So then if
not list uh meaning that it's not empty
then check then grab the first value.
Let's say that grab the first value. So
we'd have some code like this. So uh
this this not is used to like we can
negate the fact that this is going to be
empty and then this would be true and
then we can continue to access something
because that would mean it's not empty.
So not empty is a pretty standard use
case for not like to check that
something is not empty.
Uh can we also use not for checking
value in the list? Yeah, that's that's
what we're doing here to say like is it
not empty?
Okay.
All right.
Let's go to some miscellaneous
operators. So we have now some of these
are going to be incredibly useful. The
one on this page not that useful. The is
mainly because it's very rare that we
would uh that we would check these. So
is is what we call the identity operator
and this is something that um checks to
see if two references are the same.
Okay, two references are the same. um
meaning that they're referencing the
same piece of data. Um so now this is a
very interesting case where we have a
equals to a list b equals to the list.
However, when we ask the question a is b
this would actually return false. Now
that seems very counterintuitive but the
reason that's the case is because we are
creating two different references. We're
saying A equals to this list, B equals
to this list, which is a whole new piece
of data.
It's a whole new piece of data. So
therefore, we can't claim that they're
the same reference even though they're
the Now what would be true is A equals
equals B because their data is the same.
That would be true, but their their
references are different because they're
different variables, right? A and B are
different variables.
Different memory location. Exactly.
Different references. So A and A is B is
the same as checking um if ID A equals
equals IDB. Does that make sense? That's
basically checking that that logical uh
comparison if their addresses are the
same. It's the same check. So is is
basically a shorthand for doing this.
And so th those would be false because
they're going to be two different
references, A and B.
Um, however, we can use our not. So A is
not B. That's actually true because it's
the inverse of is, right? So that that
actually would invert the false and this
would be true. A is not B. That is true.
Behind the scenes, yeah, like the
memory, yeah, the location in the
computer memory is different. Yes,
because they are different variables,
different references.
Yeah, the data is equal. The data is
equal, but the references are different,
which is what this checks. You know, we
have two different names, A and B. Those
are different.
Now, take a look at this last example.
This is saying A is a list. B equals to
A. Now, remember what that does? That is
the assignment of we're saying B is the
same reference as A. That's what this
does here. The same reference. So does
it make sense to us that when we ask now
A is B. This should be true. And it is
like this is true because um
this is true because they are literally
the same reference. We're setting B
equals to A. So they are referring to
the same data. Now
um in terms of a variable reference they
so now their ids are the same
essentially their memory locations are
the same.
If you want to compare only data then uh
what operators have we looked at that
are comparison
we should think about that I mean we
just saw if we go back a couple slides
we have a bunch of operators to compare
data that's these guys right like equals
equals greater than less than so what we
could do is say does the data equal the
other data which would be something like
this equals equals operator
when two values are equal not the
reference ES. Does that make sense? This
equals equals is checking if two values
are the same, which is the the data, not
the not the reference.
Okay.
Now I want to show you a really powerful
um I want to show you a really powerful
miscellaneous operator which is going to
be uh which is going to be the in or
what's called the membership operator.
So this is an operator that checks if a
value is a member of a collection.
So this could be like uh this could be
like um you know where we have a list, a
tupil, a dictionary, just a collection
of data and we want to know is a value a
member of that collection which is
really useful for testing you know do we
have membership of something inside of
something else. So for instance, let's
say we have a list and we have a list A
equals and then we have 20 45 and 10
inside that list. So if we ask the
question
10 in A, this actually would return true
because 10 is a member of A. 10 is a
member of that list. So that is true.
Now that's useful to know. So the N
operator is a really powerful uh really
powerful operator.
Um, same same thing with not. So, we can
use not in. So, 10 not NA would be false
because there it is. We know it's a
member of A. So, that would be false. Of
course, it's NA. We can see it right
there. It's a member of that list.
Um, but if we check the different value
that's completely not inside of a, 30,
not NA, that would return true because
30 is not a member of that list. And by
the way, this in operator works for all
kinds of collections. So it would work
for a tupole, it would work for a set,
it would work for a dictionary.
Um, it would work it would work for all
kinds of collections.
Yeah, Roberto. Um, it's the fact that
there are different variables. So we may
sometimes it makes sense to have
different variables that are that maybe
they have the same value but they're
different variables altogether different
references
and the reason that is is because maybe
we have a copy of that data and then
maybe we manipulate the this one. Maybe
this is a copy of it and we manipulate
this guy and we leave this guy the same
to check the differences later.
Yes, references are tied to the variable
like A is a reference, B is a reference
even though their data that they're
pointing to is the same value.
Maybe we just have a copy of it that and
we manipulate one of those copies.
Yeah, that's why
What is the reason to check for what? If
they're equal, like as references, A is
B.
Honestly, there's not many good reasons.
Um, maybe if you want to know if
something is a copy of something else,
like you want to know that, like let's
say you're checking later down the code
and you have an A and a B and you want
to know if one is a copy of the other,
you can check and see if they're the
same reference.
That's the only reason I could think of
why you would do that. It's rarely used.
Rarely used, but it is an operator that
I wanted to show you in case you do
stumble across it um somewhere and
you're reading about Python or something
and you see the is
Like that's the only reason I can really
think of is to check if something is a
copy of another meaning it's the same
like maybe it's it's uh the same
reference
B equals to A then we could check A is B
and we know that then they're the same
reference.
Yes, the is operator is comparing
references not exact not the values.
Yes, that's true.
All right. How do we feel about this in
operator? Like the the membership
operator. Does that make sense? If
you're checking if an if a value is a
member of a collection,
that's going to be highly useful later
on. Highly useful. This one we will use
quite a bit. The is we will probably
rarely ever use, but but this one we
will definitely use.
Okay. So, I wanted to quiz you guys. Um,
what is the main difference between the
equals equals and the is operator? What
is the main difference?
We were we've just been discussing this,
so hopefully this this one is easy. Been
discussing it quite a bit.
Yeah, Roberto, now now you know the
answer. Perfect. Yeah, it is B. All you
guys answering B. Perfect. It is B. Good
job. You guys are right on top of that.
Good job. So, just wanted to point that
out. Like equals equals compares the
values. We're doing a comparison. Um is
checks if the references are the same,
which is the variable reference like the
memory location.
Yes, it is. Identity is the same as Yep,
that's what we mean. The references are
the same.
All right. So, wanted to do a short demo
on the operators on comparison, etc.
Wanted to just show off that demo um so
you can see and practice with it a
little bit. Uh so, let's hop over to
demo six inside of the lesson one. I'm
going to hop over to that.
which will be our last thing we will do
in lesson one and we'll move on to
lesson two.
Um, let me pull up demo six here.
Give me a moment.
All
right,
let me share my screen.
All right, so demo six. Uh, hopefully
you guys have access to this. This is
the last one inside of lesson one. Um,
now again, this one says try to use VS
Code. You can if you want. Again, Collab
works fine. You can use whatever you've
been using, Jupiter, Collab, VS Code,
whatever works for you. No big deal on
which one you use.
So, no worries on any of this.
Oh, yeah. So, um, F is So, yeah, this
that's a good question. What is F? So, F
tells Python to format. F is short for
format. It's basically format the the
string which we're going to print by
having some placeholders.
And um
we have variables called A and B. And
this fills in the blank of these
placeholders with whatever the values of
A and B are. So F just allows us to
format and fill in the blanks. Does that
make sense? Like wherever these um
braces are, we have a variable name
inside of it A and B. And we are um
going to fill in the blanks of those A
and B um by uh you know by just
replacing them whenever we do the print
function.
So there think of it as F is short for
format.
and we have a couple placeholders and
those will be filled in by our variables
A and B.
Okay, so this demo uh does a bunch of
operators. So it's going to do a bunch
of comparisons where we input one number
and we turn that into an integer. Input
another number, turn that into an
integer, and then do a bunch of
comparisons. So I'm going to jump over
to the notebook. I'll do that for us to
to show that off. But that's all we're
doing in in really in the beginning of
this uh demo. Um so let me jump over to
the notebook and show off that so we can
see it
but uh should be straightforward to
follow because we've done a lot of that
already.
Okay. So hopping over to notebook. Again
feel free to use whatever you want to
use. You can use collab, you can use um
uh Jupiter, you can use VS Code,
whatever you use. Um let's store a
variable as a and let's make it an
integer
and let's do enter your first number.
So we will do that.
Let's run that. So let's enter our first
number. Let's put in 10 or whatever you
want really, but I'm going to put in 10.
So that gets stored as a. Now let's do a
second number and let's do int
input um enter your second number
and let's input that.
So now I'm going to put in a second
number. Let's do 20.
So now we have a and b. Now let's do
some comparisons. So let's do print.
Um then we can do uh let's do the f
formatting like they had in there. Now
what this is going to be is we are going
to check a
um greater than or let's do yeah let's
do greater than b
um is
and then let's do uh comma a greater
than b.
Let's compare those two numbers. So now
this is doing the comparison. A greater
than b is going to compare those two
integers. What should this return? What
should a greater than b return? Should
it be true or should it be false?
Should be false. Right? So we should we
should uh display false here.
And that's what it is. 10 greater than
20 is false.
Okay, so that is false.
Let's do another one.
Let's do um let's try equals equals. So
let's say um a
equals equals to b is now what do we
think this one's going to be?
a equals equals to b.
What should this one be?
Yes, very good. This one should also be
false.
Let's run that.
And that one will be false. Very good.
Okay.
So, I think we get that. Let me give you
guys a Let me uh show you something
interesting. Let me uh go a little off
script from that demo document and let's
introduce a third number called C. Let's
do a third integer.
So enter your third number.
Let's do a third number.
Let's enter um another value of 10.
Okay. So, we have another number of 10.
Now, what I want to do is let's just do
uh multiple comparisons and do a logical
operator between them. So let's do um a
not equal to b
and
a less than c
or sorry b
less than c.
What do we think this is going to
return? This might be a little
challenging. What do you think this is
going to return?
We should use F. We can. I'm just not
printing. I'm just going to I'm just
going to run the cell. I'm I'm kind of
doing a shortcut and just print. I'm not
going to print. I'm just going to run
the cell and it should display what it
is.
Do we think it's going to be false?
Yeah, you guys are right on top of it.
Should be false. Now, let's break that
down. Why is that false? So, A not equal
to B is checking if A is not equal to B,
which is true.
A has a value of 10.
So, A has a value of 10. So, 10 is not
equal to 20. That makes sense. But 20 is
not less than 10. So this part is false.
So let's make a comment.
So the reason reason this is false
is because B is not less than C. So and
returns false.
Yes. Very good.
Now,
what if I take this code
and do this?
What's this?
Perfect. Yeah, you guys are right on top
of it. This should be true. And it is.
Now that's because the moment we have at
least one true which is going to be this
that makes sense right that is going to
be true.
Okay, one more and then we can wrap up
this demo. So I want to create a list.
I'm going to call it X and I'm going to
create a list of three numbers 10 20 30.
Okay. Now, what do you think uh this
result is?
What what should this be?
It should be true.
Very good. Should be true.
Now what should what should this be?
This is also true. Very good. This
should this should be true because this
is not going to be inside of the list.
So that makes sense. That is not true.
Now what I want you to notice is nothing
will change if I change this to a
tupole.
Nothing will change. We can still check.
So we can still check if uh 10 is a
member of this tupole and we can still
check if 40 is not a member of this
tupole. So nothing really changes,
right? It's still in membership operator
in checks is it a member of any
collection?
Okay,
very good. Any questions about the these
examples?
Any questions? Do we feel comfortable
with some of these operators? You guys
were right on top of it. It was very
impressive. You guys got those right
away.
Uh, you change print
a= b and a is b. And I entered four both
times
and I got the same result. True.
Uh so you had your A and your B. So you
you had um
you had A equals to 4
and B equals to 4 and you checked
uh you checked this.
Yes.
Yeah. So that is a little confusing and
I can understand why. So the reason this
ends up being true is because this is a
scalar data. So for scalar data it's
going to optimize in the memory to point
because four the integer four occupies
the same memory address always. But when
we create an array, when we create a
list, that is um a new object.
So yeah, that one's a little confusing,
but it's it's only because this is
scalar data that um Python kind of
knows, okay, four is the same like
integer in the me in memory always
regardless of if we're referencing it
from this from two different variables.
That's that's the reason I um yeah I
forgot to mention that example but it's
purely so the the reason this is true is
because
this is true because of scalar
data optimization. Essentially it's it's
not going to waste creating a new object
when it's just a single integer four. it
basically occupies the same memory
address
uh as as a as a reference.
But when we build a list like that is a
different object.
Okay.
Very good.
All right. So,
um, that will wrap up lesson one. What
I'm going to do is go into lesson two.
I'm going to pull up the lesson two
notes, the slides for lesson two. So, if
you have those, let's pull those up. Um,
I do encourage you now, there is a
guided practice at the end of lesson
one, and that is for your yourself. I
would encourage you as as kind of
homework between now and and the next
time we meet um to do the guided
practice for lesson one. Try that out on
your own. It is um you guys should have
access to it from your LMS and the
reference materials. There's a guided
practice. Try that. Try those out. Okay?
Try out the guided practice for lesson
one. It's just it's just some additional
practice of the things we just covered.
Okay?
Let me open up
um
lesson two.
Give me a moment here.
Okay.
So, let me share my screen.
Okay, so now we're going to move on to
looking more closely at those things
like lists, tupils, dictionaries, and
then looking at control flow with
conditional statements and loops. So
we're going to get into the fun stuff, I
would think, um that you guys may may
have been waiting for.
Okay, so we've talked about uh we just
finished talking about lesson one where
we have you know Python as a really
important thing to learn and study
because it's used all over the place
with data science and a IML. So one of
the things is we need to continue
learning about it with things like loops
things like if else statements to
control the flow of our programs and
these basic data structures.
Let's continue forward. Um so this
lesson we're going to talk about list
tupils dictionary sets uh we're going to
you know talk about the differences how
we can access data within things like
list how we can modify them um how we
can access things from tupils what are
the differences we'll review all that
one of the big topics is going to be to
um look at how we can control the flow
meaning control the flow is using like
decision logic like if this is true then
do this else do this we'll talk about
those kind of statements. We'll talk
about iteration. So how we can do loops
to repeat um statements of code that we
want to do. Um we'll also talk about
organizing our code a bit into
functions. Um which is going to be
really helpful to for our own
organization and reuse and
maintainability.
Okay.
So let's jump into it. Let's talk about
some of those data structures.
So, we've already talked about that
these aggregate data structures exist um
and they allow us to manipulate data
inside of Python. Lists, tupil, sets,
dictionaries are the main ones we're
going to focus on.
Let's start with lists. And I think
lists are going to be something we're
going to use quite a bit of throughout.
So, they're going to be a really good
place to start with. They are super
popular in Python. A lot of people um
use them to do to work with data. They
are kind of the most basic um data most
basic and useful data structure that
there there is.
So what is a list? A list is a an
ordered mutable meaning it can be
modified data structure that can hold
elements of different data types. So
it's a collection of data and there's no
requirement that all the members of the
list be the same type. In fact, you can
have different types. You can have
integers, you can have floats, you can
have strings,
um you can have even more abstract
objects be members of a list. Um so any
kind of data can live within a list. But
the big thing is that it is modifiable.
It's dynamic. You can change you can add
things to it. You can remove you can
change things. Um and it has an inherent
ordering which is nice. So you can you
can be reassured that there is some
inherent position of items and it will
maintain that order. So we can access
things based on the order like we can
access the first can access the last we
can access anything in between.
So that's nice. Um so what are some key
characteristics? So uh lists support
multiple data types. We talked about
that. There's no requirement that
they're all the same. They can have
multiple. They allow for indexing, which
we're going to talk about. This means
that we can access things based on their
index, which is their position within
the list. So, we can always access the
first thing, the last thing, anything in
between, based on its position. That's
another word for uh index because lists
have natural ordering to them which is
um really really powerful to ensure that
there's um one you know there's a first
position a second position third etc. So
lists are really nice for that.
Um
uh they are modifiable which is really
nice. So we can add things into the
list. It's very dynamic. So once we
create a list, we can throughout our
program, we can add data to it, we can
remove it, we can change. Um, lists also
allow duplicates, which may be
desirable. Like maybe we add something
into our list that already exists.
That's okay. List allow duplicates. This
is going to be different than sets. Sets
do not allow duplicates. Sets are just a
bucket of unique things. So if we added
a duplicate into a set, it would reject
it. we wouldn't have any errors, but it
just wouldn't um show up as a copy. We
would just have a set is only going to
maintain one copy of an item. It it only
allows unique items. Lists allow you can
have as many copies of data as you want
inside of a list. Um so it does allow
duplicates.
Now, we already saw in terms of syntax,
lists are um defined by brackets. So
when you see those brackets um it
defines a list and its items are
separated by uh commas.
So
we can have a list that looks like this.
Yeah, I was going to explain slice. We
have a a couple slides about slicing
coming up, but slicing just means that
we can slicing means that we can grab a
section of elements at a time from a
list. So for instance we can uh actually
let me use this example down here. A
slice would mean we can grab like these
first three slice of the of the list or
we can grab the last five elements or
whatever like this is a slice. It's just
a a subset of the list that we can grab
we can access.
So slicing just means taking a subset,
taking a smaller section of the list and
we can grab all those elements at a
time. And what that's called is within a
slice.
Yeah, like a slice of pizza. We're
taking the whole thing and we're taking
a small section of it.
So that's and actually that's going to
be possible because the list has a
natural order to it.
Does the data have to be sequential? No,
it doesn't have to be. In fact, you it
can be completely different types.
Does that make sense? Like so look at
this example down here. Like we have 10
2 5 hello. That's a valid list. You can
have different types of data in there
which isn't sequential at all
in a slice. No, it doesn't have to be.
So you can have like you can have a
slice that picks every third element
uh every other element. Um yeah, it
doesn't have to be sequential. No, it
can be customizable.
What's also nice is you can slice from
the beginning or you can slice from the
end as well. So you can you can go from
the end and slice backwards. Um or you
can go from the beginning and slice
forwards. So you can grab like every
other element from the beginning. You
can grab every other element from the
back and work your way forward and stop
at a certain point.
Slicing is very nice. Yeah. So, I'm
going to show us how to do that.
Can you slice in the middle? Yeah, you
can slice anywhere you want.
Can you slice a pizza in the middle?
Sure. Would you do that? Maybe not. But
yeah, you can slice anywhere in the
list. You can slice.
Does slicing change the original list?
No, it's just selection of a subset.
No, it just it just extracts elements.
It doesn't like permanently change it in
any way. It just gives you a view. Think
of it as like giving you a view of that
subset.
Okay, so you may be wondering when would
I ever use list? So normally you use
lists whenever you want an ordered
collection that is dynamic meaning it
may you may want to add things from it.
Um we may want to modify things from it
frequently. Um but we want something
that is dynamic and has an ordering to
it. Lists are perfect for that reason.
So they can contain data that we can add
to remove from change. So lists are very
versatile. I think most use cases with
manipulating
um a collection of items would fall
under a list. A list would be a very
good choice. Um
we can slice elements and use them.
Yeah. Yeah. You can slice you can store
a slice inside of a variable and use use
the use that resulting variable. Yeah,
for sure.
order information like
Yeah. Yeah, that's a good example. Yeah,
order like the collection of orders
would probably be in a list because it
can change. Um, but the prices would be
probably static. So, that would be more
uh suitable for a tupole. Yeah, a tupole
or maybe a dictionary. A dictionary is
probably better because you can have
like a product ID or a name that maps to
a price. Probably a dictionary would be
more appropriate for a price list. But
but yeah, tupole maybe makes sense too.
By the way, look at this list. Do you
guys see how it has different types of
data in it, right? It has so like the
first element is an integer, the next is
a string, the next is a float, the last
is a boolean. That's totally valid in
Python, which is kind of unique to
Python. Like a lot of languages don't
support that. A mixtyped array.
Don't really they don't really have that
notion of that.
All right. So, what I wanted to get to
was the positions. So, this is going to
be this is really important to pay
attention to because this is going to be
something that we will be using
throughout the program is how to access
data by its position.
Okay, which so there's another word for
that. The position um sometimes you will
hear called the index. So the index in
the list of where these where data
members live is their order like their
position amongst amongst the list.
So what's special about Python is that
it has the first position
is index zero which trips people up all
the time. The very classic trip up of
Python is that the first element of a
list is at position zero. The next
element is at position one. The next
element is at position two. On and on
and on. The last element is at position
n minus one where n is the size of the
um size of list. So, however many
elements we have in our list.
Um,
so
if you want to access the first element,
you would be looking for the element
that's at position zero. If you want to
access the second element, that's this
this name Bob, that is at position
number one or index one. So that's
something that trips up people is that
it actually starts
starts at zero
which is really important to understand
is that the positions start at zero.
Why is that? Um that's a good question.
I mean
different so so different languages
treat that differently. Um,
it's more historical reasons that it was
created that way. Uh, but I think it's
based on your like
think of it as like how far into the
list you are. So, if you're in the
beginning, that means you're basically
at at um position zero because you
haven't made any progress like
traversing the list. I think that was
the intuition.
Oh, yeah. which is kind of what Tim says
there is like yeah if you're at the very
beginning of the list you're at position
zero because you haven't made any you
haven't made any forward progress in
traversing so it's like you're at step
zero you're at the beginning
okay but okay so aside from why it was
that way does it make sense that that
the first like what I'm saying is the
position zero is the first element,
position one is the second element,
position two is the third element, and
on and on and on.
That's that's how it works in Python.
Okay.
Now,
what I'm telling you is the positions if
you were to view it going left to right.
So, in other words, in the forward
direction, we start from zero and go up
to the to the n minus one in terms of
the position. Now, what's really
convenient is the positions can also be
indexed from back to front, meaning they
can also be indexed going this way.
And what's really nice is the very last
element starts at position minus one.
The very last element on the right
starts at minus1 and then goes all the
way up to minus n.
Okay, now that's really convenient
because if we want to access the very
last element, I don't need to know how
big the list is. I just need to access
the position minus one. That guarantees
to access the very last element. The
minus one position is the very last
element.
And then it it so this is the very last
element. This is the second to last,
third to last, fourth to last, on and on
and on. And then by the time you get to
the front, it is minus n, which is like
the
uh number of elements in the list minus.
So in this case, minus 6 is the very
beginning.
Now, why why in the world do we care
about negative position? It's for that
exact reason.
we can um we can think about the
positions from the end of the list going
backward which is very powerful like if
I want to access things from the very
end I don't need to know exactly how
many there are I just need to know minus
one is the end minus two is the second
to last minus3 is the third from last um
which is pretty convenient
okay
so does that make sense on the negative
index index. It's the negative index is
going from right to left. It's from the
back to the front. Minus one is the last
element.
And then you go second to last, third to
last, right? And that is minus 2, -3,
-4.
So from right to left. Yeah. Right.
Right to left. minus 6 is actually the
first element
because it's six back from the end which
is the f which is the front. There's
only six elements.
Yeah.
So are there any question about this is
really important to understand really
important to understand because this is
how we are going to slice and access
data is based on these positions.
Uh is there syntax to get the number of
elements in a list? Yes, it's the length
function which is len. So length of list
would give you uh the number of elements
which in this case is six. So ln the
length function gives you how many
elements are in the list.
len or length. So this guy this
function.
Uh when are you counting backwards?
You're usually counting backwards when
you want to know what's at the end of
the list. So when you only care what has
been added at the end and you want to
maybe you want to get the last five
elements.
So you slice backwards from the from the
end.
That's that may be useful like maybe you
want to know like it imagine a list is
holding in your orders
and so you want the last five orders. So
you can just go from the back and go
towards the front. Min -1, -2, -3, -4,
-5
would be the last five orders. There's
going to be many scenarios where we want
to count from the end.
uh when we're manipulating data
later on when we're when we're working
with um bigger sets of data um in
something like an like a matrix, it's
going to be useful to grab like the last
five rows, last 10 rows.
So going from the end makes more sense.
So you imagine if we have like a larger
matrix of data um maybe we want to slice
out these last 10 rows in which case we
want to count from the minus like the
last row is minus one and we want to go
back towards the front.
Can you sort? Yes, you can sort. Uh I
I'll have an example of that in a
second. Yeah, you can sort.
There's a built-in sorting function in
Python that that will allow you to sort
data in a list. Yes,
there's a built-in function for that.
Okay. What I wanted to do is show you
guys an example of accessing elements
from the list. So assuming we know the
position which is that index all we have
to do is use brackets to access items of
a list. So imagine we have this list
called fruits which has some strings in
it apple banana cherry mango and we want
to access the first element. So that is
just this code here fruits bracket zero.
So the bracket tells the interpreter,
hey, I want to access something within
this list. And then all we have to do is
give it the position that we want to
access. So this is position zero, which
is going to be the first element. This
would retrieve the first element. Now,
it's not removing it. It's just
accessing it so we can view what that
is. So, it's actually going to give us a
copy of what that value is under the
hood. Basically, a copy of it. It's not
going to permanently. It's not going to
delete it. It's not going to remove it.
It's going to give us a
copy of what is at position zero. In
this case, apple.
Now if we put in a position two in that
bracket that should be remember the
indexing is 0 1 2 3
because there's four items. So the last
element is n minus one which is three.
So the item that's at position two is
going to be cherry. So this should
return the string cherry.
Okay. So we use we always use this kind
of syntax
position
to access the element that is at that
position.
Okay, pretty simple. And what's really
nice by the way about this is this the
same exact thing works for tupils
because tupils also are ordered. So
we're going to see that when we get into
tupils. But the same exact things works
where a tupole has a first item, a
second item, a third which are index
zero, 1, two, three. And we can access
things in the exact same way. So we can
access a tupole by its position as well.
Same exact thing will happen position.
So we can get the first item of a tupole
with with uh by doing um zero and we can
get the the last by doing minus one.
and on and on.
Any questions about the accessing?
Can we get position based on value? Yes.
So that is a special function called
index.
So if you do like list dot dot index
and then you pass in a value like apple,
this would return to you the index of
the first occurrence. Not every
occurrence, but just the first. So like
because we could have multiple copies of
Apple in our list, but the first time we
stumble upon Apple, this would return
this would return uh zero because apple
occurs at index zero.
So index function
uh would return um the index
which is which is the same word as
position.
Okay.
All right. Any other questions? We can
uh I think what we'll do is we can take
if unless there's any other questions we
can take a short break and then come
back and continue.
Uh, I want to talk about slicing.
No, it would return none. It would
return null. Basically none
as opening closing braces are square on
it if it's a list. No, it's actually the
same as a tupole. Uh in terms of
accessing it's the same
uh
it's it's in terms of accessing it's the
same. Um, for creating a list, yes, it
is square brackets. This creates a list.
The brackets always create a list. For
accessing elements, this is the same as
a tupole. The the square brackets for
accessing.
Actually, in most things in Python, it's
the same where we uh access data using
the square brackets. A dictionary, same
thing. We access keys by using the
square brackets.
So square brackets in Python is is
basically like an access operator.
Uh no, so list.index doesn't only work
for the first item. What I'm saying is
you can have lists that have multiple
I'm saying it returns to you the first
occurrence,
the index of the first occurrence
because we could have a copy of Apple
later on in the list, right? So there's
nothing that stops us from having a
duplicate. So, we could have another
apple down here. Like, let's say we had
another apple at the end of the list. If
I did list.index apple, it's only going
to return to me this first one, even
though there's another copy of it in the
list.
So, it only returns to you the first
occurrence.
But, you know, there's nothing stopping
me from doing list.index of banana or
cherry or whatever.
Are there data types by this for all
Unicode characters? Um, I think there's
Yeah, I think there's special strings
you can make that do Unicode,
but I'm honestly not 100% sure.
I would research that. I don't really
know. I think you can do that with
strings,
special strings.
Yeah, I don't think there's anything
like I don't think there's anything
inherently special about it that Python
can't handle. It's just you sometimes
have to like escape characters
uh
to distinguish them. But
yeah.
All right. Going back to our example,
um we had the uh going back to that
fruits list.
We can access things from the end. So
here's an example of us using the
negative index, right? So minus one
grabs the element that's at the last uh
that's the last member of the list. So
mango minus3
um would be the uh third from last,
which is going to be banana.
So negative index really handy to access
from the end of the list going backward.
Um pretty useful there.
There's an example of it.
All right.
All right. Let's talk about slicing.
Okay. Let's talk about slicing. So,
slicing allows us to extract a subset of
items from a list using a specific range
of indices. And so the syntax to do
slicing
is going to be using our brackets again,
but um we use a colon to specify where
we are starting and stopping our
positions. And not only that, but we
also use uh a colon to signal how many
we want to step by. So do we want to do
every other in which case we would step
by two. Do you want to do every third
element which would step by three? So
most slicing is going to follow um this
sort of syntax where we do a list and
then we do like a start index and then
we do colon
and then we do stop index
um and then we do colon uh step. Now
what you will see is that the the step
defaults to a step size of one meaning
we grab every element in between
starting and stop. Um so step size of
one is the default. So we actually
typically will not include the step size
unless we specifically want to get every
other which would be a step size of two
or every third or every fourth or every
fifth. Um so usually we leave off this
step size and we just we just have a
starting and a stop as part of our
slice. Um and what what that does is it
tells the interpreter to access
everything between this start and stop.
So um with one catch which is a very
important catch that trips everyone up
which is that Python um is very annoying
and that it uh when you do slicing it
allows you to include the starting
index. So uh if we start somewhere we
will guarantee that the slice will
include that but it will not include the
stopping index. it will do everything
between there up to the stop index but
not actually including what's what's at
the stop position.
So for example,
this slice that you see um on the screen
is a slice that would be starting at
position two
because remember um let me draw this
out. This is position zero. This is
position one, position two, position
three, position four, five, and six. And
so this is a slice that would start at
position two
is our start.
And meaning we're guaranteed to get 34
because we're starting there. But the
stop for this slice would have to be
here. This is our stop because we are
going to include everything in between
there. We're going to include all of
this this uh slice as part of our uh
what we can access. Um so that would be
positions 2, three, and four. So this
slice would be um basically like this
list. Um it would start at two and go to
five.
And then it technically would be a step
size of one, but remember we don't
really need to include that. So this
would really be list um two to five.
That does feel annoying. Yes, it's it's
because they you have to know where to
start and stop. And so they cho Python
chooses to to be um not inclusive of the
stopping index, but it it includes the
starting index. Um and and that's just a
choice. That's a design choice of
Python.
Okay. So the colon gives us the slice.
Um, so if you see if you see uh uh if
you see
a colon inside of a brackets, that
signals you're grabbing a collection of
items. So we So this this actually
returns a smaller list. This returns a
list of 34, 20, and 80. So it's a it's a
slice meaning we get multiple items
rather than just a single item from from
the selection.
No, 54 would not make the cut. 67
doesn't make the cut either because
remember we don't include the stopping
index.
Can you do two to four plus one? Yes,
you can do arithmetic in there and
Python will evaluate the arithmetic
first. So it will do 4 + 1 first and
determine that is five.
Yes, you can do that.
Okay. So before
before I move on, uh does the slicing
idea make sense? It is grabbing a
collection of items from the larger list
and we set it up with this syntax.
Uh what do you think? What do you think
0 to six would return? What would be
your guess?
What's included in the slice though?
It's not just 67. If you go 0 to six,
how many? Like, you should be getting
more than that. Yeah, you should get
everything, right? Exactly. You should
get all of those numbers up to
uh 54. You would not include 54. So you
would get 76, you get 12, you get 34,
20, 80, and 67.
Yes, that's true.
Yeah. Yeah. So the the step relevance is
that we can so from our slice we can
choose um our step size of how many
element like how much we want to skip
positions within that slice. So step of
one which is a default means we get
every position we go we increment by one
one position to the next to the next to
the next. A step of two
uh
a step of two would be that uh I'm going
to grab every other element from the
start. So a step of two like if I sliced
this
and changed this to a step size of two,
then that would only grab this and this
because that would step over. It would
take two steps to get to the next
element of the slice.
Whereas a step size of one is going to
grab everything
because it's going to go one index to
the next. So think of the step size as
how many positions are we incrementing?
Yes. 2 to 7 would include 54. Yes,
that's right.
2 to 7 would include 54.
Okay.
All right. Let's see some more examples.
So if we have a list like this and we
slice it from 1 to 4, the output is
going to start at index one
and go all the way up to index 4 but not
include four. So index 4 is this guy and
it's not going to include that. So it's
going to be these three elements here
would be our slice. Yep. 20 30 40. Very
good. That's what it would be.
So that's pretty useful
to be able to slice. Let me show you
another example.
Oops, don't have another. Let me go
back.
Let me show you another example with
this. Um, so what we can do is we can
actually slice backwards as well. So,
what do you Let me ask you guys this.
What do you think this slice would be?
Actually, let me erase this. Let me do
minus
or
what do you think this would return?
We can use negative index index uh
indices in our slices.
What do you think that would be?
Very good. 30 40 50. So it's going to
go. So remember -4 is the fourth from
the last. So it's going to be this is
minus4
and then this is minus one. So we're not
going to include minus one. So we should
be doing this slice here.
That should be the slice. 30 40 50 would
be that.
Okay. One other a couple other examples
I wanted to give you is that you can act
in certain special cases you can
actually leave off the starting and
stop. And what that would signal the
interpreter is that you want to go all
the way to the end or start all the way
from the beginning. So if you do
something like this,
let me show you an example. If you do
something like this and you do not
include, you leave the the start blank.
You leave the start blank and you go all
the way up to minus one. What that would
signal to the interpreter is by default
um start at the beginning. So start at
zero. Essentially start at zero. If you
leave off a slice
uh as your start that the interpreter
assumes you want to start at the
beginning.
So what do you think this slice would be
knowing that?
Yes, exactly. You guys you guys are
right on top of it. 10 to 50. Perfect.
So it's going to be everything but the
last
everything but the last would be
included in that slice.
Perfect. And so the other thing is we
can leave off the end which would signal
that we want to go all the way to the
end. So what do you guys think this is
going to be?
What would that be?
Yes. So this is So this is actually
going to include the end. So I know
that's a little counterintuitive, but
it's actually this this guarantees we
include the end. So if it's blank, it's
going to go all the way to the end,
including the end. I know that's that's
annoying. I don't blame you for thinking
it should be 50,
but it basically goes to n. Basically
goes to n, which would mean that we
remember the index is n minus one is the
last index.
So, so if it's blank, this means we
should go we should end at n
which is the length of the list meaning
that um the last element is at n minus
one. So we should include the n minus
one. So yeah that would be 20 to 60
actually sorry 30 to 60 because we uh
index two is 30. So that would be So
that slice would be this one all the way
to the end.
Good. I have one more example for you
that's really going to throw you for a
loop is this one. So what happens if we
have
this case?
This probably won't for loop, but I'll
have one more following that.
All yeah, 10 to 60 everything. Perfect.
You guys are right on top of that
because we're leaving the starting blank
meaning that we should start at the
front. We're leaving the end blank
meaning we should go all the way to the
end. So that would be everything
everything in between. So at that point
we're not really slicing anything,
right? We're not really slicing much.
We're just taking the whole list.
Now, one example I want to give you
is
what if we sliced
and we had a step size
of minus one.
Any ideas what that would do?
What does this do?
What is a step size of minus one?
So minus negative index goes from the
back, right?
So if we're step sizing minus one, what
should we be doing effectively?
So, so this is signaling we basically
want to have the whole list but step by
minus one.
Yeah. So this this would be the reverse.
So this would be the reverse list. So, I
know that seems wacky, but that actually
is a way to validly reverse a list in
Python is to index to step size by minus
one. Because what that means is you're
slicing the entire array, but you're
stepping by minus one. We know minus one
um we know minus one goes
backwards, right? Effectively, because
it starts from the end. So, step size by
minus one would be go back this way.
each element going back this way. So
that would that would effectively
reverse the list.
Yeah, there now there is a reverse
function. A list has a reverse function.
So that is in English. But this is like
this is an alternative to to reversing
minus one
step size of minus one.
Uh, no, Roberto, they're not quite the
same because remember when you do two
colon and then you leave out the blank,
that means you're you're going all the
way to the end, including the blank,
including the end.
So, -4 to minus one would be the same as
2 to
uh 2 to six or 2 to 7. Two to six.
Sorry. Two to six.
60 to 10. Yep. It would be it. So this
reverses it. Meaning this would be this
would return to 60 then 50 then 40. It's
the reverse of the list when you step
size by minus one.
I am sure.
How about slice from position one?
Slice from position one to second to
last and reverse it.
Uh what do you think that would be?
You're starting at one going to second
to last.
And then rever like we already know
what's reversing is step size minus one
that will always reverse.
What is colon colon?
Colon is is the fact that we're leaving
colon colon is not anything special.
It's the fact that we're leaving the
starting. It's just the syntax of
slicing, right? colon colon is because
we are we're leaving the starting and
stopping blank
which we can do. We're allowed to do in
slicing. So that would signal that we're
doing everything but we're going step
size minus one.
Yeah.
Okay.
Any
other any other questions?
How do we feel about slicing? Do you
feel okay with it? Are we going to we're
going to practice it more as we go
along? Uh because we're going to use
slicing quite a bit when we work with
data, but does the concept of slicing
make sense? Yeah, you need you need more
p We'll do more of it. We're going to do
slicing throughout the program.
Yes, we'll do multi-dimensional uh in
our next course.
Not right now, but in our in our data
science course, we'll do
multi-dimensional.
that help?
Yeah. Step can be a very Yeah, sure.
Sure. So, there's there's nothing that's
stopping you from, you know, there's
nothing that's stopping you from doing
like let's say x is two and then we do
um numbers
numbers and then we have uh two to to
six and then x
Yeah, that's fine. There's nothing that
would stop us from doing that.
I do you mean that I think that's I
think that's fine. There there would be
nothing wrong with that.
But you're right that it should be an
integer. If it's if it's like a float,
Python will complain. It needs to be an
integer step size and it needs to be it
needs to be uh
um in order to get any meaningful data.
We wouldn't want that step size to be
too big or like it, you know, if we pick
it to be like 20 and there's only five
elements, that's not going to make any
sense. The step size needs to be
reasonable.
All right. So, what I want to do is show
you some functions that lists have. So
probably one of the most useful
functions a list has is the ability to
append items to the list. Now this will
add items to the list and particularly
it will add it at the end. So this is
this is something we will use quite a
bit is the list.append
function.
So this will uh this will add this will
modify the list and add a new element at
the end. So append always appends to the
end. Um
and so this will uh allow us to um take
this this string cherry and now the when
we append it this list is now
permanently been changed to have cherry
at the very end. So there it is. It is
now at the back of the list and it is it
is now at the kind of end position when
we do append. So here you see the list
being really dynamic allowing us to add
elements to it through this append
function.
So append really really useful allows us
to to add we just pass in we pass in an
element inside of the append uh function
here that we want to add to the list and
it will it will go to the back of the
list.
Um, lists also have a pop function
um, which you pass in a position and it
will remove that item that's at that
position. And not only will it
permanently remove the item that's at
that position, but it will return it
back to you. So, pop is really useful if
you want to remove things um from the
from the list. um if you don't provide a
position. So if you don't provide any
index that you want to remove from and
you just do if you if you just do um pop
without any uh thing in there that will
always remove the last element by
default. So always just so if we just
did this it would remove the 40.
Can we append in a specific position?
Um, yes. You would use the insert
function and then give it the index you
want to insert into. So,
would would uh be every list has ainsert
function to to and then you put in a
position you want to add it to.
Is it common use for append and remove
during? Yes. So append is really common
to add new things to the list which
maybe we're doing like data aggregation.
We want to add things to a list and then
take the average of the list. That's
very common. Um remove. Yes. Maybe we're
working our way through a collection of
things and when we process it we want to
remove it. So we can do pop to remove it
from the list.
Yes.
Uh yes. So when so when we pop it
permanently affects the list. So um
everything gets shifted. Yes. All their
positions get shifted according to what
we removed. So like in this example um
30 is uh 30 is index um two but when we
pop it now becomes uh the last element.
So it would now be eligible to be index
minus one, right? Cuz when we remove 40,
30 is now the end of the list. Um,
for example,
so yeah, everything shifts
and you know this this example here um
pops from index two. So we would go to
index two, which is 30, and remove that.
And so 40 now shifts up to be at index 2
whereas previously it was at index 3.
Insert as well. Yep. When you insert
everything shifts. Yep.
Okay. Here's a here's a really useful
function as well. So we have the extend
function
um which would allow us to add in
multiple elements. So this is the same
as if we appended every individual item
in this collection to the list. So
extend takes a list and adds its
elements to the other list. So notice
that we have um
we have a list of colors here, red and
blue, and we're extending it with a list
of green and yellow, which will result
in the colors list now having all of
those elements. So extend is really
helpful if we want to add in multiple
pieces of data to an existing list.
a lot of questions. Um, can we get the
index and values with a print command?
Uh, yeah.
Yeah. I mean, you can use a I'm not sure
what example you have in mind, but yes.
Yes. Append is one item. Extend is is
taking an entire list and and adding all
of those elements to the existing list.
Yes. Append is for only one value at a
time. Yes. Extend is when you're adding
multiple values.
Append is one value at a time. Yes.
Uh okay. So let me ask you guys what do
you think of this? Um,
which of the following method adds a
single element at the end of the list?
So, adding a single element at the end
of the list.
Very good. It should be a Yeah, we
append. Append adds and append always
adds to the end.
Very good.
Okay. So, now we're going to have a
demo. Um, and by the way, this is um
this is a demo that uh is an existing
notebook. So, um, what you would want to
do is, especially if you're working in
collab, is take the notebook. Now, this
is within lesson two. So, we're going to
do demo one and lesson two. You would
want to take that notebook if you're
working in Collab and upload it. I'll
show you how to do that, but we're going
to do we're going to do the demo that's
inside of uh the first demo inside of
lesson two.
Um,
so let me share my screen.
Okay. So, if you're inside of Collab,
what you're going to want to do is go to
file and then upload notebook. So,
you're going to want to go to upload
notebook and then um pick the hopefully
you've downloaded the demos in which
case you have the notebook from from
lesson two. There's a bunch of notebook
files, the IP YMBs. You want to upload
um those demo those demo notebooks.
Okay. If you're working in collab, if
you're working in Jupiter, um, or you're
working in, uh, VS Code, you can just
open that file, uh, within VS Code or
Jupiter, um,
and, uh, you should be good to go from
there. So, I've I've already uh, done
that. This is this is demo one inside of
lesson two. Do you guys have access to
that notebook?
Demo one and lesson two.
There should be lesson two has a bunch
of notebooks that we're going to work
through. Um,
okay.
Very good.
So, if we run if we run this piece of
code um that's in this first set or
sorry first cell, it's going to um
create this list which has different mix
types. So this list has integers, it has
strings, it has floats, but we can
create this list. If we just hit run,
um, we now have a list. And what I want
to show you is if we were to check the
type of this my list, um, of course,
this should be a list, which it is.
Okay. Uh, thank you for uploading that.
Perfect.
Okay. So let's go ahead and access a few
elements. So we can access the first
element here. We can access the element
the fourth which would be at position
three and then the seventh which would
be at position six. We can access all of
those
and we put those into a new list here by
putting them inside of the brackets. So
that that means that we're accessing
this first one. That's the first element
of this list that we're creating. It's
25, which is here.
By the way, what do you guys think
happens if we try to if we try to use an
index that is too big for this list?
What do you think would happen? Like if
we if we tried to do if we tried to use
code that would be like um my list and
then we put in the index like 20. What
do you think would happen?
because there's definitely not 20 items
in this list.
Yeah, it'll be an error. So, let's try
running that. This will give me an error
that says it's out of range. Yeah, an
exception, right? It would be an
exception, which would say, uh, we have
an index error. Um, we're trying to use
an index that's too big for our list
essentially.
So, just pointing that out. Um, let me
make a comment there.
Um, this
uh index is out of range for our list.
So, we should get an error.
See how I'm making a comment? Making a
comment there to remind myself of why I
got this error. So, remember, comments
are useful.
All right. So, we access things and we
can uh put those inside of a list. Now,
let's do negative index. So, we know
negative -1 should give us the item
that's at the very end of the list,
which would be this 2.718.
Um, so that should be there. And then
minus 4 would be fourth from the back.
Minus 7 would be seventh from the back.
So, we can uh get those values. Not too
bad.
Then we have a slicing example.
So 2 to 7 we know as a slice. This
should um this is a slice that uh slice
that starts at index 2 and goes to index
7 but doesn't include index 7.
Right? So that should be the slice. Um,
so if we run this code, um, this would
extract everything starting at position
two, which should be the third item of
the list, all the way up to, uh,
position 7.
And one other thing I wanted to show you
is we can extract how many elements are
in the list.
So I wanted to show you guys that this
code tells us the length of the list
which would be if we did length of my
list.
Yeah. Len. So we pass that in the the
length function. We pass in my list um
which should give us 10. So there's 10
items in this list.
So just wanted to call out that there is
this length function that we can do with
the list.
Can you show printing index numbers for
the list?
Like do you mean an uh every number in
its index?
Do you mean that every number and its
index? Yeah. So the code that does that
is the enumerate function and we we
would use a loop. Um so it would be
something like for index
um value in enumerate
uh my list and then we could do um print
uh index
and value
like that
you know we haven't learned this We
haven't learned any of this yet, but
that's that's what it's doing.
Yeah.
Okay,
cool. Uh let's see. So um finally what I
wanted to show is that we can append. So
if we uh take our list this is what it
currently is
and then we append a new element we can
print out the list and you can see how
it ends up at the end. So we take that
original list and we just add a new
element at the end. Um and then this by
the way we didn't we didn't uh explain
this but remove will find that value
find the value 100 and remove it from
the list. It will find the first
occurrence.
Yeah. So remove will find the first
occurrence of this value. Pop is index
based. Remove is value based. So remove
will look for the 100 and and take that
out of the list. But pop will um be
index based. So if we do um
if we do so pop is index based. So we
could uh do my list.pop
pop and we could pass in a zero which
should remove um this will remove the
first element
and then we can uh print my list.
So that removed the 25.
Does remove all? No, I think it's just
the first occurrence.
This is the first occurrence. You'd have
to do it multiple times if you have
duplicates.
I think I have to double check that, but
I think it's just the first occurrence.
Okay. All right. So, I know we're a
minute over.
Uh, thank you guys so much. What a great
first couple of sessions. Um, I think
we're picking up this really well. So
very good job. A lot of great questions,
a lot of good um back and forth. So I
appreciate that. Hope you guys are
learning and picking up this Python as
we go along. Um we have a lot more to
cover. So uh you know next time we meet
um you know we will uh continue talking
about the other data structures. So we
have to talk about sets, tupils,
dictionaries and then we have to get
into uh loops and if else statements and
then we'll eventually work our way to
functions. So, a lot more to cover, but
we'll get there. Um, and but hopefully
you guys are learning a lot. Any
homework to do? Not formally, but I
would request that you guys work on the
guided practice for lesson one. Work on
the guided practice for lesson one if
you can.
Okay? So, go into your reference
materials, find the guided practices,
work on the lesson one guided practice
between now and our next session.
Thank you guys. Thank you so much. Uh
thank you for I know these 4hour
sessions are a lot. Appreciate your
patience. Thank you so much.
Have a great rest of your week.
>> So you might be wondering what's changed
in machine learning and why is it the
best time to get into it now. Well,
let's go back a few years. In the past,
machine learning was more about building
models based on historical data. It was
about training algorithms to predict
specific outcomes like classifying
emails, spam or not spam, predicting
house prices based on past data. But
fast forward to 2026 and the landscape
has changed dramatically. Today, machine
learning isn't just about making
predictions. It's about building systems
that can learn, adapt, and improve over
time. We're no longer creating
algorithms to just run experiments
offline. Now, machine learning systems
are integrated into real world
operations and are capable of making
decisions that impact the business
immediately. Here's an example to make
it clearer. In the past, an e-commerce
website might use machine learning to
predict what products a customer might
want based on their past purchases. Now,
the systems can constantly learn from
new customer data, continuously refining
those predictions in real time as
customers preferences are changing. The
world of machine learning has evolved
from theory to practice and this has
created a huge demand for machine
learning engineers who can build
scalable systems and make them work in
real world environments. The impact of
machine learning is now directly tied to
business outcomes and machine learning
engineers are at the center of that
transformation. You might be thinking
okay I get it machine learning is
impactful but what exactly does a
machine learning engineer do compared to
other roles in tech? That's a great
question. In the world of machine
learning, you'll hear about a few key
roles such as data scientist, machine
learning, and AI engineer. Let's break
them down so you know exactly where you
fit in. Data scientists are like the
detectives of data. They spend their
time analyzing large data sets, finding
trends, and trying to extract meaningful
insights. They build models, but their
main focus is usually on data
exploration, and experimenting with
various algorithms. They don't typically
focus on deploying those models into
production environments. Machine
learning engineers on the other hand
these are architects. They take the
models built by data scientists and
build scalable deployable systems. They
work on creating solutions that will not
only work in the short term but can also
scale to handle real world data in
massive volumes. The machine learning
engineer is responsible for ensuring
that machine learning systems are
integrated into businesses that can work
seamlessly with existing technologies.
AI engineers focus more on the
application side of things. They build
AI powered products like chatbots, voice
assistants, and real-time systems. While
their work often overlaps with ML
engineers, they are typically more
focused on the userfacing product and
how machine learning fits into it. As an
ML engineer, your primary focus is to
take models and turn them into
actionable solutions that are deployed
in real world systems. We shall now move
on to why 2026 is the right time to
enter machine learning. Now that you
know the role of an ML engineer, let's
talk about why 2026 is the perfect time
for you to jump into the field. You've
probably heard that machine learning is
a hot topic, but what does that mean for
you as someone starting out in this
field? So, I'll help you break that down
for you. First, the demand for ML
engineers has skyrocketed. The world is
full of problems that needs solving, and
machine learning has proven to be one of
the most effective tools to solve them.
From predicting customer preferences to
automating critical business functions,
machine learning is changing how
businesses operate. Secondly, the tools
used to build machine learning systems
are more accessible than ever before. In
the past, machine learning was viewed as
something experimental, something that
required a lot of effort just to set up.
But today machine learning platforms and
frameworks such as TensorFlow, PyTorch
and Scikitlearn have matured
significantly. These tools make it
easier to build and deploy models that
can scale to handle real world data. We
shall now move on to why is this the
right time for you to get started out as
a machine learning engineer. So how do
you get started? The first step is to
build a strong foundation. You might be
excited to start building models and
diving into algorithms. But before that
you need to understand the core concepts
that drive all the machine learning
systems. These include mathematics,
programming and data handling. You don't
need to be an expert in all of these
areas, but you do need to understand the
basics. Think of these as building
blocks of everything that you will need
to learn machine learning. Let's start
with mathematics. You don't need to be a
math genius, but you do need to
understand the basics. There are three
main areas of math that will help you
get started out as an ML engineer, and
those are linear algebra. This is the
study of vectors and matrices which are
used to manipulate and process data in
machine learning models. If you have
heard terms like feature vectors or
matrix operations, that's linear algebra
play. This is the core of how machine
learning algorithms operate. This is the
math behind optimizing models. You'll
use calculus to adjust the parameters of
machine learning models and minimize the
error between predicted and the actual
values. Specifically, derivatives are
used to find the best way to fit a model
into the data. Statistics. Machine
learning is all about working with
uncertaintity, and statistics will help
you make sense of it. Whether you're
dealing with probability distributions
or hypothesis testing, statistics can
help you understand patterns in data and
make decisions with uncertain
information. These mathematical concepts
will help you build a more accurate
model and optimize it effectively. Let's
move on to programming stack. If you're
new to programming, don't worry. Python
is the language that you will want to
learn. It's simple to get started with
and has a huge ecosystem of libraries
especially designed for machine
learning. The key libraries that you
will need to master include NumPy. This
library is used to handle large arrays
of matrices and data which is the
backbone of most machine learning
algorithms. If you plan on working with
large data sets, you'll be using NumPy a
lot. Pandas. This library is great for
data manipulation and analysis. You'll
use pandas to clean, organize, and
transform data, making it ready for
machine learning models. Scikitlearn.
This library provides simple, easy to
use tools for building machine learning
models. It covers everything from data
prep-processing models like regression
and classification. Once you're
comfortable with these tools, you'll be
able to start building machine learning
models and working with real world data.
Next up is SQL, which stands for
structured query language. As an ML
engineer, you will be working on lots of
data, and SQL will help you query
databases to retrieve information that
you need. You'll use SQL to extract
data, filter it, and join tables
together, making sure that you have the
right data to your models. Once you have
your data, the next step is data
wrangling. Data wrangling is a huge part
of your job, and if you master it, it
will save you a lot of time and
frustration while building models. We
shall now move on to types of machine
learning. Let's talk about the different
types of machine learning. There are
three main categories which are
supervised learning, unsupervised
learning and reinforcement learning.
Speaking of supervised learning, this is
when you have label data. You train your
model on data where the answers are
already known. The goal is for the model
to learn the relationship between inputs
and outputs so it can predict the future
outcomes. Unsupervised learning. In this
case, the model works with unlabelled
data and it tries to find hidden
patterns and groupings in the data. This
is useful for tasks like clustering or
anomaly detection. Reinforcement
learning. This type of learning involves
training an agent to make decisions by
interacting with its environment and
receiving rewards or penalties based on
its actions. This is often used in game
AI or robotics. We shall now move on to
feature engineering and evaluation.
After cleaning your data, the next step
is feature engineering. This is the
process of transforming raw data into
meaningful features that help your model
make better predictions. Once you've
engineered your features, it's time to
evaluate the model. You'll use metrics
like accuracy, precision, and recall to
see how well your model is performing.
Proper evaluation ensures that your
model is ready for real world
applications and that it can handle data
in production. We shall now move on to
the six-month learning plan in order to
become an ML engineer in 2026. So, how
do you actually get started? Here's your
six-month learning plan to guide your
journey. Firstly, focus on learning
Python, mathematics, and SQL. Dive into
machine learning algorithms and hands-on
projects. Then, learn deep learning and
work with frameworks like TensorFlow or
PyTorch. By the sixth month, you can
work on a capstone project that covers
the full pipeline from data collection
to deployment. We shall now speak about
the ML engineer tool stack. So let's
talk about the tools that you will need
as a machine learning engineer. You must
be wondering what tools should I be
learning to become a successful ML
engineer in 2026. So let's simplify
this. Git and GitHub are the first
things on your list. At first glance,
version control might seem like
something that only coders need to worry
about. However, it can be really
crucial. Why? Because version control
allows you to track every change that
you make to your code, collaborate with
teams, and roll back changes if
something goes wrong. Imagine you're
working on a huge project and you mess
something up. Without version control,
you could lose hours of work. But with
Git and GitHub, you could go back and
fix to any point and time. You might be
thinking, okay, I get that Git is
useful, but what about the tools that
can actually help me build and deploy
machine learning models? Now, that's a
great question. So let's talk about the
cloud platforms like AWS, Google Cloud,
and Azure. In the past, machine learning
models were often built and tested upon
local machines, but we quickly realized
that it wasn't scalable. These cloud
platforms give you the ability to handle
large data sets, run models on
highowered servers, and scale them as
your projects grow. Instead of relying
on your personal laptop to do all of the
heavy lifting, you can leverage the
cloud to train models much faster and
handle massive data sets without
worrying about memory limitations. We
shall now move on to the next topic
which is on MLOps life cycle. You have
now got all of your tools in place. But
how do you go from all of these tools
coming together in the MLOps life cycle.
So you must be wondering what does all
of the MLOps life cycle tools even look
like and how do I manage the process
from starting to finish? Well, here's
the thing. MLOps is like a welloiled
machine. It involves a series of stages
that ensure that your models remain
reliable, scalable, and adaptable to new
data over time. Think of it like a car
assembly line. First, you train the car
and then you deploy it into the real
world. After that, you need to monitor
it and see if it runs smoothly. If it
breaks down, you will have to retrain
it. Let's break it down into key steps.
Training your model. This is where you
create the model using your training
data. This is a very experimental phase
where you try different algorithms,
tweak parameters, and optimize the model
for better performance. Deploying it to
production. Now that your model is
ready, it's time to put it into action.
This means making the model available to
users or clients. Whether it's in a
mobile app or a web service, deployment
is where the magic happens. Monitoring
performance. Once your model is live,
you can't just forget about it. It's
like checking the car tire pressure
after it's been on the road for a while.
You need to continuously track how well
your model is doing. If it starts to
slip or underperform, it's time for
tweaks and adjustments. Lastly, we have
retraining. Over time, your model may
need to be retrained with new data to
stay accurate. This is especially true
in industries where the environment is
constantly changing like e-commerce or
finance. Restraining ensures that the
model stays relevant and continues to
provide value. You may be wondering that
this sounds like a lot of work and
that's true. That's where MLOps tools
like MLS help you streamline this
process. By automating and managing this
life cycle, you can spend less time
dealing with the back end and more time
focusing on building innovative
solutions. We shall now move on to
experiment tracking. As you dive deeper
into machine learning, you'll quickly
realize how important it is to track
your experiments. At first, this might
seem a little overwhelming. You might be
thinking, I can't just run a model and
hope for the best, right? But trust me,
tracking your experiments is one of the
best habits that you can develop early
on. Think of it like logging your
workout progress. Even if you don't
track the results, how will you know
whether you're improving or not? The
same goes for machine learning. By using
experiment tracking tools flow and
weights and biases, you can keep a
detailed record of every model and every
hyperparameter and every evaluation
metric. Now imagine you're building a
recommendation system for an online
store. You try different adjustments,
algorithms, and get different results.
With experiment tracking, you can easily
compare which configurations work best.
You can easily compare which
configurations work best and learn from
the past mistakes. You'll always know
which experiment gives you the best
results and which needs tweaking.
Tracking experiments also helps in
collaboration. If you're working with
the team, being able to see everyone's
experiments in one place makes it easier
to understand their approach and build
on each other's work. We shall now move
on to projects and portfolio. Now that
you've got your tools and workflow in
place, it's time to focus on building a
strong portfolio. You might be thinking,
how do I make my portfolio stand out to
potential employees? That's a great
question. The answer is simple. Real
world projects. Think about it.
Employers want to see what you can do in
practice. They don't just want to see
theoretical knowledge. They want to see
that you can solve problems and build
working systems that can make a
difference. So, for that reason, we have
mini project examples. At this point,
you're probably itching to start
building something yourself. Well,
hands-on projects are the best way to
solidify what you've already learned.
For example, you could create a customer
churn prediction model or a fraud
detection system. These are practical
real world projects that demonstrate
your ability to build solutions from
start to finish. Make sure you document
your projects clearly on GitHub and
always include a detailed explanation of
your approach, challenges, and results.
Remember, the goal is to show that you
understand the problem and build a
working solution. We'll now move on to
Kaggle and open source. As you continue
to build your portfolio, I highly
recommend diving into Kaggle
competitions and contributing to open
source projects. Kaggle is an amazing
platform where you can work on real
world data sets and solve problems that
companies and research institutions are
facing. Not only will you improve your
skills, but you'll also have the
opportunity to see how top data
scientists and machine learning
engineers approach similar problems.
Contributing to open-source projects is
another excellent way to showcase your
skills. It shows that you can work well
with others understanding existing
systems and contribute to the community.
Plus, it's a great way to gain
visibility and make connections with
other engineers. We shall now speak
about the rรฉsumรฉs which work for you.
Speaking about your resume, when it
comes to landing a job as a machine
learning engineer, your resume needs to
be focused on real world projects and
practical experience. So, be sure to
highlight the projects that you've
worked on, the tools you've used, and
most importantly, the impact that your
work has had. If you build a
recommendation engine that boosted
product sales by 20%, make sure that you
include that. Employers want to see how
your work contributes to solving real
business problems. Don't forget to
include links to Kaggle profile or any
other open source contributions that you
have made. Employers love seeing code
and showing them your projects is the
best way to stand out. We shall now
speak about what interviewers test in
2026. By now you must be wondering what
do employers actually look for in an ML
engineer. So in 2026 interviews are not
just about technical knowledge.
Employers just want to see how well you
can communicate your thought process and
real world problems. You'll likely face
practical tests that challenge you to
build a model or analyze data in real
time. You'll also be tested on how well
you explain your approach and justify
the decisions that you made during the
project. Prepare to discuss things like
why you choose a specific algorithm for
a problem or how you handle issues like
data imbalance or overfitting along with
what metrics you use to evaluate your
model's performance. Being able to
communicate clearly your process and how
you arrived at your solution is a skill
that will set you apart from other
candidates. We shall now speak about the
ML trends that you will need to follow.
Machine learning is evolving fast and
staying up to date is key to remaining
relevant in this field. So here are a
few trends to keep an eye on in 2026.
Generative models. These models can
generate new data based on patterns they
learn and they're used for things like
text generation, image creation, and
even music composition. AutoML automated
machine learning tools are making it
easier for non-experts to build machine
learning models. As a result, more
people will be able to contribute to
this field without needing to become
experts. Privacy first models. As
privacy concerns grow, machine learning
models are being designed to work
securely and ethically and these don't
compromise on user privacy. Staying on
top of these trends will help you remain
competitive and innovative in the ever
evolving field of machine learning. So
in conclusion, I will say consistency is
key in machine learning. You don't need
to know everything right away, but you
must stay committed to learning and
building. Start small, keep
experimenting, and keep improving. The
road to become a machine learning
engineer may seem long, but with the
right tools and mindset, you can get
there. Thank you for watching and I'll
see you in the next one. Thanks for
watching the machine learning engineer
road map for 2026. We hope that this
video gave you a clear path to becoming
a successful ML engineer. Ready to take
the next step? Explore the Simply Lance
professional certificate in AI and
machine learning in partnership with
Purdue University. Gain hands-on
experience and industry recognized
credentials to boost your career.
>> Machine learning is which is a subset of
artificial intelligence, right? That's
uh basically um machines learning from
data
in order to uh make decisions
essentially. Um, so this was a big
departure from the rules-based systems
at the time, right, that were explicitly
programmed to make decisions. So just
think of an example like a really big
kind of if this, then that, then that,
then that, and and else if this, this,
this, right? So bunch of rules that had
to be pre-programmed in order to um come
out with some final answer. Uh, with
machine learning, it's the exact
opposite of that. we're actually
training something from examples from
existing data um in order to predict
something or um
make some type of decision. Uh and so
we're going to learn about the various
ways we can do machine learning. But if
you guys remember we
um talked about some of this like the
differences and the uh basically
rules-based approaches to learning from
data approach. Um and in included in
that is going to be uh complex
unstructured data. So things like
images, text, audio. What handles those
really well is uh deep learning which we
will get to in the course after this.
But uh those are certainly in there as
learning from data even complex data.
So we had this picture uh and I think
this is kind of around where we left off
last time was uh just distinguishing
between those three terms. We see
artificial intelligence, deep learning
and machine learning kind of used
interchangeably, but this is really how
they fit in. Artificial intelligence is
kind of a broad anything mimicking human
intelligence. Um which doesn't have to
be learning from data, but uh machine
learning is part of that. And then um
one way to accomplish machine learning
is to use neural nets which is the focus
of uh deep learning. Um and so deep
learning has been has found a lot of
success especially recently with uh
those complex data types like images,
speech, text, right? So deep learning
used all over the place. Even in um
modern like generative AI, we see deep
learning used quite a bit. Um it really
anything that's using neural nets is uh
going to be deep learning.
Um
and again we'll focus on that later but
we're going to be mainly focused on
machine learning for this course
primarily machine learning that does not
use neural networks. Okay so just models
that are not necessarily neural networks
be our focus.
So in machine learning we had an example
of a game uh essentially um learning
what decisions to make uh based on the
uh kind of current um state of the
board. This could be a um you know
machine learning example that uh learns
from many previous examples. So a lot of
data around these games are used to
train these um kind of robots that can
play these games and play them at a very
high level. Um so there's been a lot of
successes actually in machine learning
and deep learning um around
uh playing games like chess or go
um using machine learning algorithms. So
pretty cool.
All right. So I think this is where we
ended. We last time we said there's a
bunch of different use cases for machine
learning. So um recommendation system is
going to be a big one and we will
actually study that uh in one of our
final lessons of this course. Um chat
bots like generative AI doing sentiment
analysis chat bots we'll study later but
those are certainly an application of
learning from data in order to uh
generate responses to text prompts
right. Um spam filtering that's a good
example like classifying an email as
spam or not spam. Um that that gets
trained from examples and uh learning
from data such as previous emails. Um
social media posts analysis is another
kind of text data um use case but you uh
can do a lot with that text like you can
predict the sentiment um you can predict
uh the category of what what the post is
talking about um those kind of things
all can be done with machine learning
>> and many other use cases not on this
list that we will uh cover
>> you know as we as as we go further.
Okay, so this is where we kind of left
off. Um, so what's doing all the hard
work here is
>> uh machine learning algorithms. So these
are things that will um these are things
that will learn from the data. So they
are uh they they are basically um
algorithms or sets of rules that uh or
mathematical rules I should say not
formal rules like in the in the sense of
a rule system but mathematical um
formulas and mathematical uh rules
essentially that help us learn from the
data. So they correlate the data to some
type of outcome. So some type of
prediction uh whether that's going to be
as we will see whether that could be
like a number like we're predicting a
price or demand or sales
um or it could be a category like is
this transaction fraud or not fraud or
what's the probability that this is
fraud um so we have different kinds of
predictions we can make with machine
learning.
Um
but uh we will study the kind of the
differences of those coming up. Um but
machine learning algorithms are really
what power they're kind of the models,
right? They're the models that help
power uh machine learning to actually
learn from data.
So we're going to spend a lot of time in
this course studying those algorithms
like the different models that we can
build and what their differences are,
what their strengths are, what their
weaknesses are. we'll we'll learn a lot
about those.
Okay. So, I guess you can imagine like
everything is so data dependent, right?
Um we're learning from data. So, uh it
makes sense that the quality of data
really really matters here in
determining how strong the model can be.
Um so you see this graph here charting
kind of the um high quality data um
versus just uh any old data but a decent
enough quantity of it. Um you can see
that performance and the performance is
measured by some evaluation metric. Um,
so think of it as uh something like an
accuracy. Like if we were predicting
fraud or not fraud, how accurate can our
model get at actually detecting fraud
um it gets better and better and better
the graph shows that the higher quality
of data that we have. So there's kind of
that there's a there's a saying in
machine learning um called garbage in
garbage out. What that means is if you
have poor data, even the best model in
the world, poor data is not going to
result in having a good model that can
be accurate and perform well. Um, so it
needs to be high quality, meaning um
there needs to be a decent amount of it
and it needs to be labeled appropriately
as we will will talk about
um and it needs to not have any, you
know, significant outliers. it needs to
be clean, not have those missing values,
all of those things. Um, you can you
have a good chance at deriving good
predictions from higher quality data
as this kind of shows.
Okay.
So, one thing we're going to learn um as
we go along is
quantity matters as well. So, not only
quality, but a decent amount of it. And
um we're going to learn those kind of
rules of thumb like how much data do I
need for certain algorithms. Um one
thing that we will see is that uh the
the basic machine learning models that
we'll study don't need as much as a
neural network would. It you know neural
networks are going to require a lot more
um than a basic machine learning model
learning model. So uh that's something
we will see as we go along. But uh this
is something we'll talk about and
discuss with each model that we study is
kind of how much data do we actually
need to produce a high quality model.
Okay, any questions uh so far?
Okay, let's talk about the different
types of machine learning that we're
going to discuss. Prim, there's going to
be two primary ones that we will study
in this course and then a couple others
that'll be a little bit more advanced
that we won't get to but worth knowing
about. Um, so there's going to be four
total that we'll study or talk about and
they'll be on this list here, which is
um supervised learning and unsupervised
learning. Now, I'd say the majority of
our focus will probably be on supervised
learning, and we'll talk about what that
means, but we'll also cover unsupervised
learning as well. And so, we'll look at
the most popular techniques in each of
these types of machine learning.
Um,
and then we'll talk about these two, but
not really study them because they're
more advanced topics. Um, that that will
be beyond the scope of what we'll do.
But uh these are going to be um
different styles of machine learning
that are going to be characterized by um
what kinds of predictions they make,
what kind of data they need and require.
Um and uh what kind of outcomes they're
actually producing. Um, so let's let's
get into each of these, but uh the the
one that we'll probably spend the
majority of our time on is going to be
supervised learning, but we will study
unsupervised learning as well. We'll
study both and we're going to talk about
we're going to define both of those um
coming up. And again, these will be a
little bit more advanced topics that we
won't spend too much time on.
Um, but but we'll discuss their
relevancy in machine learning. Um, and
give a good definition to it.
Okay.
All right. Let's start with supervised
learning. Now this is going to be uh a
term that really refers to
using examples. So using labeled
examples. So here we say labeled data to
help our model train. In other words,
help our model be able to predict guided
by specific input output pairs. So
supervised really refers to the fact
that we have answers. We have examples.
We have answers with those and we use
that collection of data to build our
model off of so that we can predict
um those kinds of things like a price,
like a category, like a spam, not spam.
in this in this slide like we would be
predicting if this shape is a square, a
triangle or a circle.
Um but but when we build a model for
that, we have data that has an answer
attached to it, right? We've talked
about this before a little bit with
labels. So there's a guide there that
can guide us towards building our model.
there's an actual every every example
has an answer and that answer is really
critical to help build our model off of.
So, um that's it's almost like you have
um a you have a bunch of exercises
in let's say like a math textbook. You
have a bunch of exercises and you have
the answers and that way you can kind of
check your work. you think about model
training um that is the really a lot of
that process of model training as we are
going to discover is um basically
checking our work against these answers
in our data in our training data.
Okay. So supervised learning is any type
of machine learning that involves
learning from labeled data in order to
predict outcomes. Okay. Predict outcomes
like now the the outcomes can be
numerical. They can be like a price,
temperature, demand, sales, revenue.
They can be numerical, but they can also
be categorical. So they can be like spam
not spam, fraud, not fraud, cancer, not
cancer. Um dog, cat, giraffe, those kind
of categories. Um we could predict
those. It's some type of outcome. Okay,
some type of outcome. The key is we're
using labeled examples to guide our
model building. That's why it's called
supervised learning.
So we know in our data we know what the
inputs are. Of course, those are going
to be think of the inputs as like all of
our columns and then we have a special
label column that represents the output
we're trying to predict. So if you think
about that housing price data, the label
could be the price. And that's something
we would build a model to predict, but
we have answers for all of our examples
in our rows. We have answers to help
guide our model building.
They help tweak our model because we
know the answer ahead of time. So
they're they're really good examples to
build our model off of.
Okay. So that's that's supervised
learning.
Uh in this example, is circle not in the
prediction because it's not part of the
test data even though it's in the
labeled data?
Um no, it just not necessarily. It just
means that like we learn against all of
these examples that have these answers
and then when we observe new examples um
we can try to predict what those would
be based on what we've seen before. So I
if there was a you know it's just a
coincidence we only have two two
examples in our test data like we could
have a circle here in which case we
would predict circle
that's fine or at least we would hope
our model would predict circle right
that's what we're hoping may or may not
get it right
um but it's it's only not there because
we only like we're just assuming that we
only have two examples we're testing
against but in reality we would probably
do a lot more than two.
It's just it's just a coincidence
really.
In reality, we would test against a lot
more data. And we're actually going to
see why we would do that. Like why would
we train our model and then kind of use
additional data um to to evaluate it?
It's actually really important that we
do that step to get a sense of how good
our model is before we take it out in
the real world. So if we apply our model
that we build on our label data to
um this kind of set of test data that we
haven't been exposed to before. It helps
give us a sense of how good is our
model. So it's test is usually used for
evaluation.
So that's something that's something
we'll study.
How do we train? Uh it depends on the
model. Um so training will be a sense uh
will be an algorithm that will um
basically update the model according to
the data. These labeled examples. Um
every model is going to be different in
exactly how it trains. So we're going to
we're going to talk about that when we
get to the individual models that we'll
study.
But uh loosely speaking, they're going
to use the data to adjust itself. Like
imagine adjust like tuning a bunch of
knobs. Um, like the best example I can
give you is we I think I did this one
last week where you have kind of a
function
that predicts the price and let's say it
has
um weights like weight one with feature
one, weight two with feature two,
weight three with feature three. So
imagine we had three input features and
we we built an answer according to that.
Essentially what we would do to train
the model is adjust these
um in order to get this correct based on
our our labeled examples.
Okay.
So that's something we're going to learn
about coming up shortly when we when we
actually dive into model. Every model is
going to be slightly different in how it
trains, but at a high level it's going
to use the training data with those
examples, right? the labeled examples to
help guide the formula essentially to
adjust to generate the proper kind of
model here. The these things are going
to be adjusted according to the data
in order to produce the correct output.
So think about these as knobs that will
turn.
Okay.
Uh which type of machine learning is
used? Uh probably supervised um which is
what we're talking about now. So
probably supervised because most people
want to
um build some type of model to predict
something. Uh so yeah, I'd say I'd say
supervise.
Yes, we're are we are definitely going
to learn how to train. Yeah, we'll see.
We'll do the code. Um, I'll tell you
about how it's done. Yeah, we're
definitely going to learn it. But what I
was saying is it's kind of on a model by
model basis.
So, I want to wait till we get into the
individual models, then we'll talk about
how they're trained.
But yeah, we'll we'll learn how to do
that.
But yeah, supervisor is used all over
the place. Even even for uh generative
models, they use supervised learning
because um like an LLM
is going to use labeled examples in
order to train, right? In order to train
how to generate responses according to
prompts. Um it needs to learn against a
lot of text examples.
So that supervised learning is what um
results in that model,
right? Learning from those labeled
examples.
Okay.
It is yeah image image uh a lot of um
yeah a lot of image processing is
supervised like object detection. So the
yolo model is an object detection model.
Yes. Um because it has to be trained
right it has to be trained on uh it has
to be trained on images
with labels such as this is what object
is in this image. This is the box around
the object.
Um yes. So if if it's if it ever uses
label data to train and build the model,
it is supervised. So YOLO is definitely
supervised and we actually we will we
will cover the YOLO model later on in in
our deep learning course. We talk about
object detection.
So we'll we'll study that.
But yeah, it's supervised
Okay. So on the slide we have some
common supervised learning algorithms
that are we will study. So all of these
we will study and understand what they
do and how they work but just giving you
some to name them. linear regression is
kind of the one I just drew out which is
the um this is the prototypical like
easiest to understand model that is kind
of the um exactly like this where we
have a weight times a feature um a
weight times a feature and then a weight
times a feature
and on and on and on. You can have as
many as you want.
um that is a linear regression. And so
that is um that's a supervised model
because we need this value here and we
need all of our inputs in order to um
actually train this model and generate
all those weights
um that that is uh that uses um labeled
examples to help tune all those knobs.
Um same with all these other models. So,
we're going to talk about decision
trees. We're going to talk about
logistic regression and and SVMs, which
are support vector machines. We'll talk
about all of those, but they're all
examples of supervised uh supervised
learning.
Okay, we'll talk about all of these.
They're all supervised because they all
require labeled examples in order to
train them and and then subsequently use
them. Okay.
Okay. So what are some use case
examples? So for for instance in uh
supervised learning we may be predicting
temperature based on yearly temperature
trends. So we would have that yearly
data as our um as our labeled examples
and those would supervise the learning
of a model that predicts temperature.
Um, same thing with predicting crop
yield based on um, seasonal crop quality
changes. So maybe we have a bunch of
features relating to crop quality. We
could predict crop yield. Um, we would
just need historical examples with those
labels, right? What the crop yield is
for each time period. Let's say we would
just need those uh, supervised examples
and we could easily build a model off of
it.
Um
uh this this last one sorting waste
based on known waste items and their
corresponding waste types. Um that's
kind of like spam. It's like filtering
basically like a spam filtering. Um so
think of it like the the shapes example.
We sorting things into squares, circles,
triangles. Um, same kind of idea here
where we have a bunch of examples on
what those um what those waste items
should uh should belong to, like what
waste bins they would go to, for
example. Um, and those could be labeled
and therefore then we could um
understand what category of waste they
belong to.
Um, same thing with spam. something is
fraud or not fraud, spam or not spam,
cancer or not cancer. All of those are
going to be supervised learning examples
because they're going to require in
order to train them, they're going to
require data that has those labels.
Okay? So, anything that has labels is
going to be supervised learning.
So, again, this is where we will spend
probably the the majority of our time is
doing supervised learning problems. ones
that we have labeled data. We're
building a model and we're going to
predict those those uh labels
essentially.
Okay, before we go to unsupervised, any
questions about uh supervised
Okay.
All right. So, supervised requires
labels
in order to have an example to go off of
to build your model. And that's because
you're predicting those kind of outcomes
like spam or not spam, cancer not
cancer. Now unsupervised learning is
completely different. It's the opposite.
So unsupervised learning is where we do
not use labels whatsoever. So we're not
using any labels at all. So it's it it
can be completely unlabeled or even if
it's labeled, we're not using labels in
any way. But um we primarily would say
it's unlabeled data. We have no guidance
because we're not using the labels in
any way. we have no guidance to um
predict anything but that's because
we're not really predicting anything in
unsupervised learning. Generally what
we're doing is looking for some
structure or pattern.
Okay, with unsupervised learning we're
looking for some structure or pattern.
So um one type of example that's very
very popular is going to be this second
one which is um identification
identification of user groups based on
similarities or commonalities. Now this
is going to be a problem basically known
as clustering
and it's a problem we will study quite a
bit. There's going to turn out to be
lots of different algorithms that can
accomplish clustering. So what
clustering attempts to do is basically
say um we have data that's like this and
then data over here and then data over
here. Let's just group these together.
So like this should be one group. This
should be one group and this should be
one group. And we can find those
structures and say okay this is group
one this is group two and this is group
three.
One 2 3. And we can basically build what
we would call clusters of data um based
on how close together the points are
kind of located in these kind of cluster
zones like these boxes I've drawn.
Okay. Now that doesn't require any label
to do which is really fascinating. So
unsupervised you don't need any label at
all to accomplish the algorithm. Um so
clustering is one good example. Um
finding outliers or anomalies is
another. So we don't necessarily have
any label of what is an outlier or what
is an anomaly. We are deriving that from
the features alone. There's no guidance.
There's no label um to doing like
outlier detection or anomaly detection.
Okay. So that's another good example.
One that's not listed on here um but is
also really important that we will study
is something known as dimensionality
reduction.
So dim reduction and what that what this
focuses on is basically compressing the
data set a bit. So we take our data and
basically compress it. Um so that but we
do it in such a way that we retain as
much information as we can. It's a very
like smart compression and what it does
is it lowers the dimension.
um dimension. Think of the dimension as
like number of columns.
Number of columns.
So imagine we had 100 columns in a data
frame. What we could do is actually
reduce that down to 10. So like 10% of
that. So we reduce it down to 10. And um
but those 10 are it's not like we
chopped out um 90 other columns. we um
smartly kind of compressed all that
information into these 10 new columns um
that are compressed versions of the
hundred that we used to have. Um so
dimensionality reduction is is another
unsupervised technique. It requires no
guidance, no label to do, but is um a
really useful technique to reduce the
size of your data if you're doing things
with it. Um so this is another one that
we will we'll study how to do it and
basically more details behind it what
the algorithms are.
Um we'll so probably those two in
unsupervised will spend the most amount
of time on clustering and dimensionality
reduction.
Uh and supervised if some data is
present but we didn't label it means in
example we had circle triangle square in
the training data we add pentagon
but we didn't label that in that case.
Uh yeah. So every um in supervised
learning, every row, think about it as
like every row in our data frame needs
to have a label
uh associated to it. It needs to have a
a column that represents the label.
So if we've never seen Pentagon before,
I can't use that as a label.
So it has to the pentagon has to exist
in the data. if I'm going to be able to
predict it,
right? So, it can't predict, right? If
we've never seen it before, we have no
examples to go off. We have no guidance.
So, how could we predict that?
Right? We can't predict it
if it's if it's in there. So if if we
have labels of Pentagon, let's say, then
yeah, we could predict Pentagon.
We could
remove. Remove what?
We wouldn't if it was talking about the
Pentagon, we wouldn't remove that. No,
let me go back to that page. We wouldn't
remove it. Um, it's just if it's not in
our labels, we're not going to be able
to predict it. So, Pentagon's a good
example here. Uh, Pentagon is not one of
our labels. So, it currently is not in
our data set as one of the labels. We
only have data that's either a triangle,
circle, or
square. We don't have pentagon. So, I
would never be able to predict pentagon.
I'll never be able to do that if I
haven't seen examples of it before.
Okay. But let's say we had that in
there.
So, we had Pentagon.
So if we had Pentagon, um we could have
an example of it in our labels
and then Yeah, we it could be then we
could predict it.
Yeah. Yeah. The the don't get worried
don't worry about the test data. So the
test data is just saying here's a new
here's a shape what is it okay that's a
square here's a shape what is it okay
that's a triangle and we could have as
many of those examples as we want in our
test data so we could have a circle and
say okay what's this should be circle
right the test data can be whatever it
whatever it wants but yeah if if we've
never seen pentagon before we're never
going to be able to predict
These are the the label data and labels
are basically the talking about the same
thing. The labels just mean what are the
categories
that are present in our data. So in this
data we only have three labels that are
present.
So the labels is are relative to our
label data, right? It's saying
what labels,
excuse me, what labels uh do we have
in our data and we only have those three
circle, triangle, square. So so Pentagon
would not be part of those labels. We
couldn't predict it.
No. So unsupervised is not going to make
a prediction. That's the big difference
with unsupervised. They're not going to
make a prediction like this. Um so
unsupervised is not going to make a
prediction. It's going to do something
different like um basically say like
these guys are similar, these are
similar, these are similar, this is a
cluster, this is a cluster, this is a
cluster. It's not going to make a
prediction. That's what supervised
learning does.
Clustering, yes, which is unsupervised.
Yes, clustering does not require any
labels. Unsupervised just means we don't
have any labels. We don't require any
labels.
So the other thing unsupervised might do
is it might say
and again without the labels it might
say that this is an outlier.
it might say that this guy is an outlier
because there's only there's only one of
those and they're not like the other. So
that that's something that um that's
something that uh unsupervised could do.
Um it it yeah and no. It kind of labels
a cluster in the sense that um it would
basically assign a number to it like
this is cluster one, this is cluster
two, this is cluster three.
It'll assign a number to it, but it's
not a very meaningful it doesn't assign
like a prediction label in in the
traditional sense of a label.
It does provide like a numerical index
for the cluster to because what we want
to know is like okay this guy has the
cluster of one. This guy belongs to
cluster one. This guy belongs to cluster
one. This guy belongs to cluster two.
This guy belongs to cluster two. Does
that make sense? So there needs to be
some like index of what cluster you
belong to.
So it's kind of like a label but not in
the traditional like prediction sense.
Okay.
Very good. So again, unsupervised, no
labels. You're doing things like
identifying clusters,
um identifying outliers, doing
dimensionality reduction. These are all
like structure and pattern oriented
things. They're not predictions of a
label. Okay? They're not which is what
we would see in supervised learning.
Okay. So an example would be that we
take we put in the data um we can group
together uh data such as images into
categories based on similarities um
which would be like those clusters. So
there's no these would be groups that we
don't have any label on ahead of time
like we don't have we don't say that
this image should belong to this this
image should belong to this we derive
that from the characteristics of the
data. Um so think like a good example is
um customer groups. So we would identify
customers based on like okay do they
have similar spending levels? How many
days do they go shopping in a week? How
much money do they spend? And we can
kind of group together customers based
on similar qualities.
Clustering will find those groups that
should exist.
um it will discover those groups based
on um the similarities in the data, but
there's no labels that that say like
this person should be in this group,
this person should be in this ahead of
time. There's no labels of that. It gets
derived during the algorithm. It's
unsupervised,
right? There's no unsupervised really
literally means no guidance. There's no
guidance to doing it. We just derive
that from the structure of the data
which is the similarities.
Okay.
All right. So,
a couple more for you. So we had um
supervised which uses the labels. We
have unsupervised which uses no labels
looking for structure. And then we have
something that's kind of in between
which is um what is known as
semiupervised learning. And this is
where you use a combination of a little
bit of label data, but most of your data
is actually unlabeled data. Um, and you
try to get some use out of that label
data in order to um build a model out of
it. And so, uh, it uses the, um, it uses
that label data to, um, generally
provide some guidance on usually what
happens with semi-supervised learning is
you use your label data to kind of
predict what the label should be for the
unlabelled data and then you can go from
there. So you can create artificial
labels on this unlabeled data and then
you can use all of it once it's all been
labeled kind of like a supervised
learning uh approach. So but but this is
semi-supervised basically refers to the
fact that you start out with most of
your data not being labeled but you do
have some labeled examples and what you
can do is basically extrapolate those
labels into the unlabeled data set and
then provide some artificial labels and
then now everything has a label you can
do supervised learning.
Okay. So, it falls kind of between um
supervised and and unsupervised.
Uh and there So, this is this is kind of
rare. Most of the time you're not going
to do that. You're actually just going
to um prefer to just start with all
label data. That's usually the preferred
approach. Most of the time you'll
actually just be doing supervised
learning, not really semi-supervised
learning. So, it's pretty rare, but um
it it could like if Yeah, it could if
the if we had a lot of examples of
Pentagon and we wanted and so they were
unlabeled and then we tried to guess
what kind of shape they were um and
provide an artificial label uh and then
um then use that whole data set to build
a model off of then then yeah, it could
it could fall into this category. Okay.
They Oh, going back to the question,
they still use some kind of label data
like age, gender. They use uh that's
those aren't those aren't really labels.
That's the features. So, yeah, they
still use the core features of the data.
They just don't have any like labels in
the traditional sense of a label. Like
you should think of a label as something
we are trying to predict.
So whether that's a price, whether
that's like a category like spam, not
spam, cancer, not cancer, it's something
we'd be interested in kind of
predicting. And so um in our data, we
would have an answer for every row. We'd
have one of our columns would be like
the the result like the outcome answer
that we're trying to predict. That's the
label.
So in unsupervised, we don't have any of
the labels.
We do have just the regular features
like gender, age, income, square
footage,
bedrooms, bathrooms, all those things.
Okay.
So, we have semi-supervised that falls
in between supervised. Now, the reason
it falls between is be is because
there's a decent amount of data that's
unlabeled. In fact, a majority of it
unlabeled. But what we can do is try to
label it. We can try to take what we
know from our existing labels and
predict an artificial label and then use
all that data together in kind of a
supervised fashion for a model down the
road.
So that's kind of what this picture uh
says is we can try to take um you know
maybe we try to infer some labels based
on we have some some labelled data here.
We have most of our data is unlabeled
and we try to supply some labels to it.
Um like maybe we have a baby's category
of teens, a tween, uh you know youth and
um adults. Um and then we try so we we
take our our labels and we try to
extrapolate those into artificial labels
for this unlabelled data so that we can
use it now because then everything has a
label at this point and then we can just
go ahead and do supervised learning from
there.
So we can do supervised from there. What
we would prefer to do and what we'll do
in this course
um is just start with supervised. We'll
just start with the labels. We won't try
to derive artificial labels usually.
We'll just start with labels.
So one example in the real world is
something like Google photos which um
whenever you take a picture it can
provide uh uh labels based on previous
uh images in your library. So it can it
can produce tags or um labels on those.
Uh generally when you take that picture
it's kind of unlabeled unless you go in
and specifically provide some tags and
some labels. But um if you don't do that
it can still it can still uh make it can
artificially create one of those based
on the other label data that you already
have.
So that's um
that's an example.
Okay.
All right. Last one in terms of machine
learning. So we have supervised, we have
unsupervised.
Uh then we had semi-supervised which is
somewhere in between a mixture of having
some unlabelled data and label data. Um
now we're going to talk about
reinforcement learning which is
completely different. Um it's it's
completely different than the other
three. It's a type of machine learning
where we uh basically learn from
interaction with the environment. And
you might ask what are we learning? We
are learning what actions to take in the
environment. Um and the way we do that
is by reinforcing
positive actions that lead to a a
reward. Um, so that's where the word
reinforcement comes from is we we
basically uh imagine like a child
that's, you know, learning from trial
and error. Like they're trying to crawl,
they're trying to walk and they keep
falling down. um eventually they learn
how to do it through trial and error and
they might get a reward
or they might um reinforce some of those
positive movements that lead them to
walk or crawl um or they might learn
from the penalties, right? They might
learn from uh some type of feedback. So
they might learn from falling down like,
"Oh, that hurts. I should uh support
myself a little bit better, right?" Or
be a little more coordinated. Um
and so they they learn from those
actions and their interaction with the
environment. Um
uh so this is a complex um algorithm
essentially uh it's it deals a lot with
um again taking actions. Usually when
you take an action something changes in
the environment um then you kind of
observe some type of feedback. So, think
about like a a board game where you're
trying to figure out what move you
should make. Or another good example is
like with a robot um trying to navigate
a maze. So, like what route should it
take? Should it move forward? Should it
move backward? Should it move left or
right? Those are different actions it
can take. Also, like a self-driving car,
should it should it turn? Should it
speed up? Should it slow down? Those are
all good examples of things that have
been trained from reinforcement
learning.
Uh yeah. So real world examples would be
like in a board game, uh a a reward
would be like if you win the game. Um or
if you like capture a piece like in
checkers or chess, that's a reward. A
penalty would be like if you lose the
game or lose one of your pieces, that
could be a a penalty.
um in a board game or sorry in like a a
robot navigation task, it could get
rewards for um moving in the right
direction
um towards the exit or like when it like
let's say you wanted to train a robot on
how to open the door and navigate a
room. Um you would penalize it for
bumping into the wall.
Um you would give it a reward for moving
usually oh like oh the algorithm
themselves usually it's like a a step
function um it's usually it's like a
discrete function that kind of is based
on the state so the reward it could be
like um like depending on the let's
let's go back to the board game example
like the reward could be like or even
the maze let's say like a navigating the
maze like getting to this let's say this
was the exit
and this was the entrance.
Then if they make it to here, they get a
numerical like if they make it to the
exit, they get a numerical reward of
like plus 100, let's say. So it's just a
number. And then if they uh like if they
bump if they go into here, like let's
say this is kind of like a death trap or
like a pit, this this would be like a
minus 100. So it could be like discrete
numerical values could be the reward. If
they're moving in the right direction
like let's say we want to encourage
going this way then we could give
smaller intermediate rewards like this
should be a plus like if you move
forward this is a plus five this is a
plus 10 this is a plus 15 if you're
moving in the wrong direction away from
the exit. Um that would be like a minus5
or a minus 10. Does that make sense? So
they're they're numerical in nature and
what you're trying to do is collect the
most reward. You're trying to get the
largest reward you can through trial and
error. So you you try this out many many
many times. You basically simulate
running through this maze many many
times. And what dictates it what
dictates like where I should go is based
on what I've observed in the past. It's
almost like you're a child remembering
like, okay, what move should I make from
this space? Like, if I'm here, if I'm
here, which way should I go? Should I go
down? Should I go right? Should I go
left? You kind of know that from
experience.
Does that make sense? Based on the
reward that I've seen in the past, like
when I've moved down, I've gotten a
higher reward than moving left or right.
Does that make sense? So, yeah, it's
it's a numerical value
as a reward.
Yeah, that's a great question. Um, how
does it differentiate rewards based on
gain and loss i.e. chess? So it's it's a
very comp complicated uh answer but
essentially every so in the chess board
you can think of the board as like every
every um
space is a state.
So I could be in this state I could be
in this state and then it's not not only
is every every uh space but where all
the other pieces are. So there's lots of
states that are possible.
Um, so
the way there's a way to quantify
essentially what's the value of taking a
certain action like moving my piece
left, moving it right, moving it up or
down um given the rest of the state. So
you're you're right, it may be
beneficial to sacrifice. Um, but we
would learn that through experience that
okay, the best move in this situation is
to sacrifice.
We would we would have to learn that
through trial and error many many many
times which is to say like okay if I'm
in this current state of the world right
all these pieces are distributed in this
way the best move for me right now in
the long run to get the most reward in
the long run is to actually sacrifice my
piece and move it right move it into
like a bad position theoretically but we
know from experience that's actually the
most long-term reward is from that
position
like moving it right may be the best
action for me. So what you learn is how
to take actions
and actions are usually like move right,
move left, move up, move down. You think
about like a self-driving car though,
that's going to be like slow down, speed
up, turn your wheel 10ยฐ,
um those kind of actions.
So the the short answer is it's there's
a calculation there that you learn what
the long-term value of every state is
every unique state
and then you're trying to basically say
what action should I take from that
state given that current state of the
world.
Okay.
And I really I really like reinforcement
learning. It's actually probably my
favorite field of machine learning.
Unfortunately, we won't be covering it
um in our main uh course. We have
offered uh electives around
reinforcement learning in the past. So,
um stay tuned. Maybe when we get to the
end of this program, uh we'll offer an
elective on it and if enough people sign
up for it, we'll we'll run it. But, um
we it's not part of our we don't really
cover reinforcement learning as part of
our main topics. It's it is an advanced
uh more advanced topic than than what
we'll cover. But, um I I really enjoy
it. Find it very fascinating.
Okay. So, all of this is kind of um
illustrating what I was saying, which is
um you think of like uh the thing that's
interacting in the environment like the
robot or the car or the human moving a
chest piece is known as the agent. It's
interacting with the environment by
taking actions which updates the state
um of of the environment. So that's
that's why you see this word state here.
This gets updated constantly every time
you take an action. Um ultimately what
reinforcement learning is trying to do
is learn the best action like what would
be the best action to take. Um
and the best action is is the one that
leads to the most long-term reward.
That's the best action. Um, so you have
to uh you have to learn what you know
what leads to a good reward by kind of
experiencing this over and over and over
through trial and error. So there's a
lot of um kind of simulation or letting
the robot try something a lot um in
order to kind of learn what's rewarding
and what's not. Think about it again
like I think a good example is like with
children, right? you kind of have to let
them try things until they learn on
their own what's what can they do and
what can they not do
what's the best actions right
so reinforcement learning has made its
way into other places so I I said like a
good example is self-driving cars or ro
robotics a lot of reinforce
reinforcement learning is used there one
place it's found its way into recently
is recommendation systems have kind of
merged with reinforcement learning
learning. Um, and this is because you
you can imagine there's kind of a
built-in reward for you clicking on a
video and kind of watching it.
Um, so that kind of reinforces that
recommendation and then uh that's where
um you can then kind of recommend a
similar thing and see if that's
rewarding and generates a click or
generates some view time or watch time
or whatever. Um so reinforcement
learning has found its way into a lot of
areas. Um recommendations being one of
them because it's just natural for the
idea of like what um should I recommend
next to generate the most reward. In
this case the reward is kind of
correlated to did they click on it or
not or did they how long did they watch
for longer it's more rewarding.
um those kind of things but uh place
places where reinforcement learning have
been used I said self-driving cars um
games so uh one of the most famous
examples if you want to look it up is
the um Alph Go this was in 2016 um the
Alph Go uh algorithm was a reinforcement
learning bot that beat um some of the
world's best Go players which go if
you're not familiar Go is a um board
game
that is a little bit more uh complex
than chess. It has more more uh it's a
larger board um more pieces to it. Um
but they there was a reinforcement
learning powered bot that actually um
learned how to play the game so
effective it could beat um world kind of
masters at the games was pretty amazing.
Um that's the alpha go and that was by
deep mind Google and deep mind in 2016.
That was pretty that was only in 10
years ago not that long.
Um so certain uh we said recommendation
uh even autocorrect um learning to
predict like what is the best correction
uh to generate a reward which would be
like you accept that correction or you
reject it would be a penalty. Um so
reinforced learning has been adapted to
these kind of problems very
successfully. Let's take a look at the
packages that we will use throughout. So
um of course we will rely on these three
which we've already relied on to do a
lot of things like numpy to do numerical
manipulations and calculations.
Uh mapplot lib to do any plotting and
not only mapp but maybe seabour as well.
both of those to do plotting. Um, pandas
is a big one because
that's where all of our data is going to
be manipulated and prepped before it
goes into modeling.
So, all of that stuff we learned from
pandis is definitely going to be applied
here in this course uh as we actually
build models. Um, so of course like
these old ones that we've been working
with quite a bit um still going to be
useful here in the modeling stage. Um,
mainly for different reasons though,
mostly to get our data prepared to do
some type of modeling or maybe to
visualize it before we do modeling to
get a sense of what it looks like, those
kind of things.
Um,
sci is sometimes useful for certain uh
um processing like in unsupervised
learning. We'll actually use scyp a
little bit to do dimensionality
reduction or help us do that. Um so
scypi will be used here and there and
we've seen it before with hypothesis
testing we use scypi like the test and z
test came from there. Um some of the
unsupervised learning stuff will come
out of there but the package we will use
by far the most in this course is going
to be scikitlearn
which is here. Um and we've already seen
a little bit about scikitlearn in terms
of its pre-processing capability. So we
use the uh minmax scaler and the
standard scaler from there from the
pre-processing module in scikitlearn.
But it has um many different models
built into it that we can use to help uh
do our training and predictions. Um so
it's a incredibly useful machine
learning library. It is the industry
standard machine learning library. Um if
you're going to do anything in machine
learning, it would be expected that you
know how to use scikitlearn.
Now what's really lucky about that is
that scikitlearn is a really easy
package to get used to. Nearly
everything we do in scikitlearn will
mostly follow the same pattern and so um
the code will be extremely simple. They
did a great job with that package of
making things really user friendly,
really simple. Um, it's a really
fantastic package and we're going to get
a lot of practice with it uh as we go
along. Every model we build will
essentially be from scikitlearn
and not only like the models but um
doing the training, doing the
predictions and then doing the
evaluation will all come from different
uh scikitlearn u modules. So that'll be
really nice and we'll get um good
exposure to that package throughout the
course. So if anything will come away
from this course as um scikitlearn uh uh
experts that'll be very nice. So this is
this will be the new one for us learn
but we'll get a lot of practice with it.
Okay.
All right. So just to recap that lesson
before we move on to lesson three. Um we
talked about machine learning as
learning from data um which is included
underneath the AI umbrella but deep
learning is also included under machine
learning because it's still learning
from data but it's learning using neural
networks.
Um we talked about the four different
types of machine learning. We had
supervised, unsupervised,
semi-supervised and reinforcement. So
those are the the different types of
machine learning that are out there. Um
and then we talked about some of the pi
Python packages uh that we will use. The
main one being scikitlearn and of course
we'll use our older like pandas to
manipulate our data and get uh pass it
into our model training etc. But
scikitlearn will be uh our goto for
anything machine learning.
All right. So, I have some questions for
you guys, some checks.
So, let me know in the chat. What do you
guys think? Uh, which of the following
best describes machine learning?
Which choice do you think makes the best
is the best for this?
Very good. Very good. I see I see a lot
of choices for A and A would be the
correct choice. So machine learning is
definitely um a a subset of AI. It's
underneath that AI umbrella, but of
course we're learning from experience
and of course that experience is
recorded in the data um without being
explicitly programmed. Uh so it's the
exact opposite of BNC. We're definitely
not learning from rules and it's
definitely not just used for image and
speech speech recognition. It can be
used for many other things beyond those.
So yeah, A is the best choice there.
What we say here?
Okay. What do you guys think about this?
Which example illustrates the use of
machine learning to enhance customer
experience in an ecommerce company?
In other words, what would be some what
would be some uh typical use cases of
machine learning?
Good. So I think uh C is going to be the
best answer here. Definitely C. So it's
using machine learning to do uh fraud
transactions. So so that would be a
prediction probably a supervised
learning, right? If if this is fraud or
not fraud. Um, and then maybe some
customer behavior uh that might be
unsupervised. So maybe grouping together
customers uh clustering them based on
their data like their shopping behavior
and characteristics. Um that that might
be unsupervised but either way it's
machine learning.
Okay.
Okay. Final one. What distinguishes deep
learning from machine learning in
artificial intelligence? So what's
unique about deep learning?
Oh, very good. Yep. So, deep learning
uses neural networks as so you guys are
right on top of that. Neural deep
learning uses neural nets. That's what
makes it unique. So, machine learning
would be part A. Machine learning is
focused on learning from data.
Underneath of that is learning from data
using neural networks which is what uh
deep learning is.
Very good.
All right. Let's go to lesson three.
And lesson 3 has two notebooks. We're
going to be starting with 3.1.
So, you'll want to open up that
notebook. I'm going to go over to it
now. Give you a moment to open that up.
So, we're going to open the 3.1
notebook. Um, there's two of them. We'll
see how far if we can get into the
second one today. probably will.
Um, but we're going to do the uh we're
going to start with 3.1 notebook. Do you
guys have this notebook? Should be in
your materials for for this course.
Let me give you a moment to open that
one.
Do you guys have it?
Okay.
All right. So, we're going to start by
talking about uh supervised learning.
um in our machine learning journey. So
remember we're going to talk about uh
supervised and unsupervised after we do
supervised
um and there's going to be a lot to
cover with supervised mainly because um
there are uh two different types of
problems we can tackle uh which will be
uh we'll talk about in a moment
predicting different kinds of values. Um
but let's talk about the kind of what
we're hoping to learn here which is um
talk about the different kinds of
problems that we'll study which are
these these categories of supervised
learning. Um those two categories are
going to be called classification and
regression. We'll talk about those and
their differences and then talk about
some applications and some uh example
algorithms
and that's just within this notebook. Um
3.2 two we'll get into uh regression in
particular um which will be uh very very
interesting. Okay. So that'll be our
first models that we'll build will be
over there in 3.2.
Okay. So if you guys remember um
supervised learning is where we learn
from labeled data. So we have input and
outputs in our in our data set. Um and
so you train a model on this data that
includes input features and
corresponding outputs that are that are
the labels, right? So um the goal is to
learn a relationship between the input
and the output. Of course, that's what
any model is trying to do. Um, and what
this allows us to do is then take that
model and use it to make predictions on
never-beforeseen
uh data. Right? So then we have a
predictive model out of that that we can
use um going forward on new examples. Um
so
remember we will have in our data a
bunch of features which are columns and
then generally one of those columns will
be the label that we're trying to
predict.
And our model is going to try to learn
some type of relationship between those
inputs and the output label. So the
output label could be like fraud not
fraud, cancer, not cancer, uh a price, a
temperature, those kind of things.
So let's talk about that. inside of um
supervised learning there are two
different types of learning that we can
do and they're really based on the label
or sometimes that label is known as the
target that we're trying to predict. Um
and depending on that type we get these
two different categories of learning or
two different types of learning. One is
known as regression. So that's generally
when we are predicting something that is
continuous or something that is a
numerical.
So numerical
numerical value. So think of price,
think of temperature, think of revenue.
We're trying to predict something like
that. Um versus something that is
categorical. So that the predicting
something categorical would be like
fraud, not fraud, spam, not spam. um
those are discrete categories and the
problem of predicting categories is is
known as classification because we're
trying to classify examples as belonging
to one category or another.
So we have these two main types of
supervised learning problems. we have
regression and we have classification
and they're going to be handled slightly
differently um for many reasons that
we're going to uncover. Um one of the
primary reasons is that of course we're
predicting something that's continuous
in the regression case versus something
discreet. So the models have to be
slightly different to account for that.
Um but then a step beyond that is the
evaluation has to be different too. Um I
kind of alluded to this last week, but
when you're predicting a regression,
it's very very difficult to to get the
exact numerical answer. So um generally
we don't care about that. Um generally
we don't care about getting it exactly
uh we don't care about getting it
exactly right.
um we just care about getting it um
we're just we care about getting it
nearby, getting it close enough. Um
whereas classification, we do care about
getting it exactly right because it's a
discrete category. So we're going to be
able to evaluate that a little bit
differently to say did we get the answer
right or wrong. Regression is going to
be did we get close? Um because it's we
assume it's going to be nearly
impossible to predict a continuous
number. Um, that's very hard to do.
Okay.
So, any questions on
uh that?
Any questions on those two differences?
Let me give you some examples. Maybe
it'll it'll help too.
So, again, the classification is going
to be predicting uh something that's
categorical. regression is going to be
predicting something that is continuous.
So think about trying to predict the
price of a house based on those other
features we talked about before like
square footage, bedrooms, bathrooms, all
of those things we predict the price.
That would be a regression problem
because the price is a continuous value.
Let's take a look at an example here.
Um, imagine we were trying to uh predict
the temperature tomorrow. That's going
to be a regression problem, a a
supervised learning kind of regression
problem because we're trying to predict
a numerical temperature.
Okay? And versus a category like a
discrete category would be this would be
a classification. So this is a
regression on the left. This is a
classification
on the right. Classification
um because we are um predicting one of
two categories. Is it just hot or cold?
Now, we're not saying exactly where that
threshold is on what's hot or cold. That
would be a decision on on what we want
to what our discrete categories actually
mean.
But, um we only have two choices, hot or
cold.
versus predicting the entire temperature
which would be um a numerical prediction
of some exact number. Right? So that'd
be a regression and then on the right
would be a classification.
Um now again why is this so different?
You can see the types of predictions
we're making are completely different.
One's a number, one's a category. But
again with the valuation it's like if
the if the true answer in our labels was
84
and we predicted 83 that's a pretty good
result. That's still pretty close.
That's pretty close to this. So from an
evaluation perspective that's pretty
good. Um whereas like if I predicted
cold and it's actually hot that's that's
a wrong answer. So they're evaluated
slightly different.
Um, and that's something we're going to
see as we talk about evaluation of our
models once we build them is depending
on if it's classification regression,
there's going to be different ways of
evaluating them.
You can kind of see why it's very
difficult to say, okay, we got exactly
84 when it could be any number. Our
model is going to be predicting a
number. That's really hard to pin down
an exact floatingoint number. So, the
best we can do is kind of say, how close
did I get? Like, this would be a worse
answer. If I got something all the way
down here, that's a really long distance
to here. That's bad. That's a bad
prediction. But if I get something
really close, that's better, right?
That's a decent prediction because it's
pretty close,
right?
Of course, being perfect would be
getting exactly right, but that would be
nearly impossible to do.
Okay.
All right. Any questions on this?
Does it make sense on regression versus
classification? We're going to use those
words quite a bit as we go along. So,
regression predicting that continuous
value. Classification predicting a
category.
And they're going to be um different
models that do that
different models being used for
regression versus different models being
used for classification.
All right, let's talk about supervised
learning uh applications here. So just
to name a few, we have HR operations. is
imaginary recruiter tasked with finding
the best candidates. Um so supervised
learning can help by um rejecting or
accepting candidates. Now this is
something that happens quite a bit even
today. Um and that it's kind of like uh
how recommendations happen like this
this resume should be um recommended
this should not um from a whole pool of
applications. Um so there's those kind
of use cases of of um predicting a
category that would be like a
classification. Should we should we
accept or reject the the candidate?
Um finance you see this all the time
with things like risk and loan
approvals.
Um you can uh predict the the the
category of like if the if the loan if
we should accept or reject the loan
application.
um you know that would be a
classification.
Um what's interesting about
classifications by the way so it says
here like we can predict the likelihood
of a of a loan being repaid
um is a lot of classifications um we we
say that they predict a category but
under the hood they can actually predict
a probability and we turn that
probability into a category. So, um, you
know, like we could say what's we could
say the likelihood of her loan being
repaid is very low. Let's say it's less
than 50% probability. Um, then we could
label this as reject,
right? We could label that as a
rejection. Um, if it's greater than 50%.
Then we could label this as accept. So
we can set a threshold there
and say okay truly we're predicting a
prob like our model spits out a
probability but we turn that into a
category by saying should we accept if
it's less than 50% we should reject if
it's greater than we should accept.
Okay, so that's something we will see
with some of our classification models
is that they actually produce a
probability and we turn that probability
into a category label
um by by doing something simple like
this putting a threshold on it um for
the for the category.
So finances is used all over the place.
Not only just loans like fraud, we
talked about fraud, not fraud. That
would be a classification.
Um predicting sales revenue, that would
be a regression, right? What is the
revenue going to be in the next two
quarters? That's going to be a
regression problem.
Uh emails like spam, not spam, that's
going to be a classification.
um that's going to operate on the that's
going to take the text input and predict
if this email is a spam or not spam.
That's going to be a uh supervised
learning problem, but it's going to be a
classification problem,
right? Uh manufacturing supervised
learning is used to inspect and uh
quality and classify products in
different grades. For example, a factory
might use a model to check for defects.
So this actually something that happens
is you look at images of products as
they go through the assembly line and
you can take a look at those images and
predict if it's a high quality, low
quality, medium quality. Um so they can
be this is a classification, right?
They're going into different categories
of quality. Um so it's much much like a
manual kind of intervention by some uh
QA or quality control uh specialist.
Okay. But that's a classification.
So in the maritime industry, supervised
learning can be used to predict current,
so current level
um and that can be used to forecast uh
supply and demand. Um so those would be
like regression models that are used to
predict um kind of like temperature, but
in this case like title levels.
We talked about fraud already, so that's
there. Um, that would be a
classification.
Okay,
any questions on these uh examples?
Of course, there's many more. Um
recommendation is kind of like a
supervised learning problem uh where you
are
taking examples of things that people
have viewed in the past or or reviewed
in the past and using that to predict
what they would want to watch in the
future. Um so recommendation is
supervised learning. Um and it's like a
classification, you know, trying to
predict um uh certain number of
categories of of uh shows or movies that
you would want to watch. Um
and that's something that we will study
in the future. Recommend we'll we'll
have a whole lesson dedicated to
recommendation as well.
All right.
So when it comes down to the uh actual
models themselves, so there's going to
be lots of different models that we are
going to cover. Um and they are um going
to be different in their purpose and
kind of their uh what kinds of problems
they're used for. Um and uh their their
how they actually train is going to be
different. Um, but at a high level,
they're all trying to do the same thing,
which is learn some sort of relationship
between the input data and the and the
label, right? That's really what they're
trying to do because they're all
supervised. They're they have those
labels, trying to build some
relationship there. Um, they just do it
differently.
And what we're going to study is the
pros and cons of a lot of these models,
like when would I use one of them, when
would I use another. Um, so we'll try to
talk about that as we go along. Um, but
they're all trying to learn some
relationship between the input features
and the output, right? So we have to
keep that in mind. They're trying to
model that relationship. They just do it
in different ways. Okay? So as we go
along and learn about new models, um, we
will learn the details. will learn the
ins and outs um and those pros and cons,
but they're no matter what, they're all
trying to
uh learn that relationship, right? And
be able to make predictions on new data.
Okay,
so here's a list of models that we will
cover and work on throughout the uh the
sessions that we have.
um we're not going to do them all in one
one sitting, but um the first one that
we're going to start with and that we'll
cover today is going to be linear
regression.
So we will cover linear regression and
then we'll cover the rest of these guys
mostly in the context of uh
classification.
So, um, what's interesting is some of
these guys can actually be used for both
regression and classification as long as
you make, um, certain adjustments to
them. They have variations that can be
used to do classification and regression
is very interesting. Um, but we're going
to start with linear regression today
and then work our way through the rest
of these models when we do um, we're
going to do a separate lesson four on
classification. And so these all these
guys will come from lesson four.
Um and then uh we will do this guy in
lesson three in the 3.2 notebook. We'll
do all about linear regression.
Yeah. I so logistic regression is a
classification. Um which is kind of
strange that its name is regression but
it's doing a classification. But the the
reason is that the logistic regression
um computes a probability. So it does a
regression to predict a number but that
number is actually a probability. So it
it produces a result that's between it
produces a probability that's between um
obviously uh zero and one.
So it uh and then we take that
probability and we turn it into a
category
like a spam not spam fraud not fraud.
Um but so so logistic regression is kind
of special. It's sort of like a
regression but it's predicting a very
specific type of value which is a
probability. So for for that reason it's
a classification uh algorithm primarily.
So we'll study that one in lesson four.
Uh but yeah, that's that's why it's
under that kind of umbrella of
classification is because it's it's
producing a probability as its main
output which we can then turn into a
category as long as we interpret that
probability as um in the right way. Uh
like the probability of spam,
probability of not spam.
Okay.
Okay. So, let me focus on um
let me focus on linear regression. I'm
not going to go through all of these
other use cases because we haven't
learned these models yet. Um so, I don't
think they're good. Uh I don't think
it's good to read about them yet until
we've covered them. So, once we cover
them in lesson four, I'll come back and
describe these examples to you guys and
we'll see why it makes sense. But I
think for linear regression um which is
what we'll cover next, let me talk about
that example. So a prototypical example
would be like predicting the house
prices that we've seen in that house
price data set.
So um if we wanted to uh if we wanted to
predict um if we wanted to estimate the
market value of a house so the price
um we could do that by using the
features such as number of bedrooms,
square footage, location, age of the
property. Um and you know then when a
new when a new house comes on the market
we could estimate what the price should
be based on those features. So linear
regression is a good one to predict the
price like a housing price. Um and we'll
actually practice that in the next uh
notebook.
So we'll we'll uh and then all these
other now there's descriptions of these
other models but again we haven't
covered these guys yet. So I don't want
to really go through those until we get
to those models. So we get to those I'll
come back and mention the example.
Uh can K andN be used for clustering?
No. So um the clustering model is going
to be different. It's going to be uh K
means
K means that's the primary clustering
model. Not K nearest neighbors. K
nearest neighbors is used for uh it can
be used for regression. It can be used
for classification.
So we'll we'll talk about K andN which
is the K nearest neighbors in lesson
four.
It sounds really similar. Yeah, it
sounds really similar but K means is a
clustering algorithm that's that's
slightly different
different uh there's no labels used at
all. This K nearest neighbors is a is a
supervised learning algorithm. It uses
uh labels.
Good. Any any other questions so far?
Okay.
So that being said, let's move on to the
3.2 notebook.
Let's move on to that which will be our
um first discussion around uh
regression. So going into supervised
learning and regression. Give you guys a
moment to pull up this notebook.
But yeah, you want to pull up the 3.2.
We'll do this one next. So we'll focus
in. So our plan is to do regression
first and then we'll talk about
classification in lesson four
which we will cover all those other
models which you you could use for
classification uh on that list. But then
we're going to talk about linear
regression uh first.
All right. So we have a a big agenda.
This is a big notebook um to go through
a lot of material here surrounding
regression. So we're we're going to
start with linear regression and see um
how we actually perform it, what that
model is doing. Um which we've kind of
seen the idea of it a little bit
already, so it should be somewhat
familiar. Um and then we'll talk about
how to adapt that linear regression idea
to um nonlinear what's called nonlinear
regression which is going to be using
like polomial uh features. We'll talk
about how to do that. Um and then a big
big big topic for us is going to be
evaluating the model. So it'll be it'll
be quite easy to actually build it.
building the model will be really easy
but evaluating and interpreting that
will be uh a lot of interesting work
there um because we want to know what
the performance of that model is once we
have it built right we want to know how
good of a model is it is it worth using
or do we need to retrain it or get new
data or change the model up to talk
about that um how do you determine what
to do based on that performance
um and then we'll We'll talk about here
um a couple things. We may not get to
this today, but regularization
which is used to boost the performance
uh in certain situations um whenever the
model is kind of performing um poorly
against test data even though it
performs pretty well on training data.
In that scenario, you can use offshoots
of linear regression that do some uh
what's called regularization. We'll talk
about that.
Um, and then we'll talk about
hyperparameter tuning, uh, generally as
a strategy, which is something you
generally do want to do when you're
training machine learning models. Um, so
again, these two we may not get to
today, but um quite a quite a lot to get
to be prior to that mainly centered
around evaluation and building linear
regression.
Okay, so pretty cool. we'll get to our
first kind of model here. This linear
regression
to start with.
Okay,
so let's start with uh linear regression
here. Um, and really what linear
regression is attempting to do and I
want to show you this in this picture is
draw this line sometimes what is known
as the line of best fit. So this is our
model that kind of goes through the data
and it's generally a good predictor
um because if you give me um features uh
if you give me new features and let's
say they are let's say you give me a
feature that's right here.
So you say, okay, I have a feature
that's this value on the x- axis. Then I
know all I have to do is plug that into
my line equation, and I will generate a
a value that's like right here.
Okay, that's pretty that's on that line
at that input. And that's going to be my
prediction for what the output variable
should be. It's just going to be
something on that line. And what you can
see is this line is a decent estimate
for this data because it slices through
this pretty evenly. So it's a good guess
as to what the output should be given
any one of these inputs. It's a it's a
good estimator this line. And so our
goal building a linear regression is to
kind of build the equation of this line.
So we want this equation.
Equation of this line
is going to be our model.
Yes, it's going to look just like that.
MX plus B or yeah, MX plus C. It's going
to look exactly like that. uh except
that it's going to be more than just MX
because we have um generally more than
one feature. So you think of X as a
feature um it will be more than just MX.
It will generally be like uh it'll
generally look like this
and then plus maybe some bias here plus
an intercept. Yeah, it'll generally look
like that. So, yeah, you're exactly
right. MX plus B is the right idea.
Exactly right.
It'll generally look like that.
Nonlinear, it can be adapted to
nonlinear. Yeah. If we transform, we're
going to talk about that. If we
transform all of our features in a
nonlinear way, um we can apply linear
regression to it. Yes. And and that
would be a nonlinear regression. So yes,
we can do nonlinear things too.
We'll talk about that.
Okay. So linear regression again is the
art or science I should say not really
art but it is an exact science of
finding the equation of this line that
fits through this data. Um now why one
thing you should be thinking about is
why is this line a good predictor and
the argument is that if you take a look
at this distance from these blue points
so let's say these blue points are
actual data points this line is going to
be found such that it minimizes this
distance
from the points to actually I should
draw it this way from the points to the
line. So, we want this distance to be um
actually I should draw it that way. This
way. We want this distance to be kind of
at a minimum. So, it would be bad to
draw a line all the way out here because
then that's a lot of distance, right?
So, and that would be a lot of error um
contributed from not being able to
predict those points in our data set
very well. Um which is our training
data. That's why we have labels, right?
that that guide us in building this
line. Um so our goal is to build that
line especially so that this error or
this distance can be as minimum as
possible. Right? Which are all these
distances from these points to the line.
We want those to be as minimum as
possible. So our goal is to find this
equation.
So we're going to build a model that's
going to find this equation.
of the line
um such that our error
is minimal.
And what is the error? The error is the
distance
of our data points
to to
the line that we build. So essentially
what we'll do in order to train this
will be to adjust the parameters or the
or in that like I think is really good
you brought up the MX plus C. Basically
the M and the C will adjust. So we
adjust those accordingly to make this
distance as small as possible.
Okay? To minimize that distance as much
as possible.
Okay.
So um where is regression used? We've
already seen some examples. Here's some
more uh advertising like predicting
sales, predicting um oil and uh oil
production and demand. Those are like
forecast those are regression problems.
Um retail like demand forecasting for
inventory. Um healthcare predicting um
uh the levels of certain um uh blood
markers or you know something like that.
um real estate predicting prices based
on those uh talked about like square
footage, bedrooms, bathrooms, those
things. So regression is used again
whenever we want to predict a number a
numerical output um that's a regression
problem.
So this kind of regression we're talking
about here is generally
um known as uh a when that equation is
linear that is known as a linear
regression. So go back to that picture
when we have a when that equation of the
line that we find is a linear equation
meaning that it is exactly the form I've
been telling you. So it's it's something
like um weight time feature
plus weight time feature
plus weight time feature
and then maybe some intercept um term
like some some bias term there.
Um this is a linear equation because all
of the features are to the single power.
So it's a linear power and this is a
linear combination of features with with
those different weights. So this is a
linear model
because it is uh it's what in math we
would call this a linear equation right
everything is to the first power. It
resembles mx plus b. It is a linear
equation or linear model. Um so when we
talk about linear regression that is a
regression model so we're predicting
some continuous target that assumes we
are model our model is formed from this
kind of equation a linear equation.
So this is going to be our our model for
a linear
uh regression.
Okay.
And so when you when you train a linear
regression, your goal is to learn these
weights so that you can plug in um you
can plug in any one of your uh input
features and you um can generate a
prediction. You can which is going to be
something on that line, right? It's
going to be a value that's sitting here
on this line.
We put in all of our features and we end
up there somewhere on that line.
This output.
Okay.
Okay. Let me pause there. Any questions
on the linear model here or why it's
called linear regression?
Okay. And by the way in these notes um
this bullet point here where it says it
uses the least squares criterion to
estimate the coefficients that is
exactly what I said earlier with the
distance. So the distance is based on
the square
of this this quantity like how far away
you are from the line is based on this
square distance here and here and here
and here. So what we're trying to do is
find the least distance or least squares
which is that minimum distance. So
that's how we find all of these weights
is from minimize. We basically tune them
enough using our labels. So here's our
label which is the y. We basically plug
in our data and tune those enough to
minimize the error. It's it's a it's an
optimization problem, right? We we're
trying to find the minimum of this
quantity which is that best fit line.
Okay.
So we have linear regression
um and we can do a simple linear
regression that only has one feature. So
if it only has one feature that's
exactly the so if there's only one input
feature sometimes that is known as um
simple regression or simple linear
regression and there's only basically
there's only one feature. So one
independent variable is the feature.
There's only one feature. And so this
equation resembles the
exact equation that you guys just put in
there, which is um mx plus b,
right? It resembles exactly that. um
we're just using different symbols for
those like beta beta 0 and beta 1 but um
basically exactly that simple line
there's only one feature. So and that's
because that line is going to um that
line is going to be generated uh
according to that equation. So here's
kind of what it looks like.
This is the best fit line through all of
these blue dots. This is something we're
going to be able to build. we're going
to be able to build that equation um
pretty easily in scikitlearn.
So we'll be able to find that um and it
won't be too hard. So this line will be
um y = beta 0 plus beta 1. So some
weight beta 1 times the only feature we
have x1.
Okay. So in this case um we would be
predicting sales. So sales would be the
value basically the label that we're
trying to predict and the feature that
we're putting in is uh I think it's the
number of TV expenses. Yep. TV expenses
which is on the x- axis. So there's one
feature which is um TV expense.
So um on this graph this would be this
would be our model.
Okay that would be our model. We only
have one feature and we have um these
two weights. We have an intercept B 0
and or beta 0 and then a one weight
which gets applied to that one feature
beta 1. And so our model would have
certain value for beta 0 and a certain
value for beta 1. That's what get that's
these guys get learned
learned during
model
training.
Okay. So those are what get learned
during our model training and they get
learned by a a a least what's called a
lease squares algorithm that is trying
to minimize that distance. It tries to
tweak beta 0 beta 1 to minimize this
distance of this line
um this line
to all of these points
trying to minimize this.
So imagine taking a line and kind of
moving it around and turning its its
slope, its angle um to try to find that
best fit,
which reduces that error the most.
Right? That's kind of what we're doing.
Uh can I explain? Yeah. So uh sales is
in dollars and and TV expense
um
uh
TV actually I think it's the other way
around. I think the sales is actually a
quantity. So I this is number of sales
that we have and TV expense is um I
think I think it's in dollars. So how
much money how much expense um did we
put into the into the product and then
this is how many sales did we have of
that product.
So I think it's the other way around.
But what this what this graph is showing
is the blue points are our actual data
points. Okay. So so we have a collection
like we have a data frame that has so
imagine we had a data frame that has the
uh true values.
So it has the um TV expenses. Um, it has
points that are like one. So, it has
points that are like 120 and then the
sale sales could be like 700
700 units, let's say. And then it has um
so this is just our data set, right?
This would be like in a data frame that
we have. And then we had ones that were
um 50 and then this could be um this
could be 400, let's say. And on and on
and on, right? So this is our data and
this data is plotted in the blue. So
these are these blue points here,
right? So these are the blue points here
and the red points are is our model. So
we built a linear regression model
um where we are putting in some values.
We're putting in some e fake x values
here and generating some predictions
which is this line
this linear uh regression line.
Right? And that line is derived from
this data. Right? It gets learned from
this supervised uh examples.
Does that make sense?
That line is derived from the data. It's
actually um learned from like the line
of best fit is learned from that data
and the actual data is in the blue.
So you can see we're trying to build
this such that this distance is kind of
a minimum.
So it's an optimal fit
to balance out these distances.
So it's just plotting. So it's just
building that relationship between the
input and output. like when the when the
expenses are higher, um we seem to have
more sales.
Uh what's perpendicular like the
distance? This should be this should be
perpendicular because it's a distance
here.
Is that what you mean? Like the distance
from the real points to the line. Yeah,
that should be perpendicular
because it's it's a it's a distance
formula.
Okay.
All right. So, more generally now do do
we usually have one feature? No. So
generally we expand this to the more
general case where we have more than one
feature like what we see in the housing
data right where we could predict a
price but we have many different inputs
like bedrooms, bathrooms, square footage
etc.
So more broadly
instead of simple linear regression we
have what's known as multiple linear
linear regression which means we have
multiple variables or multiple features.
Um so this is exactly the equation I've
been talking about. Um so we just extend
that that one into many features. So
which is this case and then a intercept
term which is uh um there as sometimes
known as the bias. Um
but this is the intercept term to kind
of orient the line to start out in the
right place. Um and uh but this is the
um this is the equation that we would be
building the model. This is our model
essentially, right? This is the equation
we would be learning.
Intercept is like a constant. Yeah. So
if if all of the features were zero, um
this is what our our data would be. This
is what our result would be. If
basically if this was zero, this was
zero, this was zero, it would reduce to
this as the prediction. Yeah. It's like
a constant. Yes.
So in in geometry, the intercept is
actually really important because it it
orients where your line should start. So
it orients like so so these values are
kind of like the slope. They orient the
tilt of it. Like should it be tilted
like this or should it be more sloped?
But the intercept orients where it
should start like vertically like should
it start all the way up here? Should it
start more down here?
Um, that's what the intercept kind of
tells us.
Okay, so this is the situation. This is
going to be our linear regression model
that we will be building most of the
time because we will have again these
are all going to be features.
So this is some feature the X this is
some feature this is some feature
X1 etc. These are all features and what
gets learned during the training are
these coefficients. So all of these
coefficients including the beta 0ero um
will get learned. So these will get
learned
um from our data right they get learned
they will be trained from our data um in
order and and how do they get trained
it's from reducing that distance we try
to get that line of best fit by tweaking
those betas enough to uh until we reach
a minimum distance but there's there's
an algorithm behind that um that that
scikitlearn will run for us to find that
best fit Um, so we don't need to do that
manually, but that's that's the process
is basically tweaking those weights to
end up with that line of best fit. So in
higher dimensions, instead of a line,
you get more of what's called a plane
here. Um, which kind of looks like this.
So the best fit is actually this plane
where all um, it kind of dissects all
these points just like that um, in
higher dimensions. So this is uh instead
of a line you get this in in three
dimensions you get this plane like this
but it's still it's like a line of best
it's just a more general line of best
fit. It's still the same idea. Um we're
still trying to um come up with the best
coefficients to minimize that distance
from our from our points to the line.
Although in higher dimensions it's no
longer a line. It's more like a plane
like this. So you're trying to minimize
this distance from here down to the
plane
here up to the plane
in higher dimensions. So I want you to
keep in mind what we're trying to do
before we go into the code because the
code's going to make it seem really
really simple and that's because
scikitlearn is great and that's what it
does.
But we should realize that there's
something really complex going on which
is again finding the best value of these
weights
that minimizes the distance of this line
to the data points that we have. So
there's an algorithm there that will
keep trying to make adjustments to this
based on those distances. So it's going
to use those distances as a guide to
kind of tweak them to find the one that
results in the lowest amount of
distance. So we keep making tweaks, keep
making tweaks, keep making tweaks and
eventually we try to find we converge to
the set of weights that gives us that
best fitting line. Um and and there's an
algorithm there that occurs. Now luckily
that gets abstracted for us a bit behind
um scikitlearn
um finding that best fit. So there'll be
a function that we use in scikitlearn
when we build the model that will go
ahead and find the best weights for us
and that's then we now have our optimal
model right that then we can just plug
in different values of these features
and generate a prediction which is going
to be this uh result right so so that's
what we're ultimately trying to do is uh
train the model which will uh find all
those optimal weights and then uh we can
predict with it which would be plugging
in different feature values to to
generate a prediction.
Okay,
so let's see how that happens. It's
actually going to be super easy um with
scikitlearn.
So uh in this scenario we have um we're
going to import our pandas because we're
going to load our data from that. Um, so
of course we need some data to work
with. So we're going to load this uh
CSV.
Um, I
uh so I was not actually able to find
this CSV for this example, but I mean
that's okay because we'll do some we'll
do other examples where we'll work with
the data. If you happen to have it, um,
great. I didn't see it in in my files.
So just have to take the word for it
that these are the this is that TV and
sales columns here um from this data
set.
Okay. Um as an example. So um just to
see how it's fit um what we're going to
do and this is going to be a very
standard process for us for building a
model. These steps are going to be very
very standard for us which is going to
be first of all splitting the features
away from the label. That's the first
step that we always will take. So if you
take a look at this code, it's taking
all rows but only the first column.
Okay, so it's extracting all the
features from the dataf frame um which
happen to be which is just the first the
first column uh which is the TV uh
column right just that column there and
our target variable which is our label.
So our target variable aka the label um
is the second column, right? It's that
that sales column.
Um and so our first step here, let me
call that out here. First step is to
always split apart
features from labels.
Okay, so we put all those features into
a data frame called X and we have all of
our labels into technically a series but
uh sort of like a data frame, right? Um
called Y, which is just the um which is
just the uh uh labels. So that's just
the TV values. Um now you're going to
see why we do that. It's because we need
um our our features and labels split
apart to put them into the model
building function. It expects our
independent variables or our features to
be separated from our answers or our
labels that guide the model building.
That's the first thing you got to do is
separate those.
Okay, so this code will separate those
out into a capital X and a lowercase Y.
And that's actually pretty industry
standard notation. Whenever you split
apart all your features, usually you put
them into a data frame called capital X
and then you have a lowercase Y to
represent your labels. That's actually
pretty standard.
So it's pretty standard that um X
represents
features
and
Y represents labels
label column
whatever our label column is in this
case it is the sales because we're going
to be predicting sales
using the TV column the TV quant uh
expense quantity.
Yeah. So what it so the assignment is
that we are um the assignment is that we
are
uh we are um splitting apart our data.
So that when we first read in the data
um it is a data frame right that has two
columns TV and sales.
Oh perfect thank you Tim. I will I will
go ahead and so if we look at this data
it only has those two columns right it
only has those two columns. Okay. So
what we're doing with this is we are
splitting apart
our our independent variable our
features. So this this X will contain
our features
and Y will contain
our label.
Does that make sense? We're splitting
this data apart. So, we're only grabbing
that first column here to be our
features. And then we're we're grabbing
the second column, which is the sales,
because we're going to predict the
sales. This is our label. We're going to
we're going to build a model to predict
the sales given the TV input, TV expense
input. So, the first thing we have to do
is split apart the features and the
label.
Okay, that's the first step we usually
will take. And the reason we have to do
that um just to reiterate the reason we
have to do that is because our model
will expect our our data features to be
separate from the label. We will pass
those in separately.
X is TV. It's the first column
because we're using right. It's the it's
all rows but the first column
which is TV.
Why is sales? This is what we're
predicting.
We are predicting the sales given the TV
expense value.
Yeah. Which is why we split it into So
this is the second column, right? The
index one column.
Uh you just put in read CSV and pass in
the URL.
So you could So exactly the code that
was up earlier from temp
um you just do this
and then data equals ddread CSV URL.
So we split our data into X and Y here.
All right. Now, one other step that
we're going to take that's a very very
critical step and you're gonna we're
going to see this step over and over and
over and over again. So, splitting apart
into X and Y will become we'll do that
over and over and over and over again.
Not only that, but doing this next step
which is what's called a train test
split. Now, let me show you what the
train test split does. It takes our data
And it's going to split apart our data
that we have, our X and our Y data. It's
going to split it apart into a
percentage that will be used to train
the data
and then a percentage that will be used
to test. Now, why would we want to do
that? It's mainly so we can do
evaluation. So we build the model over
here and then we test it on data that
has not seen before. So we reserve a
percentage of the data to be used for
test. Usually this this data is um
somewhere between uh 20 to 30%.
So somewhere between 20 to 30% of the
original data. So that means the
majority of it is used for training. So
the majority of the of that X and Y over
here is going to be between 70 to 80%.
Will generally be used for for uh for
training. Okay. So somewhere between 20
to 30 the industry standard is some
anywhere in between there. Um a lot of
people like to use 30%, some people like
to use 20%. Um anything in that range is
acceptable. um we will I think we
generally will favor like 30%.
Um to be used for testing but um the the
point is we don't we don't want to mix
those together. We want those to be
separated out so that we can have a fair
evaluation, right? We want to train our
data on this train our model on this
data and then see how well it performs
on this data that it has never seen
before.
Right? So in order to have data it's
never seen before, we're going to take
our X and our Y and we're going to split
it using this function called train test
split that will do this kind of
splitting for us. Okay. So scikitlearn
has a function called train test split
that will go ahead and we're going to
pass our x and our y and we'll pass in a
percentage like 30% that we want to
split out into a test set and then the
remainder of that the 70% will be used
for training the model.
Okay.
So what we're going to get let me redraw
that. So, what we're going to get out of
this for the train test split is we're
going to we're going to have an X and a
Y per
training and test. So, we're going to
get now we're going to get an X train
and a Y train.
So, we're going to get training features
and training labels. And then we're
going to get test features
to plug into our model and and test
answers or test labels
to do evaluation because what we should
be able to do is build the model over
here and then apply the model on this
data. Meaning we can take these features
and plug it into our model and then see
what answers we get and compare those
answers to this testing data. Right? We
should be able to do that to generate an
evaluation.
Okay? Now you may be wondering why do we
do any of that? What's the purpose of
that?
Evaluating it on this test data gives us
a good sense of will our model
generalize to new examples. Right? If it
performs pretty well on this data,
that's a good signal like when it's
performing pretty well on data it's
never seen before, that's a good
indicator that it's going to perform
pretty well when we use it on brand new
examples
um in the future.
Right. So that's a that's why we do this
evaluation on this data that it has not
seen before. It's going to see this
training data, right? We're going to
train the model on that data. But that
model will never be exposed to this test
data until we do the evaluation
and and generate some metrics to see how
good is this performing
and does it have a good chance of
generalizing to never before seen
examples which is what we want right
because we're going to use this model in
the real world. It's going to be being
used on new examples that it hasn't seen
before. We want it to perform well. So,
this is kind of our test, our
evaluation.
Okay. Any questions on the We're going
to do this in a moment. I'll show you
what it looks like in the code, but any
conceptually, any questions on the train
test split idea? It's a very very
important idea that we um basically use
part of the data to train it and then
another part of it to evaluate. It's
very important we do that. By the way,
this has a term um this in machine
learning this is called cross
validation
because we are using one data set to
train the model and then we're cross
over we're crossing that over into
another data set to validate it which is
the uh the the testing set.
So this is called cross validation. Um
there's actually many ways to do cross
validation. That's something we'll
study. This is a very simple way of
doing cross validation. There's more
complex ways. You can take your data and
you can actually divide it into many
sections
and basically train it against most of
these and evaluate it against one at a
time and then rotate. So that's another
way to do cross validation. We're going
to study that. Um but this is the this
is the simplest way to do it here.
Okay.
So let me show you what you get when you
use train test split. So uh we're going
to import from sklearn.
We're uh from the model selection
module. Now we haven't used this before.
This is our first time using it. But
here's our model selection. We're going
to import this train test split function
and we're going to use it on our X and Y
and we're going to set a test size of
30% which is which is.3. So our test
size
is 30%.
Converted to decimal
right converted to.3. So that means
we're reserving 30% for that test set.
Um you can set a random state. Now
that's completely optional. Um the
random state
is for reproducibility
because what the train test split is
going to do is it's actually going to
shuffle the data and then split it apart
into the 7030.
So um yes, the seed. Exactly. It's like
a seed. So it's it's saying like when
you do that shuffling every time I run
this notebook I'm going to get the same
result but it's going to be random the
first it's going to be random but I'm
gonna be able to reproduce that
randomness with that random state. Yes,
it is like a seed.
Uh it's you can choose any number to be
your your um your random state. It 42
isn't important. You could choose zero.
You could choose one. Um, you could
choose any positive integer. Um, 42 is
kind of like the uh industry standard.
It's it's you'd have to look it up why
it is. Um, apparently 42 is a special
number. Um,
in in kind of the history of development
of this stuff, there's nothing really
special about 42. You could choose a
random you could choose a random seed to
be uh zero. That's fine. It it doesn't
really it doesn't really matter.
Um you just want you could choose it to
be uh one, two, three. Um you could
choose it to be 15. You can choose it to
be anything you want it to be. It's
really so that your your shuffling is
consistent. Every time you run this
notebook, you get the same shuffle
result. So I'm always going to get the
same rows in these splits.
Hitch. There it is. I knew it was from
something.
Yeah. So 42 is kind of like a
it's it's just used ubiquitously
uh you know as kind of a um paying
tribute to the Hitchhiker's Guide to the
Galaxy, but it's no it's there's nothing
that special about 42. It doesn't it's
not going to change our result or
anything.
It's just so that this train set split
is going to shuffle our data and split
it apart into 7030.
You just want to set this to something
so that you get a cons every time we run
this notebook, we get a consistent
shuffle.
And so the data in these sets
are uh consistent. That's all.
Okay. But do you guys see how we pass in
our X and our Y and we generate four we
generate four different data uh
quantities here which is we generate
training features, test features,
training labels and test labels because
again we are generating these four
different we're generating data on these
two different sets. A training set and a
test set. So we have training features,
training label,
and then test features, test label.
Okay, that's why it's so important to
split apart our data into the X and the
Y. We need those split apart in order
for this part to work.
So by the way, these two steps we will
always do for any model we build. We'll
generally do X and Y and then train test
split in order to generate the data that
we will use for building our model.
Okay. So this this data here is going to
be what we actually use to guide the
training of our model. So it's
definitely supervised, right? Linear
regression
um we we will use that
Okay, so we haven't built the model yet.
We're just getting our data split apart
and ready for the training. We haven't
actually built our model yet, right?
That'll be coming up uh in a moment.
But this is getting our data ready. We
started with our data frame. We split it
apart into uh an x and a y. And we split
that into a train test split. And um you
know then we can uh then we can go ahead
and um pass in to our model training
which we'll do in a moment.
Um you that's a good question. You could
run so what you could do is you could
run
um should we import numpy? Let's see.
We did. Okay. You could run the average
on the um you could check the MP mean on
the X train and see how it compares to
um
see how it compares to X.
So you could you could do that and see
what the average of this feature is um
compared to the average of the original.
They may not be perfect because we are
taking a reduced data set size. So I
don't think there's really any good
there's not like a one-sizefits-all
validation we can do because we're
taking a random shuffle and taking a
percent. We're taking 70% of the data
out. So we're not guaranteed to maintain
the same statistics. We can see if
they're close.
Um but does that make sense? Like we're
not guaranteed to get the same stats
because we're taking a slice of it.
We're taking 70%.
So it's not guaranteed to to to
be the same distribution really.
Delete that.
Uh is it a good practice? Yes, it is.
It is. Uh 30% is the industry standard.
Anything between 20 to 30. So 0.2.25.3
any of those are acceptable. It's really
up to you. Um I mostly see 30%.
Mo I think.3 is is a good good practice
to use for sure.
Um I did explain random state. Uh random
state is so that you get consistent
shuffling. Um you can set this to any
integer that you want it to be. It it
doesn't really matter. Um you can set it
to uh 100, you can set it to 10, you can
set it to 15. Um it just ensures because
what this split will do is it will
shuffle the data first. It'll shuffle
the rows and then um split it apart into
the into the train and test sets. So you
set the random state so that the next
time you run this you get the same
consistent shuffling. That's the only
that's the only thing it it helps you
with because it is randomized but when
you set a random state um it's so that
like if you run it again you'll get the
same shuffling.
You'll get the same the shuffling
matters because it it it uh dictates
what ends up in in these sets.
Okay.
All right. So let's see let's do let's
build the model
um and let me show you how easy this is
going to be to build the model and this
is really how it's going to be for every
single scikitlearn model will basically
look the exact same for training it
which is what's going to make it really
really nice. So the first thing we have
to do is import our model. So from
scikitlearn we're going to be using a
linear from the linear model package or
the linear model module I should say
within sklearn we're going to be
importing the linear regression
and we're going to create an instance of
the linear regression here.
Okay, so linear regression and look how
easy this is going to be. Nearly all
nearly all sklearn models use
ffit function to train.
So every one of them, no matter which
one we use, like the decision tree, like
the um logistic regression, any of those
like we use for classification that are
going to be coming up in lesson four,
they're all going to look the same in
terms of it's going to run.
Which is um scikitlearn's
uh generic function for training your
model. So this will execute the training
once we run this code. And what that
again the linear regression training is
going to do that least squares distance
procedure or algorithm to try to find
the right weights. It's trying to find
those weights that minimize that squared
distance uh from our line that it's
trying to build to the data.
And what I want you to notice is what we
put into the ffit. See how we put in the
training data where we put in the
training features and we put in the
training labels. Now this is supervised.
So of course we put in the labels,
right? Of course we put in these labels
here and of course we put in our
features here. So we're putting in all
of our examples from our training split
into this ffit which is going to train
the model uh so that we can we can use
it for prediction.
Okay, it's really fast. If I run this,
it's going to be pretty much instant.
Pretty much instantly it gets trained.
And you can see here we now have a
linear regression. you can see in this
little box. Um, and it and this
information says that it has been
fitted. So, it's now ready to be used,
right? So, we now that's it. We've
trained our model. We tr That's how easy
that was. We did fit. Now, what we
should realize is there's a lot of work
going on behind the scenes of this ffit.
Okay, there's a lot of work being done
there to do the least squares algorithm
and find those weights and and create
that line of best fit. Right? So there
there's a lot of work being going on
there that's going on there behind the
scenes, but scikitlearn is abstracting
it away for us. Right? And all we have
to do is fit when we're using this code.
Really easy. Really easy. Fit. And there
we go. We've trained our linear
regression model.
And by the way, if you want to see what
the coefficients are, you can actually
extract them if you do so if you take
your lin regression and you do um
coefficients like this
coeff with a with an underscore. So this
gives us the trained
weights coefficients
also known as the coefficients right.
Um so if you run this you can see uh
right now we have this coefficient here
um which is the only coefficient we had
on our feature. So we only had one
feature coefficient there.
And we can take a look at our intercept
which is this.
So this gives us the train weights
and so we can look at the intercept we
can look at the the the coefficient. Um
so obviously if we have multiple
features our model has many features
it's going to have more values in that
coefficient but the intercept is just
the single value 7.23
and then the coefficient
is 0.046. So that's the weight that gets
learned.
Is there a size limit? No, not really.
There's no size limit. Um,
no, you can use as much data as you
want.
There's really no size limit other than
what like what you can fit in memory.
I'd say that's the only limit is
basically what the amount of data that
can fit in memory.
Okay.
All right. Were you guys able to run
this? Were you guys able to run the
linear regression ffit?
Okay, perfect.
Perfect. You are Okay, great. Great.
So, we have a model and we can use it to
predict. Um, and so that's actually what
we're going to do next. If we go down
here, um we're going to have a function
that's going to um build a scatter plot
of our original test data.
Um so we're going to have our test data
here.
Um,
and we're going to then take our uh
we're going to take our training data
and plot we're going to use the uh this
data versus our sales predictions. So
you can see we're going to you this is
how by the way this is how you use the
scikitlearn model to predict. You have a
fit to train it and look at the function
you use to predict. It's literally just
called predict. That's how easy it is.
and you pass in your data, all your
features into this predict and it
generates a prediction for every row. So
every row in these features in this data
frame um will end up with a prediction
using our model. So what we're going to
do is plot our training date uh features
against the predicted sales to see how
good of a fit that really was.
Okay. to see to see the regression fit.
Okay. And so there's the regression fit.
We have all of our test data here
plotted in the green. We have our blue,
which is our um we have our our blue,
which is our uh um training data line
that we built our model on. So that's a
pretty decent fit. Um, and then our test
data is here. We just plotted in the
green scatter. But the thing I want you
to see is this prediction, right? We we
were able to generate some predictions
on that training um by running our
predict function with our model. Now,
this model has been trained. So, we've
already fit it and now we're using it to
predict, right? And so, we're predicting
the sales and plotting that on the
y-axis.
So the sales are we're using the
predicted sales there which is our blue
line. So this is our line of best fit.
So this is our model prediction.
This is our model predictions. Right?
You can see it's a pretty decent uh
line, right? Pretty decent line of best
fit. Of course, there's some error here
like there, you know, it's not perfect,
but it it does a decent job of being a
best fit line.
Okay, so look how easy that was to
just to recap this to fit our model was
a linear regression.fit. And of course,
we're going to do more examples. So no
worries uh on that. We're going to see
this many many many times throughout
this notebook. But we have linear
regression.fit to train it. And then we
have linear regression.predict
to and we pass in our features and that
generates a predicted output.
Right. So what this is actually doing is
is computing this quantity.
We could do either.
We could do either. Um, so we could do,
so one thing we could do is plot uh, so
we could swap it out. We, we could do
either one. It doesn't, it's not a big
deal to do the training set. We could
do, so we could plot X test and then we
could plot linear regression X test.
So it's it's a similar line. Um it's
just different input features, but the
line is going to be the same. Just
different inputs,
but the coefficients are the same,
right? It's the same line. It's just we
generate different outputs.
So yeah, you could do either one.
This is this is honestly this is
probably better. I see what you're
saying. This is probably better because
this is the line of best fit through
this data. So that probably makes sense
to do to do predict on the test set.
Agreed on that. Probably makes about
most sense.
But you could do either one.
Yeah, I think that would be the most I
think that makes the most sense is for
it to be on the same one just to
validate. So like we could do we could
do training here and then train and
train just to see how that data lines
up.
Really, what we're trying to do is have
our scattered data and then our line of
best fit on the same plot. That's all
we're trying to do, right? So, yeah, I
think I think they should be the same.
I think that makes sense.
These values
or which values do you want to see?
Yeah, we could uh we could generate
those if we just do um let's go down
here. So the the line values
um are going to be uh the prediction.
So, um the the uh test
predictions
equals um
test predictions equals linear
regression.predict x test and then we
could uh we could print out our test
predictions.
Yeah. So, we can see what those actual
values are on our uh on the test set.
Yeah.
Um, we will do that. Yeah. So, you
thought we were checking how well our
data was trained. We will do that. Yes.
We haven't learned how to evaluate this
yet. We're going to talk about that
coming up next. Yeah. We will do that.
We just haven't learned how to do proper
evaluation
of a regression model.
But yeah, it's something we're going to
talk about for sure
and see how to do in our code.
Okay.
All right. Any other uh questions on
this example?
Again big takeaways
fit to train it and then predict to use
it
predict on the features to use the model
and make predictions with it.
So here is an example we we made all the
predictions. This these are all the
values that are on that line.
These are all our predictions and notice
they this is a truly regression right?
These are all floatingoint values. Um,
so this is definitely a regression,
right?
Okay.
Uh, that's a good question. Um,
I'm not sure if there is
If there's like a verbose
there's not really no there's not really
a verbose you can I mean you can look at
the source code if you really want to
see you can view the source code to see
um how it's done I can tell you I mean
so generally linear regression is done
in two ways either you use a formula um
to to solve the optimization problem of
minimizing like this this distance from
the points to to the line. Um,
or you use something called gradient
descent, which is how a lot of these
things do it is they iterate through a
bunch of different iterations where they
update these weights according to um a
certain uh basically a gradient of the
the error function. The error function
in this case is the is the squared
distance from the line to the uh to to
the points.
So uh we can compute the gradient of
that and do um gradient descent. So if
you really want to look into it, I would
do some research on like linear
regression gradient descent.
Okay, linear regression gradient descent
to see how that's uh how that's being
done. Yeah, it it's it's a pretty simple
procedure. Um, again, you have the the
notion is that you want to minimize
minimize the loss or the error. Uh, in
this case, the loss is the square
distance. So, it's like um there's like
a it's a formula. It's like a sum of a
square distance from your prediction
um or your label sorry to your model
which is the beta 0 um plus beta 1 x1
plus beta 2 x2
etc like your model and then squared. So
this squared this is the squared
distance here and you're minimizing this
guy which is like a calculus problem.
You you find you basically find the this
is this is I'm getting so far into the
weeds of this, but this is like a
parabola and you work your way No, no,
you're good. It's it's it's a good
question. Um you work your way down to
the minimum of it. Does that make sense?
Like you're working your way down here
and you do that through a descent
process, like a descent iteration.
Um
so
that's how these are found.
Um, but you don't see that happening in
the background. But if you look at the
source code, it I guarantee you it would
be it's either going to be this or
they're going to use the they're going
to use a a a matrix formula to basically
solve an equation um that involves this
basically the derivative of this set
equal to zero and you find the minimum.
Either way, you're finding the minimum
of this.
Okay. But yeah, I don't think Psycharn
has like a uh maybe there's some type of
verbose flag you can look for.
I don't think they have that though. Not
that I've seen.
All right.
So I have uh an important um concept to
talk about next which is going to be uh
called overfitting and underfitting
um which is a really important concept
that's related to the training and test
data we just split apart to do
evaluation.
And um essentially the the issue with
machine learning is that it's not
perfect and it can struggle in different
ways. And the two ways that it primarily
struggles is going to be overfitting and
underfitting. So overfitting is a
situation where the model basically
memorizes the training data so well that
it's it fails to generalize to new
examples. So what we see with
overfitting is this exact sign here
where we have really good performance on
the training data. So when so when we do
that train test split we see a really
good accuracy or really low error on the
training data but it does not perform
anywhere near that on that test data
split. So what that means is that the
model is overfitting to the training
data. it's basically memorizing it and
it's not able to generalize very well.
Now, why does that happen? It's usually
because the model is way too complex.
And that means generally you need to do
something to reduce the complexity.
Either you need to use a simpler model
or you need to use some type of
technique to mitigate overfitting. And
we're going to we're going to study some
of those techniques coming up in this
notebook. Uh we might not get to it
today, but we're going to study
particularly what can we do to prevent
overfitting because overfitting is the
more common issue with machine learning
models. They tend to do so well at
learning from data that they pick up on
small details and patterns in the
training examples that they're exposed
to. They don't do a great job at
generalizing to new examples. they can
struggle with that. So that's
overfitting is struggling to generalize
to new examples, but you do really well
on your training data. So it appears
like you have a good model, but it it's
not able to go and make predictions on
test data very well, which means we
would not want to use that model in the
real world, right? Because it's not able
to generalize outside of what it's
already seen. And that's not a good
thing if we're trying to use it for real
world examples, right?
So overfitting is a real issue. Um you
see it all the time. I've seen it many
many times in the real world, real
industry uh work that I've done.
Overfitting is a is a challenge for a
lot of machine learning models. And so
we need some techniques to overcome
overfitting and we're going to study
some of those uh coming up shortly.
Um, one of the things that we can do,
one of the one of the things that we can
do to detect overfitting is exactly what
we just did, which is you split apart
your data into training and testing so
that you have a chance to do an
evaluation to see if you're even
overfitting in the first place. You want
to see that performance be consistent
from train to test, right? You want to
see consistency. What you don't want to
see is performance that drops off on the
test data. It's much worse. You don't
want to see that. That means that your
model is overfit uh to your training
data and it's not going to perform well
in the real world.
Okay. So, we're going to have a couple
ways to uh overcome that. Talk about
that. Um now, the opposite can actually
happen as well, which is called
underfitting.
And underfitting
refers to the fact that a model is too
simple and it actually just performs
poorly across the board. So if we see
poor performance on the training and
testing data, that's a good signal that
the model's underfit and that means it's
too simple usually and you should try
using something more complex. Um, so the
best way to combat underfitting is to
use a more complex model. And as we go
through and learn about the models,
we're going to learn about which ones
are simple and which ones are complex.
So we're going to have a scale of kind
of complexity. And if you're
underfitting, you want to bump up to the
to a more complex model. If you're if
you're overfitting, one way of combating
that is to actually go down to something
more simple. Go the opposite way to
something simpler. So we need to learn
right now we've only learned linear
regression
but we will learn other models you know
in the future and we'll we'll talk about
uh their complexity and how they're
related to each other.
Okay, but these are two issues we see
just to draw that out again is if we
have a train test split where we have
7030 split let's say and we perform
really well over here but we go to apply
that model over here and it fails it's
accuracy drops off significantly more
error that's that's definitely
overfitting which is not good
right and then underfitting is just not
performing well in either case so even
on the training data itself self your
your accuracy is not very good. So
you're not really learning effectively.
You're underfitting your model. So
that's that's um underfitting case.
Okay.
All right. Now the issue is that it can
be very difficult to balance these two
and get it correct. That's what makes
machine learning a little bit
challenging is getting this balance
correct of simplicity and complexity. So
you don't want to be overly complex that
you overfit, but you don't want to be
overly simple that you underfit and
you're not able to learn effectively. So
there's a bit of a tradeoff there. And
this trade-off is typically known in the
community as bias variance trade-off. Um
in which case, uh it's basically like a
complexity simplicity trade-off. It's
another word for that. Um,
and so, uh, it's it's thought that, um,
if you, uh, if you have very, um, if you
have a situation where you're able to
fit the training data very well, you
risk not being able to generalize. In
other words, you risk overfitting, and
it's hard to um, it's hard to combat
that in a way. Um, and um, on the
reverse side, if you have something
really simple, um, you risk not learning
enough. Even if you're trying to combat
that overfitting, you risk not learning
enough and your model just doesn't
perform as well as it could. So, there's
a bit of a trade-off there of trying to
find the right balance between something
complex enough to learn, but something
not overly complex that it's going to
not generalize to new data. That's the
challenge. Um, like I said, we are going
to have techniques to overcome this. So
luckily there are things to basically
overcome this trade-off and um and help
us along the way so that we don't
overfit. They basically prevent
overfitting
um and allow us to use complex enough
models um that that won't be overfit.
This is in the um this was in our uh
lesson 3.2 notebook. So you want to pull
that one back up. We were working on
Monday.
Um, and just to recap this a little bit,
remember we were building a linear
regression, I wanted to recap some of
the steps we took there, um, that we
will be doing over and over again. And
really the same kind of steps, uh, that
we do here, we'll do in a lot of our
model building. Pretty much all of our
model building um, that we do, whether
it's regression or classification,
doesn't really matter. um we'll still be
doing a lot of these steps which are um
remember first we split apart our data
into kind of a features and a label
uh x and y and the reason that's
important is because um the model
training uses the features and the label
um to help train the model right they
use those separately um so we want to
split those apart whenever we can and so
we have usually Uh it's a good practice
to call your features capital X and your
labels lowercase Y. And what we do with
that is remember we immediately split
that into what we called a training and
a test set. And the picture we had for
that was something like this
where we had about 70% of the data
we used to train the model against and
then the other 30% of the data we use to
test the model against. Meaning that we
build a model over here and we apply it
to this set over here um to make
predictions. And then the that's where
the supervised learning really comes
into play, right? is on this test set.
We already have the answers. We already
have the label. And so we can apply our
model to this to the features over here.
Predict uh what the the label should be
and compare that. We can get a a metric,
right, that compares how close we are in
our prediction to the actual values. Um
and that was some of our performance
metrics. I'll recap some of those that
kind of measure that distance away from
our predictions to what the actual label
is. Um, but remember we had this train
test split function which helps us split
apart our features and our labels into
these uh four sets of data. So we have
our training features, our testing
features and then our training labels
and our testing labels. So we have all
of those and um really these two guys
are going to be used to train the model.
That's why they're called underscore
train. They're going to be used to train
that model and then the then we're going
to predict on these set of features and
then com use those predictions to
compare to this set of labels right
that's on the test test set. Um and you
notice here our test size is set to 30%.
Um, that's a pretty standard number.
Anywhere between like 20 to 30% is
pretty standard. Um, we'll typically
use.3, but it could be 02. Anywhere in
between is fine.
Okay, so we had that. Hopefully that uh
we remember that from Monday.
So we had a train and a test set. And
then building the model was actually
really really easy. Once you have those
train and test sets, um, we just import
our model object. So from uh scikitlearn
sklearn
um linear model uh module from that
package we import the linear regression
model and then we do um linear
regression.fit
and we pass in our features and our
labels and this is again this is where
that supervised learning is really
coming into play because we're passing
in these labels.
That's really what makes this work,
right? We need those labels to help
guide the model to make those updates.
If you guys remember, the model is
something that looks like this.
So, this was a bunch of different
coefficients
um times the features,
however many we have. Um, and so these
labels are really taking the place of
this and they're helping us um make the
correct updates to these to these
coefficients or sometimes we call them
weights. Um, these B 0, B1, B2. Um, we
find out what the optimal one is to get
the best fit, right? To get the line of
best fit. Um, that's what the model
training when we call this fit. That's
really what it's doing in the background
is finding all those coefficients,
right, to end up with the line of best
fit that has the lowest amount of error.
Okay, so hopefully that makes sense.
That's just a fit um to train our
models. And that's really going to be um
the case for
uh pretty much every single model that
we uh train with scikitlearn. It's
pretty much going to be a fit. we pass
in our training uh features and our
training labels.
Okay, so we had that and this was the
visualization of that where we had our
test points kind of scattered and we see
our line of best fit is the one that
goes through there with that minimal
error. That's that's the whole goal.
Pretty decent predictor.
Okay. And then we talked about
overfitting, underfitting. So just to
recap this, overfitting is the concept
of our model basically memorizing our
training data. It performs really well
on that training set, but it is not able
to generalize outside of that. So it
performs poorly on the test set or data
that it's never seen before. Um, and
that's overfitting. So the reason that
it overfits is generally the model is
too complex and it needs to be um it
needs to be simplified a bit. And one of
the things we're going to do today is
see a couple of ways we can alter the
linear regression model um if we are
overfitting to prevent overfitting. Um
so there's going to be ways to handle
this. Um and so we're going to explore
some of those today.
Uh underfitting is kind of the reverse
of that. Remember it's where the model
is not learning enough. So the
performance is poor even on the training
data. It's not good on the test data
either. Um that is a sign that the model
is probably too simple and maybe we
should use something more complex like
go from a linear regression maybe use a
polomial regression. Um or maybe use an
entirely different model altogether. Um,
if we're underfitting, our performance
is poor, it's a good signal we should
try something else. Um,
okay.
So, we talked about those
and one of the things we also talked
about was evaluations. If you guys
remember, we had different metrics that
we could compute to get a gauge of how
good our model is actually performing.
Um, one of those was MSE, which is this
mean squared error function. Um so we
did this example during class last time
on Monday um where we uh were able to
generate the mean squared error. That's
one of our metrics. And we can see what
the mean squared error is on the
training set and see what it is on the
test set by um just passing in our um
training predictions and our training
labels, our test predictions and our
test labels. pass those into this mean
squared error function and it computes
the MSE and that's that's a helpful
function from the scikitlearn metrics
um package um or module I should say and
we'll be using that quite a bit to do
you know evaluation of of especially of
regression right mean squared error is
pretty is probably the most common uh
performance metric we can have and if
you guys remember what it's really doing
is measuring these distances So mean
squared error is kind of like the
average distance away from our our
points to the actual um to the
predictions which the predictions are
all on this line. Um so it's like
measuring on average how how much error
do we have on average right? Um, and the
idea is the closer to zero the better.
Generally means that the distance away
from our prediction to our points is
pretty low. The closer to zero it is.
Um, which is pretty desirable.
So a low MSE is kind of what we're
looking for. Um, closer to zero the
better. And so um if one model has if
one model has um a low lower MSE than
another, it's it's a better performing
model, right? It has less error.
Okay. And then we also looked at the R R
squared or sometimes known as R2 um
score. Um this is another metric that we
could use that measures the the
variability
um of uh the predictions and if our
model is capturing that variability um
well um and so R squ is has a range of 0
to one one is better that means the
model is capturing the the changes in in
the um output it um our predictions
follow along with those same changes um
so they're pretty close um so closer to
one would be a better score. So we have
those kind of metrics. So like on this
data um this would this would show that
this model was underfitting remember
because this
mean this MSE was bad and this MSE was
bad.
Um and what we should think of these in
the units of what our labels are. um
especially if we take the square root of
this the RMSSE that was another metric
we had um the square root of this is
actually in the exact units that we um
have for our labels. So uh in this
example this was the um this was the the
units or the sales versus the TV
products, right? Um and so this would
indicate that on average if we take the
square root of this um
in the square root of this um we have uh
um we're on average about 11 sales units
off squared. So if we take the square
roo of that um it's somewhere around 3
to four um somewhere in between three
and four units off. And this is as well.
Um, and because both of these are still
not close to zero, um, this would be
under fit. And this shows that as well.
This isn't that close to one. It's
decent, but it's not, um, not that close
to one. So, we would say, and
performance is poor on both training and
test sets. That's the key indicator of
underfitting. It's poor on both.
Yeah, exactly. High MSE correlates to
underfitting. Yes. Yes. And it what's
key is it's high MSE on both on both the
training and the test sets.
If you have a high MSE on your test set
but a low MSE on your training set,
that's overfitting, right? Where it's
not generalizing from the training set
to the test data that it hasn't seen
before. That's overfitting. So the key
is high MSE on both sets.
All right. So we talked about that. Um
we did polomial regression last time. So
that was um doing
that was uh making a curved graph um by
transforming the features into polomial
features and then doing linear
regression with that. So you guys
remember from Monday we did this where
um we took our features and uh
transformed them according to this
polomial features from scikitlearn. So
we can go all the way up to degree
whatever degree we want. So we put in
four here but there's nothing special
about four really. This is just testing
it out. um and we generate the the
polomial features and we can fit a
linear regression on those polomial
features and we get a slightly better
model, right? Um it fits the data a
little bit better than just a straight
line. This curved line with the polomial
features um performs a little bit better
and we could see that with the MSE,
right? or we could evaluate the MSE of
this um and it would be lower.
It would be lower than the curve line.
And so that's something we could do. Um
we would just have to pass in these test
predictions, the training predictions
and then the the test labels and
training labels and pass those into the
mean squared error function and we could
compute that, right? Wouldn't be hard to
do.
All right. And then finally where we
left off um you know is on our
performance metrics. So we talked about
mean squared error. That's that average
distance away from the labels to our
predictions. Um and we take the square
root of that. It's it's basically
measuring the same thing but it's the
square root of it is um more
interpretable because it's in the same
units as our label.
um mean absolute error is is the average
distance of the absolute value. So it's
not the squared distance formula like a
uklidian distance but it is a absolute
value. So it's a little bit um less
sensitive to outliers. They don't get
magnified as much. Um but it's not
typically used as much as a mean squared
error would be with regression. um we
talked about the last time because um
the distance formula or that distance is
actually what's used to train the model.
So it's a more natural um fit for a
performance metric for it.
All right. And then we had R square. We
just talked about that closer to zero
would be um worse. Closer to one would
be better. That means that the model
explains um all the variability in the
in the predictions. Uh it captures those
predictions um closely to the labels
um very well. So uh one would be better.
Closer to one would be better.
All right. So that's where we left off.
Um we're gonna pick up from there with
cross validation. um we've actually
already seen one method of cross
validation. So we're going to study um
we're going to kind of recap that and
and then um talk about cross validation
in general um and look at some more
sophisticated techniques of it um coming
up next. But before I do that, any
questions about anything we've covered
um to this point in in the recap or
anything from Monday? Any questions on
that?
All right. So let's talk about uh cross
validation. Um now this term cross
validation refers to a technique that
evaluates performance. And what it does
is it divides our data into essentially
um training and test sets which we've
kind of already seen. And then we are
able to train a model on on the training
set, evaluate it on the test set. And
that's where that's where we get the
name cross validation because we're
crossing over our model from one batch
of data used to train it over to another
set of data used to validate those
predictions. Um, and there's actually
different ways to do cross validation.
So cross validation is a bit of an
umbrella term for multiple ways to do
that. We've already seen one way of
doing that um which I'm going to scroll
down to is um known as a hold out cross
validation. So that's um what we've been
doing so far. So this is just um
generating a train and a test set
train um split.
Um that's the that's what's known as the
hold out cross validation method. Um and
and this is exactly what we've been
doing so far, which is you split your
data into some type of split, usually
7030,
um of a train and test
and then you um train your model on this
section of data and then apply it to
this to evaluate performance. Right? So
that's that's what's known as the hold
out method. Um it is uh you know
relatively simple. It's pretty fast to
do. Um, but there are more robust ways
to try to divide up our data a little
bit uh more evenly. Instead of just
having one split, we can actually do
many splits, which is the idea of um the
next kind of cross validation I'll
cover. But hold out method is one that
we've already studied. It's the most
basic type of cross validation you can
have. Um so hold out this is the most
basic
and we we've already been we've already
been uh working with this type. Okay.
So we've we've already seen hold out
method. Let me uh explain to you a more
sophisticated method a little bit more
advanced of a cross validation um which
is known as Kfold cross validation. So
this is um going to be a little bit more
advanced of a technique but this is the
idea of kfold is that you take your data
set
and you split it into k number of what
are called splits or folds. So you take
your data and you let's say it was let's
say k equals 5. So we have five splits
here.
Okay. So let's say k equals 5. We have
five splits. So what we're going to do
is we're going to we're going to train
our model on K minus one of those folds.
So if K was five, we had five splits.
We're going to take our model and train
it on four out of five of those uh
splits. So let's say it's these four.
We'll train it on these four.
Okay. And then what we do is the one
split that's left over, we will we will
test our model against that split. So
we'll test here.
Okay. Now, this sounds very similar to
the hold out method where we're doing a
train test split, but it's a little bit
this kful cross validation a little bit
more sophisticated because we repeat
this process that I just mentioned over
and over for all combinations of the
splits. So then what we'll do, this is
just one trial that we'll do it again,
but this time we will pick um four
different splits. So, this time we might
pick,
let me do blue. This time we might pick
this one, this one,
um,
this one,
and this one.
And then those four we will train our
data on. And then we will test against
this one. Okay. And we'll do we'll
repeat this
repeat for all combos of the folds.
Okay. So we'll repeat that. So
essentially what we're doing is rotating
through. Every time we rotate through
one of the folds is going to be left out
as a test set. Now this is a little bit
more robust than just a train test
split, right? because we are exposing
our model to more of the data in in
doing this, right? Because we're going
to split it evenly into five or 10
splits. Those are pretty common um
number of folds to use. 10 or five. Um
those are the ones I've most commonly
seen. Um but we're going to by rotating
through which folds are being used for
training, which ones being left out. um
we are exposing our our model to more of
the data this way than just doing a
single train test split. Right? So now
what do we do with with the results is
every time we do this we we generate um
an MSE let's say or some type of
performance metric. So let's say we
generate an MSE from this guy
we generate an MSE from this version and
we generate an MSE for all combos.
each combo we generate MSE and then what
we do is we average
the metrics
or the in this case uh if we use MSE we
would average those together. So every
time we do a fold combination and we
keep four of them for training, one for
test and we rotate through all those
combinations, we are going to generate
an MSE for every combination
then we're just going to average those
MSSE's to get a final. So the final MSE
of cross val of this K-fold.
So the final metric
is just the average of the uh
performance on all of the fold
combinations. Okay. So our final MSE, we
just average all those MSE from all of
our combinations.
Okay.
Now, what's the advantage to doing this?
It's way more robust of a estimate of
the of the performance of the model
because we're exposing it to all
basically all of our data, right? We're
getting a sense of how it performs
across all those different folds. Um
rather than just doing a single train
test split, which is a bit it's basic,
it works, but it's a bit basic. Um so
this is more robust estimate of the
performance.
Now, what's the drawback to doing this
is that it's more intensive. So, if you
have a lot of data, this is going to be
pretty expensive to do because you're
going to have to especially you have a
high number of folds, right? You're
going to have to divide your data into k
number of folds and you're going to have
to do this over and over again. Um, and
if it's a large data set, it might take
your model a long time to train. It's
going to be a little bit more uh
computationally intense than if we just
did a train test split.
Okay, we just did a single like 7030
split. We only do that once. We only
train the model once, right? We train it
on the 70, apply it to the 30% test data
and evaluate performance that way. Um,
so we're only really using the model and
training the model once, but in this
kfold, we're going to do it um, you
know, k number of times essentially
or I should say one for every
combination that we have to work through
of of all the folds.
Okay.
All right. Does that make sense? Any any
questions on kf fold cross validation?
So k K is an important uh number here.
It it's how many folds how many splits
do you have? A typical value for K is
going to be somewhere like five or 10.
So 10 folds or five folds. Those are
pretty pretty standard
from what from what I've seen.
But does the does the concept make sense
or is there any questions on it on in
terms of um you're always going to leave
one fold out. You're going to split it
up into K number of folds. Always leave
one out. Train on the rest of it.
Evaluate on that one that gets left out
and then rotate those through. And
you're going to do that for every
combination and average all those
metrics.
And by the way, there's going to be an
easy function in scikitlearn that will
do this for us. So managing all these
combinations will be really easy. It's
actually just built into scikitlearn. So
we don't have to um we don't have to do
this all by hand. Okay, this will be in
scikitlearn. It'll handle doing all
these combinations of folds for us and
computing the average metric will be
really easy. So um
we don't have to worry about that. We're
going to see an example of this coming
up shortly.
All right, of kfold cross validation,
but this is a this is a really widely
used technique. And again, like the
purpose, you may be wondering like
what's the purpose ultimately of doing
this? It's to get a sense of if our
model is going to perform well on new
data. That's really what we want to
know. Like is the model going to perform
well when I start to use it on new data
that it's never seen before? And this
kffold is a decent indicator of that
because we are varying which data it
sees across many different folds. Right?
So it's a it's kind of a good um proxy
to exposing it to different kinds of
data each time and seeing how it
performs.
All right? Because we're working our way
through each one of the folds. There's
always going to be one fold left out.
We're going to change which fold gets
left out each time. And um that's sort
of mimicking the idea of we're going to
apply our model to new data and see how
it performs. And it's it's new data
every fold.
um how we know which model is best suits
for which scenario because we have Yeah,
that's a good question. Um,
so my we're going to learn this as we go
along because we haven't covered all the
models yet, but generally the best
advice I can give on that is
you you generally want to start as
simple as you can get and then if it's
not performing well then work your way
up to something more complex.
So we are going to have models that are
simpler. We're going to have models that
are more complex. The rule of thumb is
to start with the most simple model that
works.
So you're usually going to have the same
ones that you're going to try in the
beginning. And linear regression is a
very simple model. It's usually the
first one you want to try for regression
because it's the simplest.
Um, and for classification, we're going
to have a similar like logistic
regression is the simplest kind of
classification model we could have. So
usually want to start with that and then
if it underfits like if we see it's
producing a lot of error then we work
our way up to a more sophisticated
model.
So um that's the way we that's the way
it should usually go is simple to
complex it based on their performance.
So we evaluate it and then we can repeat
the process. If it's not performing well
we can try something different that's
more complex if it's underfitting.
Uh this is a good question. Does a model
reset after training each K minus one
fold? Um yeah, it's essentially like a
blank model every time uh every fold. So
um we imagine like you have a brand you
have a fresh model every um k minus one
combination. Yes.
And the reason the reason it has to be
that way is because you don't want the
other folds influencing the model that
like on on the next combination. You
don't want the previous combination to
influence the results on the next one,
right? Um you want it to be a fresh
evaluation on every combination of
folds.
Okay.
All right. So, let me describe to you a
variation on what we just um talked
about with the K-fold. So, there's
another cross validation known as
stratified K-fold. And um this is the
same exact procedure as k-fold except
that when we this is used for
classification.
Um so when we do classification
uh we want to make sure that the
different categories are going to be um
split amongst those folds in a
proportional way. So we don't what we
don't want to happen is um when we split
apart the data. So, let's say we have
let's say we're predicting um spam not
spam. What we don't want to have happen
when we do our splits is we don't want
to have all of the spams end up in one
fold and then every other fold has no
spam, no spam, no spam, no spam, right?
That's not very good. Um because if we
if we train against all these guys, we
have no shot at predicting spam when
they've never seen spam before. So
stratify kffold is is used in
classification
and it's to um it's to make our splits
ensure that they have basically a
balanced number of categories for each
split. Um so that we don't end up with
certain splits with way more spams than
not spams. Um so we we do what's called
stratifying where we make sure the
proportions are balanced across each uh
split. So this is only really useful in
classification, not really necessary in
regression because we're predicting a
value. But if we were predicting a
category,
like in classification like fraud, not
fraud, we don't want to do the split and
have every single fraud example um by
bad luck in our shuffling and split end
up in one split and every other um every
other split has no examples of fraud.
Right? Right. So we want to stratify
this to spread out those um frauds
against all the other splits. Um so uh
again um scikitlearn will take care of
that for you. Um but if you're doing
classification and you have an
imbalanced data set um you you really
want to make sure you stratify kfold. um
imbalanced meaning that you have a a um
different number. Like if you're doing
fraud, not fraud, you have way more not
frauds than frauds. Um when where that
category is imbalanced,
you want to make sure it's balanced
across all your splits.
Um so this is this is useful in
classification only, not really
regression, which is what we're talking
about right now. Um but it's just a
variation on this that ensures when we
do those folds um the data is
distributed evenly amongst those folds
as much as we can. The labels are I
should say.
Okay. So that's stratified kfold. It's
the same same procedure once we have our
splits. It's the same where we do k
minus one of them. We train test on that
last fold um and then rotate through all
the folds and and average all the
metrics. the same exact procedure. It's
just the splitting itself um is going to
be balanced in a stratified kfold.
Okay, so hold out we've already talked
about um is just doing a single train
test split. We've talked about that. One
more variation that is a bit of an
extreme version of Kfold. So it's
actually the same process as Kfold, but
it's an extreme version is if you set K
equal to the number of data points. So
you basically are um this is a really
really extreme kfold where you um
basically are training on all the data.
Um so you're training on all the data
except one point and then you test
against that one point. Um now why would
you ever do this? Um it's mainly so for
this reason here. it's to um maximize
the amount of training data that your
model gets exposed to because instead of
just doing instead of just doing five
splits
um which would be like
you know these four folds are going to
be used and then we um test against one
fold um we're essentially going to use
99% of the data right one point is going
to be left out 99% of the data gets used
to train um and then we're always going
to leave out one point and and the issue
is we're actually going to do that over
and over and over again and rotate that
one point to cover the whole data set.
So we're going to train on 99% leave one
that one point out
and then rotate through every
combination of points until we've left
out every single point and then average
all those together. Um so this is a this
is an extreme kffold. Again the number
of folds is actually equal to the number
of data points in this case. So we have
every point is its own fold and we train
on everything but one test on that one.
This gets you the maximum size of your
training data because you're basically
going to have every point but one used
in the training.
This gets you the maximum size. However,
it gets you the maximum uh expense
especially for large data sets. This is
going to be usually you're not going to
use this um especially for large data
sets because it's just too extreme. It's
going to take you a really long time to
work through every single point being
left out. Um it's just going to take a
while to do.
So for that reason, the leave one out um
that that's why it's called leave one
out because it's you're leaving one out
every single time. Um is rarely used. I
I don't really see it used that often,
but it is an extreme version of K-fold
cross validation.
Okay. But rarely ever actually used. I
think the the ones that get used the
most are definitely the hold out method
with just a regular train test split. Um
and then uh the other one that gets used
quite a bit is is Kfold
or stratified K-fold if you're if you're
doing classification, but certainly
K-fold in the in a regression case.
Okay.
All right. Um, we're going to do an
example with these guys. So, we'll do
that next. Um, with with the different
cross validation techniques. Um, but any
questions on what they are doing
conceptually before we actually do the
code example?
Okay,
very good.
All right, so let's see some examples.
Um let's go into our code and build a
model and do the different cross
validation techniques on it. Um you're
going to see it's actually going to be
really easy to do and we it sounds
complex like doing the kfold and leaving
one out and testing it sounds kind of
complex but I promise you scikitlearn
makes it really easy to do. Um
and so uh we won't need to do too much
besides just use the right uh tools from
scikitlearn. Uh so we're going to we're
going to see that. Um so here we have
some imports. The um primary uh thing
that's a little bit new for us is going
to be these um different kinds of cross
validation techniques. So we have our
kfold, we have our stratified kfold,
leave one out um which are those
different cross validation techniques.
Um these are going to be used in
combination with this cross val score
which is going to keep track of the
different um metrics and then average
them
uh while we do one of these um cross
validation techniques. So this guy gets
used in combination with one of these to
um as as we're going to see in the code
uh to average those metrics. um doing
the different folds, right? Perform
doing performance against the different
folds. Okay. And then of course we need
a model
using linear regression. That's that's
the one we've studied so far. Um and
then we have just a regular metrics if
we want to compute those. Um using maybe
just hold out, right? And hold out um
which which is just a regular train test
split. Um we could use these guys to
evaluate performance.
But in a more sophisticated K-fold style
of cross audition, we're going to use
this to evaluate the the performance.
Okay, let's see.
So, we're going to be working with this
housing with ocean proximity data. Um,
you guys should have this one. Uh, so
you guys should have this one. So, if
you want to follow along and run it
yourself, um, you can load that one in.
Um, I want to make sure that I have it.
Let me pull that one in. So, it should
be this guy.
I'm going to load that in so I can make
sure I run it with you guys.
Um,
so let me run this.
Do you guys have that data?
The housing with ocean proximity?
It's another it's another housing data
set. Um,
but it it's a little bit different than
the ones we've seen before. It has a a
special feature for how close it is to
the ocean, the different locations.
So, it looks kind of like this. If we
load it in and do our head, which is
usually what we do, right, we can see um
we can see that it's got these features.
So, it's got uh uh bedrooms, total
rooms,
um it's got uh median age. Now, this is
this is looks a little strange for total
rooms and um uh bedrooms and population,
etc., but it's um
it's it's got those uh it's got those
because it's representing an entire
neighborhood. So, it's an entire
neighborhood and we're looking at this
um this is actually going to be our
label is this median house value for the
entire neighborhood. So, what's that
median value uh in the neighborhood? And
this is the total number of bedrooms,
total number of rooms, um population,
households. So, how many houses are
there? Um median income. And of course,
these are scaled. So these are um likely
times you know uh thousands
um
but um that's our data. We could
describe it.
So we can see the average age median age
um which sounds a little um weird but
that's it's because again this is the
median of data within a neighborhood. Um
so the average of those is about 28 or
29. Um we have
u
total bedrooms. The we can look at the
min. There's some data that only has
one. So it's likely only one house in
there. Um which is what this represents.
There's only one house. So there there
is some neighborhood that only has one
house. Um, and we see the median, um, we
see the minimum, uh, median house values
there. And then the maximum down here,
um, is a pretty big number.
6,000 households is the largest that we
have in any any one of these
neighborhoods.
Okay. So, just a little bit of
description of the data.
Okay. So then we can run.info. So this
is um let me ask you guys, were you able
to load this? Were you able to run this?
If you're following along, were you able
to
load it and take a look at dot head.
Okay, great. Great.
Okay, so we're able to load that and
then look at dot head. Perfect. Um
okay.
Um and then we run describe which gives
us that uh usual kind of statistical
description. Uh so we can see some
interesting stats about those.
What do you guys notice about the info?
Anything interesting that we see from
there?
Is there any missing data?
Any features that have missing data? Can
we see
object? Yeah, object type usually is
string. If it's an object type, that
usually means string. Python when we
read it into pandas it usually is just a
string.
So that that makes sense like we have
mostly numerical features but then we
have a this ocean proximity which is a
string.
Yeah. Total bedrooms has nles. That's
right. Because you can see here this
does not equal the number of uh rows
that we have. So, this is the number of
rows. There's about 20,000 rows. That's
a good size data set, right? 20,000
rows. That's decent. Um, we're
definitely missing some data here for
sure. Um, we could count how much we're
missing exactly by running this is NATO
sum. Um,
and so we see that total bedrooms is
missing about 200 uh 200 rows are
missing total bedroom uh value.
Okay. And then one thing I wanted to
look at is yes, this is a string. So
what remember what we can do with those?
That's a categorical.
So ocean proximity
is a categorical
string
feature.
So we can take a look at its value
counts, which is usually a good idea to
take a look and see what possible values
that feature could be. So if we look at
our
um what are we calling this? Housing
data.
Housing data
ocean
proximity
dot value counts.
So here's the different types that that
one can be. So there's some
neighborhoods that are less than 1 hour
from the ocean. There's some that are
inland. There's some that are near the
ocean. There's some that are near a bay.
There's even five of them that are on an
island. So, these are the different
values of the ocean proximity. So,
remember, you can always do that. If you
see a string feature, you can always
take a look at what its um categories
are. And it looks like most things are
less than 1 hour from the ocean, but
it's kind of evenly distributed here. Um
otherwise
very few islands.
But as you can imagine like this feature
is probably going to be important for
determining um what the value is, right?
Probably going to be important.
Okay. So, um, we need to deal with these
NLES. If we're going to build a model,
right? So, um, this is all of our
typical data prep. If we want to build a
model, we're going to have to deal with
these NLES. What do you guys think we
should do with the NLES? What would you
what do you think for total bedrooms?
What do you think is a good strategy to
do? Keep in mind, we have 20,000 points,
20,000 rows I should say, and about 200
of them are null.
Right. So about 200 are null. Um so what
do you what do you guys think would be
like a good strategy to deal with those
nles in that case?
average. We can't ignore it because we
can't ignore that column.
We can't ignore the whole column. So,
something needs to go there.
Probably don't want to make it zero.
I think average is a decent average is a
decent idea. Probably don't want to make
it zero because um that would indicate
that there's no bedrooms and yet we
still have a bunch of total rooms. So,
it probably doesn't make sense to do
zero.
Average, I think average could be a
decent one.
Now, in this example, what we're
actually going to do is we're
rows.
We're actually going to drop the rows al
together. Now, why are we doing that?
It's because we have so much data and
only 200 of them are null.
Okay, only 200 of them are null. So,
we're actually just going to drop the
rows. Now, that's a choice.
Um, that's a choice, right? Is that we
could fill in with the average like you
guys are suggesting. What we're actually
going to do is just drop the rows. It It
makes up less. It makes up about 1% of
the whole data. So it's not that much of
it is missing. We can drop those rows.
So that's actually what we're going to
do here is we remove all the roles with
the NLES by doing drop NA. So this just
drops them. So those rows are cut out.
Um, it's arguable that we could replace
it's arguable that we could just replace
it with something and I think you guys
have good thoughts which is the average
a default
um assume total bedrooms. We could we
could try that. Yeah.
Assign a value based on comparable home
value. Yes, you could do that too.
That's a good strategy is to look at the
other rows that are similar to it and
fill in a value. That's absolutely fair.
Um, in this example, we're actually just
going to drop those rows,
but I think that's totally um totally
valid.
This is a choice.
We could fill NA with different values
such as the average,
total bedrooms,
um, derive a value, etc. So, we could
derive something, which I think Brent,
you have a good suggestion. That's a
good suggestion. Um, we could derive
something like that, uh, and fill in the
blank, and that's I think that's totally
valid. Um, we could take the average of
the um bedrooms. Uh, I meant total rooms
here. Sorry, total rooms. Um, we could
fill in we could fill it in with the
total rooms for that category um or for
that row. Um, many options. In this
case, we're actually just going to drop
those rows because they make up such a
small percentage relative to the 20,000
rows that we have. It's about 1%, right?
200 rows is about 1% of 20,000.
So, we're just going to drop them. But
that's a choice. We don't have to drop
them. We could fill in with something.
Um, and if we did that, we would use
fill na rather than drop NA, right?
Uh after dropping the rows, how many? So
it's just so after we drop the rows, um
after we drop the rows, it's just going
to be we still have all our other rows
are intact, right? So if we look at this
now,
we now have um slightly uh slightly less
entries.
So now we have this this many um rather
than rather than this many,
right? We dropped those 200
But they're all filled in. Yeah, they're
So all the other columns are still
filled in. We're just we're we're
cutting out the whole row. So if you
think about our data set, um we have all
these rows and all these columns. What
we're doing is like if there's a null
here, we're just we're just getting rid
of that whole row, right? And so we
still have all the other rows intact.
Uh, we can drop them because we have a
good sample size. Yes,
that's exactly right, Ronald. Yep, we
can drop them because we have we have
20,000 rows and only 200 are missing
values. So, that's totally fine.
Uh, drop a removes all rows that has any
null. Yes, that's true. It it will go
ahead and just drop any row where
there's any null, no matter what column
it's in. Yes,
index. Yeah, the index is not getting
reset. Um that's true. So um what we
what you can always do is you can reset
the index. So, um, if you want to, it's
optional. We we're not really going to
use the index for anything that
important, right? But what we could do
is, uh, reset index.
Uh,
we could do that, right? Which will
reset it.
So now now it gets reset.
But um let me actually I don't I don't
really want to do that. I'm going to
reset this.
Um,
yeah, we could do that.
Okay.
So now importantly there should be uh no
missing data of this of this new one
where we've dropped NAS right. So now
this is good. If you now the reason we
had to do this is because if we try to
build a linear regression and we have
nles in there um the the issue is like
how do you build a model where you have
something like this
and these are null like what do how do
you multiply a number by a null?
Um, we can't really do that, right?
We can't really do that. So, um,
so therefore, uh, we need to get rid of
NLES like the the null's not really
going to work in there. So, uh, we need
to get rid of them for linear regression
to to really have a chance to work,
right? To train it and be able to use
it.
You got to get rid of those nles.
All right,
any questions so far? So, we haven't
done any modeling yet. We're doing some
We're doing some data preparation before
we get to the modeling. And we haven't
done any cross validation yet. We
haven't set that up. We're just doing
our data preparation before we get to
the modeling. Right. So, we've dropped
some NAS. We've checked it. Um, we're
going to do one more prep step, which is
to um change that ocean proximity
feature into something numerical because
again, how do you build a model where
you're inserting a string into those
like beta 1, beta 2, beta 3 times of
features? You can't really do that when
it's a string. Um, so what we're going
to do, and I'm going to get rid of this
because I don't think we really need
that. um is we are going to uh run this
get dummies function which is our um our
get dummies function is our usual one to
uh our git dummies one is our usual one
to um
uh get our one hot encoding. So this is
our uh one hot encoding here.
We now are going to have data that's
like this, right? So we have ocean. So
So by the way, this prefix
um this prefix is OP, which which is
short for ocean proximity, right? So we
have ocean proximity uh less than 1 hour
from the ocean, ocean proximity inland,
ocean proximity island, near bay, near
ocean. So these first five rows are near
the bay. Um so they have a one there and
a zero in the other spots. So this is
good. This one hot encodes that feature
into these numerical uh values,
right?
Were you guys able to run that one? The
get dummies
So the reason that Yeah, that's a great
question. How did it go ocean proximity?
It's because um that is the only uh
string feature we have. That's the only
one we have. So it it's going to look
for any non-numericals and one hot
encode those however many however many
there are. So whatever objects we have
which are strings, it's going to
automatically oneh hot encode those.
Yeah, we could have right we could have
went here and did Right. We could have
done ocean
proximity,
but we only have one of those features.
So, it's just going to do that to the
whole data frame
uh on that one feature. So, what we're
going to do is um go ahead and split it
into an X and a Y um which the X is
always what includes our features. The Y
is what we are trying to predict, which
is the label. Now, um, in order to
separate those out, what we're going to
do is assign X to be the variable that
is, um, our data frame minus this median
house value column. So what this is
doing is um uh it's not permanently
dropping because we're not uh dropping
it in place but it is returning us a
copy of the data frame with the median
house value column left out right it's
dropped. So this is this is uh something
we want to do because that will the rest
of it will contain our features, right?
So um this will temporarily or I should
say return a copy of the DF with um
median house value
dropped,
right? Median house value dropped. Um so
we go ahead and drop that one. Uh now
remember it's not permanent. It's just
giving us uh the remainder of it which
is this housing data dropping this and
it's assigning that to x and then we're
taking the actual median house value
column from the original data and
assigning that to y. So this is going to
be our labels
right. So this is what we are trying to
predict.
Okay, so that is our Y and that's always
how it is. X is our features, Y is our
labels. Um, hopefully that makes sense.
What this is doing is this is going to
get rid of that label column and
everything else will be our features.
And then this will get rid of this will
just assign the label column to Y.
All right. And then what we can do is
pass X and Y into our train test split
function. And this will generate the
hold out set. So if we want to do the
hold out cross validation, this is how
we would do it is we would split the
data into X-ray, X test, Y train, Y test
um using train test split. So this is
what we did last time. This would be
this would be for hold out cross
validation,
right? where we are uh uh just have that
one one set for testing and one set for
uh one set for training one test one set
for testing I should say right so this
is pretty standard train test split um
we pass in that x we pass in the y we
use a 30% test size is pretty standard
and random state so that we get the
consistent shuffling if we were to run
this multiple times um we we get that uh
consistent randomization
Okay,
so we have that and so now our X train
is a percentage um of the data frame of
the 20,000 uh rows and the X test is uh
30% of that. So it's only about 6,000
rows which is what um the shape of that
is.
Yeah. X so X is our features. So we're
we're putting all of our data in that is
our features into X. And so the the um
most efficient way of doing that is um
the most efficient way of doing that is
to
uh just take our data and drop the
median house value column because that's
our label column. So we just remove
that. The rest of the data is our
features. So that's what that's what
this X is, right? It's all of our
feature data. All of our columns that is
not the label column essentially is what
that's doing. And then Y is our label
column from our original data.
Right? Y is our label column. And so
this this will um contain all of our
labels which is the median house value.
X X X contains every column but the one
we're gonna so we we ultimately decide
that but X contains um X is everything
that is not our dependent variable which
is what we're predicting. So we're
removing what we are trying to predict
from X. X should be everything else.
That's always how it's going to be. X is
X is always going to be all of those
independent variables that we're using
to predict the median house value. So we
are going to predict the median house
value. We need to remove it from X.
So we're we're taking everything but
that column.
So it's the whole data frame. It's the
whole data frame minus this one column
with just the dependent variable. Right.
Exactly right. Removing the dependent
variable and keeping all the
independence. That's exactly right.
Exactly right. So think about it in
terms of the model. Let's go back to the
features. Right. Think about it in terms
of the model. We are trying to predict
this this value. We're building a model
to try to predict this. So we are going
to make sure X is everything but this
right. So this is actually just Y.
That's our label. That's our dependent
variable. Right? That's Y. Everything
else is belongs to X. Everything else
belongs to X including all of these.
Right? We choose this one to be Y
because we're building a model to
predict that. That's our label.
All right. So, we have our we use X and
Y to do our train test split. So, we
have our our training features and our
test features and then our training
label and test labels here. Um, pretty
standard there.
Um, okay. So, this is what's new is if
we want to do K-fold uh validation, what
we're going to do is create a kfold
object. So, we have this kffold from
scikitlearn that we already imported. we
are going to create a kfold um where we
are going to specify how many folds we
want. So that is the in uh inslits
parameter as this says um this is going
to be uh uh in this case we're going to
do 10 folds. That's pretty standard. So
I think the typical number of folds that
I've seen and I've worked with in my in
my career is usually five or 10.
Five or 10 folds is the standard.
Okay. So, we're doing 10 folds in this
case and we're setting a random state
because we're going to do shuffling. So,
in order to produce those folds, we're
going to shuffle the data first and then
split it into five folds, right? So,
this this kfold object is going to
manage creating these splits for us,
right? These even splits. I know I I
didn't draw it even, but um it's going
to manage these five folds for us and
it's going to shuffle the data and
assign them to these different folds and
we're and then what we're going to do is
use those to do our training. We're
going to execute the cross validation
using this k-fold object.
Okay, so we create the kffold
um we initialize our model as well. So,
of course, in order to train something
uh in the Kfolds, we're going to need a
model. In this case, we're using linear
regression, right? Which is which is the
model we've been studying so far. So,
you have a linear regression. Um now,
look how easy it's going to be in order
to execute cross validation. All we need
to do is um all we need to do is create
a crossfile score function.
um or I should say use the cross val
score function from scikitlearn. So we
use that with the model we want to
train. So our model goes first. So
that's the linear regression object.
Then our data. So our extra our features
and our label for our training.
And then um let me skip over this for a
second. I'll explain what this is in a
second. Um but then we are using uh the
cross validation technique is our kfold.
So this is where our k-fold object goes
in the CV parameter which is cross
validation. So what cross validation
strategy are you using? We're using
kfold and the kfold we're using is this
one we defined up here KF. So we're
putting that right here for this. And
then um in jobs um allows us to
parallelize this. So if we set it to
negative one, that's the that that's the
default. Um it will do it will actually
train across the different combinations
in parallel. Um which speeds it up. So
you want to you want to keep this to
negative one if you can. So um now let
me describe the scoring. So what this
means is we put in our metric here. Um
and so you can put mean absolute error,
you can put in mean squared error. Um
those are the two that we can use. And
um the reason we it has a negative in
front of it is because we want to find
the one that has the lowest score.
That's going to be our best model is the
one that has the lowest score. So we
take the absolute value.
I'm sorry. we take the abs the the the
metric and we take the negative of it um
because the highest scoring one is going
to be the closest to zero. Um so it's
just a we use the we use the negative of
the of the metric um because on the
number line like the the highest um
scoring one should be the least um or I
should say the maximum negative that we
can get. that's going to be closest to
this to zero. So if here's zero, this
would be like -1 is better than -10,
right? So something that scores um the
maximum negative uh absolute error would
be closest to zero
and something that has more is going to
be on this side.
So this is only the reason we need this
is only just to keep track of the scores
of each individual um fold. Okay.
So the one so the reason we can do that
is at the end we can kind of see which
which combination performed the best. Um
it's going to be the one that has the
highest uh highest value of the negative
which is closest to zero.
That's just a convention.
Yeah, it's just because um it's because
the cross validation is looking to
maximize the metric. So, whatever has
the best score
um whatever has the best score is
considered the best uh performance. Um
but we are using uh something where
lower is better. So we we take the
negative and like the the highest
negative would be closest to zero,
right? The highest negative is going to
be closest to zero.
So that so it's it's just because like
we want the lower score to be the best.
The lowest score should be the best.
So we take the negative of it. Um and so
something that is more negative is going
to be worse. Yeah, that's the reason.
So something that's down this way is
going to be worse. Okay, so it runs this
and what you can see is if we actually
print this out, if we print out our
k-fold scores, what we should get is 10
different scores.
And you can see um we have 10 different
uh scores here which are all negative
because we're taking the negative of the
absolute of the mean absolute error. Um
so what we would be looking for here is
um we want to take the average of these
scores but take the absolute value of
them to get the best performance. So
this is capturing like this is the score
on the first fold combination. This is
the score on the second fold
combination. This is the score on the
third fold combination and on and on and
on. And these are the absolute errors.
Okay, these are the absolute errors. Um
so if we take a look at computing the uh
average, which by the way, we don't need
this import because we're using the
numpy average. So that's fine. um we can
take the absolute value of those um and
take a look at the average MSE
or sorry MAE. Now I want you to think
about this this uh average performance.
So this is our performance right here on
the cross validation.
This is our average
MAE
across all of our fold combinations. So
that's a that's an indicator of our
performance, right? um for the cross
validation.
Now, what are the units of our original
uh the original median value? They're
already in the thousands, right? So, if
we go to that feature, they're already
in these hundreds of thousands. So, this
is not a very good error. It's it's kind
of high, right? because it's in this is
49,000.
Um that's that's how far away we are in
absolute value on average is $49,000 um
dollars on the median value. That's not
very good. So this score
this score is
um not very good. So this model is not
performing that well and we can see that
by comparing this error to our actual uh
data. So this is right around 50,000
and our median uh house values are in
the hundreds of thousands. So on average
we're 50,000 off when we make a
prediction. That's a significant amount,
right? It's a significant amount on
average um when our when our data is in
about the hundreds of thousands here.
So we are um we have a significant
amount of error 50,000 relative to the h
to our units that our our data is in.
Right? Um so this score is not very
good. Um
and so we see that from the cross
validation. So look how easy the cross
val is. Again we just do cross val
score. We put in our model. We put in
our data. We put in our cross validation
uh strategy here which is k-fold and we
can generate these metrics across all
the fold combinations. So it's this
function is taking care of rotating
those and doing every combo with just
the 10 different combinations here of
the of the folds.
10 different instances where you have
you know 10 different folds are the ones
that are left out for evaluation.
Um so it's managing that for us using
this data right using this training data
here. Um and we uh we generate these um
generate these scores.
Okay. So that's kfold. It's not hard to
do. All you have to do is um just use a
cross file score. And we could change
this to mean squared error. That's you
know we could do that too. That'd be
pretty easy. Um, so that'd be no issue.
We just happen to be using the absolute
error here. Of course, we could use
squared error.
Were you guys able to get this to run
kfold scores?
It produces an array of 10 10 different
scores, which should make sense because
those are these are the um we're
splitting our data into 10 different
folds,
right?
10 different folds and leaving one out
to do our evaluation on. So the one that
gets left out every time is what's
producing these scores. So it's 10
different ones get left out when we
rotate through all the combinations.
And so we average these scores
and we get this amount of we get about
50,000 in error on average.
Um, what do you think would be what do
you think would be acceptable? So, if
our if we're predicting the price, like
if we're a real estate agent and we're
predicting these prices and they
typically are
Yeah, close to zero would be great.
That'd be fantastic. Closer to zero
would be better. The average is um
206,000.
So 50,000 is a decent percentage of
that. Um so you know you can compute it
as a percentage right? So 50,000 is a
decent percentage of that. Um probably
you want this to be less than 20,000
would be about 10% error. 20,000
right? So maybe like 30,000 somewhere in
there.
Yeah. 10% would be 5% error. 10,000
would be 5% error. That's true. That's
true. So that would be that would be
much better. So being closer to zero,
like the smaller the better, of course.
Of course. Um but yeah, I would say an
acceptable percentage of error is
probably 20%.
Probably 20%, which would be um like
40,000 or less would probably be
acceptable.
Usually when we usually when you build
models um 80% accuracy is usually uh
considered decent.
Usually considered decent
80%. So I'd say 40,000 or less would be
kind of ideal.
Does that make sense to answer the
question?
That's a good question. What value is
acceptable? I think probably less than
40,000 would be ideal. That's right
around 20% error.
All right, so that's kfold. Um let's do
just a regular hold out now. So this is
just using our training and test data.
Um doing model.fit and calculating an
MSE on the test data. So this is this is
just the um hold out strategy here where
we just have um this is less robust but
it's a lot quicker to do and easier to
set up. Right? So um this is using the
hold out strategy. So just a regular
um train test split.
Are we going to rebuild the model? No,
not necessarily. There's some things we
could do most likely. And like one thing
we did not do was scale our features.
Remember I said that's a pretty
important thing to do is to scale our
features. We did not do that. So that
would be an enhancement to this that
we're going to So I I actually do think
we'll do that later. Yes. So I think we
will actually do that now that I'm
thinking about it. Yes. One of the
things we can do is scale these features
using like a minmax scaler, a standard
scaler. that's actually going to help us
um that's going to help us do better
predictions.
So that that's one thing we could do. Um
but yeah, we will we'll try to see if we
can get better.
It should help it. Yeah, usually you
want to scale you want to scale the
data. That's something we didn't do in
our preparation step. We did a lot of
the things we should do. We removed nles
and we did one hot encoding to the
proximity feature like this one. Um
those are good to do but we didn't scale
any of these other we didn't scale any
of the features right we didn't scale
any of them. Um it you it will have an
effect. It usually when we scale it
it'll be a better model.
It'll it'll learn a little bit better if
we can scale the data. Um so that way
like these
um like ages aren't you know drastically
different than like in scale then total
bedrooms or income
uh those kind of things. So we usually
want these to be in a similar scale
range.
So we'll we will I think we'll scale
them coming up in a bit and it should
help the model.
We've talked about that before, right?
scaling usually is a good idea to do
when you're prepping your data for
modeling.
No, you want to you want to scale your
test data as well. You're going to do
both. You're going to scale your
training data, you're going to scale it.
So, that's actually a good point you
bring up is any transformations you do
on your training to build your model,
you should also do on your test set so
you get an applesto apples comparison.
You should always do the same
transformations.
Yes.
Would scaling data impact K? Yeah, it
could it could make it better. It could
uh yeah, it should impact it. We should
get a better model. So when we do the
different folds, we'll get different
we'll get better scores. Yeah, it it
will impact
uh yeah, if they're so that's a good
point. If they're going to use our
model, then yes, they have to scale the
data as well. If they're gonna if we
build the model on the assumption that
the input is scaled, then yes, they have
to also scale their data when they're
using it with our model. That's true.
I mean, not really. I'll show you why.
There's something that's actually going
to make it easier um that that will
automate doing the scaling for them. So,
they don't they don't have to do the
scaling manually. it'll just it'll
happen automatically when they use the
model. I'm going to show you something
that's going to automate that which is
going to be called a pipeline.
So that part will be automated and they
won't have to do that. So it won't be
heavy on the user. No, in theory it is,
but
has a really helpful tool to make it
easy to do that. So I'm going to I'm
going to show us that um later on in the
notebook.
No, the data data is not for a single
house. It's for like a neighborhood. So
there's a certain number of households
in the neighborhood and this is the
we're predicting the median house value
of that neighborhood.
Yeah. So there's a there's certain
number of households. There's there's
like an a median income, a population,
certain number of people that live
there. Um proximity generally of where
that location is. It also has a latitude
and longitude.
So,
and a median age in that neighborhood.
So, yeah, it's not just a single house.
Okay, let's go back to this was the hold
out strategy. So, this is a lot simpler.
This is just model.fit, right? This is
just model.fit on the training uh data.
And then we um can predict on the test
features and generate test predictions.
And then we can compute our error on
those um we can compute our error
amongst the test predictions and our
test uh label. So that's our useful mean
squared error function, right? To to
compute the MSE. Um let's see what the
MSE is. So MSE is right here.
Um now what we can do is we can take the
MSE
and we can take the square root of it.
So let's actually do that. Let's um do
MP. square root of the
um test
MSE
and we get um 67 we get 67,000.
So that's pretty high on this. So when
we just now look at the difference of
that, right? When we just do a train
test split, um
when we just do a train test split, we
get a worse score because it's not as
it's not as robust, right? We're not
showing that to many of the other uh
folds. So we get a lot more error this
way on the test data.
So this is um actually worse performance
just doing the train test split.
This is a really higher.
Yeah, we can. We can. I'm going to I'm
going to show us how to how the scaling
will be done automatically. Yes, we can.
Um there's there's a really easy tool to
do that will scale it automatically.
It's going to be later in this notebook.
I'll show us it.
All right. So, just to recap this, this
is fitting the model.
This is fitting the model. This is
making the predictions, right?
Model.predict.
So, this is making the predictions. And
then this is calculating the error, the
mean squared error, which is looking at
our test labels versus our test
predictions, right? And this is
computing the distance, the average
distance away from these values to these
values,
right?
And then we can also compute the R squar
R R 2 and we see that it's not a very
good R squared. 65 uh is not a very
great model
um because it closer to one would be
better. So this is still this is not
very good.
We know that we knew that from the cross
file score. But this is just doing um
this is just doing a hold out uh where
we do a train and test split, right? So
it's a little bit simpler, but it's not
quite as robust. Um
it's not quite as robust as the cross
valve, but it works. Um it's, you know,
we can do hold out. Um,
we can do hold out uh to to quickly
evaluate a model and see if we need to
make any adjustments.
It's a little bit quicker to run.
Okay. And any questions on it? Does it
make sense what we're doing here?
Model.fit to train it predict to get our
predictions. Um, this is pretty
standard, right? To train is the
model.fit it and then to use the model
to predict. We predict on the test
features.
Um, so this is passing on on all of our
features into this model to generate
predictions for every row. That's
something I also want to point out that
may be a little bit confusing is this is
a data frame. So we're passing in a
bunch of rows of features with columns,
right? So um, we're passing in a bunch
of data that looks like this. And what
we're doing is essentially making a
prediction for every row. So this will
generate a prediction. This row will
generate a prediction. This row will
generate a prediction. And on and on and
on. So this this predict will predict
for every row. And so we end up with
this collection of predictions here for
each row. And we're comparing those to
the labels that we have for those rows
from our from our supervised learning,
right? From our data set.
So that's truly supervised learning,
right? We have the examples and we're
comparing those to what our model is
predicting to to get our performance.
All right.
So let's uh let's try the other just so
you can see it. The leave one out. Now
the leave one out cross validation is
going to actually work the same way
where we put in the leave one out um
strategy inside of the cross file score.
Now here we don't need to specify how
many folds there are because we know how
many there going to be. It's going to be
the number of data points, right? So
which is actually going to be quite
large because there's 20,000 rows. So
this is going to be extremely
uh extremely um intensive because we are
doing um you know 20,000 examples and
leaving one example out to be our
validation and then um doing that across
every 20,000 uh examples.
So we could do it though just to see how
it works. Um we have this again leave
one out. We generate our cross file
score from our model our data and then
same scoring that we had before and but
this time we change our cross file to be
instead of our k-fold object we have our
leave one out object which is this
um and then we could run this. We can
compute our average uh across the all
the folds. Now this is going to be a lot
bigger of an array. It's going to be a
20,000 size array and we're going to
compute the average across it.
So, let's do that. It's going to take a
moment because there's lots. So, if you
notice it when you run, it's going to
take a little bit of time to run because
it's running across all 20,000 examples
and leaving one out. So, you have 20,000
and then one left out to uh test
against. So, it's quite intensive. You
can see it's taking a lot more time.
It's still running. It's taking a while.
Okay, just let that run. Still running.
So, if you guys try running this, it's
going to take a little bit of time.
Hopefully that makes sense why it's
taking so long, right? It's because it's
instead of doing 10 folds, it's it's
putting every data point but one is the
training set and then iterating through
all 20,000 points.
This takes a while to do.
Let's see what our
RAM our memory is a little increased.
Okay,
still running. That's okay. I'll let it
run.
Come back when it's finished.
Yeah, exactly. This is a This is for
This is giving us a performance
evaluation. This is like the average
error across all of our uh different
folds. Um now this is the extreme case
where we have the number of folds equals
the number of points.
Right? So it's an extreme case but yes
it's just like kfold. It's giving us
that performance estimate.
Okay. It's about the same right. This is
still around 50,000.
Not much difference, right? Still right
around there. But look how much longer
it took. That took 2 minutes to run. The
other one was pretty instant, right? So
this this took about 2 minutes to run.
So um definitely uh
yeah, definitely don't want to run this
uh too often. I think that it's
generally preferred to do k-fold if
you're going to do cross validation.
Generally want to do k-fold or just the
regular hold out train test split. Uh
generally better than doing leave one
out. It's just going to take too long
and um it results in about the same kind
of score as the kfold.
Okay,
any questions about um the cross
validation that we just did.
Okay,
good. And as it says here that the
stratified kfold is usually used for
classification. Again, we're not doing
classification yet. That's in going to
be in lesson four. So, we don't need to
worry too much about that. Just for
regression, um regular kfold is
preferred, right? Because we don't need
to um worry about distributing
categories amongst our folds uh in any
regression problems.
And as we see the error is kind of high.
Um there's going to be some things we
can do to improve that which will be uh
later on we'll learn about some more
advanced models. This signals that the
performance is bad. We probably need a
more complex model. Um one thing we
could try before we try a complex model
is to do scaling. We will try to do
scaling. I'm going to show us how we can
do that coming up um in a in a nice
streamlined fashion. Um, but uh outside
of that, if we still had bad
performance, we would likely need to use
a more advanced model. And we'll learn
about more advanced models uh in the
next lesson. And what's great is some of
those advanced models can actually be
used for regression. So they have
variations that can be used for both
classification and regression, which is
pretty cool. So I'll point those out
when we get to them. Um, okay.
So what I want to talk about now is a
way we can combat overfitting. So if we
have overfitting which remember that is
the case where the uh the we see good
performance on the training data but
then um it doesn't generalize over to
the test data. We get poor performance
on the test data. Um there's there's a
drop off there. Um that would signal
overfitting.
overfitting
and one way of um combating overfitting
is to do something called regularization
which we're going to talk about next. So
the key idea in regularization
is to
change our uh the change the way we
train. Essentially, what we're going to
do is modify our training
uh error function or sometimes called
the objective function or loss function.
We're going to change that to add a
penalty to penalize excessive complex
complexity. Essentially the the way that
we're going to penalize is by making
sure the size of the coefficients
doesn't grow too much which should
mitigate overfitting because remember in
linear regression what we are learning
are the coefficients right we're
learning the beta 0 the beta 1 the beta
2 and on and on however many betas there
are beta n we're learning all of those
guys um through the regression error
function we're trying to minimize that
error function. That's how it trains. We
talked about that on Monday.
Um so what we're going to do is um
basically penalize the these guys
growing too big and making sure we kind
of keep them small so that no one
coefficient has a dominant uh effect on
the model. And this should help with
overfitting and complexity. should make
the model simpler because all the
coefficients are going to be encouraged
to be smaller. They're not going to grow
too big. Um and this this has the effect
of making the model so basically make
the model simpler.
Make the model simpler is what these
regularization techniques are
essentially trying to achieve is is
remove complexity, make them a little
bit simpler, make these coefficients
smaller so that you can generalize a bit
better and and prevent overfitting. So
we want to prevent
uh overfitting,
right, is what we want to do. Um so
there's going to be a penalty and I'll
show you where that penalty gets added
and kind of what it looks like.
Um but uh to control the level of that
penalty we are actually going to
introduce another parameter to our model
um called alpha.
Alpha is going to scale the penalty. So
if alpha is really high that imposes a
stronger penalty on the coefficients um
which will make the model a lot simpler.
So the higher the alpha the simpler the
model we will get and we the the risk
with that is we actually underfit. So if
alpha is too big we may underfit the
training data
um a bit too much because it will make
the model way too simple. Um and I again
I'll show you what this means
mathematically in a moment. Um but on
the other hand if we have a lower alpha
this will have a lower penalty. it's a
weaker penalty term and that'll lead to
a model that is um a bit more complex.
Um which could um risk some level of
overfitting. Um so there's so there's
still the risk of overfitting if you
have a low alpha. And of course if alpha
goes all the way to zero, there's no
penalty at all. So you're back to your
original linear regression. Um which
could risk a lot of overfitting,
right? So you you generally want to pick
an alpha um effectively and actually
we're going to see h what's the best way
to pick alpha. Um we're actually going
to learn how to do that. I'm going to
show us how doing some tuning techniques
to pick what alpha should be. Um but um
a a pretty industry standard alpha that
most people default to is alpha equals
to one. So just just one which signals
that there should be some penalty. we
just have alpha equal to one is a
standard penalty. We don't want it to be
too high. We don't want it to be too
low. Like we don't want it to be a
fraction. Um but a penalty of one is
usually uh good enough.
Okay, I'm going to show you where that
comes into play in a moment.
Um but the whole purpose of doing this
is to mitigate overfitting, right? Um
that's what and and doing this penalty
is is called regularization. So adding
so going beyond just regular linear
regression adding this extra penalty to
to the training process um to penalize
large weights large coefficients
um is known as regularization.
Okay. Um and there's two common
penalties that are added. Um so there's
actually two different variations on the
penalty. Um we're going to study both of
them and um they're they're known as
lasso. So if you take linear regression
and add a particular type of penalty,
it's known as lasso. If you add another
type of penalty, it's known as ridge
regression. We're going to study both of
those and what their differences are.
But these are the primary two
uh regularization tech uh models that
are used um to take a regular both of
these take regular linear regression and
just modify the training process a
little bit in different ways. Two
different ways. um using that alpha
um to penalize the terms in slightly
different mathematical ways. So we're
going to learn about these two guys.
Lasso regression there. Both of these
are just offshoots of linear regression.
So underlying model is still linear
regression. It just adds different types
of penalties to the training process.
So both of these are still in the family
of linear regression. In fact, in um in
scikitlearn, they both come from they
both are still from the linear model
family in inside of the linear model
module, which is where linear regression
comes from. So there's still linear
regression. They just have different
styles of penalties added to them. Um
which we're going to see.
Okay, so just to recap that
regularization is the process of adding
a penalty to the training to discourage
complexity. In this case, we're going to
discourage large coefficients.
And um this should help prevent
overfitting.
And so uh these are going to lead us to
two different offshoots of linear
regression that have two different
penalties.
lasso and ridge regression, which we're
going to uh study next,
but they they function the same way as
linear regression. They will just have
different penalty terms added onto their
training process um to discourage
uh discourage um again those large
weights.
Okay, any questions about regularization
before we first look at our we're going
to look at our first uh variation on on
our first regularization technique which
is going to be called lasso regression.
Okay, let's look at lasso regression. So
what is lasso regression? It's actually
lasso is short for least absolute
shrinkage and selection operator
regression. Um and this will function by
adding a particular penalty to the
linear regression model. So again it's
based on linear regression. That's the
underlying model. It's just that during
the training process we are going to um
add a penalty which has the effect of
shrinkage of the weights. That's why
it's called shrinkage. It encourages
smaller weights through that penalty and
it also will shrink some of them so much
that they'll become zero and so it has
has an effect of kind of selection which
means that some of them get wiped out to
zero
and this means that whatever is left
over is kind of what's selected as our
features because the other ones will
have zero weight applied to them. So
this penalty will really favor small
weights um and penalize really large
weights. In fact, it will favor small
weights so much that some of them will
actually um be shrunk to zero um during
the training process. And the ones that
are left over are the ones that um are
the ones that are what we call selected
because they are the ones that remain in
in the training um after the other ones
get uh coefficients of zero. Um now when
you make some of the coefficients zero,
you are inherently making the model
simpler, right? There's less features
involved in the prediction that or less
features that have an effect on the
prediction. So this definitely makes the
model simpler. This lasso, this
shrinkage and selection uh process makes
makes the model simpler for sure. Um
and this is supposed to reduce
overfitting, right? If you make the
model simpler, it's not as complex. It
has less of a chance of memorizing
training data and not generalizing over
to test data. So our whole goal with uh
regularization is to make our model
better at generalization right over to
test data from the original training
data.
Um so how does this happen? We have to
go back to the
uh training process. If you guys
remember I I wrote out this equation a
little bit earlier which is the
distance. This is the sum of squared
distance between our labels and our
prediction.
This is basically the mean squared error
uh calculation that we're trying to
reduce when we build our model using the
training data. Um so this is just in
standard linear regression. This is the
um uh sum of squares uh distance right
so this is this is what the model is
trying to minimize when it learns these
coefficients.
So when it learns these coefficients,
it's trying to minimize this guy.
Minimize. It's trying to find the betas
that minimize this quantity.
Mathematically, that's what it's doing.
Um, and there's there's a algorithm that
will discover what the best betas are
that actually minimize this. That gives
us the line of best fit, right? That's
what we've been talking about for
regression.
Now in regularization
here's by the way here is that same
thing but we've just inserted our model
for the predictions. This is our model
just a fancy way of writing down our
model right it's the beta 0 plus all of
these betas. So beta 1 x1 plus beta 2 x2
plus on and on and on. Right? That's
that's what this uh means if you're
unfamiliar with the sigma notation. It
just means sum. So it's the sum of all
these guys or this term. Um so this is
this here is just a regular linear
regression
uh training regular linear regression
training. So we the training process
solves for these parameters right it
solves for these weights. We discover
what those are by minimizing this
quantity. That's the whole training
process. Um but when we do lasso
we add a penalty which is this
here is our penalty.
So basically um we take our linear
regression training which is this and we
add on a penalty which is this and you
can see exactly what this penalty when
when you minimize this penalty it's when
these weights are small. So this
encourages
So minimizing this quantity encourages
small weights
encourages small betas
beta I
right you or in this case beta j sorry
this encourages small beta js uh because
we want this thing to be minimized
minimize
So um what's going to make this minimal
is of course the line of best fit and
small weights right are going to make
are going to bring this error down the
most.
So um and here's our alpha right here's
our alpha. So you can encourage a higher
penalty with a larger alpha or a lower
penalty. If alpha equals zero
what happens to that term? It just goes
away. So if alpha equals zero, there's
no penalty and we're back to uh we're
back to regular linear regression.
We just have regular linear regression
because we have no penalty at that point
when alpha equals zero. So the smaller
alpha is, the less penalty we're
enforcing in in the regularization.
Okay.
Now what happens is in reality when you
train with lasso. So this is lasso is
this particular penalty. This is called
the lasso penalty
or sometimes um people call this the L1
penalty.
Um L1 just comes from the fact that this
is the first power or absolute value. Um
so it's not a squared penalty. It's a
single uh single power penalty
um there. But when you add this lasso
penalty, what can happen is it c it does
because the because you're minimizing
this, it does encourage some of these
weights to become zero.
So some if you're really trying to get
the lowest quantity of this,
the lower the better.
What makes this thing lower is of course
if some of these go away. If some of
these go to zero then that of course
will lower this as much as we as much as
possible. Right? So what happens during
the training is some of these
coefficients actually they're encouraged
to be small because of this penalty. But
some of them will actually becomes will
actually become zero um in order to get
the best model the best fit. Some of
these will actually get so small that
they'll basically become zero. And that
means that that that feature basically
has no effect anymore. It's it's been
the model has been simplified, right?
That feature no longer really has an
effect.
So just to call out the alpha again um
if alpha zero some co uh basically you
have your linear regression you're back
to linear regression because alpha 0 is
just wiping this out and you're back to
linear regression. Um if alpha is
infinity now if alpha is infinity that's
an extreme. So if alpha is infinity the
only way to make this minimize is if all
your coefficients are zero. If every
beta is zero, then this will lower the
the error as as much as possible. So you
basically have no model. So if all
coefficients are zero, you have no model
and that's useless. So you don't want
your penalty, you don't want your alpha
to be huge is what this is saying. You
also don't want your alpha to be small.
You're basically back to linear
regression. So you want something in
between. Um and the typical typical
value is alpha equals 1.
typical is alpha equals 1
to have some level of penalty there. So
just a regular kind of regular penalty
term.
But we are actually going to have a way
to test and evaluate which alphas are
the best.
Um,
basically you can, yeah, you can have a
you can have a penalty that's close to
zero. You can get rid of this if just a
regular linear regression performs
pretty well. You can basically have no
penalty in that case.
Yeah. So near zero or like it could be
that adding a little bit of penalty
actually helps the overfitting and it
could be really small. One thing that
we're basically going to do is have a
strategy to try out different alphas,
try different alphas
and evaluate performance
and then we can decide which. So that's
what we're going to do is have a
strategy to just plug in different
alphas, generate the like train the
model, and then see what its performance
is and see if those alphas are good.
What what which alpha is the best? We
can evaluate that
because we can train the model and see
what it performance is,
right?
Yeah. Yeah. So we'll do that. We'll
practice that.
Okay, great. Any other questions about
this lasso regression? So, remember this
is linear regression here. This is the
this is how you're training to find the
betas in linear regression. So, this is
just linear regression uh um training
function there.
We're adding a penalty which is this is
the lasso penalty
lasso penalty there right we're adding
that this is known as regularization
and the goal of regularization is to
prevent overfitting so you add a penalty
here this makes the model simpler which
prevents overfitting it helps you
generalize better when it's simpler Any
questions conceptually on this? We're
going to do a code example with it
coming up, but any questions on this?
Uh yeah, you you so that's the thing,
Ronald, is you may be willing to
sacrifice some accuracy in order to
generalize to unseen data because
remember that's what we're really trying
to get after is we may be willing to
sacrifice some accuracy on this training
data in order to have it perform better
on the test data, right? We may be
willing to do that. That's a willing
that's an okay sacrifice
as long like if if it generalizes
better. That's what we want. That's what
we're trying to do here is add a
penalty, make the model simpler and help
it generalize better to new and unseen
data. Right?
That's that picture I've been using with
the with the um train and test split.
Where is the square?
So in the model there's no square. So
remember the model is the model is this
um equation uh that has no squares in
it. Right? It's beta 0 plus beta 1 x1
plus beta 2 x2 plus beta n xn.
That's the that's the linear regression
model. This is the now this this is the
model but this is the equation that
helps us train and find the betas. This
is how this is what we find the betas
with. So we'll continue. Um we were
talking about the lasso regression which
uh adds it takes linear regression right
which is this optimization and adds in a
penalty um scaled by the alpha. Um, and
what that does in order to minimize this
whole thing, it encourages these to be
small uh as possible. Um, which makes
the model simpler, right? The weights
don't get overly big and complex. Um,
they they tend to stay small. In fact,
some of them can even go all the way to
zero. Um, which makes the model even
more simpler,
right? Um, so let's practice uh using it
in code. It's actually really easy to
use. It's going to be essentially the
same uh style and and code as linear
regression except we are um just going
to have to uh put in our alpha parameter
um when we use the lasso. So here we are
um from the linear model family right
which makes sense. It's a linear
regression offshoot that has this
penalty in it during the training. um we
are grabbing our lasso regression. Um it
also has a version of the lasso that
we're going to take a look at that is
used for cross validation which is
really um convenient as well. So it has
a cross validation lasso which is a
really convenient um combination of
basically cross val score and lasso um
all in one. So it actually is really
nice to use that way. Um so we'll take a
look at that example. Um, but we are
importing it. The main thing is going to
be the lasso model here. Um, we're going
to be using a different data set for
this one. So, not the ocean uh data, but
this hitters data, which is a baseball
data set. Um, so it has 322 rows um with
20 different columns and it looks like
this. So, you want to download that one.
Um, hopefully you guys have access to
that one.
Um,
so I will upload it into
this.
So give me a moment.
There's that. And then we can run this.
Okay. So we are displaying the data and
so it has um the the hitters names and
then it has a bunch of different
statistics. These are all baseball
statistics.
Um, if you're unfamiliar with with them,
that's okay. It's not a big deal. Um,
but just different baseball stats here.
Okay. Were you guys able to load that?
Um, if you're following along, were you
able to load that? You should have
access to this data. The hitters CSV.
This is the one we're going to use for
the lasso model
to build a lasso model.
Yeah.
Okay. Able to load that one. Perfect.
Okay. So, able to load that one. Um, and
we take a look at the the head. Um, so
we're actually going to uh drop this
unnamed column because we don't care
about their name. it's actually just the
batter's name, which is not going to be
useful in modeling. Um, so and remember
that's generally true like an ID, a user
ID, like a customer ID, a name, that's
usually not going to be useful in any
kind of modeling. So we're actually just
going to drop that uh column and we're
going to do it in place.
And access equals 1 means we're dropping
that column. Um, so we're going to drop
that and we should no longer have that
column. And we have all of these guys
now. So you want to run that. This will
drop that. Um this will drop drops the
column in place.
Um and now we can see we have uh all we
have this data where um we have this
data where it's now removed. So this
that column is now gone and now we have
these guys. Um, do you notice anything
about this
from the info?
Looks like we have a couple categorical
features, a few of them, league and
division
and new league. What do you notice about
this?
Nolles. Yep. So, there's definitely some
missing data there um that we're going
to have to deal with.
So, it looks like there are uh there are
59.
Um there are 59. Now, we could we the
alternative to doing that is we could uh
we could just use our usual code where
we do df.is is uh is null.
Um and then we do uh dot sum to total
those up across our different columns.
And we can see that uh we have 59 of
those in this salary column. That's this
is the standard way of doing that,
right?
Standard way of doing that. And we have
so we have 59 of those
59 of those. So, we have to deal with
it. Any ideas on how to deal with it?
Any ideas on how to deal with it? This
is now This is 59 out of 300.
So,
what do you guys think about that? It's
a little bit different than 200 out of
20,000. A little bit different. We have
We have about 60 out of 300.
There's a decent amount.
Any ideas on how to handle this one?
Replace. Yep, we should replace. What do
you think we should replace with?
It's a float. It's a floating point uh
value.
By the way, something unique about this
that's a little different than usual,
too, is that the uh this is actually the
column we're going to use as our label.
So, we're actually going to predict the
salary based on the uh based on the um
rest of the features. So, we definitely
need to fill in these nles, right?
because they're actually going to be the
labels
and we're missing some labels uh in our
data. We we definitely need to fill them
in. Yeah. So, we're going to replace
them.
All right. So, we'll we will replace
them down below. That's going to be
coming up. Uh we'll come back and
replace them. um before we replace them,
we're actually going to get our uh one
hot encodings for those three different
um features we have. Um so we do uh get
dummies with this. Now um of course we
don't need to do this if we just so this
code we don't need to do if we just pass
in the dtype here
um which is uh then we don't need to do
this. So we can comment this out.
Um so now what I want you guys to notice
is this is the alternative to what we
did before where we are purposely just
doing these columns not the whole data
frame but just doing these columns and
then we can um concatenate those these
one hot encodings. We're going to
concatenate back to the data frame.
Right? So if we do our dummies and then
do dummies.info info. Um, we can see
that we end up with six new columns. And
in fact, we can do dummies.head
and take a look at what those are.
Right? So, these are league A, league,
uh, N, division E, W, division W, new
league A, new league N.
Okay.
So, um these are uh these are our one
hot encodings for these three different
features which are strings, right? So,
those those features were strings. If
you go back up, those were our only
string features we had. So, we've one
hot encoded those so we can use them in
our model. What we need to do is just
concatenate this back to our data frame.
Right? So, we just need to concatenate
it back into our data.
Okay. So, what we're going to do then is
we're going to grab um we're going to
grab Y as our salary. And of course,
we're going to fill nles on that Y
coming up shortly. But we're going to
grab Y as our salary and X new. Now
before building a full X, we're going to
take a look at X numerical as our data
frame minus these columns. The reason
we're doing minus those is because we
are going to concatenate our dummy
variables back into this that are going
to replace these guys. So we're going to
replace these anyways with our one hot
encodings. We don't want the strings. So
we're going to get rid of those. And
we're also going to get rid of the
salary because that's going to be part
of our that's just the label. So we
don't want that in the X, the eventual
X.
Are you guys able to run this one?
Hope I'm not going too fast. You guys
able to run this? And does it make
sense? What we're doing is we're putting
our labels in Y, which is what we
usually do. So we're going to predict
the salary
and we're getting ready to build the X.
But before we first want to get rid of
those one hot the the strings. This is
getting rid of the strings
and this is getting rid of the label and
that's going to be part of our features.
What we need to do is build our final X
by concatenating our dummies with this.
Do you guys see that? We're going to
concatenate our dummies with this to
build our final X.
But but prior to doing that, we need to
get rid of these string columns here. So
we're dropping those
dropping those from the uh data frame uh
and getting a numerical uh x numerical
here.
You can see the columns of that are just
these guys here. So the the results we
need to concatenate our we need to
concatenate this guy um into this and
then that'll be our full x all of our
features.
Okay. So you can see x is going to be
pd.con
of this with our dummies.
This with our dummies. And um
uh instead of doing this, I'm actually
going to do the full dummies. We don't
need to um pick just a few columns.
We're actually going to do our full
dummies here and um do x equals 1. Now,
the reason that's the case is because um
this will get rid of one column per
feature and basically assume that if you
have a if you have a zero, the other one
should be a one. If you have a one, the
other one should be a zero. Um so it
basically makes that assumption because
we only have two of them. Um so whenever
there's a one, the other should be zero.
Um, so you can get away with just having
these three, but um I think it makes
more sense to just have to have the full
dummies,
but by process of elimination, you can
get away with just using two of them
because anytime you have a zero, the
other one should be the other feature
would have been would have been a one,
right? And vice versa, when there's a
one, the other feature would have been a
zero.
So we do that one.
And you can see all of our uh all of our
one hot encoding features end up back in
there.
So this is the code that I want you guys
to run. I think it makes more sense. It
follows along what we've been doing.
um which will concatenate our dummies
back to our features here to build out
our full X. So now X is all of our
features. Um remember X
X X contains all of our features
now.
So X contains all of our features and so
we have all of this now.
Okay. Were you guys able to run this
one?
D. We have y, we have x. We still need
to deal with the nles in y. So that
something we still need to deal with.
But hopefully you have this. Now
all of these are numerical.
So that should be good with the model.
That's one thing about X is you should
you our X should have all numerical
features, right? Because it's going to
go into a model to to learn those betas.
So it needs to have all numerical
features,
right? These are going to be all
numerical, which makes sense. We change
we did one hot encoding to change all
those guys to numerical.
Sorry, I'm scrolling down.
Okay, we do fill in the nator. Okay.
Okay.
Any questions so far? So, we're just
getting our data ready. We haven't
applied the lasso yet, but we're just
doing some prep. Now, hopefully you guys
recognize th these are some standard
steps that we're taking when we do our
modeling. We have to do these data prep
steps. They're necessary. And so, if it
seems like it's a lot of work, that's
because it is. It is work that you do to
prepare your data to get ready for
modeling. You have to do that. Okay.
So, we're doing that here. Um, now we're
going to do our train test split because
we're just going to do uh we're going to
do hold out here. So, we're doing a
train test split with about with a test
size of about 0.25. So, again, anywhere
between 02 to.3 would be okay.
Um, so uh it's our choice. We could do
02. We could do 3. We could do anywhere
in between there. We're doing 0.25.
That's fine. Um, that's okay. So, we we
build our train test split right there.
Um, so pretty pretty simple and we've
seen that a bunch of times with our X
and our Y data frames. There we have our
train test split.
Okay.
Um, now what we're going to do is do our
our scaling. So, we're we didn't do this
last time, but we're going to do this
now as uh because we should get in the
habit of doing that. Um is um we're
going to um go ahead and scale our
features and we're going to use the
standard scaler here uh to do that
scaling. Okay. Now, we could use minmax
scaler that's fine, too. We're just
going to use the standard scaler here.
Um and remember we are going to uh um
use the standard scaler from sklearn and
we're going to transform our features uh
uh according to our um according to our
training data. So we have our
pre-processing standard scaler here. So
we import that guy and then we um build
our standard scaler and fit it on the
training data only on the numerical
features. Um so that's which is going to
be uh all of these guys. So we're doing
the scaling on all of these guys. Now,
something to note is that we are not
scaling all of these one hot encodings
mainly because it doesn't make sense to
scale those really. They're zero or one.
They don't need to be scaled, right?
They're already zero and one. So,
they're they don't need to be even if we
were doing minmax scaling, it's going to
put them between zero and one. It
wouldn't affect it really, right? So,
these one hot encoding features, we're
not going to scale because they're
they're always going to be zero or one.
There's no need to scale them really.
Um, but we're going to scale all the
other features here that are floats.
So that's these guys here. These
numerical features we're going to scale.
Okay.
Don't need to we don't really need to
scale the one hot encoding. Uh, it's
pretty much already scaled.
Oh, you should change that. Um, go back
and rerun go back and rerun this. But
make sure you have your data type as int
here.
Make sure you add that in there to
change that over to integer and rerun
that and then rerun the rerun the
concatenation.
So make sure you run this
and then u make sure you rerun this and
rerun the concatenation part which is uh
this
Okay. So, we go ahead and fit the um
scaler to this data and then we're going
to transform our training features,
those numerical features. Um we're and
then we're going to uh transform these
features uh uh the test features in the
same way. So we're going to perform the
same transformation from the scaler on
the test data. So that's something
really important I want to note here is
that we always scale both the training
and test data. We always scale both. Of
course, we're going to train the model
on the training data. Um, but we are
going to also test it on the testing
data and it also needs to be scaled
because our model that we build is going
to assume scaled features. The
coefficients that it learns are going to
be assuming scaled features.
So we need to also scale our test data
in the same way. So we're doing that as
well.
So, we scale that. And now we have our
uh training and testing features have
been scaled.
No, we haven't replaced. We're going to
do that. We have not yet. We're going to
do that coming up in a minute. Yeah, we
haven't done that. Um, it is it is the
label. We definitely need to replace
NLES. We just haven't done it yet
because it's not in the features and
we're doing all of our uh uh
pre-processing to our pre-processing to
our features.
Yeah. So, we're definitely we need to
we're going to in a minute.
Okay. So, if you look at the data now,
it's all been scaled. So, these are all
um zcores. These are all on a much
better scale now. Um, and these are we
still have our one hot encoding features
which are zero or one. So this scaling
should lead to a better model than if we
didn't scale. So scaling is really
important. We can see that here.
Okay.
Now, um, let me ask you guys, were you
able to run the scaling? Are you caught
up to here? If you're following along,
were you able to run the scaling?
Okay, great. Great.
Awesome.
Okay. So, uh what we're going to do now
is replace nulls in the uh replace nles
by calculating the median of the data.
So, what I want you to notice is that we
are taking the NLES now this is um this
is on purpose is we are purposely taking
the NLES um out of the median
calculation. So we're skipping the NLES
when we compute the median because we
don't want those NLES to affect the
median calculation.
Um so we compute a median salary here
and then we fill our NLES with the
median salary um from the training data.
So this is our choice. This is a choice
um to use the median and it's also a
choice to use the training set median
for both train and test. What we could
have done, this is an alternative that
we could have done is use the entire
column and then um use the median of all
of the data to replace. That's really up
to us. Um this is one way of doing it.
We could have done before we did the
split. We could have um filled in with
the median earlier. We chose to do it
here mainly because it doesn't affect
the features. So, we could have done
this earlier and did it before we did
the split and filled the NAS. Um really
doesn't it's doesn't matter that much
which way we do it. Um but we do need to
fill in NLES. We cannot have those be
null when we when we put it into our
model. So some way we need to fill in
NLES. Um and so in this strategy we're
filling in our Y train um with the
median salary from our training data.
And same with this we're filling in with
the median salary of the training data
as well. But that's a choice f we could
fill in with the mean with the average.
Um we could fill in with the we could do
it with all the data together before we
split it. we could have filled in with
all of the the median across the whole
data set. Um either one works. You can
do it either way, but we we did it um
later here to show that it doesn't
really affect the features. So, we can
choose when we do it, right? It doesn't
affect the features at all. So, we can
do all of our pre-processing on the
features and then do our label uh
filling and nulls um if if we have them.
uh x numerical. Um make sure you're
running uh this
uh x numerical was defined here
when we split it apart um from
uh when we dropped these columns here.
So make sure you're running this. This
is x numerical
gets defined there.
So, go back up to uh this cell
where we split apart the y and we and we
have the x here x numerical.
Make sure you run this.
Make sure you run this. And then you can
run these. Then you run this to build x.
All right.
Are we up to here with this filling in
the labels?
Uh because then we can build our model
once we're up to here. We've scaled
everything. We filled in our NLES.
We've gotten one hot encoding.
Yeah, it is. That's why you know that's
why we spend a lot of uh time on model
on data preparation with pandas, right?
That's why we did all that pandas work
for sure. Yes, there is a lot of work
before we can build a model.
Yes, the mo do you guys notice that like
the modeling is relatively easy. It's
just a fit and predict. The modeling is
actually really easy. It's all the other
work that's that's more involved, right?
more code.
The modeling itself is really easy.
It's just it's just one line of like
ffit.
Yeah, pretty easy to do.
And then you do evaluation which is a
couple lines.
Yep. There's these are all the these are
the common steps. All these steps we're
doing are very very prototypical in
model building is you let's just go back
through this to see what we did right we
imported our data
um we analy we dropped this name column
because it's not useful to us so we
dropped that um we filled in the nles
eventually um but you know if there were
any nles in our features we would have
to deal with those as well by replacing
them or dropping the rows like we did
earlier Um
and then we do one hot encoding because
of course we can't have any string
columns going in our models. We got a
one hot encode.
Um we uh then build our X and Y by
concatenating the one hot encoded back
to the numerical features.
Then we train test split. Right? That's
pretty common. Or we could do cross
validation either way. Um the K full
cross validation. Then we scale. So, we
didn't do this last time, but this is
something we should get in the habit of
is scaling um our features. So, we do
that and now we're ready to model. So,
now we're ready to model. Um so, that's
this part.
Okay. So, let's do the model. Um the
model's actually uh pretty easy to do.
So, we're going to use a lasso. So, we
have a lasso model here. Notice what
we're setting our alpha to. So the big
parameter we really need ignore this
iterations. We actually don't really
need the we don't really need that
parameter. Um so just ignore it for the
moment. But the big one that we're
setting here is the alpha. So when we
did linear regression, we didn't need
any parameters to go inside the linear
regression object. We didn't need any
parameters, right? Because there are
really no parameters of it. But for
lasso, the important one is the alpha.
And so we need to know what to set alpha
to. Um let's start with alpha equals 1.
That's a good starting place. So a
typical um starting point
for alpha
um is uh is one. So that's a typical
starting point. And so we can set alpha
equals to one. This max iterations is
the the parameter that governs the
training process because it is
iterative. So if for some reason we we
can't converge to the right betas and
we've run it for 10,000 steps once we
pass 10,000 steps, uh it will stop and
just give us the betas at that point.
But it will likely never hit this
number. It'll converge before then. So
um we don't really need to um specify
it. So, I'm actually just going to get
rid of it. Um, it's not really a big
deal. It should converge before then.
Um, but if if we want to set like a
maximum step size in the optimization,
we definitely could there. Uh, but not
concerned about that too much. But
here's our lasso. And then we're just
going to do a fit on our data. So, look
how easy that is. Just like a linear
regression. Lasso.fit,
right? So, we do fit. Um,
oh, I didn't run this. I'm sorry. I got
to run this. Okay. Actually, that's a
good example of what happens when you
don't when you have nulls, right? So, it
says our our null contains nan. That's
because I didn't run this. But now, that
should be filled in. Now, we should be
able to run this. Okay, perfect. So it
runs.
Okay. So you can see what the intercept
is. Um this is one of our coefficients,
right? The intercept is 457. And look
now what's really interesting about the
coefficients is look at what some of the
coefficients are.
Some of them are actually zero, which is
really So some of them ended up being at
zero, which is very very interesting.
that means that those features get
cancelceled out and they're basically
not part of the model which is really
interesting. Um so we have all these
coefficients and some of them are zero.
Yeah, negative0 is just because of the
convergence like they started out
negative and worked their way up to
zero. it. Negative zero really just
means zero, but they just were coming
from they were like small negatives and
ended up at zero
during the training process. They were
negative at one point and ended up zero.
Um
so yeah, negative 0 just obviously means
zero. Um it's still still zero there.
So what's interesting is some of these
features ended up uh being zero which
you don't usually see in a linear
regression. So if we were to train this
using a linear regression we typically
wouldn't see that but some of these turn
out to be zero because again we're
encouraging those betas to be small.
we're encouraging them to be uh small
and so um you know what happens is some
of them can be shrunk all the way down
to zero meaning those features don't
contribute that's a really simple model
at that point right so we've taken
something complex that includes all of
these features and actually reduced it
into something simple that only includes
these features
right
so that's what it does um now we need to
evaluate this to see how good of a model
it is. But that's what this is saying
here in this text is that um a positive
uh coefficient indicates that as the
independent variable increases the
dependent variable also increases.
Negative coefficient means as the
independent variable increases dependent
decreases because it's reducing the
value. Um and lasso is known for feature
selection by shrinking some of them to
zero effectively removing those
variables from the model from the
equation right
um
so that's what happens
some of them end up being zero
were you guys able to run this this
lasso uh fit which is the training of
the lasso
No, it doesn't ensure there's no
overfit, but it helps with overfitting.
It's supposed to help by making the
model simpler. And this is definitely a
simpler model because it's removing some
of the features from the model
essentially, right? Because some of the
features aren't going to contribute.
It's a simpler model.
It doesn't it doesn't mean there's not
going to be any overfitting, but it
helps prevent it. That's what it's
designed to do, help prevent it.
Yeah. So, higher coefficient. Yes. The
higher coefficient means it's a more
important feature towards the
prediction.
Yes. That's what it means for sure. The
higher the magnitude, the more of a
contributor towards that prediction. Uh
it is. Yes.
And it's not just it's it could be
higher positive or negative there. Like
a higher negative is also a pretty big
factor,
right? So So you want to think about it
in terms of absolute value.
does not guarantee but helps. Yes, it
doesn't guarantee it but it's designed
to help overfitting, help prevent it.
Yes, absolutely.
Okay.
So let's do some evaluation. Um so let's
do in this case we are going to do our
predict
Oh, yeah. I'm not sure why that's the
case.
Interesting.
We could try increasing the um max
iterations.
Okay, that's why. Yeah. So then you get
that result with the with the higher max
iterations.
It doesn't get cut off there.
I think that's why you probably left
this in there,
which is fine. You get about the same
numbers.
Yeah.
All right. Let's evaluate this. So,
we're going to to to do evaluation. I
want you guys to see again, we should
get in the habit of doing evaluation,
which is taking our model and predicting
on the training and predicting on the
test sets, right? So we predict on the
train set and calculate our MSE
and we um calculate our R2 score um or R
squar score I should say. Uh but again
the MSE is the one we're really going to
use mostly. Um but we calculate so we do
our predictions and then we compare that
into our mean squared error with our
labels
and we uh go ahead and do the same thing
with the test. Right? So we do uh
lasso.predict
on our test features and we go ahead and
compare that with the test labels. And
so what we're doing there is generating
our MSE.
So we we take a look at our MSE and we
get uh 84,000
MSE. Um and so of course we could take
the um what we could do with that is
take a look at the um MSE on the uh we
could do um MP. Square root
and do the square root of the MSE test.
and we get um 340. So this would be in
the units of our label. So, if we go
back and look at our label um for some
of those um
so uh we are in 300s and our data is
like right around the 500. So, of
course, if we describe this um we could
see what the statistics are of it. So,
we could do df.escribe describe and
generate that. But that doesn't look
like a very good error, right? If these
are in the 400s, um that's that's not a
very good error.
So again, it's not a very great model.
But one thing I want you to see is that
it's it's not overfitting.
Um if anything, it's actually
underfitting, which is what this kind of
um MSE suggests, right? because our
error here is 84 uh excuse me 84,000.
Um
our our area here is 84,000,
excuse me. And on the test set it's
116,000.
Um so these two errors are both bad. So
it's not overfitting. This is actually
underfitting. So it's not overfitting,
it's actually underfitting. Um, and so
that's the risk with something like
lasso is that it's making the model a
bit too simple and we actually risk
underfitting, which is what happens. We
have too much error across both the
training and the test set. Overfitting
is when we do we have really good
performance on the training set, but bad
performance on the test set. We're not
overfitting.
um we are uh underfitting because our
performance is not good either way. Even
this R squar is pretty low. It's not
even at 50%.
Okay, so that's so we we do the
evaluation and again the evaluation just
comes down to making predictions and
computing our error amongst those
predictions to our labels. That's always
what the uh evaluation is going to be
for MSE.
What's the ideal MSE? What do you think
it should be? What is So, think about it
like this. The MSE represents the
average distance between our predictions
and the labels.
So, if we're getting it right all the
time, what's that distance going to be
if we're always right? What's our
distance from what's our distance from
our predictions to our labels going to
be if we're always getting it right?
Zero. Yeah, there's not going to be any
distance. It's going to be right. It's
going to be perfectly aligned, right?
There's going to be no distance there.
So, yeah, an ideal MSE is zero.
That's an ideal MSE.
So, anything close to like the smaller
the better for MSE. The smaller the
better. Um, for this R squared, uh, it's
it's a scale between 0 to one where one
is the best. So, one would be perfectly
aligned predictions. Um, so, and again,
this this is we actually multiply by 100
to get uh because it's it's a number
between 0 and one. So we get about 47%
which is not good.
Okay.
All right. Any questions on this
evaluation?
All right. I want to show you something
which is
Yeah, this that's true. the scale of it
matters on the data because we should be
you should always interpret your MSE in
the scale of
um your your labels because your labels
like in this case our labels um you know
we could take uh for example we could
easily let's actually do that let's take
the average
let's take the average of our labels on
the training data
and and we can see what those are. Um,
so the average is 500,
right? The average is 500. And look at
what our uh square root of our MSE is,
which is in the same units as our
original. Um, so we have uh quite a bit
of error. 340 when our units are right
around 500.
So that's quite a bit of error.
Yeah, MSSE of zero means our our uh our
predictions are nearly identical to the
test labels. Yes, that's what MSSE of
zero means. There's zero distance.
So closer to zero, the better.
But we talked about it as you you really
so the rule of thumb should be what is
your RMSSE as a percentage of your
typical value. So your typical value is
in the 500s. Our our RMSSE is 340.
That's just really high. That's over
like 60% of that value.
So that's just a lot. That's too much
error. What we would love this RMSSE to
be is under 20% of the typical value. So
that means on average we are 20% or less
off in our prediction. That would be
good. That would be pretty good. That
means we're like 80% accurate,
right? That'd be pretty ideal. So you
got to think about it in terms of this
RMSSE which is in the same units as your
labels.
This is the
RMSSE
which is in the same units as the
labels.
So and then to interpret this we have
340
is compared to
typical
um salary unit of 500
right so this is uh quite a bit when the
typical value is 500 and we are off on
average by 340 units
that's so much relative to the typical
value
that's just too. That's a lot of error.
That's not a very good model, right?
It's underfitting. It's definitely
underfitting.
Yeah. So, that's a great question. What
should we do from here? So, um because
we're underfitting
um we should use a more complex model.
So uh we're going to learn about those
in lesson four, but we should use
something different. This linear
regression is still too basic. Even with
lasso, it's still too basic.
Yeah, we're underfitting because we But
it could also be we're underfitting with
a regular linear regression. We should
test that out. Um, and maybe it would be
an exercise for you guys um to test that
out yourself. It shouldn't be hard to
do. Um, you already have all the data
scaled. You So, do you see how you would
do that? You would just come in here and
build a linear regression rather than a
lasso and dofit. And then you would
evaluate it the same way with a predict.
It's really easy to do that. And then we
can compare that um to to this. It
shouldn't be that hard to do that,
right?
And something you guys could do for
sure. Um,
is build the linear regression and
actually compare it and see what kind of
difference it makes. I mean, we honestly
we could do it ourselves. We could do it
right now. Maybe it's worth trying that.
So, let's build a linear regression
for comparison.
So we have our linear regression
uh is linear regression and then we do
ffit linear regression.fit fit
right so so this will train it um and
then we can evaluate it so lin mse is
mean squared error
and then we can do our um let's do our
training let's do the training and then
um let's predict
actually let me do that here
uh y prediction
train
linear
equals um linear regression.predict
and then we're going to predict on our
training features.
Okay, do you guys see what I'm doing?
I'm building a linear regression for
comparison.
I'm doing fit here to train it and then
I'm making some predictions on the
training set and we're going to evaluate
those. I'm going to replace that here
with y prred
uh train
linear. So these predictions
Okay. So, if you guys want this code, I
can paste it in.
So, let's see what the RMSSE for just a
linear model is.
It's a little bit better. It's better
for sure.
So 289 is better than this 340. It's
better. It's getting closer to zero.
It's still underfitting though,
right? And that's just on the training
set. Let's look at the Let's do the same
thing, but on
Let's change this. Let's swap this out
for um test
And then let's do test.
And then let's do test
test.
And then
test test.
Okay. Okay, so this is producing test
predictions on the test set.
We are generating an MSE test
and then we're doing MSE test
which is using the test labels and our
test predictions and then we take the
square root of that for RMSSE and then
we're going to generate that. So it's
still under fit. I mean this is still
high. This is still high um on the test
set and versus on the training set. So,
it's still pretty high. Um, even the
basic linear regression is under is
still underfitting. Still underfitting,
right? Even without the lasso,
which is lasso is supposed to help with
overfitting. It's definitely not
overfitting. Um, it's definitely
underfitting, but this is a signal that
it's kind of overfitting because this is
performing better on the training data
and then it gets worse on the test data.
Definitely gets worse, right?
Did you guys follow?
I'm just running this above I'm running
this above this. It doesn't matter where
you put it. We could uh we could move it
down.
We could move it down to I just ran I
just picked a new cell right here and
ran it. But we could move it actually
let's do that. Let's move it down
to
after the lasso evaluation.
Okay. So I just moved it there.
And then let's move
this down.
So I just put it here after the um after
this. So this is the um this is
basically the objective function right
of the training process. So during the
algorithm that runs when we call ffit in
scikitlearn it's going to find these
betas right it's actually going to learn
what these best betas are for our model.
Um this is our model here, right? It's
the combination of betas times our
features um plus an intercept beta. Uh
so that's our model. But um we penalize
those large uh weights in absolute value
by um adding a penalty term like this um
where alpha is some level of penalty
that we want to provide. Usually alpha
equals 1 is okay. But um actually what
we're going to learn uh to finish out
this section is there's going to be a
systematic way we can test out different
alphas um that represent the level of
penalty we want to uh apply to lasso or
even ridge
uh regression. So that was the lasso and
um if you guys remember using it was
super easy. Uh we worked through this
problem with this um baseball data um
and we had uh
let's see scrolling down we um split out
our numerical data and we did uh we one
hot encoded our our categorical data
combined it back together. Hopefully
that um rings a bell there. Um and we
actually scaled our data which is pretty
standard to do is we do some type of
scaling to our features especially our
numerical features right want to scale
those in some way whether it's minmax
scale or standard scaler um want to do
that and so we did that for this example
and then we um ran the lasso regression
which is pretty easy to use. You just
use the lasso object and you pick an
alpha here. Um, again, we are going to
have a way to test out different alphas
that could be candidates and we can see
which one's the best. Um, so I'm going
to show us that today coming up shortly.
But that was that was the lasso. If you
guys remember, we did that. Um, this it
we compared that to a basic linear
regression which is just this pretty
straightforward just a fit and then
predict and then we can generate mean
squed error. Um, still not a very good
mean squared error on this data, it's
still fairly large. Um, so it's still
not, no matter which model we use, it's
still not very good, but at least we can
practice doing that comparison. That's
what we did last time. We did this on
Wednesday.
Um
and then
we saw that the effect of different
alphas we had a lasso um
we had a lasso uh cross validation
example here. So beyond just using a
regular lasso model that um scikitlearn
has a lasso cv which allows you to try
out different alphas uh with cross
validation and um figure out what the
best alpha is. Um, now we're actually
going to have a different strategy
that'll instead of just picking random
ones, we can actually um supply multiple
parameters that we may want to test. Um,
as many as the models may support. And
in some more complex models, we'll have
more than one parameter like lasso only
has the alpha. Um, technically it also
has this max iterations, but really the
only one that matters is this alpha.
Other models have many more
hyperparameters that we can um uh change
and so we want a way to systematically
test out those different combinations
and to see which one leads to the best
uh version of that model. Let's say the
best results. So um we're going to
explore that coming up. So we had lasso.
Um now this is where we ended last time.
We had ridge regression. If you guys
remember, this one is just a slightly
different penalty. Um,
it takes the it I drew it out for us. It
takes the same penalty we had before.
So, it has that um residual sum of
squares error, which is the main one we
use for linear regression, but it has a
penalty with an alpha. And then it has
the sum of the beta squares
beta i squares. So it penalizes it has a
penalty but it penalizes slightly
differently where it uses the square not
the absolute value. That's the ridge
regression. And this has the similar
effect of you don't in order to minimize
this right because our goal in training
a model was to minimize this thing
minimize this um quantity and find the
best betas that minimize this. Um so
generally yes you want to encourage
lower values but the um once you get
values that are a fraction if you square
them they actually get smaller. Um, so,
uh, it's it's not, um, it's not
necessary to shrink them all the way to
zero. They will get smaller as soon as
they're kind of below one. Um, so they
don't encourage it to completely go
away, uh, like the absolute value does.
It's just slightly different
minimization. Um, so what we see with
the ridge is we don't see the features
kind of get wiped out completely like we
do with a lasso. and lasso they get
encouraged to be um to become zero
because that's kind of the only way to
minimize an absolute value. But with
squares they can keep getting smaller
and smaller and smaller um fractions and
they don't have to become zero. It's not
as harsh of a of a penalty.
Um so uh the ridge was easy to use as
well. Um and it also has an alpha that
we can set. So, it's literally the same
exact code, just a different model, just
slightly different penalty, and it
results in different coefficients. You
notice that none of them are exactly
zero. Like with the lasso, you can get
ones that are exactly zero. We don't see
that with the ridge. You remember that.
Um, so we we s pointed out that last
time. Notice the coefficients aren't
zero. Um, and then we can evaluate it.
So we did our MSE calculation which is a
pretty standard thing where we use our
model to predict on a training set,
predict on a test set, evaluate those um
by computing the metric like the mean
squared error and we can see if we're
overfitting underfitting. This is
definitely the same kind of story we've
seen with all these models is
underfitting because the error is so big
across both sets
across training and test. So it's it's
definitely underfitting.
Um
and same thing as lasso, it has a cross
validation uh variation on it that
allows you to try out different alphas
and um do different folds. So 10 folds,
five folds, whatever, and compute the um
try to find the best alpha that way.
Okay.
All right.
Any questions on this so far from last
time from reviewing that a little bit?
Hopefully that uh hopefully that is
jogging your memory a little bit on
ridge and lasso. Um you know where we're
going to pick it up today is to finish
out this lesson with one more model
which is going to be a combination of
ridge and lasso. So you can actually
combine them together
um in a linear fashion those penalties.
So you can actually have both penalties,
the absolute value and the square. And
when you have both penalties um that's a
special model called the elastic net uh
regression or elastic net model. Um so
this is a combination of lasso and ridge
together. So you have lasso, you have
ridge and then you have elastic net
which combines both of those penalties.
Um let me show you the equation.
So here is the uh so here is the the
model. This is the same that we've
always had. This is our usual u model
fitting for linear. This is a basic
linear regression um loss function or
objective function that we're trying to
minimize to find the betas. Notice how
we have both of our penalties though
this time. So instead of just having one
of the penalties, we actually have both.
So we have the lasso penalty
and then we have the ridge penalty here.
So we actually use both of them and um
try to find a balance of minimizing
those two uh those two penalties.
Okay. And notice how they instead of
just a single alpha, we kind of have a
balance on both of them.
So, we can actually weight the lasso one
more. We can weight the ridge one more.
We can weight them the same. Uh we can
um change that around as much as we
want. So, they have two different
weights there um that they could be.
Um now what happens in reality is uh
we're going to see this in the model is
that um usually what happens is these
get combined into a fraction. So there's
usually a ratio of lambda 1 to lambda 2
and this is known as the um this is
sometimes known as the L1 ratio
and this is a this is a a parameter
inside the model that we'll be able to
set um along with alpha. So we'll be
able to set an alpha and then this
ratio. Um the idea is is that um the
ratio will uh allow us to control which
one is more dominant. So if this number
is bigger the um this lasso penalty will
will be weighted more. If this ratio is
smaller if it's less than one for
example that means that the um ridge
regression is more uh dominant. Um but
the so we'll have this we'll have really
this and this at our disposal and alpha
is um
alpha is kind of like a a you can think
of it as a scale that is um so lambda 1
kind of like lambda 1 plus lambda 2 um
combined to equal alpha.
So it's like our total level of penalty
um our total level of penalty and we can
set that equal to one. We can set it
equal to whatever we want. Um and so
these will be in this ratio and there'll
be a total level of penalty that we can
apply. So the model will actually use
these two parameters when we when we do
it. But that's how they're that's how
they're all related.
Okay. So ridge uses both penalties.
That's the only difference between lasso
or sorry elastic net uses both
penalties. Um so one thing I want you to
notice is that uh if we um if we want we
could set this L1 ratio all the way to
zero
um which uh if we do that um the only
way this L1 ratio could be zero would be
if lambda 1 is zero. So it would just
revert back to ridge regression. So it
complet if if this is zero this will
wipe out this term and we'll be back to
ridge if the L1 ratio is zero.
Okay.
All right. So we have a elastic net
model. Um now it's used the exact same
way as we did the other models in the
code. So we have elastic net um uh from
the scikitlearn linear model family just
exactly where we had linear regression
lasso ridge all of those came from this
linear model um elastic net also comes
from there and then the cross validation
version also comes from there um so
let's see so when we build our model
it's going to be um very very simple
easy stuff because it's the same code
that we always have um we just use the
elastic net. We set an alpha alpha
equals 1 is pretty standard um just like
it is in in the last one ridge that's
industry standard is one and then an L
L1 ratio of.5
that's pretty standard as well. What the
L1 ratio.5 is is kind of a um
uh kind of a that means that the lambda
1 to lambda 2 ratio is 1/2. Um, so
that's that's a pretty standard uh ratio
as well, but again, we could set this
equal to one and they'd be kind of
equally weighted. Um, L1 ratio of a half
means that the uh ridge regard the the
ridge penalty is a little bit more
weighted uh in that in that situation.
Okay.
So uh once we have this model um we can
do ffit and we can run that on our
training data and we can um get we can
figure out what our parameters are like
our coefficients and our intercepts. Our
model will have that but more
importantly we can use our model to
predict right so we can predict on the
test set. Um let me go back and load our
data and actually run this.
So, we're going to be using the same
data that we did for
uh lasso,
which is the I'm scrolling back up so I
can load it. It's the baseball data
here.
Um,
just run it from there.
It's this hitters.csv. So, hopefully you
have that one.
Let me load this.
Okay, so we loaded that and then that
should load.
Drop that unnamed column.
We will get our dummies
and then concatenate those split
scale. I'm just rerunning things. I'm
rerunning things so we can see our model
one more time.
So rerun that. Take a look at that. That
looks good. and then
fill in the NLES on the on those.
Okay. So, we should be able to run our
uh elastic net now.
Okay. So, let's import that and then
let's build our model. So, there we go.
We build our model and the intercept is
that. Now, of course, we can look at our
coefficients. Let's look at that.
Look at our coefficients. So remember
the coefficients are the uh betas. These
are our betas that are in our model. Um
so we can take a look at those. Now um
they're it's somewhere in between. It's
not a full lasso where we're going to
see some of these be zero. It's not a
full ridge. Um so the coefficients we
get are different. They're somewhere in
between there. Those two models that
we've already built. So not quite the
same um somewhere in between there.
Um and then we can use our model to make
predictions and and compute the MSE
uh or the RMSSE I should say as well. So
we can take the mean squared error, pass
that into the square root and compute
the RMSSE. So still pretty bad. Um this
is right around that 300 range of what
we've gotten for our other RMSSE. So,
it's not like elastic net is any better
than those other like linear or lasso or
ridge. And that's not surprising because
it's just adding those extra penalties.
We don't expect it to magically get
better. It's actually a more complex
um when we add when we add those in,
we're actually reducing it and making it
simpler. And we need something more
complex, I should say. So, we're making
it simpler um by by making penalizing
our weights a little bit more. And so,
it's still not a good fit. That's not
really surprising, right? It's still not
really a great fit.
And we can we can even double check
that. We know our RMSSE is pretty bad.
Um but we can double check it with this
R2 score. And it's, you know, still not
good. Remember, a one would be really
good. Um that'd be like a perfect linear
model. This is um still pretty bad.
Okay, so as we said, the alpha controls
the overall strength. Um so the higher
the alpha, the more overall penalty
we're supplying, which makes the model
simpler. Um uh but the L1 um ratio
determines the mix or that ratio of the
lambdas, the lasso to the ridge. Um if
you have it be um exactly zero, you you
revert all the way back to um if you if
you put it at zero, you revert all the
way back to ridge. One would be all the
way to pure lassos. Somewhere in
between, like one half is is good.
Okay,
so this is another example of trying out
different values of alpha in the CV to
see which one works. Now again, I'm
going to show us in a minute a
systematic way to do this, but this is
just trying out um different alphas that
we set up in this uh in this um
uh range. So we have different uh values
between minus2 and two um
logarithmically.
Um so these are uh logarithm values that
are between this between minus2 and two
and we choose a 100 different alphas and
then we choose a 100 different um L1
ratios between 0.01 and one and we run
that we run this um cross validation
with 10folds. So this is quite a bit.
So, we're doing 10 folds and we're
trying out a hundred different um
options. Uh every time we do an option,
we're trying out 10 folds to evaluate
it. So, it's going to take a minute to
run.
It's still running here. But again, what
this is doing is trying out different
alphas and it's it's going to do a cross
validation. And you guys remember the
t-fold cross validation is where we take
our data and we divide it into 10 folds
and then we um train on nine of those
and then test on the remaining fold and
then we rotate all the folds 10 times.
and that we average those mean squared
error metrics together um against those
10 different uh fold options to generate
a basically like an average performance
for that value of alpha. And we're doing
that a 100 times for all these different
100 alphas that there are and 100
different L1 ratios that we're trying
with them.
So that's quite a bit of processing but
uh it did finish.
So we can see what our best alpha is and
our best one ratio. So we get the best
alpha is this best one ratio is this. Um
and therefore we can uh build a model
with those with just these two guys as
the alpha and the L1 and um see how that
performs.
We build that model and then we predict
on the test set and we generate the
RMSSE. It's just a little bit better.
It's still not It's just a little bit
better, but it's still not good, right?
It's still 338. It is just way too big.
Remember, this is RMSSE, so it's in the
units of our uh target variable. So,
it's in the units of if we go back to
our data, actually, I could just print
it out here.
um this RMSSE.
If I just do this, we could take a look
at um DF
or I could look at Y test
and you can see some of these values.
These are these salary values in the
hundreds, right? Some of them are in the
thousands. Um but an error of like 338
is just too big. That's a really big
error. That means we would be off by an
average of 300 when our our values if we
just do the mean
um
is only 550 as on average is 550 but we
have this amount of error on average um
so that's just a way too big of a
proportion of error right it's not a
very good model and again we can verify
that by looking at this R2 for.
So if we go down here,
still not very good.
Here's some of our coefficients. So
remember, you can always take your
coefficients and line them up to your
your data columns. Uh so that you can
get a sense of what coefficient belongs
with what feature. So that's all we're
doing here is just creating a series
where those coefficients instead of just
printing out the coefficients, we're
actually lining them up to the columns.
So this tells us um remember the larger
it is the more influence it kind of has
on the on the final result. Um either
way, so like this has a big negative
influence. um this has a large positive
influence.
Okay, let me pause there. Any questions
about the
elastic net model?
This is a really this model is a really
good one to use when you are building a
linear regression and it's performing
well but it's overfitting. This is a
really good one to use because you can
balance
lasso and ridge you can get the best of
both worlds. So the the main strategy is
if you are using a linear regression and
you see overfitting
um meaning that it's performing decently
so on the training set
it's performing okay but then on the
test set like you know it's it's not
underfitting. it's performing pretty
well on the training set, but then on
the test set it's um performance is much
worse. That's overfitting. If you're
overfitting, then this is a great model
to use because we can try basically by
by rotating through different alphas and
different L1 ratios, we can try out
different strengths of penalty and
different variations on lasso and ridge
together. This is a really good model to
to use for those overfitting cases where
linear regression is doing decently. Um,
but it's overfitting,
right? So far, we haven't ran that case
because so far, no matter what model
we've used, it's always underfit. So,
anytime we have those underfitting
cases, it signals that we should likely
just use a more complex model. And we
haven't learned about those yet.
um we will coming up in lesson four, but
um that's for this data. That's ultim
ultimately what we'd want to do is
probably use a more advanced model
because it's underfitting um just using
a linear regression and and then using
the the overfitting variations of linear
regression like lasso ridge and elastic
net.
Okay.
Any questions on this on elastic then
the TV? Yeah, we Yeah, I think I have
it. I can share it with you.
I said that and now I can't find it. I
thought I had it.
I don't have it. I thought I had it, but
I don't.
If anyone does have that one.
Yeah, I'll look one more time. I thought
I had that one.
Um,
yeah, it's not in there. I had it. Let
me see.
Yeah, I don't have it either. I thought
I had it in here.
Yeah, I don't have that one.
I don't have that one. I'll have to find
it. Uh I have this marketing data. I
don't think this is the same one.
I have this marketing data. I don't
think that's the right one, but you can
take a look at it.
No, we're using So, for this example,
we're using the same hitters data set
that we used earlier for lasso.
No, that's an earlier one.
That's from the uh very beginning of the
notebook. So that's the that's from this
one.
Oh, this Oh, this is where it is. Sorry.
This is where it is. You can find it
here.
That's right. It was from a URL.
It was used in the very beginning of the
notebook.
And we did we did this.
Okay.
That's right. That's why I didn't have
it downloaded.
Okay.
All right. Any other questions on the
elastic before I move I'm going to move
on to uh finding those a systematic way
to find the best hyperparameters.
Um, I'm going to show you a couple
strategies to doing that. Um, so far
we've just ran CV with some random
choices. Um, I'm going to show you a
better, more systematic approach. That's
kind of the industry standard for doing
tuning. Um, so I'm going to I'm going to
show you that next, but any questions on
the elastic net?
Okay. And again like you know
scikitlearn makes it really easy for you
guys
because it just behaves the same way as
any other model. You use the object and
then you do ffit and predict right? So
the ffit is going to train it um and the
predict is going to allow you to use
that model to predict. It's it's super
easy that way. Every scikitlearn model
is like that dofitit and predict. So it
provides a really simple way to use
basically every model.
Okay,
let's talk about let's finish up this
lesson with a couple things. Um, one of
those things is going to be
hyperparameter tuning. So what is this?
The hyperparameter tuning is a
systematic way to find the best
parameters in a machine learning model.
So a lot of machine learning models have
what are called hyperparameters.
These are not the betas that we learn
during the training that's learned from
the data. These are settings that we set
ahead of time like the alpha. That's a
perfect example like alpha L1 ratio in
in the elastic net. We set those up
ahead of time and depending on what we
pick for those we get different
performance, right? And so what we
really need is a systematic way to find
the best settings for those
hyperparameters as we are training our
models. Um the the the main like idea
behind this process though is going to
be to systematically try out different
combinations as many as we want to try.
And so we're we're basically going to
have a strategy for tuning that is going
to exhaust all the combinations of those
hyperparameters that we want to try
until we find the one that performs the
best. Um and and that strategy is known
as grid search. Um and essentially what
it does is it sets up a grid um where
which is basically like a matrix to say
okay which parameters do you want to
try? I want to try um alpha and I want
to try L1 ratio
um L1 ratio like let's say I want to try
these two. So we set these up in a grid
where we say, "Okay, I want to try this
value. I want to try this value. I want
to try this value. This one, this one,
this one, and on and as many as we want
to try." So we could set up set those up
systematically like a linear um a
linearly spaced like I want to try every
alpha between between 0 and 10 spaced by
one um whatever. You know, we can set up
different ranges of those, but that's
going to be in this grid. And then the
L1 ratio, same thing. We can try out
different values of these that we want
to try. Let's say there's many of those.
Um maybe every um tenth between 0 to one
I want to try out. Um so you set up your
parameters and you can set up as many as
you want in the grid. And then
essentially what you're going to do to
do grid search is you're going to work
your way through every combination of
those. So you're going to try out this
combo. You're going to try out this
combo. You're going to try out this
combo.
this combo. So the first value of alpha
with every possible L1 ratio, then go to
the next, try out the next value of
alpha with every L1 ratio, and on and on
and on. So we're going to try
all combos
in the grid.
We're going to try all combos and we're
going to find the lowest MSE
combination. find lowest
MSE
combo.
So whatever leads to the best model is
going to be the um parameters that are
that are deemed to be the best. And the
idea is once we have found those we know
that we can use we can go ahead and
train a model with those best alpha and
len ratio and on and on and on.
Yeah, when you get an So this goes back
to the error. Remember that for a
regression,
the error is this measurement of how far
off we are, right? So if we have a bunch
of points and we draw we fit a line
through there, the the MSE is measuring
this distance, right? So what do you
think is a good distance? Like if our
model is perfect,
what's the best distance from our
predictions to the actual points? Zero.
Yes. So the lower the better. The lower
the better. Um so for an R RMSSE, the
lower the closer to zero the better.
However, the RMSSE can be it's its units
are interpreted in the units of our
target.
So what is deemed to be good is relative
to our target. Like let's say our target
is in the thousands. Like it averages in
the thousands. If we produce an MSE of
50 or sorry an RMSSE of 50, that's
pretty good, right? Because our units
are in the thousands
and we're only on average we are off by
50 units,
right? Our distance away is about 50
units. That's pretty good. So the RMSSE
is relative to your target variable.
Does that make sense? Yeah. It depends
on the target. It depends on what you're
trying to predict.
So that's why we got RMSSE that were in
the 300s for those hitters, but the
average was the average of the target
was in the 500s. So that's a really bad
proportion of error relative to the
average target value. Right? If our
RMSSE was 300,
but the target was sitting in the 500s,
that's just too much error. Way too much
error, right? That's just too big of a
value. Um, our predictions are just off
way too much
in terms of that distance. So, this
would be this was a bad model. It was
underfit.
We know that from the the RMSSE. So
yeah, the RMSSE closer to zero, no
matter what is good,
zero is being perfect. Um, but it to
know what's good, you need to know what
your target is on average and then think
of this as kind of a ratio to that
average target. I think that's the best
way to think about it.
Okay, so going back to this grid idea is
so the grid is just basically laying out
all possible parameter combinations and
trying them all out by fitting and
predicting until and generating an a
metric like an MSE
until we find the one with the lowest
MSE. So find the lowest MSE combination
and that will be the best
that will be the best combo and then if
we once we know that best combo we can
use that we can use that alpha we can
use that L1 ratio and use that model
going forward we can we can use those
parameters in our model so this strategy
it has a name it's known as grid search
so it is a hyperparameter tuning process
that tries out all combinations S.
So what's the what's the U benefit to
this is that we get to test out a lot of
different combo combos of those
parameters like the alpha and L1. So we
can be confident what the best model is,
right? So we can pick the alpha and L1
perfectly because we're trying out a
bunch of different combinations on the
data to see which one's the best. What's
the downside?
It's expensive, right? It's an
exhaustive search. So if you have many
different parameters and you're trying
out many different combinations, it can
get exponentially
expensive
to perform this search. Okay, so grid
search is great except for the fact that
it can be expensive if you have many
parameters with with very wide ranges
that you're searching over because that
that's a lot of combinations you have to
test, right? And especially if you have
a lot of data, that's going to be
expensive
um to do.
So, uh we're going to practice doing
grid search, but that is that's the pro
and con. The pro is that we get to try
out all these combinations and see which
one's the best. The downside is it can
be expensive to do that if you have a
lot of parameters um that you want to
tune for your model um and you have very
uh many different choices that you're
trying to evaluate for those and it just
creates a really big um collection of
combinations that you have to try out,
right? Um that's the only downside to
grid search.
Now on the opposite end of the spectrum
of that is a randomized search or random
search and this will basically just um
do a sampling of those parameters from
um kind of fixed uh specified
distribution. So essentially what you do
is similarly you define your range. So
you say I want to look at alphas um
between zero sorry between let's say
yeah 0 to 10. I want to look at a bunch
of different alphas. Um, and I want to
look at a bunch of different L1 ratios
that are between 0ero to one.
0 to one. And um, what we do is we say,
okay, I'm going to restrict only testing
20, 30, 40 times. I'm not going to do
all possible combinations. I'm just
going to randomly sample something in
this range and randomly sample something
in this range. And so and I'm going to
perform that experiment a fixed number
of times. So let's say I set the uh
sampling where I'm only going to do um
20 evaluations.
And so 20 times we're going to pick a
combo randomly. So I'm going to pick an
alpha and I'm going to pick an L1 ratio.
L1 ratio.
And um we are we are just going to uh
sample those randomly from this range.
Um and we're going to use those and test
those out and then it's but otherwise
it's the same as grid search. Whatever
is the lowest MSE
um so whatever is the lowest MSE is the
best.
So we evaluate those. We sample we train
the model. Evaluate it. Whatever is the
lowest MSE
is the best is the best combo. Now,
what's the benefit to this is it's a
much more controlled experiment in the
sense that we um aren't going to iterate
through every possible combination in
the grid. We're we basically set up a
fixed number of times we're going to try
out stuff.
The risk to doing this is that you're
not you're not exploring all
combinations, right? Because you're
randomly sampling, you may get unlucky
and you may not stumble into the best.
You you can make um samples and figure
out what's the best amongst your
samples, but you may not be covering all
the combinations. Does that make sense?
The grid search is going to try every
combo. The random search is going to
randomly sample those combos.
So, it's not going to try every single
one. It's going to try a limited number,
however many you set up. Now, if you set
that number really, really, really high.
Now, you're starting to approach a grid
search because now you're sampling so
many of those combos that you basically
are trying them all at that point,
right? Um,
so, so that's the way the random search.
So by the way, both of these use cross
validation in the sense that when you
evaluate accommodation, you're actually
doing it with cross validation. So when
you do an evaluation, you're going to do
probably 10 or five folds where you
split your data and then you test it on
the rest of the folds and evaluate or
train it on the rest of the folds,
evaluate it on one of them and generate
an average MSE to get your evaluation.
So every evaluation is using cross
validation.
That's why that's and hopefully you can
see why this would be so expensive for a
really big grid, right? Because you're
trying out many different combinations
and every combination is going to do a
cross validation procedure. So it's
going to train 10 times and test against
10 different folds and average those
together. it's going to be a pretty
expensive operation
for a really big grid, right, of of
parameters.
Um, but these are the two kind of
systematic approaches we have at trying
out different hyperparameters. Remember
those those things are called
hyperparameters. These are those choices
that we have before we train our model.
Um, those choices we have that affect
the performance of the model like the
alphas, the L1 ratios, those kind of
things. um we have control over what
they're going to be. This is a
systematic approach to find out what the
best
uh value of those parameters is going to
be on our data.
Right?
Okay. So before we practice this, we're
going to practice with grid search
first. Um
any questions?
Uh, I don't know if it has a built-in
That's a good question. By time limit, I
don't know if it has a built-in way of
doing it, but you could certainly set up
like a a a loop um to like to wrap
around. Do you know what I mean? Like
you could set up a loop where you check
the time if it's if if the time elapsed
as you're doing a search if the time
elapsed is greater than the the time
limit then you can kind of break early.
Um so it's not hard to implement that
but I don't know if it has that built
in. I don't think it does
because I don't think it really cares
how long every evaluation takes. It's
just going to exhaust all those
especially in a grid search.
But um yeah, I there's probably a way to
manually kind of set up a time time
loop.
So hyperparameters are um settings that
we have on the model itself. And a
really good example of this is like the
alpha and L1 ratio in the in the elastic
net. So they're not things that we um
learn from the data directly like the
betas in the model like those get
trained directly by doing the um least
squares process right um by doing that
gradient descent and all that
optimization.
Um so these are not learned from that.
They're actually set ahead of time. And
so what we're saying is the best way to
understand the effects of those is to
try out different combinations of those
until we land on the best one. Right? So
hyperparameters are those options we
have in the model like the alpha like
the alpha and l1 ratio in the uh elastic
net. Many models have hyperparameters.
Um we're actually going to see that in
in future models that we study. they
have options that you can set that
affect their performance.
And so this this is just a strategy to
evaluate those different options to see
which one's the best.
Yeah. So again, hyperparameters, those
are settings on the model itself um that
affect the performance of it.
And basically we have the two two
strategies here. We can set up an
exhaustive grid and search through all
of those until we find the lowest MSE uh
option or we can randomly sample
potential options, try them out and see
which one's the lowest as well. And do
that a fixed number of times. Um
sort of like a fixed number of trials
almost. um which has a risk of not
trying out every option but but
hopefully you try out enough that you've
explored the space a bit and you get
some quality choices there but no
guarantees right no guarantees you try
everything which is what a grid search
will do it will try everything
okay now luckily per usual scikitlearn
has something to manage this process for
us in terms of grid search. Um so in
that way we will not need to manage this
process ourselves. We can just rely on
scikitlearn. And so if you're doing
hyperparameter tuning um this is going
to come from the model selection module
inside of sklearn. So we're going to
import from from sklearn the model
selection module. We have our grid
search cross validation.
Okay, that's what the CV stands for.
grid search cross validation. So, this
is going to do that grid search
strategy. Um, we're going to set it up
with our dictionary essentially of
choices. So, we're going to say, hey,
here's the alphas I want to try. Here's
the L1 ratios I want to try. Um, and
here's my other settings like uh how
many folds I want to use, what my random
state is for the shuffling. So, we'll
set all that up. Um,
and then we'll just run the grid search.
And then what should come out of that is
the best options for our parameters from
the grid and then we can use those going
forward in the we can build a model with
those best options right so we're really
doing some evaluation here of what is
going to be those best alphas those best
0 to1 ratios on our data set right and
the only way to really know that is to
evaluate them because they're not things
that are learned during the training
hopefully that makes sense right they're
not things that we learn directly from
training. There are things that we have
to set and then kind of evaluate and see
how they affect things.
Okay, so we have grid search CV. That's
going to be our primary um tool to do
the evaluations of the different
hyperparameter options.
Grid search CV. Um we're going to set up
our cross validation uh object here. Now
I want you to pay attention to this is
that um it's a slightly different
version than the kfold we had earlier.
So we've used k-fold before with a
certain number of folds. This would be
10 folds and we can set a random state
for the shuffling um that happens in the
folds.
But this is actually a slight different
variation on it where it is a repeated
kfold where we do three repeated trials.
Now why would we do that? It's to be
extra extra extra careful with the
shuffling.
So this what this means is we do three
different shuffles. So we do kfold, we
actually repeat it three times with
three different shufflings. That's all
that means. So the repeated kfold is
actually a bit beyond the just basic
kfold. What basic kfold will do will
we'll will shuffle and then do our
splits into 10 splits and then train on
nine of those. test on the other one and
rotate through all the splits.
We're actually going to do that process
three different times with three
different shuffles. So this and we're
going to average 30 results instead of
just 10. So repeated kfold is just going
above and beyond to do extra to repeat
the kfold three different times. In this
case only three. We could do more,
but um now is that necessary to do? You
could argue not necessarily. Um but it
just provides extra robustness
uh beyond just our single shuffle and
then split and then rotation of those
folds, right? We're doing it actually
three different shuffles. Um so we're
repeating our kfold three times uh for
every now is the thing is we're doing
that for every evaluation. So it is
going to be more expensive than just a
basic K-fold.
So we have three different K-fold trials
that we're doing essentially.
Okay, hopefully that makes sense. This
is the repeated K-fold. We haven't
really seen that before. we've only
worked with the Kfold, which would get
rid of this repeats option and only have
uh 10 splits in a random state for the
for the single shuffle that we do. So,
we can um recreate that same shuffle
every time. Um but now we're actually
going to do three random shuffles, uh
three different trials. So, one shuffle
creates the and then create the 10
splits, evaluate, then go back and do
another shuffle, another new 10 splits.
So, one thing that should be um clear is
that we get different splits every time
because we're going to shuffle once,
right? We're going to shuffle once and
generate our splits
and then we're going to shuffle again,
generate these splits which are going to
be different and then shuffle one more
time for for three different times,
right? And then get get these splits and
then we're going to get 10 metrics here,
10 metrics here, 10 metrics here, and
then average all of those together.
So, it's a bit more just going up extra
above and beyond for a K-fold. Okay.
All right. So, here comes the fun of
when you do grid search. Now, the grid
is actually just a dictionary. It's a
Python dictionary where you declare what
your parameters are going to be inside
the dictionary and you set up a range of
values that you're go or a list. It can
be a list. It can be a range
but some declaration of what you are
going to test and evaluate inside of
your grid search. So the grid is
initialized as an empty dictionary.
And then what we do is we say okay in my
grid I want to check different alphas.
So we're going to add a collection of
alphas in here that we're going to test.
So let me make a comment there. We add a
add a range of alphas to test. And this
range is a this is just like the Python
range. Um
this is just like a Python range um uh
operator here where this is going to be
uh every so it's going to be um every
uh value between
zero and one um uh steps with a step
size
of 0.1. So, it's going to try a bunch of
different alphas um between uh zero and
0.1
sorry 0 and one stepping by 0.1. So,
it's going to try zero.1
2.3 point 4.5 6 right all the way up to
one.
So, that's what this will do. And it's a
numpy range. So, it's just all those
decimals between 0 to one.
It you can use either that's valid.
Yeah, you can do you can do that to
create a dictionary or you can use the
keyword um dict. You can use either one.
Either one works.
Whatever whatever you want to use.
They're the same.
Yeah. The the reason people prefer
dictionary is because um sets are
created with the same braces.
So it it makes it clear what you're
creating as a dictionary. If you use if
you use this, that's the only advantage
is it's just plainly obvious what you're
making. Uh because technically you can
make a set with the curly braces as
well.
Yeah,
no worries. Um okay, so we have our
alphas here. So what I want you to
notice is that we are going to try out
different alphas and we are that's the
only parameter we are going to try in
our ridge regression. So we're going to
we're going to try ridge but just try
different alphas in the in this range um
in our grid search. So the grid search
CV takes in a model. It takes in our
grid dictionary which is really
critical. We need that dictionary to
declare what we're going to try.
um we need a scoring to say to find the
best. Now remember it uses the negative
to find the lowest which is going to be
the the least negative option.
Um otherwise it wouldn't um just based
on the optimization it would look for
the highest value. Um so the highest
would be closest to zero in this
situation. Um so we use negative and
again we could use squared error. It's
using absolute. We could use um squared
uh either either one works.
Um more typical would probably be
squared error, but um absolute is fine.
Here's where we have our repeated kfold.
So we pass in our um how we're doing CV.
That can be it can be a kfold object. It
can actually just be an integer, which
is say I just want to do 10 splits or
five splits um to to do every
evaluation. But these are the bare
minimum that you need. Just really the
model and the grid and your CV. Um what
metric you're using to evaluate what's
going to be the best. And then this end
jobs is to parallelize. If you have it
set to minus one, it's going to it's
going to try out all the grid options in
parallel. Um which is nice. It's going
to help speed up the overall search.
Okay. So let me mark that down as n
jobs equals minus one.
tries out the combos in parallel.
So in this situation, we actually don't
have more than one parameter. We only
have the alpha. So we're really just
going to be systematically working our
way through every alpha and evaluating
which one's the best right with this.
And notice that in order to use this
grid search, all we have to do is call
search.fit. So it works kind of like
every other model does, right? It's the
grid search.fit.
And we pass in our data.
And we um once we're once this prints
out the results, you get a results
object um which has a best score and
then a dictionary with your best
parameters. So, whatever your best grid
member was or grid members, um it prints
that out and you can So, for from that,
we can um grab our best alpha, which
which let's confirm what that ends up
being.
Oops. We need to import repeated kfold.
So we'll import that.
Oh, I didn't. Let's do from
sklearn.linear
model import ridge.
Okay.
Okay. So, it completed the search and
what we found is this is the best score
is 238 for the mean absolute error and
the best alpha that we got was 0.9. So,
the best alpha that worked here, the one
that gave us the best score was actually
0.9 as the alpha. So what it did is it
tried out everything between this range
and 0.9 was the best. So it did cross
validation, tried out every single combo
in our grid.
So if we want we could actually print
out
print our grid so we can see
what our combinations were.
So, it tried out all of these guys and
the best one that we had was 0.9.
Okay, so pretty cool how that works. And
if we had other parameters, like if we
were doing a elastic net, we could add
those into our dictionary and it would
do all combinations of those. So if we
did um so for instance to add to our
grid we could do grid
um L1 ratio
this would be for like an elastic net
right now the ridge regression by itself
doesn't have an L1 ratio parameter but
just as an example um we could try out
different ranges um similar range
different one um maybe an exact list
whatever we want to do. So this is going
to try out different ones between 0ero
to one
as well. And so it's going to try out
every combination of these from this
grid.
Okay, if we did that. But again, this
the ridge regression doesn't have an L1
ratio. The elastic net does. So that the
ridge regression only has an alpha to as
a hyperparameter. So we're only testing
out that one.
Okay. So that's grid search CV.
Pretty useful. This is pretty useful in
doing parameter tuning again when you
want to try out ranges of different
values and you can evaluate those to see
which one is your best and then we can
use that best going forward. So we can
for instance this is what this code does
below it is it fetches the best. Um you
can do it this way or you can do it um
the alternative is to do results.b best
params
and then you can just grab it like this
alpha.
Either way you can do get or like this
um and it this is just a dictionary,
right? And you can grab your alpha. So
that's the 0.9 um and we can pass that
alpha into the ridge regression and go
back and refit it to our data um and
then use that model going forward. So
the grid search really just evaluates
those different options, tells you
what's the best according to this score,
right?
And you should, by the way, you should
interpret this score in the positive
sense. It's only negative because we're
purposely making it negative to find out
what the lowest option is, right?
Because the lower is the better. So we
we purposely make it negative to make it
whatever is the least negative is the
winner. Um more negative is worse.
So it's really positive version of it is
the is the true result for the error. Um
and they are a tool from scikitlearn to
put together your model with your
pre-processing steps. So they kind of
get automated together. Um and they
combine everything into kind of a
streamline process. You're going to see
what that looks like, but it's a really
nice um feature of scikitlearn. Um why
would we care about pipelines? They help
organize our code um so that we ensure
that we basically always run the
pre-processing steps before we train and
use a model to with the predictions. Um,
so it bundles those steps together,
minimizes the risk of forgetting a step
because one of the things that can
happen is when you do pre-processing, if
you're doing it on the training set, you
have to do it on new test data as well
when you put it through your model
because your model is training against
that pre-processed data.
So in order to make sure you never
forget that, you can bundle it all
together in a pipeline which is going to
make things really really easy to use
and and make sure that those steps
happen in a repeatable way. Um and it
makes things easier to uh deploy that
model as well because everything is
together in one pipeline. So in the in
the industry, I've seen this a lot. Um
you know, people will do their initial
exploration steps and initial model
building. They may not use pipelines
right away, but as they found their
model, um they'll generally move it into
a pipeline and all their steps into a
pipeline so that it's uh easier to work
with um when you're when you're
deploying it, actually using it uh in in
the real world. Um, so here's what a
pipeline generally looks like. It's from
scikitlearn. It's this pipeline object.
Um, and the pipeline is made up of steps
that we're going to see that that are
various um uh basically um kinds of
pre-processing we've seen before like a
scaler or um filling in missing values.
Those kind of things we can put here in
the steps which is basically a list. um
steps is just going to be a list of
scikitlearn functions that we can apply
to data. One of those being a model. Um
and then whenever we use the pipeline,
it's basically um you know, it's going
to be something like pipeline.fit
or pipeline.predict.
So the pipeline kind of behaves like a
model. It's just going to contain many
more steps than that like the
pre-processing steps we've worked with
before. Um, and it also has some
capabilities for caching. So you can
like uh cache some of the data in
memory. Um, so that if you're reusing
the predictions, it kind of goes faster.
Um, so there's some options for that
too. I'm not too concerned about that at
this stage, but the main thing is going
to be filling out our steps and then
using the pipeline.
Okay.
Um, so some important bits of
information about the pipeline is that
it is going to be a sequence of data
transformations that will have at the
very end of the pipeline the model
because of course we're going to do
transformations and then train a model
or predict with a model. So every
Oh, can you guys hear me? Okay,
not able to hear me. Thanks for letting
me know. Can you guys were you able to
hear me so far?
Okay. Make sure. Yeah, it might be on
your internet or your your uh Yeah, it
seems like seems like it's good. So,
no, you can't hear me. Check your
volume. Check your headphones if you're
wearing headphones. Oh, no issues. Okay,
perfect.
Okay. Yeah, local internet issue. Yeah.
Okay.
Always let me know. always let me know
cuz it could be the case that it is me.
So, um always always make sure to let me
know. Um but sounds like yeah, you may
want to check on that. Um
so, okay. What I was saying is every
pipeline's going to have a uh a sequence
of steps that go first and then the
model at the end. Um so, the order
really matters. Um
uh so the order matters in the sense
that we want our transformations to go
first. Things like scaling, things like
filling in missing values, we want those
to be first and then we want our uh
model to be last because we want those
transformations to happen prior to
training or prior to prediction. So
usually what you'll see in these
pipelines is a model at the end, right?
a model that's going to be at the end of
the pipeline because we want basically
our processing steps then our training
or our processing steps then our
predictions. Um so everything in the
pipeline though is going to be from
scikitlearn. Uh that's how it gets
automated in the sense that all of those
things are going to have fit and
transform functions built into them so
the pipeline can use them. Uh, and then
the last step is going to be a model
that has a fit and a predict. So it's
pretty standard that the last part of
the pipeline is just going to be a
model. Um,
uh,
so we can um, as we do more modeling,
we're going to play around with the
pipelines quite a bit and see how we can
change up some of the parameters. like
if we want to change a model's parameter
um we can actually adjust it to do
things like uh grid search or cross
validation. So um we're going to see
some examples of some pipelines but for
right now mostly what we're going to see
is how to build one and then how to use
one. And then as we get into lesson
four, we'll get some more practice with
pipelines cuz we're going to start using
them quite a bit uh to build our models
rather than do manual steps uh all the
manual pre-processing
um and then kind of building a model
from there. We'll just include all of it
together in a pipeline.
Okay, so the example we're going to do
is with this housing with ocean
proximity. So we've actually looked at
this data set before. Um so we have uh
this ocean proximity data set that has
the feature of like how close it is to
the ocean like the bay or the less than
1 hour. Remember we had that and it had
the median house value for different
neighborhoods. Um so we're going to work
with that one again. Let me make sure I
have that one uploaded.
You guys should have this one. It should
be in your uh data sets.
Um, I'll I can upload it here in case
you don't have it though.
Does this use multi-threading? I think
it does. Yeah, I think in order to do it
can do uh um I think it can do
processing in parallel for some of the
pipeline steps. Um, now does it use that
all the time? Not necessarily because
some of it is sequential in nature where
you have to do one step and then you do
the next step and then you do the next
step. So it's not like you can do them
in parallel.
Um in terms of the like you need to know
the output of one step to compute the
the output of the next step. Um so it
can but it it doesn't always lend itself
well. The thing that will use
multi-threading is is like the training
process could be parallelized
like the fit um can be for some models
it can be parallelized not every model
it so long answer is or the short answer
is that it depends
depends on what kind of transforms
you're doing and what kind of model
you're using if you can really take
advantage of
Okay. So, we load our data here and take
a look at that. Um, do you guys have
this data set? Are you able to load it
in? If you're following along, are you
able to load it?
Okay.
And and again, we've worked with this
data before, so hopefully it's somewhat
familiar. Remember, every row represents
a neighborhood and it has a we're going
to end up trying to predict this median
house value as our target um variable,
our dependent variable. Um and we're
going to use the rest of these features.
Remember that um this feature is in
particular going to need to be one hot
encoded,
right? It's going to be one hot encoded
because it is currently a string and we
need to turn that into a numerical
feature which is the one hot encoded
feature. So we're going to have to do
that but we're going to do that as part
of our pipeline.
Okay. So we'll be able to include that
in our pipeline steps uh to to do one
hot encoding which is nice.
All right. So we're going to split apart
our data um as we normally do. So we're
going to uh create our feature uh data
frame which is everything but this
median house value. So we go ahead and
drop that column and then our target is
the median house value. So it is just
that column here. Pretty standard. Um
and then we're going to train test split
and um split it into 30%
uh test data. And again random state you
can choose whatever you want to be. that
just affects the shuffling. Um, so
whatever doesn't really matter what it
is. It's just so that when you rerun
this, you get the same result in the in
the shuffle.
Okay, so we have our train and our test.
So you want to make sure you run those.
All right, so what we're going to do is
take a look at our data
and see if we have any null values. Um
if you guys remember this data actually
did have null values. You can see it
here in this this guy and exactly how
many there are is from this the sum. So
we have um 162 nles in in this data. Uh
and this is just a training data. So of
course you know the test data could have
that in there as well. Um so that's
something we're going to want to make
sure we fill in the blanks on any data
set we use whether we're using the
training or test set. Um, like if we're
doing training, we want to make sure
that gets filled in. If we're doing
predictions with the test set, want to
make sure that gets filled in. Um, so we
we should be doing that. Um, now
what we're going to do is use this data
to help uh train our pipeline or or use
with our pipeline. We need to construct
our pipeline. So far, we've just split
apart our data. We haven't done anything
with our processing steps in our model
yet. Um so roughly
it this should be the flow of our
pipeline. What should happen is we
should be doing some type of feature
scaling
um some type of uh feature um
manipulation. So that could be
engineering, that could be um that could
be uh doing the one hot encoding. Um so
extracting new features like one hot
encoding,
one hot encoding. Um we are going to be
doing that and and by the way, this is
split up into this is when we use our
pipeline for training.
Um it's going to look like this where we
do our scaling, we do one hot encoding,
um we have our model here. Um, so that
could be a linear regression, that could
be a lasso, that could be a ridge, it
could be elastic net. Whatever model we
end up using is going to be last in the
pipeline. And we're going to run this
pipeline. Ultimately, we're going to run
pipeline.fit,
right? We're going to run a fit function
and we get a fitted model as the result
of this pipeline.
Then when we use it when we use our
model for prediction,
we use our model for prediction in this
lower part. It's the same pipeline, same
exact pipeline, but it's this model has
now been trained.
So we now have a trained model here. So
the great thing about the pipeline is
it's the same this is the same pipeline
that we're using here. So it's just
going to it's going to repeat those same
transformations. It's going to do our
scaling. It's going to do our one hot
encoding. It's going to use our model
and it's going to generate predictions
and generate uh we can we can do
predictions. We can do evaluation like
in a cross validation. Um we can use it
however we want to use it. Uh but notice
that the pipeline makes it consistent
between training and test. We're using
the exact same transformations
and the model is last. It's it's either
being trained or it's being used for
prediction, but it's last. Our
transformations are upfront, which are
things like our scaling, things like our
one hot encoding, right? Those happen
first. No matter what data we put
through there, we put our training data
through there, we put our test data
through there, they're going to go
through the same steps,
right?
So that's that's the design of the
pipeline. That's what it's supposed to
do. So our job is to create those steps.
So we need to create those relevant
steps and then put them together into
this pipeline. Okay. So that's going to
be the code we're going to see coming up
is we're going to build out these steps
and then put them together into the
pipeline.
Um any questions on this diagram? Does
it make sense what we're trying to do
with this pipeline? We want to have
repeatable steps during the training,
during a prediction process.
Okay.
All right.
All right. So, um, a couple of things
we're going to need is, uh, to first of
all, let's jot down what steps we're
actually going to do. We're going to
need to deal with missing values. So,
we're going to fill in we're going to
need a pre-processing pre-processing
step that fills in any nulls. We always
need that, right? So, if there's nles,
we're going to fill them in somehow.
We're going to define how we do that in
our in our step. Um and we also need to
one hot encode and we need to scale
right those are pretty standard steps
that we've dealt with whenever we're
building these models right so pretty
standard things fill in nles one hot
encode any categorical data whatever
however much we have and then go ahead
and um standardize which is the scaling
so this this just is the same word for
scaling our numeric features so we're
going to we're going to define Windows.
Um, so that's why we're going to go
ahead and import from pre-processing.
We're going to import our scaler. Um,
again, we could use minmax scaler here.
We're going to use standard scaler. Um,
but we could use minmax. Um, we have our
one hot encoder here. Now, usually when
we do oneh hot encoding, we use pd.get
dummies. This does the same thing as
that, but because we're going to be
building a pipeline, we actually want
the scikitlearn version of git dummies.
So this is the scikitlearn version of
git dummies here. And it and we have to
use that version in the pipeline because
everything in the pipeline needs to be
an sklearn object. It needs to be an
sklearn tool or object.
So um instead of using pandis get
dummies we're using one hot encoder
which is does the same thing. Okay. In
fact it just this basically just uses
pd.get dummies um under the hood.
Okay. So it just uses that uh anyways.
It's just code that builds on builds on
that.
Now what's really nice here is we're
also going to use from sklearn.impute
impute. We're going to use a simple
imper now what this is is an automated
way to fill in missing values. So this
is a fancy way of basically doing the
the fill na on a data frame. So simple
imputer um we are going to basically
fill in the blanks. What we're going to
do when we create this object is give it
a strategy of how to fill in blanks.
Should you use the average? Should you
use the median? Should you use the max?
Should you use the min? should use a
default value. We're going to tell it
what to do in this object.
Okay. So, we're going to we're so we're
going to use this as our automated tool
for filling in missing values. So,
that's really nice. It has so this is
going to be a critical part of our
pipeline an imputer that's going to fill
in missing values.
So, we have that.
Yeah. Coding to reduce coding. Exactly.
Uh we have our pipeline now. So we have
our pipeline. So our pipeline is going
to hold everything. So we need the
pipeline object um to hold everything
and that comes from sklearn.pipeline.
Um so everything's going to actually go
into a pipeline object. We're going to
see how that looks. Um and finally we're
going to from skarn.compose we're going
to use a column transformer. The reason
we're going to do this is because we are
going to specify for some columns like
the numerical features we should be
scaling
for some columns like the categorical
features we should be one hot encoding.
So the column transformer will allow us
to map different transformations to
different sections of columns which is
really useful. So this is actually going
to be a critical part of our pipeline to
apply to make sure we only apply this to
numerical features and only apply this
to categorical features. Right? So this
column transformer will help us um to to
apply pre-processing to particular
columns. Um like that ocean proximity is
the only one that really needs this but
every other column is going to need this
all the numerical features.
So, we're going to use this column
transformer. And again, we're going to
see how this looks, but just trying to
give you an idea of why we're importing
all these things.
Okay. So, let's import those.
Uh, this mentions about the column
transformer. We just talked about it. It
allows us to have a particular column or
group of columns get the right
transformation. So again, uh, looking
ahead to our pipeline, the numerical
features are the ones that are going to
need scaling, but the categorical
features are the ones that are going to
need one hot encoding. However many
categoricals there are. In this case,
there's really only one, which is that
ocean proximity. Go back to our data.
Um, you can even see that in the info,
there's just that one. Um, and we see
that here, right? Just this one string
column that should be one hot encoded.
All these other guys should be scaled.
Right? They should all be uh uh standard
scaled.
So this will allow us to specify those
distinctions.
All right. So let's get started building
our pipeline. So this is going to be
really cool. We're going to build out
the pipeline. Um let's extract our
numerical data and our categorical data.
Now this is a really neat way of doing
that that I'm not sure we've seen
before.
Um so what this does is we'll take our
data frame particular our training data
frame and select our data
that's what this select dtypes does is
select data from it um which includes
only the object type columns so only the
object types. Now what's that?
The object type is the string right? So
this should select only this column
because it's in the include.
We go here include only object types in
the result. And so this should only have
our one categorical column which is
ocean proximity. So, housing cat is
going to have a reference to our uh it's
going to be a list that has a a
basically just our ocean proximity
feature because this select dtypes will
make sure we only pick object types and
um
grab those columns. So this is a way to
neatly grab um our categorical features
here by including the object types. Now
on the flip side we can exclude object
types and get everything else. So this
is going to be all other columns which
is excluding the object. So this is
excluding this meaning we should get all
of our numerical features that way. So
this will be all of our numericals
by excluding the object type and this
will be our housing num which is short
for numerical. So this excludes
the uh object type meaning all numerical
features
right all numerical features there.
Okay.
So, if we were to uh let's double check
this. Let's sanity check this. If we
were to print out the housing
cat, um this should be just the ocean
proximity feature, which it is. So, just
that one. If we were to print out the
housing num, this should be all the
numerical features, which are all these
guys. So it's just a reference to those
columns so that we can uh use those
later when we're mapping uh this
transform needs to go to this column
like the one hot encoding needs to go to
this column and the scaling needs to go
to these columns right so we have those
uh names of those columns already at our
disposal. So, we're just doing that.
And this is just a
simple check.
Uh, are you guys able to run this?
If you're following along, let me pause
there. Make sure I'm not going too fast.
Uh it so the the issue with a specific
data type like that is none of these are
ants. They're actually all floats. So we
did float. I think that should work. But
yes, that's the idea.
Great. I'm glad to hear that right there
with me. Great. Glad to hear that.
Okay. So, we have our columns picked out
here, which we're going to use later.
Okay.
All right. So, let's go ahead and build
out our steps for each of these types.
So, um for our numerical features, let's
build out our pipeline steps. So what
we're going to do is build out a
numerical pipeline. And it's going to be
a pipeline with a list
of tupils. And the reason these are
tupils is because every tupil has a
name. So here this is a name that we can
it can be whatever we want it to be. So
we're calling it imputer. We could call
it anything we want. We could call it
fill in the blanks. We could call it
null filling. Call it whatever you want.
We're calling it imputer because that's
that's a pretty um easy name for it. An
accurate name to what it's doing. Um but
the important thing is after the name
you give it, you put in the scikitlearn
object that you are going to use to
operate on your data. So in this case,
we're using a simple impery
of median. Now that's a choice. We could
use a strategy of mean, max. Um, we
could provide it a constant default
value. But what this means is we are
going to fill any blanks we find in
those columns with the median value of
that column. That's the strategy for the
computer. So that's pretty cool. This is
kind of an automated way to fill in the
blanks using for any column using its
median,
right? And so we could change that. We
could put mean here or max or min or
whatever. Um
but we are filling in the blank on any
column with its median. And the reason
this works is because we are going to
apply this pipeline only to these
numerical features. So that is fine.
We're we're not going to apply it to the
categorical features. We're going to
apply it to only those numerical. So it
should have a median value, right? So
that that's totally fine. So we're going
to now look at how we're constructing
the steps. We have a list of tupils.
Here's one tupole
which is the imputer with a simple imper
of strategy median. And then we can have
as many tupils as we want which
represent processing steps. So every let
me write that down. Every tupil
represents
a pre-processing
step on our data.
Okay, so we have an imputer step named
imputer and the reason it has a name is
just so you can reference it in the
pipeline if you need to. So you so it
has like a a reference name um that you
give it. Um but this is the more
important part is the actual scikitlearn
object that's doing the processing. So
in this case a simple computer but
notice that we have a secondary step
which is our scaling. Now this makes
sense. This is something we should be
doing to our features is we should be
scaling them. So here we we say okay
let's fill in any blanks first.
By the way order
matters.
So, and what I mean by that is the
simple imputer
is before the scaler. Now, that's
important because what that means is we
should be filling in any blanks before
we attempt scaling.
So, that order actually matters. We're
going to fill in blanks first in this
list. That's first. We're going to fill
in blanks. Then we are going to scale
right then we scale which makes sense
right so we we fill in blanks first then
we apply the scaler to scale our
features so those are our two steps
so so pretty simple um we are building
out our two steps now this is just one
piece of the puzzle we are going to put
this pipeline together with our one hot
encoding that's going to be coming up
next and build out our final pipeline.
But this is um a a pipeline that has two
steps that will actually be used with a
larger pipeline coming up where we we do
one hot encoding to our categoricals and
then we put a model in there at the end
to train and and use for prediction. So
um pipelines can actually be composed is
is uh something to realize there is that
we can have a pipeline that contains a
few steps. We can have another pipeline
over here that contains a few steps and
we can actually um kind of put them
together into a final pipeline that has
both pipelines uh kind of merged
together. Okay. So we're going to see
that coming up when we construct our
final one. Our final one, as you can
imagine, needs to handle this mapping of
basically saying, let's do one hot
encoding to these guys and then do this
pipeline here to these numerical
features. That's what our final pipeline
needs to handle. And it will. We're
going to build that out.
But let me pause here. Um, were you guys
able to run this? Are you with me on
this this pipeline here?
Does that make sense? Those two steps
one is filling in blanks with a median
whatever column. So where so this is
this is what's so amazing about this is
this is going to automatically search
for nulls and if you come across a
column with a null, it's going to use
the median of that column
to fill in the blank, right? To fill in
those nles.
Okay,
great. Glad to hear. Glad to hear.
Okay.
All right. So we are going to now um put
this together with a column transformer
to basically say what steps are going to
be mapped to what columns.
Um so now you can see what we're doing
here is using the column transformer
which is going to be a list of tupils
again. So this is another um list of
tupils.
But the important thing is um
each tupil
has a name
followed by so it has a name uh which
again is is generic. You can say
whatever you want it to be. So here
we're kind of shortening this to
numerical. This is short for
categorical. But the important thing is
it's followed by a pipeline
slashstep
followed by a pipeline slashstep
um followed by a uh followed by a list
of columns that it applies to. So you
can see that pattern here. What we're
saying is we're going to apply that
numerical pipeline we just defined. So
this is saved in a numerical pipeline
object here. We're going to apply that
to those numerical features. So this is
that list
of numerical features here. So that's
how we do the mapping. We have a tupil
here that says okay apply these steps to
these columns.
Those go together in that tupil, right?
Apply these steps to this uh these
columns. And then apply this step. Now
what is the step? This is a one hot
encoder
which is going to uh uh encode um those
features and it's going to uh ignore um
basically nulls for now. That's a choice
but it's going to ignore um uh basically
ignore nles and and uh skip over them
for now. We now we know there's no NLES
because we already did an is NA from
before and we know there's not any NLES
in that ocean proximity. So this isn't
going to be an issue. But that's what
that would do.
But we have a one hot encoder here which
we're going to apply to our categorical
features. Now of course that's just the
ocean proximity feature but that but
again you see the pattern in the tupole
is apply this transform which is the one
hot encoding to this column apply these
numerical transforms which is a whole
pipeline. So it's two steps in a
pipeline of um
uh an imputer and a scaler are going to
be applied to this
really nice. So those are going to be
all together in this column transformer
and that is our way to signal that for
these numerical features use these
steps. For our categorical features use
this step and and you know if we had
more than one step we were applying to
categorical we could build a pipeline
for the categorical and it would and do
the same thing. We have more than one
step here and so it's good practice when
you have more than one step to just put
that in a pipeline because we have more
than one step. We'll just put that in
this list inside of the pipeline and we
can map that pipeline to those features.
Here we only have one step. So it's okay
to just put that there um and apply that
to the categorical features. But if we
had more than one step um it would be
good practice to put that in a pipeline
which is what we do here. Right? This
pipeline is being mapped to these
features. This step is being applied to
this feature.
Okay,
how about that? Are you guys able to run
that one? Does that make sense what we
have set up so far? So, we're almost
there. We almost have our final
pipeline. We have our pre-processing
basically done to say our numerical
features should be processed with that
other pipeline and our categorical
features should be one hot encoded.
We're getting close. The only thing
we're really missing here is a model.
The only thing we're really missing is
to have our final model training
pipeline is to actually include a model
which should come at the end.
Right? So it should we should be doing
these steps first
then doing modeling which we know right
we we've done that uh many times. We've
done our pre-processing and then we do
our modeling.
Any questions on that?
Okay.
Fantastic.
All right.
So, if we wanted to uh see if we wanted
to test this so far, um we could. So we
could run the pre-processing and
actually run a fit transform on our data
and this will um basically apply that
pipeline to the data. Now this would be
a sanity check. This is a good this is a
good kind of um this is a good sanity
check that our pre-processing
works. So it's doing what we expected to
do. It's not our final pipeline because
we don't have our model in there yet.
But this is just to ensure that all of
the features are kind of behaving as we
expect. So we can uh we can do that and
we can take a look at the um results.
This looks pretty good. This all of our
numerical features ended up scaled
which is pretty good. And we have one
hot encoded features for that ocean
proximity over here.
Okay. So this looks pretty this looks
reasonable of those steps being applied
to the right columns. But this is a good
kind of sanity check to just run our fit
transform on our data to ensure those
steps are actually happening and they
are. You can see here the result of the
scaling and the uh the one hot encoding.
So that that all looks pretty
reasonable,
right?
And uh what we should also do is make
sure there are no nulls in this which
there shouldn't be because we did the
imper. So we should be doing uh is na
dot
sum
and there is no nulls anymore. So that
looks pretty good right? Those got
filled in uh by doing our steps. our
pipeline steps executed really nicely on
our training data um and and we were off
and running. And there's nothing unique
about the training data. We could do
this to our test data as well
and verify that those steps are running
and they would, right? There's nothing
really that special about running it on
the training data. Um it should also
work on the test features as well and it
does. You can check that for yourself.
Okay.
All right. So, that's pretty cool. We
can uh verify all that's working.
Any questions on that?
We're almost there with our full
pipeline. This this is this is not the
full pipeline, but this is something
that will run during our full pipeline.
Of course, our features are going to be
transformed according to those steps and
then it will be uh put into our model to
either predict or train with. Um
so let's do that. Let's actually build
out our final uh model here. So it's
actually going to be really easy to do.
All we need to do is um put in our
model. So here we're going to import the
ridge model here. Now, we could use any
we could use linear regression, we could
use lasso, we could use elastic net. Um,
we're just going to use ridge um uh um
just to test it out. And um we are going
to uh now put in a final pipeline. So,
we're going to use our pipeline. And so,
we're going to create a new one here and
map our pre-processing to our
pre-processing that we've already built.
So this is a column transformer that
already has all of our steps. And then
notice what comes after it is just the
model. Now that's pretty pretty basic,
but it makes sense that it should come
after that model. Um and of course this
is a generic name. We could we can name
it whatever we want to. Um model ridge
is pretty reasonable um to because it is
a ridge uh regression. But uh of course
we could we could change that.
Okay. So that builds out our uh final um
pipeline. So now we have a pipeline. And
what's great about that is this signals
that all of these steps should be
completed prior to doing anything with
this model. So all of those processing
steps are going to run and then we're
going to do ffit or predict and that. So
that's really great. It ensures that all
those steps are running together every
single time we call.predict with this
with this model. So we're just going to
use the pipeline in place of the model
to ensure that all those steps are
running together. And this is our this
is kind of our final pipeline that we
would use uh with like something like
ffit or predict.
So let me make that a note of that. Now
we can use this final pipeline just like
a regular model i.e. pipeline.fit
or pipeline.predict.
So we could use it in ei in either
fashion uh to to train the pipeline
would be this guy and then use the
pipeline to predict would be this. And
what we should realize is under the
hood, these steps are running first and
then we train it or these steps run
first then we use it for prediction.
Okay,
questions on that. Does that make sense
on this final pipeline here? It's just
now it it's really cool because we have
a pipeline
made up of a of a pipeline really,
right? a pipeline made up of a pipeline.
But that's scikitlearn allows you to do
that to compose pipelines in this way.
That's is pretty uh pretty uh normal
there.
Okay.
What I want to show you is we can
actually use this pipeline in a grid
search. So that's pretty amazing. We can
use this pipeline in any way we can use
a mo like a regular model. It's just
that now our pre-processing steps have
kind of been packaged together with our
model to ensure that they always run
anytime we do any processing with this
model. Um so for instance we can do a
grid search just like we did with a
regular with with just a model right
with just this. Um we can do the same
thing with the whole pipeline. Um, so
the only catch is that you want to make
sure in your grid you name things in the
appropriate way inside of your your uh
keys in your dictionary. So uh for
instance um inside of the grid uh we're
going to set up the alpha that would be
used with this ridge regression by
referencing its name. So this is model
ridge is this is the name of the model
inside of the pipeline. So you want to
make sure that goes first.
And then what scikitlearn does is it
recognizes parameters that belong with
this model by using a double underscore.
So the so you have underscore alpha um
here. So the double
uh underscore
signals a parameter
belonging to model ridge in the in the
pipeline.
Okay, so we have a model ridge is just a
reference to the model in our pipeline.
That's the one we're going to test out
these parameters with. and
underscore_pha is just a way to say this
alpha belongs to this model. Okay, it
belongs so it's going to be used with
that model in our pipeline. Um otherwise
it's going to work exactly the same way.
It's just we need to line up this naming
convention of of scikitlearn.
You just have to reference this to
whatever name you provided here and then
underscore parameter. So L1 ratio alpha
whatever right would go there.
Okay. So there is a range from 0.1 to2
uh step size of 0.1
um and then we do our grid search CV. So
this is exactly the same setup as we had
before. It's just that our model is now
the pipeline. So our pipeline is going
in there. Um we have our grid going in
there. we have our scoring is the same,
you know, negative absolute error. Um,
we're using five-fold cross validation
and we're parallelizing that search. Um,
so we're going to search through these
alphas and uh basically fit this to our
um data and find the best um find the
best alpha.
So, it's going to try out all those
combinations and try to come up with the
best alpha.
So looks like the best alpha was 0.1 for
the ridge.
Okay. Is the best. So then um if we
wanted to we could uh then predict using
the model um which would be doing
something like this. Um, and we could
also go back and do something like so we
could
now use um this param. So we could do
model
um equals ridge
and then we could put in our alpha.
Um, alpha is our results, our best
parameters, and then we get that model
ridge alpha. And then we just rebuild
our our pipeline
equals um pipeline and then we uh put in
this new model here. So we could do
this. This would be going back and just
um putting in our best alpha here for
this model and then uh ensuring that's
part of our our pipeline. So we're just
overwriting that pipeline with the best
model there
to get the best model in our pipeline.
Okay,
so that's all this is doing is just
initializing a new um let me actually I
can put this code in here.
This is actually just getting this is
just getting a model with the best alpha
and then reinserting that into our our
uh we're just overwriting our final
pipeline there with the best model that
we have.
So pretty cool that pipeline can be used
basically exactly like a model, right?
It's it's going right here in the grid
search and being used uh entirely like a
basic model. So we do ffit
um and that allows us to use it. We
could do predict we could even do
pipeline.predict once we we could go
back and do final pipeline.fit
um with this and then final
pipeline.predict with this and evaluate
Okay,
so pretty cool that pipeline can be used
uh basically exactly like how a model
would be any way we'd use a model.f
model.pred predict we can use a
pipeline.
So grid search is for instance something
that can use a model in there. Um but
instead of just a model we're ensuring
we have our pre-processing steps kind of
bundled with that model in this
pipeline.
Any
questions on
uh this example so far?
Were you guys able to run it up to here?
Were you able to run the grid search?
Okay, great.
Okay.
Okay. So, this is this is uh just
showing you what's actually happening
underneath the hood is uh you know,
we're doing some scaling. We're doing
some one hot encoding. Um
and we're doing some uh we're doing a
model here. And that's all part of our
pipeline. Um, and then we can use the
pipeline however we want. So for
example, I know it's not here, but for
an example, we could use um once we do
once we have this final pipeline um we
can can use the final um
pipeline to predict. So we can do um
predictions
equals final
pipeline.predict
and then we can pass in our test data.
Now what happens on this is once we have
ran our our pipeline.fit we have a
trained pipeline and then when we run
this final pipeline.predict uh this data
is going to be transformed.
It's going to go through those
transformation steps and then we would
apply our model to it at the end uh to
to make those predictions and then we
can evaluate those predictions which is
what we're doing kind of here
right.
Okay.
All right. So in conclusion uh we have
gone through a lot of stuff here. Um,
we've gone through regression, we've
done the regularization on regression.
So hopefully we have a good foundation
on regression. Um, what we're going to
do in a little bit is actually do some
additional practice with regression on a
new problem. We're going to do a
capstone problem and do some additional
regression work with that. Um, so we'll
do that next. Um but the other thing we
learned is how to evaluate the
regression using things like mean
squared error, RMSSE which is square
root of that. Um which is which is
really cool. So we have a sense of that
error which is our distance from our
prediction to the actual value. That's
always what these uh that's always what
these things are doing like this, right?
This mean absolute error metric from
scikitlearn is computing the average
distance from these predictions to these
test labels that we have right those
actual values. Um and that gives us a
sense on average how far away are our
predictions
um to see how good of a model that we
have, right? And we should be evaluating
that error generally
um against the scale of our targets to
see, you know,
uh how far off we typically are.
Okay. Any questions at all on this
lesson on regression? Uh anything we
covered up to this point?
We're going to do some more practice
with the next. We'll do the capstone. So
we get So we just do some more
regression problems.
Yeah, it's a that's another bad score.
It's a little bit hard to interpret this
though because it's m ae. Um so one
thing we could do is is compute mean
squared error and then take the square
root of it to get the RMSSE which is a
much better uh evaluation metric in
terms of our target. Um so we could
actually run that. Uh if we go back here
and um we could generate for instance we
could generate the MSE which is the mean
squared
error
and it's it's the same exact function uh
of using our predictions
um
and then we could just print that out
mean squared error.
So we have mean squared error and then
what we can do is let's take the um MP.
Square root of that.
So that way we can generate the RMSSE.
So yeah that I mean that's pretty bad.
That's uh pretty bad. Uh now let's let's
go back and look at our
uh data though. So let's take a look at
the average for our Y. Um remember one
thing we should be doing is taking a
look at um what our uh let's take a look
at y test mean
to get an average value. So the average
value is in the 200,000s. So
this isn't this isn't awful. This is
70,000. It's still a decent amount of
error. It's not as bad as the models we
have before though, right? This is an
average
median price of the house is in the
26,000 range and our error is off by
like 70,000,
right?
So, it's not good. Um, but it's not
hor like as bad as the it's not as
horrible as we've seen so far. Right.
This is a little bit better of a model.
A little bit better. closer to zero
would be better, right? Um but the
smaller the better. But uh remember this
is the um these even the mean absolute
error is is technically in similar units
as the as the uh um
as the target. So 50,000 60,000 here
70,000 it's still a decent amount of
error in terms of 200,000.
Uh so far we only come up with models
and test their accuracy with available
data. We haven't used a model to make
completely new predictions on No, we
haven't done that. Uh except we know how
to do that. Um it would so to make
predictions on new data would be exactly
how we're making them on our available
data because we actually do that all the
time. If we go back down to our model
building,
um it's it looks just like this, right?
where we take so for instance we do
predictions all the time on test data
that was never involved in the training.
So it's it's as if this data mimics new
data that we've never seen before. So if
we had new raw data it would just it
would be the same exact process. the new
now with our pipeline it makes it a
little bit easier because with the
pipeline
um the raw data will go through those
transformations which it should right
the raw data should because if it's
missing data it needs to be filled in if
it has categoricals it needs to be one
hot encoded so that's the purpose of the
pipeline actually is to make sure that
if we're dealing with raw data um those
steps can happen on the data before it
goes into the model. Right? So we so
that's kind of the purpose of the
pipeline
is to ensure that we run those steps
ahead of using it using a model with it.
But but ultimately that's how it uh any
scikitlearn model is going to be doing
the predict even if it's a pipeline
right it's going to be uh we just go
back down here it's going to be um
predict it's always going to be that on
new data
yeah
okay
really good question uh what I wanted to
do next was do some practice um I wanted
to go over to the capstone session five.
So, in this course, we have some more
capstone sessions. So, if you have a
moment, you want to pull up those
capstone session materials, the
incremental capstone session materials.
Um, we're going to be doing session five
today. So, this is just um remember it's
just extra practice that we do after
we've covered some concepts. So we are
going to do some regression practice now
that we've uh covered regression um
pretty fully and then um this will this
will be good practice before we head
into lesson four on classification. So I
just want to do this practice now while
it's fresh while the material is kind of
fresh in our in our minds. Um do you
guys have the capstone materials? Do you
know where to get it in your LMS? It's
in your LMS and the resources the uh
capstone materials you want to download
that so you can get the the data um and
the the slides for the instructions
right the or PDF I think for you guys
but uh let me ask you do you have those
we want to pull up session five if you
have it
thank you I was just going to share that
appreciate that yeah so this is going to
be session Question five. Um, you're
also going to want the data that we're
going to use with this, which is going
to be the, uh, bike rental data set.
I'll share that with you guys now.
So, we're going to be using this bike
rentals data set for this uh, for this
capstone. Um, it's the one we're going
to build a regression model off of.
Okay. So, you should have that one from
the the capstones data sets as well.
All right. So, let's go through this.
Um, we are going to be doing uh machine
learning here. So, we're going to be
doing uh so we're talking about that
kind of example product that Aura
product that was in our original
capstone. Um, and in order to do uh to
to to um make decisions, it's going to
have to build some models. Um, in this
case, it's going to be doing some bike
rental modeling. um which is the data
set we have. So we're moving away from
that healthcare data set going into this
bike rental data set as an example of
the capabilities here. Um so we're going
to do this first capstone. Um after we
do classification, we'll do this
practice. Um after we do unsupervised
learning, we'll do this practice on
clustering. And then after we do
recommendation uh which is the last
lesson um we'll come back and do
practice with building a recommendation
engine. Okay, but we're going to do this
one today. Um and then we will uh do
these other capstones as we go along. So
this will be session six, session seven,
and session 8
uh later on in the course. Okay.
Um,
okay. A little bit about the data. So,
uh, in this capstone, we're going to be
working with a, um, a shop, like a
retail shop that rents out bikes. And
they have data, um, on a per day basis
with, um, actually on a per hour basis
on the number of bikes that they rented
in every hour. Um, so maybe one hour
they rented out 20 bikes, another hour
they rented out 30 bikes. Um, another
hour they rented out 15. So they have
that data here in in the CSV. Um, they
have other kinds of data like the
environmental data like what the
temperature was at that hour, the
humidity, if there was snowfall if it's
a holiday, the wind, the visibility, the
due point, um, solar radiation,
rainfall, uh, what season it was. um
what uh what day it was like
um the functional is like if it's uh I
believe it's like if it's a weekend or
weekday um which would be uh
nonfunctional
um so
based on those features we have a bunch
of tasks okay so we have um based on the
uh rented by count hour of the day um
temperature, humidity, wind speed,
rainfall, and whatever other features
we're that are in the data set. We're
actually going to build a model to
predict the bike count required for
every hour to have a stable supply of
rented bikes. So, our goal is actually
going to be to predict the bike rental
count per hour. Um and uh we are going
to do that using all the features we
have at our disposal. um like mostly
those environmental features and what
day it is, those kind of things. Um
so we're going to load our data. We're
going to do our usual check. So check
for any NLES, handle those missing nles.
Um we're actually going to practice
converting our date because we actually
do have things based on a date here. So
we can convert it over to a datetime
object, extract different uh features
from that um like the month or the day
of the week. Um we're going to check uh
correlation using the heat map. We're
going to do some plots, some very basic
plots. The focus is going to be on the
modeling. So, I probably won't spend too
much time on the plots today, but um
there's some plot tasks in here like the
uh the the histogram of the bike count,
the histogram of the numerical features.
Um box plot of the bikes against the
categoricals.
Um so we can do some plots like that
from Seabor for instance. Uh Seabor
category plot of rented bike count
against features like hour, holiday,
rainfall, snowfall. Um, so we can see
how that stacks up against like
different hours of the day, different
holidays, rainfall, different weather
events. Um, then we're going to do then
we're going to start building our model,
right? So encode our categorical
features. Um, identify target variable
and do the split and then do scaling and
do three different models. So we're
actually going to build a linear
regression. We're going to build a lasso
regression and build a ridge regression
for the hourly bike count. and we're
going to see which model performs the
best. Now,
we could and should um build this into a
pipeline. So, that could be something we
practice. Um but this initial
instructions actually doesn't require
doing that, but I think it's really good
practice to um build our model. So, we
could use git dummies. Like it says
here, hint to use git dummies. We could
do that and build our model that way,
but more practical would be doing the
steps we did towards the end of lesson
three, which is um actually putting
everything together into a pipeline. All
right, that'd be more practical and then
fitting the pipeline and predicting with
it um for evaluation.
So, uh we'll do that. I think we'll do
that instead because I think that'll be
more practical. The the pipelines are
really uh useful. So the things we'll
have to do when we build our pipeline
will be um making sure we handle the
missing values. So we'll want that imper
in there for numerical features. We'll
want our one hot encoding very similar
pipeline to the one we built earlier. Um
and then we'll want to basically map
those to the right columns using the
column transformer. And then um we'll
have our pipeline ready to go for
training and prediction.
Right.
Okay. So, those are going to be our uh
steps. Any questions on this before we
kind of get started on it.
Okay. So, let me go over to the
notebooks and let me actually do a new
notebook.
And I'm going to name this uh
capstone
session five.
Okay.
So, I'm going to come back here and
reference the steps here. Okay. So the
steps are to load our data set and
basically check for nles.
So we should be pretty adept at doing
that. We're going to import pandas as
pd.
And I need to make sure the data sets
available. So I need to
uh load that here. So we're going to use
our bike rental.
Florida bike rentals.
And then look at the first five rows.
Uh, I got a decoder error.
UTF8 codec can decode by in
one moment.
I think we have an error in the data
set. Was anybody able to get this to
run?
Hopefully they
or did you get the same error as me?
giving me the same error.
I think we need to set a
encoding
having other issues. What other issues?
Sorry, I think this data is kind of
corrupted.
I may need to open it externally.
Okay, let me open it. There may be just
a bad character that needs to be
removed.
Okay,
let me try a different Let me try
something real quick.
I think the data is needs to be updated.
Oh, saved it as the wrong file.
One moment.
Okay, let me try uploading this.
Okay, there. That worked better. So, let
me give you the data. I think it was uh
yeah, I think it was that the degree
code was giving it some issues. So, I
actually just removed it.
You can do that or you could just work
with this one.
I removed the temperature had a strange
like degree symbol that wasn't being
parsed.
So, I just removed that in this data
set. This one should work. The one I
just sent you guys should work. Or yeah,
I guess you could try the CP1252
with the original data. See if that
works. Did that work for you?
Okay, it works with that encoding. So
yeah, you could use that encoding or
uh
remove that temperature degree which is
what I did from that. So I use that
other data set.
Yeah, that opens the data but
it's not it's not in a data frame.
We just we want it in a dataf frame to
work with our models and doing all of
our Yeah. Like that opens the file. If
it just if it was a text file that's
fine, but it's not in a data frame. Want
in a structured data frame so we could
use it.
Okay. So, one of those methodologies
hopefully works. So you can either work
with the file I sent and read it like
this or you can use the encoding.
Did you guys were other people able to
open it once they change the encoding?
Okay.
Okay, very good.
Let's see what we have. Let's do
df.info.
Let's see what we have, which is pretty
standard step to do once we first load
in some data. Um, so we have about 14
columns here. We have a date and then we
have um which is an object right now but
we're actually going to convert that
over to a datetime object in a minute.
Um we have our bike count which is what
we want to ultimately predict. This is
the target variable in our data.
Remember that's going to be the problem
is to predict that hourly bike count. Um
the hour of the day that we are
producing that bike count um is a
feature. the temperature which is in
degrees Celsius. Uh I I removed that
from there but that's it was in Celsius.
Um
and then we have a bunch of different
features which are uh a kind of boolean
like yes no holiday no holiday the
season. So these guys are going to be
good um candidates for
uh these are going to be good candidates
for um doing our uh one hot encoding.
Right? These are probably the three that
we should pick to do uh some
transformations to for our one encoding.
Okay.
Everything else is pretty numerical
though, so those should be fine to keep
those. But they would they're just going
to be good candidates for scaling,
right? Really good candidates for
scaling.
All right, let's go back to here.
And let's go back to the task.
So we loaded the data set. Um we're we
are going to check for nulls. It looks
like there's actually not any nles. So
there may not be any uh imputing that we
really need to do for this. Um but we
can check. So we can use is na or is
null and then do the sum.
Um looks like we don't have any. So that
that's good. There's no nulls.
That's pretty good. So we don't have to
worry about filling in any blanks
really.
Um but you know that would normally um
that would be an important part of our
pipeline right is filling in any nles.
Looks like we don't have to worry about
that here.
So that's good.
So we can
uh
we can extract we can do the date uh
extraction. And now one thing that we're
going to do is uh create purposely for
this we're going to create a weekend or
weekday feature. Okay, weekday or
weekend which should be really easy to
do from the day of the week. um which we
should be able to extract from this uh
from the date.
Let's go back to Were you by the way,
were you guys able to run this?
Just want to make sure everyone's with
me. Checking if there's any nulls. We
did that.
Okay.
All right. So, we're following along
there. We checked if there's any nles.
Let's do um our conversion. So, let's do
um df
uh date.
And this is going to be um
PD.2
date time and then df
date.
All right. And then let's sanity check
that that it got converted by doing
info.
Oh, we might need this format.
Um, maybe we should do
Okay, so that actually worked. Let's
see. Let's double check the
date.
Okay, so that worked. It extracted it
into
uh it extracted it into the right dates.
Some weird encoding with this
So it goes all the way up to 2018
from
the very beginning data is in 2017 the
beginning of the year right. So it go
and then the tail is all the way in
2018.
This this extracts the date.
If you guys want to run that
you guys able to run this
to extract the date. Okay, perfect. So,
this extracts the date. By the way, the
reason we need this is because our date
formats are not uniform. Uh they were
actually in different encodings. So some
of them had the year uh the string was
in a slightly different format where the
year was last or the year was first. So
if we do format mixed, it kind of
rearrang it kind of puts it in a uniform
arrangement with the year. It's year,
month, date, but it parses that out um
correctly. Uh so we have year, month,
date in there.
Um but the this strings were in a mixed
format. So we uh put that argument in
there to handle that case.
Okay. So the reason we're going to do
that is that we should be able to create
a day of the week uh feature. Um so
let's actually do that and add it to our
data frame. So we should be able to
create um day of week
And we should be able to extract um
the DF
date. And then we do our usual DT dot um
day
of week.
Okay. So, and then let's see what that
does. So, if we add that, let's actually
see um let's see what our new features
are.
So we add that it should go onto the end
of the data frame as the day of the week
is a numerical day of the week. So this
this is the uh um
this looks like a th or no a Wednesday.
I think that's the third day.
Sunday Monday being uh Monday being
actually this would be a Thursday. I
think Monday would be zero.
This would be Thursday.
Okay. So, we extract that day of the
week.
So, it's just we're adding a new column
called day of week, which extracts the
day of the week from the date feature,
the datetime feature, uh the day of the
week.
And we're verifying that here. It's now
a new feature called day of the week.
which is a number.
Does that make sense what this is doing?
Yeah, it's a Thursday. So, I think yeah,
Monday is zero
and Sunday is six. So,
um, Monday is
zero, Sunday is
six.
Yeah. So, this should be a Thursday.
This first five rows is a Thursday. And,
and by the way, the data um, this is all
on the Thursday. This is a different
hours of the day.
Different hours of the day. Um, and we
the the thing that we don't know about
this is the necessarily the time zone.
So, it may seem strange that like this
is zero, which is kind of like midnight.
Um, it could be it could be in a
different time zone. So, we don't know
that uh necessarily, but the there is um
you know 250 bikes, 200 bikes, 173. It
starts to decrease over these hours.
Just kind of notice that.
Okay. So, by the way, uh we should be
able to create a new feature. So let's
create
the weekend feature
which is um we can create using
uh weekend
and we can take our DF
um day of week
and then we can just um
say is this um greater than or equal to
uh greater than or equal to five
cuz that would be five or six. And we're
going to
um put this as
actually. Let's just leave it like that.
That should be fine.
No, let's change it as type
uh int.
So, let's see this feature.
Let's see if this works.
So this is not a weekend uh because it's
a Thursday, right? So anything bigger
than five would be bigger than or equal
to five would be five or six which would
be Saturday, Sunday. Um so we know it's
not a weekend day. This is uh zero.
Okay.
So just creating those features there
and I can paste these in. Does that make
sense what what I just did?
Any questions on that code? And this is
important. If you don't have this um it
this will be a true or a false, but we
want it to be a zero or a one. So when
it's actually true, this should be a
one. When it's false, it'll be a zero.
So, we want it to be an integer rather
than a true false. So, that's why I have
this part. That's why I did this here
to make sure it's um make sure it's an
integer.
Good.
Okay,
so we're pretty much uh doing that and
we convert it to day and extract day. Um
let's
check our heat map.
Let's do that.
So let's import Seabour
as SNS.
So, we're going to do our heat map next.
Let me uh make a note of that. So, we're
going to do
heat mapap
heat map. Um,
so we're going to do uh SNS
heat map
and then let's do uh df.correlation.
And let's make sure we do numeric
um only
equals to true
and then let's do annotate
equals to true
so that we get those uh correlation
values that are displayed on the heat
map. So what this is going to do is
create our heat map with our correlation
matrix um where we're making sure we
only do the numerical features of course
when we do uh the heat map.
Let's generate that. Um okay so we have
some
uh pretty mild um correlations. Now the
ones that are so the ones that are
correlated are the temperature and
dupoint temperature. Those are pretty
correlated.
Um so I think we could argue that we
should drop one of those. Probably just
the dupoint temperature we could drop.
Um that's a really high correlation
right right here.
Let me draw it in red. That's a this
this one here
which is the Dupoint temperature against
the regular temperature. That's super
high. 0.91
is nearly a onetoone correlation.
So, that's a good candidate to be
dropped. Uh, one of those guys, I would
argue probably the Dupoint temperature
we could get rid of dropping. Um, and
just keep the regular temperature
because they're nearly identical. Um, if
you go down here,
day of the week and weekend are
correlated. Um, and that makes sense.
That's a pretty strong correlation
because of course if it's depending on
what day of the week it is, it is the
weekend or not. So we could go and we
could go ahead and drop the day of the
week column if we wanted to because
we've already derived the weekend uh
feature which is a simpler feature. Is
it the weekend or is it not the weekend?
Um so we could probably drop day of week
and be okay. That's a pretty strong
correlation. Otherwise, it's all pretty
weak. I don't see any other strong
correlations
uh necessarily.
Um so
that seems pretty reasonable is that we
could get rid of we could get rid of
this one and we could get rid of day of
week and probably be okay,
right? Those are pretty strong
correlations. is 08 and 0.91.
Pretty strong.
Okay. Were you guys able to run that one
and see the the uh heat map? And does
that make sense based on what I'm
saying?
So, we want to run the heat map
and pass in that correlation. And we
want to make sure we turn this to true
to only do the numerical features. And
then this to true to show the value
on the uh on the heat map.
able to run that one.
Great.
Okay.
Um, that's a good question. Any
recommendation on how to choose between
the two? Uh,
not really. I think
I don't think it really matters. If
they're correlated to each other, then
uh including one of them,
um, including one of them should give
you the same information as the other,
especially if they're really correlated.
So, it doesn't really matter too much.
Um
the way that I would choose is to think
about like I'll give you an example in
this day of week versus weekend. Um we
could I would argue drop day of week
because the weekend is a simpler
feature. It's only is zero or one and
that's directly derived from day of
week.
So it has less um complexity to it. it's
a little bit simpler of a feature and I
think that makes it easier to work with.
Um,
but the truth is that uh we could do
both options and try them out and
evaluate the results, right? So, well to
be to be truly thorough, what we could
do is build a model where we've dropped
this one and do the evaluation and then
go back and build a model where we've
dropped this one and do the evaluation.
Right? So that's the proper way to do it
is to actually just build both models
with each one dropped and see which
performs better.
Um otherwise I I tend to prefer to go
for the simplicity whatever one has kind
of a lower range.
But the truth is if they're if they're
really strongly correlated, it's not
going to matter too much. Uh because
they're going to give you the same
information, right? Uh because they're
so strongly correlated. Like if I in
this data, like if I know the
temperature, I pretty much know what the
Dupoint temperature is going to be.
They're so correlated.
So it doesn't it doesn't really matter
which one I drop
but I yeah prefer to go for the
simplicity.
Okay.
Um going back to let's let's do this
plot now of the distribution of the
rented bike count. So that should be
taking a look at the
um
the rented
bike count and then just doing a
histogram
and we can see what that distribution
is. Most of it is it less than 250.
Um but there are some values that are
really high, right? There are some days
that turn out to be over 3,000 into the
3500 range. Um, that's quite a bit, but
most of the days are stacked over here
in this like 250 bucket. So, by far
that's the most. And then it kind of
decreases from there. Most values are in
that range and then uh it kind of
declines.
So that should just be this simple
histogram here.
So this is the
histogram.
That should be a simple one to build.
Were
you guys able to run that one?
Sweet.
Okay, pretty basic. Just showing us how
it's distributed. And we kind of noticed
that most of it is in the 250 bucket or
below. But there's a good amount that's
out there beyond like in the 500,
7500,000
um
all the way up to there's some days that
register with a 3,000 and above, right?
3500.
So,
in fact, we could look, we didn't do
this, but it might be worth doing is we
could look at the describe because we
never looked at the maximum.
Um,
so for the bike count, there are some
days that are zero
and the maximum is 3500. 3556
and the median is around 500 bikes.
Um
the average around 700.
So that's just our usual describe
All right, let's see what else. Uh,
plot the histogram of all numerical
features. So, uh, let's do that.
So uh luckily we have a shortcut to do
this. If you guys remember we have our
SNS um pair plot and we can pass in our
uh our data
is our df right so we can pass that in.
Um so this is kind of a nice this is a
nice thing to run that will um generate
the uh the scatter plots of all the
features against each other kind of like
the correlation but at the same time
produce the histograms on the diagonal.
Right? So this should be a nice plot to
to see all the histograms
uh in one kind of uh grid.
That's just this one.
Okay, so pretty this is obviously quite
a bit of data, but this is all the
features against each other. Um, look at
this. I mean, a couple things you see
right away is look at the ones that are
really highly correlated like the
temperature against the Dupoint
temperature. Do you guys see this strong
correlation here?
So that's that's a very indicative of a
very strong correlation, right? That's
the temperature against the Dupoint
temperature. That kind of makes sense
that it's uh that was the 0.91
correlation. So of course it's like
that.
Of course it looks like that, right? Uh
just a very strong correlation on the on
the x-axis or sorry on the diagonal is
all the histograms. They're a little bit
zoomed out so difficult to see. Um but
um we can get a sense of how some of
these are distributed like the first one
uh
sorry this first one
which is the bite count we already did
um temperature we can see how that's
distributed this third one it's kind of
evenly distributed
left of due versus temperature this one
that one is the hours
versus the due point. So, it it's kind
of evenly spaced out. Um, which would
probably be which would make sense
because this data is across two years.
So, you're going to get a lot of
seasonal data in there, right? So, it's
going to it's going to vary across like
seasons. Yeah. So, it kind of looks it's
very evenly spread out.
Okay. Were you able to get Parpot to
run?
Takes a moment to run.
takes a moment, but it produces all of
these plots, including all the
histograms.
And again, if we wanted to zoom in on
any one particular histogram, we could
do that. We just have to basically copy
and paste this code and swap out this
feature. Just swap out this feature and
we can get a zoomed in uh plot of any
one of those uh features like the
temperature
um or visibility or whatever it is.
So, what I want to do, uh, I think we'll
take one more break here. Um, and then
what we'll do is we'll come back and
just finish up our practice. Um, I'm
going to start building the model. I
know it wants us to do some additional
plotting. Um, but I want to get to
mainly the plotting elements, including
doing the pipeline one more time. Um, so
we're going to do that. Um, we're going
to practice building out the pipeline in
a similar fashion to exactly how we
built the pipeline earlier and then do
we're going to build our models and
we're going to practice that coming up.
I'll probably leave the plotting to you
guys to do as kind of homework if you
want to do that. Um, just because I want
to get to the to the modeling. Hello and
welcome to machine learning tutorial
part one. This is part one of a machine
learning series put on by SimplyLearn.
My name is Richard Kersner. I'm with the
SimplyLearn team. That's
www.simplearn.com.
Get certified, get ahead. What's in it
for you today? Well, we'll start off
with a brief explanation of why machine
learning and what is machine learning.
And then we'll get into a few of the
types of machine learning. machine
learning algorithms, linear regression,
decision trees, support vector machine,
and finally, we'll do a use case where
we're going to classify whether a recipe
is of a cupcake or a muffin using the
SPM or the support vector machine.
Sounds like a delicious way to explore
machine learning. So, why machine
learning? Why do we even care about
having these computers come up and be
able to do all these new things for us?
Well, because machines can now drive
your car for you. still very in the
infant stage but it's just exploding as
we see with uh Google's Whimo and then
Uber had their program which
unfortunately crashed. They know that
this is huge. This is going to be the
huge industry to change our whole
transportation infrastructure. Machine
learning is now used to detect over 50
eye diseases. Do you know how amazing
that is to have a computer that doublech
checkcks for the doctor for things they
might miss? That's just huge in the
health industry. pretty soon they
actually do already have that with in
some areas where maybe not for eyes but
for other diseases where they're using
the camera on your phone to help
pre-diagnose before you go in and see
the doctor. And because the machine can
now unlock your phone with your face, I
mean, that's just cool having it being
able to identify your face or your voice
and be able to turn stuff on and off for
you depending on where you're at and
what you need. Talk about an ultimate
automation our world we live in. And as
we dig in deeper, we have a nice example
of Facebook. As you can see here, they
have the Facebook post with Halloween.
Comment yes if you want it order here.
Nobody likes spam posts on Facebook that
annoy them into interacting with likes,
shares, comments, and other actions. I
remember the original ones were all if
you don't click on here, you will have
bad luck or some kind of fear factor.
Well, this is a huge thing in a social
media when people are getting spammed.
And so this tactic known as engagement
bait takes advantage of Facebook's
newsfeed algorithm by choosing
engagement in order to get the greater
reach. To eliminate engagement bait, the
company reviewed and categorized
hundreds of thousands of posts to train
a machine learning model that detects
different types of engagement bait. So
in this case, we have we're using
Facebook, but this is of course across
all the different social media. they
have different tools are building and
the Facebook scroll gif will be replaced
kind of like a virus coming in there and
notices that there's a certain setup
with Facebook and it's able to replace
it and they have like vote baiting react
baiting share baiting they have all
these different these are kind of
general titles but there certainly are a
lot of way of baiting you to go in there
and click on something so they fed all
this this data was fed into the machine
and then they have the new post the new
post comes up that takes over part of
the Facebook setup up and that's what
you're looking at. You're looking at
this new post that's replaced like a
virus has replaced that. So what
Facebook did to eliminate this is they
start scanning for keywords and phrases
like this and checks the click-through
rate. So it starts looking for people
who are clicking through it without even
looking at it or clicking through it and
it's not something that normally would
be clicked through. Once Facebook has
scanned for these keywords and phrases,
it is now able to identify the spam
coming in and this makes your life
easier. So you're not getting spammed.
It's not like walking through an airport
and in a lot of countries you have like
hundreds of people trying to sell you
time share. Come join us. Sign up for
this. Eliminates that annoyingness. So
now you can just enjoy your Facebook and
your cat pictures. Or maybe it's your
family pictures. Mine is family.
Certainly people like their cat pictures
too. Another good example is Google's
Deep Mind project Alph Go. A computer
program that plays a board game Go has
defeated the world's number one go
player and I hope I say his name right.
Kijiji the ultimate go challenge game a
three of three was on May 27th 2017 so
that was just last year that this
happened and what makes this so
important is that you know go is just is
a game so it's not like you're driving a
car or something in our real world but
they are using games to learn how to get
the machine learning program to learn
they want it to learn how to learn and
that is a huge step a lot of this is
still in its infant stage as far as
development
as we saw what happened with the as I
referred to earlier the Uber cars. They
lost their whole division because they
jumped ahead too fast. So still an
infant stage, but boy is this like the
beginning of just an amazing world that
is automated in ways we can't even
imagine what tomorrow's going to look
like. We've looked at a lot of examples
of machine learning. So let's see if we
can give a little bit more of a concrete
definition. What is machine learning?
Machine learning is the science of
making computers learn and act like
humans by feeding data and information
without being explicitly programmed. And
we see here we have a nice little
diagram where we have our ordinary
system, your computer nowadays, you can
even run a lot of the stuff on a cell
phone because cell phones have advanced
so much. And then with artificial
intelligence and machine learning, it
now takes the data and it learns from
what happened before and then it
predicts what's going to come next. And
then really the biggest part right now
in machine learning that's going on is
it improves on that. How do we find a
new solution? So we go from descriptive
where it's learning about stuff and
understanding how it fits together to
predicting what it's going to do to
postcripting coming up with a new
solution. And when we're working on
machine learning, there's a number of
different diagrams that people have
posted for what steps to go through. A
lot of it might be very domain specific.
So if you're working on photo
identification versus language versus
medical or physics, some of these are
switched around a little bit or new
things are put in. They're very specific
to the domain. This is kind of a very
general diagram. First, you want to
define your objective. Very important to
know what it is you're wanting to
predict. Then you're going to be
collecting the data. So once you've
defined an objective, you need to
collect the data that matches. You spend
a lot of time in data science collecting
data and the next step preparing the
data. You got to make sure that your
data is clean going in. There's the old
saying, bad data in, bad answer out or
bad data out. And then once you've gone
through and we've cleaned all this stuff
coming in, then you're going to select
the algorithm. Which algorithm are you
going to use? You're going to train that
algorithm. In this case, I think we're
going to be working with SVM, the
support vector machine. Then you have to
test the model. Does this model work? Is
this a valid model for what we're doing?
And then once you've tested it, you want
to run your prediction. You want to run
your prediction or your choice or
whatever output it's going to come up
with. And then once everything is set
and you've done lots of testing, then
you want to go ahead and deploy the
model. And remember I said domain
specific. This is very general as far as
the scope of doing something. A lot of
models you get halfway through and you
realize that your data is missing
something and you have to go collect new
data because you've run a test in here
someplace along the line. You're saying,
"Hey, I'm not really getting the answers
I need." So, there's a lot of things
that are domain specific that become
part of this model. This is a very
general model, but it's a very good
model to start with. And we do have some
basic divisions of what machine learning
does that's important to know. For
instance, do you want to predict a
category? Well, if you're categorizing
thing, that's classification. For
instance, whether the stock price will
increase or decrease. So in other words,
I'm looking for a yes no answer. Is it
going up or is it going down? And in
that case, we'd actually say, is it
going up? True. If it's not going up,
it's false, meaning it's going down.
This way, it's a yes, no. 01. Do you
want to predict a quantity? That's
regression. So remember, we just did
classification. Now we're looking at
regression. These are the two major
divisions in what data is doing. For
instance, predicting the age of a person
based on the height, weight, health, and
other factors. So based on these
different factors, you might guess how
old a person is. And then there are a
lot of domain specific things like do
you want to detect an anomaly? That's
anomaly detection. This is actually very
popular right now. For instance, you
want to detect money withdrawal
anomalies. You want to know when
someone's making a withdrawal that might
not be their own account. We've actually
brought this up because this is really
big right now. If you're predicting the
stock whether to buy stock or not, you
want to be able to know if what's going
on in the stock market is an anomaly,
use a different prediction model because
something else is going on. You got to
pull out new information in there or is
this just the norm? I'm going to get my
normal return on my money invested. So
being able to detect anomalies is very
big in data science these days. Another
question that comes up which is on what
we call untrained data is do you want to
discover structure in unexplored data
and that's called clustering. For
instance, finding groups of customers
with similar behavior given a large
database of customer data containing
their demographics and past buying
records. And in this case, we might
notice that anybody who's wearing
certain set of shoes goes shopping at
certain stores or whatever it is. are
going to make certain purchases. By
having that information, it helps us to
market or group people together. So then
we can now explore that group and find
out what it is we want to market to them
if you're in the marketing world. And
that might also work in just about any
arena. You might want to group people
together whether they're uh based on
their different areas and investments
and financial background, whether you're
going to give them a loan or not. before
you even start looking at whether
they're a valid customer for the bank,
you might want to look at all these
different areas and group them together
based on unknown data. So, you're not
you don't know what the data is going to
tell you, but you want to cluster people
together that come together. Let's take
a quick detour for quiz time. Oh, my
favorite. So, we're going to have a
couple questions here under quiz time
and um we'll be posting the answers in
these part two of this tutorial. So,
let's go ahead and take a look at these
quiz times questions and hopefully
you'll get them all right and it'll get
you thinking about how to process data
and what's going on. Can you tell what's
happening in the following cases? Of
course, you're sitting there with your
cup of coffee and you have your checkbox
and your pen trying to figure out what's
your next step in your data science
analysis. So, the first one is grouping
documents into different categories
based on the topic and content of each
document. Very big these days. you know,
you have legal documents, you have uh
maybe it's a sports group documents,
maybe you're analyzing newspaper
postings, but certainly having that
automated is a huge thing in today's
world. B, identifying handwritten digits
in images correctly. So, we want to know
whether uh they're writing an A or
capital A, B, C, what are they writing
out in their hand digit, their
handwriting. C behavior of a website
indicating that the site is not working
as designed. D, predicting salary of an
individual based on his or her years of
experience with HR hiring uh setup
there. So stay tuned for part two. We'll
go ahead and answer these questions when
we get to the part two of this tutorial
or you can just simply write at the
bottom and send a note to SimplyLearn
and they'll follow up with you on it.
Back to our regular content. Now these
last few bring us into the next topic
which is another way of dividing our
types of machine learning and that is
with supervised unsupervised
and reinforcement learning. Supervised
learning is a method used to enable
machines to classify predict objects,
problems or situations based on labeled
data fed to the machine. And in here you
see we have a jumble of data with
circles, triangles and squares. And we
label them. We have what's a circle,
what's a triangle, what's a square and
we have our model training and it trains
it. So we know the answer. Very
important when you're doing supervised
learning, you already know the answer to
a lot of your information coming in. So
you have a huge group of data coming in
and then you have new data coming in. So
we've trained our model. The model now
knows the difference between a circle, a
square, a triangle. And now that we've
trained it, we can send in in this case
a square and a circle goes in and it
predicts that the top one's a square and
the next one's a circle. And you can see
that this is uh being able to predict
whether someone's going to default on a
loan because I was talking about banks
earlier. Supervised learning on stock
market whether you're going to make
money or not. That's always important.
And if you are looking to make a fortune
in the stock market, keep in mind it is
very difficult to get all the data
correct on the stock market. It is very
uh it fluctuates in ways you really hard
to predict. So it's quite a roller
coaster ride. If you're running machine
learning on the stock market, you start
realizing you really have to dig for new
data. So we have supervised learning.
And if you have supervised, we need
unsupervised learning. In unsupervised
learning, machine learning model finds
the hidden pattern in an unlabeled data.
So in this case, instead of telling it
what the circle is and what a triangle
is and what a square is, it goes in
there, looks at them, and says for
whatever reason, it groups them
together. Maybe it'll group it by the
number of corners. And it notices that a
number of them all have three corners, a
number of them all have four corners,
and a number of them all have no
corners. And it's able to filter those
through and group them together. We
talked about that earlier with looking
at a group of people who are out
shopping. We want to group them together
to find out what they have in common.
And of course, once you understand what
people have in common, maybe you have
one of them who's a customer at your
store, or you have five of them are
customer at your store, and they have a
lot in common with five others who are
not customers at your store. How do you
market to those five who aren't
customers at your store yet? They fit
the demographs of who's going to shop
there, and you'd like them to shop at
your store, not the one next door. Of
course, this is a simplified version.
You can see very easily the difference
between a triangle and a circle, which
is might not be so easy in marketing.
Reinforcement learning. Reinforcement
learning is an important type of machine
learning where an agent learns how to
behave in an environment by performing
actions and seeing the result. And we
have here where the in this case a baby.
It's actually great that they used an
infant for this slide because the
reinforcement learning is very much in
its infant stages. But it's also
probably the biggest machine learning
demand out there right now or in the
future. It's going to be coming up over
the next few years is reinforcement
learning and how to make that work for
us. And you can see here where we have
our action. In the action in this one,
it goes into the fire. Hopefully, the
baby didn't it's just a little candle,
not a giant fire pit like it looks like
here. When the baby comes out and the
new state is the baby is sad and crying
because they got burned on the fire. And
then maybe they take another action. The
baby's called the agent because it's the
one taking the actions. And in this
case, they didn't go into the fire. They
went a different direction. And now the
baby's happy and laughing and playing.
Reinforcement learning is very easy to
understand because that's how as humans
that's one of the ways we learn. We
learn whether it is you burn yourself on
the stove, don't do that anymore. Don't
touch the stove. In the big picture,
being able to have machine learning
programming or an AI be able to do this
is huge because now we're starting to
learn how to learn. That's a big jump in
the world of computer and machine
learning. And we're going to go back and
just kind of go back over supervised
versus unsupervised learning.
Understanding this is huge because this
is going to come up in any project
you're working on. We have in supervised
learning, we have labeled data. We have
direct feedback. So someone's already
gone in there and said, "Yes, that's a
triangle. No, that's not a triangle."
And then you predict an outcome. So you
have a nice prediction. This is this
this new set of data is coming in and we
know what it's going to be. And then
with unsupervised trading, it's not
labeled. So we really don't know what it
is. There's no feedback. So, we're not
telling it whether it's right or wrong.
We're not telling it whether it's a
triangle or a square. We're not telling
it to go left or right. All we do is
we're finding hidden structure in the
data, grouping the data together to find
out what connects to each other. And
then you can use these together. So,
imagine you have an image and you're not
sure what you're looking for. So, you go
in and you have the unstructured data.
Find all these things that are connected
together and then somebody looks at
those and labels them. Now you can take
that label data and program something to
predict what's in the picture. So you
can see how they go back and forth and
you can start connecting all these
different tools together to make a
bigger picture. There are many
interesting machine learning algorithms.
Let's have a look at a few of them.
Hopefully this gave you a little flavor
of what's out there and these are some
of the most important ones that are
currently being used. We'll take a look
at linear regression, decision tree and
the support vector machine. Let's start
with a closer look at linear regression.
Linear regression is perhaps one of the
most well-known and well understood
algorithms in statistics and machine
learning. Linear regression is a linear
model. For example, a model that assumes
a linear relationship between the input
variables x and the single output
variable y. And you'll see this if you
remember from your algebra classes, y =
mx + c. Imagine we are predicting
distance traveled y from speed x. Our
linear regression model representation
for this problem would be y = m * x + c
or distance = m * speed + c where m is
the coefficient and c is the y
intercept. And we're going to look at
two different variations of this. First,
we're going to start with time is
constant. And you can see we have a
bicyclist. He's got a safety gear on,
thank goodness. Speed equals 10
meters/s. And so over a certain amount
of time, his distance equals 36 km. We
have a second bicyclist who's going
twice the speed or 20 m/s. And you can
guess if he's going twice the speed and
time is a constant, then he's going to
go twice the distance. And that's easy
to compute. 36 * 2, you get 72 km. And
so if you had the question of how fast
would somebody going three times that
speed or 30 m/s is, you can easily
compute the distance in our head. We can
do that without needing a computer, but
we want to do this for more complicated
data. So, it's kind of nice to compare
the two. But, let's just take a look at
that and what that looks like in a
graph. So, in a linear regression model,
we have our distance to the speed and we
have our m equals the ve slope of the
line. And we'll notice that the line has
a plus slope. And as the speed
increases, distance also increases.
Hence, the variables have a positive
relationship. And so your speed of the
person which equals y = mx plus c
distance traveled in a fixed interval of
time. And we could very easily compute
either following the line or just
knowing it's 3 * 10 m/s that this is
roughly 102 km distance that this third
bicycle has traveled. One of the key
definitions on here is positive
relationship. So the slope of the line
is positive. As distance increase so
does speed increase. Let's take a look
at our second example where we put
distance is a constant. So we have speed
equals 10 m/s. They have a certain
distance to go and it takes him 100
seconds to travel that distance. And we
have our second bicyclist who's still
doing 20 m/s. Since he's going twice the
speed, we can guess he'll cover the
distance in about half the time, 50
seconds. And of course, you could
probably guess on the third one, 100
divided by 30 since he's going three
times the speed. You can easily guess
that this is 33.3333
seconds time. We put that into a linear
regression model or a graph. If the
distance is assumed to be constant,
let's see the relationship between speed
and time. And as time goes up, the
amount of speed to go that same distance
goes down. So now your m equals a minus
v slope of the line. As the speed
increases, time decreases. Hence, the
variable has a negative relationship.
Again, there's our definition. positive
relationship and negative relationship
dependent on the slope of the line and
with a simple formula like this um and
even a significant amount of data. Let's
uh see what the mathematical
implementation of linear regression and
we'll take this data. So suppose we have
this data set where we have xyx= 1 2 3 4
5 standard series and the y value is 3
22 43. When we take that and we go ahead
and plot these points on a graph, you
can see there's kind of a nice
scattering and you could probably
eyeball a line through the middle of it.
But we're going to calculate that exact
line for linear regression. And the
first thing we do is we come up here and
we have the mean of Xi. And remember
mean is basically the average. So we
added five plus 4 plus 3 plus 2 plus 1
and divide by five. And that simply
comes out as three. And then we'll do
the same for y. We'll go ahead and add
up all those numbers and divide by five.
And we end up with a mean value of y of
i equals 2.8 where the x i references
it's an average or means value. And the
yi also equals a means value of y. And
when we plot that, you'll see that we
can put in the y= 2.8 and the x= 3 in
there on our graph. We kind of gave it a
little different color so you could sort
it out with the dashed lines on it. And
it's important to note that when we do
the linear regression, the linear
regression model should go through that
dot. Now, let's find our regression
equation to find the best fit line.
Remember, we go ahead and take our y= mx
plus c. So, we're looking for m and c.
So, to find this equation for our data,
we need to find our slope of m and our
coefficient of c. And we have y = mx + c
where m equals the sum of x - x average
* y - y average or y means and x means
over the sum of x - x means squared.
That's how we get the slope of the value
of the line. And we can easily do that
by creating some columns here. We have
xy. Computers are really good about
iterating through data. And so we can
easily compute this and fill in a graph
of data. And in our graph you can easily
see that if we have our x value of 1 and
if you remember the x i or the means
value is 3. 1 - 3 equals a -2 and 2 - 3
= a -1 so on and so forth. And we can
easily fill in the column of x - x i y -
yi. And then from those we can compute x
- x i^ 2 and x - x i * y - yi. And you
can guess it that the next step is to go
ahead and sum the different columns for
the answers we need. So we get a total
of 10 for our x - x i^2 and a total of 2
for x - x i * y - yi. And we plug those
in, we get 2/10, which equals2. So now
we know the slope of our line equals2.
So we can calculate the value of c.
That'd be the next step is we need to
know where it crosses the y ais. And if
you remember, I mentioned earlier that
the linear regression line has to pass
through the means value, the one that we
showed earlier. We can just flip back up
there to that graph. And you can see
right here, there's our means value,
which is 3 x= 3 and y= 2.8. And since we
know that value, we can simply plug that
into our formula. Y =2x + c. So we plug
that in, we get 2.8 8 =2 * 3 + C. And
you can just solve for C. So now we know
that our coefficient equals 2.2. And
once we have all that, we can go ahead
and plot our regression line. Y =2 * X +
2.2. And then from this equation, we can
compute new values. So let's predict the
values of Y using X= 1 2 3 4 5 and plot
the points. Remember the 1 2 3 4 5 was
our original x values. So now we're
going to see what y thinks they are, not
what they actually are. And we plug
those in, we get y of designated with y
of p. You can see that x= 1 = 2.4, x= 2=
2.6, and so on and so on. So we have our
y predicted values of what we think it's
going to be when we plug those numbers
in. And when we plot the predicted
values along with the actual values, we
can see the difference. And this is one
of the things that's very important with
linear regression in any of these models
is to understand the error. And so we
can calculate the error on all of our
different values. And you can see over
here we plotted um x and y and y
predict. And we draw a little line so
you can sort of see what the error looks
like there between the different points.
So our goal is to reduce this error. We
want to minimize that error value on our
linear regression model. Minimizing the
distance. There are lots of ways to
minimize the distance between the line
and the data points like sum of squared
errors, sum of absolute errors, root
mean square error, etc. We keep moving
this line through the data points to
make sure the best fit line has the
least squared distance between the data
points and the regression line. So to
recap with a very simple linear
regression model, we first figure out
the formula of our line through the
middle and then we slowly adjust the
line to minimize the error. Keep in mind
this is a very simple formula. The math
gets even though the math is very much
the same, it gets much more complex as
we add in different dimensions. So this
is only two dimensions. Y equals MX + C.
But you can take that out to X ZQ all
the different features in there and they
can plot a linear regression model on
all of those using the different
formulas to minimize the error. Let's go
ahead and take a look at decision trees.
A very different way to solve problems
in the linear regression model. Decision
tree is a treeshaped algorithm used to
determine a course of action. Each
branch of a tree represents a possible
decision, occurrence, or reaction. We
have data which tells us if it is a good
day to play golf. And if we were to open
this data up in a general spreadsheet,
you can see we have the outlook, whether
it's rainy, overcast, sunny,
temperature, hot, mild, cool, humidity,
windy, and did I like to play golf that
day? Yes or no. So, we're taking a
census. And certainly, I wouldn't want a
computer telling me when I should go
play golf or not. But you could imagine
if you got up in the night before,
you're trying to plan your day and it
comes up and says, "Tomorrow would be a
good day for golf for you in the morning
and not a good day in the afternoon or
something like that." This becomes very
beneficial and we see this in a lot of
applications coming out now where it
gives you suggestions and lets you know
what what would uh fit the match for you
for the next day or the next purchase or
the next uh whatever you know next mail
out in this case is tomorrow a good day
for playing golf based on the weather
coming in. And so we come up and let's
uh determine if you should play golf
when the day is sunny and windy. So we
found out the forecast tomorrow is going
to be sunny and windy. And suppose we
draw our tree like this. We're going to
have our humidity. And then we have our
normal, which is uh if it's if you have
a normal humidity, you're going to go
play golf. And if the humidity is really
high, then we look at the outlook. And
if the outlook is sunny, overcast, or
rainy, it's going to change what you
choose to do. So if you know that it's a
very high humidity and it's sunny,
you're probably not going to play golf
cuz you're going to be out there
miserable, fighting off the mosquitoes
that are out joining you to play golf
with you. Maybe if it's rainy, you
probably don't want to play in the rain.
But if it's slightly overcast and you
get just the right shadow, that's a good
day to play golf and be outside out on
the green. Now, in this example, you can
probably make your own tree pretty
easily cuz it's a very simple set of
data going in. But the question is, how
do you know what to split? Where do you
split your data? What if this is much
more complicated data where it's not
something that you would particularly
understand? like studying cancer, they
take about 36 measurements of the
cancerous cells and then each one of
those measurements represents how
bulbous it is, how extended it is, how
sharp the edges are, something that as a
human we would have no understanding of.
So how do we decide how to split that
data up and is that the right decision
tree? But so that's a question that's
going to come up. Is this the right
decision tree? For that we should
calculate entropy and information gain.
Two important vocabulary words there are
the entropy and the information gain.
Entropy. Entropy is a measure of
randomness or impurity in the data set.
Entropy should be low. So we want the
chaos to be as low as possible. We don't
want to look at it and be confused by
the images or what's going on there with
mixed data. And the information gain, it
is a measure of decrease in entropy
after the data set is split. Also known
as entropy reduction. information gain
should be high. So we want our
information that we get out of the split
to be as high as possible. Let's take a
look at entropy from the mathematical
side. In this case, we're going to
denote entropy as I of P of and N where
P is the probability that you're going
to play a game of golf and N is the
probability where you're not going to
play the game of golf. Now, you don't
really have to memorize these formulas.
There's a few of them out there
depending on what you're working with.
But it's important to note that this is
where this formula is coming from. So
when you see it, you're not lost when
you're running your programming, unless
you're building your own decision tree
code in the back. And we simply have a
log 2 of p + n minus n / p + n * the log
squar of n of p plus n. But let's break
that down and see what actually looks
like when we're computing that from the
computer script side. Entropy of a
target class of the data set is the
whole entropy. So we have entropy play
golf. And we look at this. If we go back
to the data, you can simply count how
many yeses and no in our complete data
set for playing golf days. In our
complete set, we find we have five days
we did play golf and nine days we did
not play golf. And so our I equals, if
you add those together, 9 + 5 is 14. And
so our I equals 5 over 14 and 9 over 14.
That's our PNN values that we plug into
that formula. And you can go 5 over
14=.36.
9 over4=64.
And when you do the whole equation, you
get the -.36
log^ 2 of.36 minus.64 log of
64. And we get a set value. We get 94.
So we now have a full entropy value for
the whole set of data that we're working
with. And we want to make that entropy
go down. And just like we calculated the
entropy out for the whole set, we can
also calculate entropy for playing golf
and the outlook. Is it going to be
overcast or rainy or sunny? And so we
look at the entropy. We have P of sunny
times E of three of two. And that just
comes out how many sunny days yes and
how many sunny days no over the total,
which is five. Don't forget to put the
we'll divide that five out later on.
equals P overcast = 4 comma 0 plus rainy
= 2a 3 and then when you do the whole
setup we have 5 over4 remember I said
there was a total of five 5 over 14 *
the i of 3 of 2 + 4 over 14 * the 4 0
and 514 over i of 23 and so we can now
compute the entropy of just the part
that has to do with the forecast and we
get 693 similar We can calculate the
entropy of other predictors like
temperature, humidity and wind. And so
we look at the gain outlook. How much
are we going to gain from this entropy
play golf minus entropy play golf
outlook? And we can take the original
0.94 for the whole set minus the entropy
of just the rainy day and temperature
and we end up with a gain of.247.
So this is our information gain.
Remember we define entropy and we define
information gain. The higher the
information gain, the lower the entropy,
the better. The information gain of the
other three attributes can be calculated
in the same way. So we have our gain for
temperature equals 0.029.
We have our gain for humidity
equals.152.
And our gain for a windy day equals
0048. And if you do a quick comparison,
you'll see the 247 is the greatest gain
of information. So that's the split we
want. Now let's build the decision tree.
So, we have the outlook. Is it going to
be sunny, overcast, or rainy? That's our
first split because that gives us the
most information gain. And we can
continue to go down the tree using the
different information gains with the
largest information. We can continue
down the nodes of the tree where we
choose the attribute with the largest
information gain as the root node and
then continue to split each subnode with
the largest information gain that we can
compute. And although it's a little bit
of a tongue twister to say all that, you
can see that it's a very easy to view
visual model. We have our outlook. We
split it three different directions. If
the outlook is overcast, we're going to
play. And then we can split those
further down if we want. So if the over
outlook is sunny, but then it's also
windy. If it's uh windy, we're not going
to play. If it's uh not windy, we'll
play. So, we can easily build a nice
decision tree to guess what we would
like to do tomorrow and give us a nice
recommendation for the day. So, we want
to know if it's a good day to play golf
when it's sunny and windy. Remember the
original question that came out,
tomorrow's weather report is sunny and
windy. You can see by going down the
tree, we go outlook sunny, outlook
windy. We're not going to play golf
tomorrow. So, our little smartwatch pops
up and says, I'm sorry, tomorrow's not a
good day for golf. It's going to be
sunny and windy. And if you're a huge
golf fan, you might go, "Uh oh, it's not
a good day to play golf." We can go in
and watch a golf game at home. So, we'll
sit in front of the TV instead of being
out playing golf in the wind. Now that
we looked at our decision tree, let's
look at the third one of our algorithms
we're investigating. Support vector
machine. Support vector machine is a
widely used classification algorithm.
The idea of support vector machine is
simple. The algorithm creates a
separation line which divides the
classes in the best possible manner. For
example, dog or cat, disease or no
disease. Suppose we have a labeled
sample data which tells height and
weight of males and females. A new data
point arrives and we want to know
whether it's going to be a male or a
female. So we start by drawing a line.
We draw decision lines. But if we
consider decision line one, then we will
classify the individual as a male. And
if we consider decision line two, then
it'll be a female. So you can see this
person kind of lies in the middle of the
two groups. So it's a little confusing
trying to figure out which line they
should be under. We need to know which
line divides the classes correctly. But
how the goal is to choose a hyper plane
and that is one of the key words they
use when we talk about support vector
machines. Choose a hyper plane with the
greatest possible margin between the
decision line and the nearest point
within the training set. So you can see
here we have our support vector. We have
the two nearest points to it and we draw
a line between those two points. And the
distance margin is the distance between
the hyper plane and the nearest data
point from either set. So we actually
have a value and it should be equal
distant between the two points that
we're comparing it to. When we draw the
hyperplanes, we observe that line one
has a maximum distance. So we observe
that line one has a maximum distance
margin. So we'll classify the new data
point correctly. And our result on this
one is going to be that the new data
point is MEL. One of the reasons we call
it a hyper plane versus a line is that a
lot of times we're not looking at just
weight and height. We might be looking
at 36 different features or dimensions.
And so when we cut it with a hyper
plane, it's more of a three-dimensional
cut in the data, multi-dimensional that
cuts the data a certain way. And each
plane continues to cut it down until we
get the best fit or match. Let's
understand this with the help of an
example. Problem statement. You always
start with a problem statement when
you're going to put some code together.
We're going to do some coding now.
Classifying muffin and cupcake recipes
using support vector machines. So the
cupcake versus the muffin. Let's have a
look at our data set. And we have the
different recipes here. We have a muffin
recipe that has so much flour. I'm not
sure what measurement 55 is in, but it
has 55, maybe it's ounces, but it has a
certain amount of flour, certain amount
of milk, sugar, butter, egg, baking
powder, vanilla, and salt. And so based
on these measurements, we want to guess
whether we're making a muffin or a
cupcake. And you can see in this one, we
don't have just two features. We don't
just have height and weight as we did
before between the male and female. In
here, we have a number of features. In
fact, in this, we're looking at eight
different features to guess whether it's
a muffin or a cupcake. What's the
difference between a muffin and a
cupcake? Turns out muffins have more
flour, while cupcakes have more butter
and sugar. So, basically, the cupcakes a
little bit more of a dessert, where the
muffin's a little bit more of a fancy
bread. But how do we do that in Python?
How do we code that to go through
recipes and figure out what the recipe
is? And I really just want to say
cupcakes versus muffins like some big
professional wrestling thing. Before we
start in our cupcakes versus muffins, we
are going to be working in Python.
There's many versions of Python, many
different editors. That is one of the
strengths and weaknesses of Python is it
just has so much stuff attached to it.
It's one of the more popular data
science programming packages you can
use. In this case, we're going to go
ahead and use Anaconda in Jupyter
Notebook. The Anaconda Navigator has all
kinds of fun tools. Once you're into the
Anaconda Navigator, you can change
environments. I actually have a number
of environments on here. We'll be using
Python 36 environment. So, this is in
Python version 36. Although, it doesn't
matter too much which version you use. I
usually try to stay with the 3x because
they're current unless you have a
project that's very specifically in
version 2x 27 I think is usually what
most people use in the version two. And
then once we're in our um Jupiter
notebook editor, I can go up and create
a new file and we'll just jump in here.
In this case, we're doing SPM muffin
versus cupcake. And then let's start
with our packages for data analysis.
And we almost always use a couple
there's a few very standard packages we
use. We use import oops import
numpy
that's for number python. They usually
denote it as np that's very comma that's
very common. And then we're going to
import pandas as pd. And numpy deals
with number arrays. There's a lot of
cool things you can do with the numpy uh
setup as far as multiplying all the
values in an array in a numpy array data
array. Pandas I can't remember if we're
using it actually in this data set. I
think we do as an import it makes a nice
data frame. And the difference between a
data frame and a numpy array is that a
data frame is more like your Excel
spreadsheet. You have columns, you have
indexes. So you have different ways of
referencing it easily viewing it. And
there's additional features you can run
on a data frame. And pandas kind of sits
on numpy. So they you need them both in
there. And then finally, we're working
with the support vector machine. So from
sklearn, we're going to use the sklearn
model. Import SVM support vector
machine.
And then as a data scientist, you should
always try to visualize your data. Some
data obviously is too complicated or
doesn't make any sense to the human. But
if it's possible, it's good to take a
second look at it so that you can
actually see what you're doing. Now, for
that, we're going to use two packages.
We're going to import mapplot
library.pipplot as plt. Again, very
common. And we're going to import seabor
as sns. And we'll go ahead and set the
font scale in the SNS right in our
import line. That's what this U
semicolon followed by a line of data.
We're going to set the SNS. And these
are great because the the seabour sits
on top of map plot library just like
pandas sits on numpy. So it adds a lot
more features and uses and control.
We're obviously not going to get into
mattplot library and seabour. It' be its
own tutorial. We're really just focusing
on the SVM, the support vector machine
from sklearn. And since we're in Jupyter
notebook, uh we have to add a special
line in here for our mattplot library.
And that's your percentage sign or amber
sign mattplot library in line. Now, if
you're doing this in just a straight
code project, a lot of times I use like
Notepad++
and I'll run it from there. You don't
have to have that line in there because
it'll just pop up as its own window on
your computer depending on how your
computer's set up because we're running
this in the Jupyter notebook as a
browser setup. This tells it to display
all of our graphics right below on the
page. So that's what that line is for.
Remember the first time I ran this, I
didn't know that and I had to go look
that up years ago. It's quite a
headache. So mattplot library inline is
just because we're running this on the
web setup and we can go ahead and run
this. make sure all our modules are in.
They're all imported, which is great. If
you don't have them import, you'll need
to go ahead and pip. Use the pip or
however you do it. There's a lot of
other install packages out there,
although pip is the most common. And you
have to make sure these are all
installed on your Python setup. The next
step, of course, is we got to look at
the data. You can't run a model for
predicting data if you don't have actual
data. So, to do that, let me go ahead
and open this up and take a look. And we
have our uh cupcakes versus muffins. and
it's a CSV file or CSV meaning that it's
commaepparated variable
and it's going to open it up in a nice
uh spreadsheet for me. And you can see
up here we have the type we have muffin
muffin muffin cupcake cupcake cupcake
and then it's broken up into flour,
milk, sugar, butter, egg, baking powder,
vanilla and salt. So we can do is we can
go ahead and look at this data also in
our Python.
Let us create a variable recipes equals
we're going to use our pandas module
read CSV. Remember is a commaepparated
variable
and the file name happened to be
cupcakes versus muffins. Oops, I got
double brackets there.
Do it this way.
There we go. cupcakes versus muffins.
Because the program I loaded or the the
place I saved this particular Python
program is in the same folder, we can
get by with just the file name. But
remember, if you're storing it in a
different location, you have to also put
down the full path on there.
And then because we're in pandas, we're
going to go ahead and you can actually
in line you can do this, but let me do
the full print. You can just type in
recipes.head head in the Jupyter
notebook. But if you're running in code
in a different script, you'd need to go
ahead and type out the whole print
recipes.
And Pandanda's knows that's going to do
the first five lines of data. And if we
flip back on over to the spreadsheet
where we opened up our CSV file,
uh you can see where it starts on line
two. This one calls it zero. And then 2
3 4 5 6 is going to match. Go and close
that out because we don't need that
anymore. And it always starts at zero.
And these are it automatically indexes
it since we didn't tell it to use an
index in here. So that's the index
number for the left hand side. And it
automatically took the top row as
labels. So pandas using it to read a CSV
is just really slick and fast. One of
the reasons we love our pandas, not just
because they're cute and cuddly teddy
bears.
And let's go ahead and plot our data.
And I'm not going to plot all of it. I'm
just going to plot the uh sugar and
flour. Now, obviously, you can see where
they get really complicated if we have
tons of different features. And so,
you'll break them up and maybe look at
just two of them at a time to see how
they connect.
And to plot them, we're going to go
ahead and use Seabor. So, that's our
SNS. And the command for that is SNS.LM
plot. And then the two different
variables I'm going to plot is flour and
sugar.
Data equals recipes. The hue equals
type. And this is a lot of fun because
it knows that this is pandas coming in.
So this is one of the powerful things
about pandas mixed with seabor and doing
graphing. And then we're going to use a
pallet set one. There's a lot of
different sets in there. You can go look
them up for seabor. We do a regular fit
regular equals false. So, we're not
really trying to fit anything. And it's
a scatter KWS.
A lot of these settings you can look up
in Seabor. Half of these you could
probably leave off when you run them.
Somebody played with this and found out
that these were the best settings for
doing a Seabor plot. And let's go ahead
and run that. And because it does it in
line, it just puts it right on the page.
And you can see right here that just
based on sugar and flour alone, there's
a definite split. And we use these
models because you can actually look at
it and say, "Hey, if I drew a line right
between the middle of the blue dots and
the red dots, we'd be able to do an SVM
and and a hyper plane right there in the
middle.
Then the next step is to format or
pre-process
our data.
And we're going to break that up into
two parts.
We need a type label. And remember,
we're going to decide whether it's a
muffin or a cupcake. Well, a computer
doesn't know muffin or cupcake. It knows
zero and one. So, what we're going to do
is we're going to create a type label.
And from this we'll create a numpy array
nump where and this is where we can do
some logic. We take our recipes from our
panda and wherever type equals muffin
it's going to be zero. And then if it
doesn't equal muffin which is cupcakes
it's going to be one. So we create our
type label. This is the answer. So when
we're doing our training model remember
we have to have a a training data. This
is what we're going to train it with. Is
that it's zero or one? it's a muffin or
it's not.
And then we're going to create our
recipe features.
And if you remember correctly from right
up here, the first column is type.
So we really don't need the type column
because that's our muffin or cupcake.
And in pandas, we can easily sort that
out.
We take our value recipes
columns. That's a pandas function built
into pandas.
values converting them to values. So
it's just the column titles going across
the top and we don't want the first one.
So what we do is since it always starts
at zero, we want one
colon till the end.
And then we want to go ahead and make
this a list. And this converts it to a
list of strings.
And then we can go ahead and just take a
look and see what we're looking at for
the features. Make sure it looks right.
Me go ahead and run that.
And I forgot the S on recipes. So, we'll
go ahead and add the S in there and then
run that. And we can see we have flour,
milk, sugar, butter, egg, baking powder,
vanilla, and salt. And that matches what
we have up here, right? Where we printed
out everything but the type. So, we have
our features and we have our label.
Now, the recipe features is just the
titles of the columns. We actually need
the ingredients.
And at this point, we have a couple
options. One, we could run it over all
the ingredients.
And when you're doing this, usually you
do. But for our example, we want to
limit it so you can easily see what's
going on because if we did all the
ingredients, we have, you know, that's
what, um, seven, eight different
hyperplanes that would be built into it.
We only want to look at one. So you can
see what the SVM is doing.
And so we'll take our recipes and we'll
do just flour and sugar. Again, you can
replace that with your recipe features
and do all of them, but we're going to
do just flour and sugar. And we're going
to convert that to values. We don't need
to make a list out of it because it's
not string values. These are actual
values on there. And we can go ahead and
just print
ingredients. And you can see what that
looks like.
Uh, and so we have just the nanoflower
and sugar, just the two sets of plots.
And just for fun, let's go ahead and
take this over here and take our recipe
features.
And so if we decided to use all the
recipe features, you'll see that it
makes a nice column of different data.
So it just strips out all the labels and
everything. We just have just the
values. But because we want to be able
to view this easily in a plot later on,
we'll go ahead and take that and just do
flour and sugar.
And we'll run that. And you'll see it's
just the two columns.
So the next step is to go ahead and fit
our model.
We'll go ahead and just call it model.
And it's a SVM. We're using a package
called SVC.
In this case, we're going to go ahead
and set the kernel equals linear. So,
it's using a specific setup on there.
And if we go to the reference on their
website for the SVM,
you'll see that there's about there's
eight of them here. Three of them are
for regression.
Three are for classification. The SVC,
support vector classification, is
probably one of the most commonly used.
And then there's also one for detecting
outliers and another one that has to do
with something a little bit more
specific on the model. But SVC and SVR
are the two most commonly used standing
for support vector classifier and
support vector regression. Remember
regression is an actual value, a float
value or whatever you're trying to work
on. And SBC is a classifier. So it's a
yes, no, true, false.
But for this we want to know 01 muffin
cupcake. If we go ahead and create our
model and once we have our model
created, we're going to do model.fit.
And this is very common, especially in
the sklearn. All their models are
followed with the fit command.
And what we put into the fit, what we're
training with it is we're putting in the
ingredients, which in this case we
limited to just flour and sugar, and the
type label. Is it a muffin or cupcake?
Now, in more complicated data science
series, you'd want to split into, we
won't get into that today, where you
split it into training data and test
data. And they even do something where
they split it into thirds, where a third
is used for where you switch between
which one's training and test. There's
all kinds of things go into that. It
gets very complicated when you get to
the higher end. Not overly complicated,
just an extra step, which we're not
going to do today because this is a very
simple set of data.
And let's go ahead and run this. And now
we have our model fit. And uh I got an
error here. So let me fix that real
quick. It's capital SBC. It turns out
I did it lowercase.
Support vector
classifier. There we go. Let's go ahead
and run that. And you'll see it comes up
with all this information that it prints
out automatically. These are the
defaults of the model. You notice that
we changed the kernel to linear. And
there's our kernel linear on the
printout. And there's other different
settings you can mess with.
We're going to just to leave that alone
for right now. For this, we don't really
need to mess with any of those.
So, next we're going to dig a little bit
into our newly trained model. And we're
going to do this so we can show you on a
graph.
And let's go ahead and get the
separating.
and we're going to say uh we're going to
use a W for our variable on here and
we're going to do model.coreeficient_0.
So what the heck is that? Again, we're
digging into the model. So we've already
got a prediction and a train. This is a
math behind it that we're looking at
right now. And so the w is going to
represent two different coefficients.
And if you remember, we had y = mx + c.
So these coefficients are connected to
that but in two-dimensional it's a
plane.
We don't want to spend too much time on
this because you can get lost in the
confusion of the math. So if you're a
math wiz this is great. You can go
through here and you'll see that we have
a= minus w of 0 over w of 1. Remember
there's two different values there. And
that's basically the slope that we're
generating.
And then we're going to build an xx.
What is xx? We're going to set it up to
a numpy array. There's our np line
space. So we're creating a line
of values between 30 and 60. So it just
creates a set of numbers for x. And then
if you remember correctly, we have our
formula y equals the slope * x
plus the intercept. Well, to make this
work, we can do this as y
equals the slope times each value in
that array. That's the neat thing about
numpy. So, when I do a * xx, which is a
whole numpy array of values, it
multiplies a across all of them. And
then it takes those same values and we
subtract the model intercept. That's
your uh we had mx plus c. So, that'd be
the c from the formula y mx plus c.
And that's where all these numbers come
from. A little bit confusing because
it's digging out of these different
arrays. And then what we want to do is
we're going to take this and we're going
to go ahead and plot it. So plot the
parallels to separating hyper plane that
pass through the support vectors. And so
we're going to create B equals a model
support vectors. Pulling our support
vectors out there. Here's our y, which
we now know is a set of data. And we
have uh we're going to create y down = a
* xx + b1 - a * b 0. And then model
support vector b is going to be set that
to a new value the minus1 setup. And y y
up = a * xx + b1 - a * b 0. And we can
go ahead and just run this to load these
variables up. If you wanted to know
understand a little bit more of what's
going on, you can see if we print
y, let me just run that. You can see
it's an array. This is a line. It's
going to have in this case between 30
and 60. So there's going to be 30
variables in here. And the same thing
with y y up y y y y y y y y y y y y y y
y y y y y y y y y y y y y y y y y y y y
y y y y y y y y y y y y y y y y y y y y
y y y y y y y y y y y y y y y y y y y y
y y y y y y down and we'll we'll plot
those in just a minute on a graph so you
can see what those look like.
Just go ahead and delete that out of
here and run that. So, it loads up the
variables. Nice clean slate. I'm just
going to copy this from before. Remember
this? Our SNS, our Seabor plot, LM plot,
flower, sugar. And I'll just go and run
that real quick so you can see what
remember what that looks like. It's just
a straight graph on there. And then one
of the neat things is because Seabour
sits on top of piplot,
we can do the piplot for the line going
through. And that is simply plt.plot
And that's our xx and y are two
corresponding values xy. And then
somebody played with this to figure out
that the line width equals 2 and the
color black would look nice. So let's go
ahead and run this whole thing with the
pi plot on there. And you can see when
we do this, it's just doing flour and
sugar on here.
Corresponding line between the sugar and
the flour and the muffin versus cupcake.
Um, and then we generated the support
vectors, the y down and y up. So let's
take a look and see what that looks
like.
So we'll do our plot.
And again, this is all against xx,
our x value, but this time we have y
down.
And let's do something a little fun with
this. We can put in a k dash dash. That
just tells it to make it a dotted line.
And if we're going to do the down one,
we also want to do the up one. So here's
our y
up. And when we run that, it adds both
sets of line. And so here's our support.
And this is what you expect. You expect
these two lines to go through the
nearest data point. So the dash lines go
through the nearest muffin and the
nearest cupcake when it's plotting it.
And then your SVM goes right down the
middle. So it gives it a nice split in
our data. And you can see how easy it is
to see based just on sugar and flour
which one's a muffin or a cupcake.
Let's go ahead and create a function
to predict
muffin or cupcake.
I've got my uh recipes. I pulled off the
um internet and I want to see the
difference between a
muffin or a cupcake. And so we need a
function to push that through. And uh we
create a function with deaf. And let's
call it muffin or cupcake. And remember,
we're just doing flour and sugar today.
We're not doing all the ingredients. And
that actually is a pretty good split.
You really don't need all the
ingredients to know it's flour and
sugar. And let's go ahead and do an if
else statement. So if model predict
is of flower and sugar equals zero. So
we take our model and we do run a
predict. It's very common in sklearn
where you have a predict. You put the
data in and it's going to return a
value. In this case if it equals zero
then print you're looking at a muffin
recipe. Else if it's not zero that means
it's one and you're looking at a cupcake
recipe. That's pretty straightforward
for
function or def for definition. Deaf is
how you do that in Python. And of
course, if you're going to create a
function, you should run something in
it. And so, let's run a cupcake. And
we're going to send it values 50 and 20.
A muffin or a cupcake. I don't know what
it is. And let's run this and just see
what it gives us. It says, "Oh, it's a
muffin. You're looking at a muffin
recipe." So, it very easily predicts
whether we're looking at a muffin or a
cupcake recipe. Let's plot this. There
we go. Plot this on the graph so we can
see what that actually looks like. And
I'm just going to copy and paste it from
below where we plotting all the points
in there.
So, this is nothing different than we
did before. If I run it, you'll see it
has all the points and the lines on
there. And what we want to do is we want
to add another point. And we'll do
pltot.
And if you remember correctly, we did
for our test we did 50
and 20. And then somebody went in here
and decided we'll do yo for yellow or
it's kind of a orangeish yellow color is
going to come out. Marker size nine.
Those are settings you can play with.
Somebody else played with them to come
up with the right setup so it looks
good. And you can see there it is
graphed clearly a muffin.
In this case in cupcakes versus muffins,
the muffin has won. And if you'd like to
do your own muffin cupcake contender
series, you certainly can send a note
down below and the team at SimplyLearn
will send you over the data they use for
the muffin and cupcake. And that's true
of any of the data. We didn't actually
run a plot on it earlier. We had men
versus women. You can also request that
information to run it on your data
setup. So you can test that out.
So to go back over our setup, we went
ahead for our support vector machine
code. We did a predict 40 parts flour,
20 parts sugar. I think it was different
than the one we did whether it's a
muffin or a cupcake. Hence, we have
built a classifier using SVM which is
able to classify if a recipe is of a
cupcake or a muffin. Which wraps up our
cupcake versus muffin. So the key
takeaways, what is machine learning? We
discussed that with some of the
different aspects of machine learning on
there. We went into types of machine
learning. If you memorize we have
supervised, unsupervised and
reinforcement learning. We discussed
regression line or best fit and we did
the building a decision tree and what
the logic is behind that. And finally we
did classification using SVM support
vector machine and we did the code in
there. Today we are diving into machine
learning, the technology behind things
like Netflix recommendations, CD, and
even the face unlock of your phone.
Machine learning helps devices get
smarter by learning from data and
predicting what we might like or need.
And here's why machine learning is huge
for your career. Right now, machine
learning jobs are among the fastest
growing roles worldwide. Companies in
every industry, tech, healthcare,
finance, and more, are looking for
people with machine learning skills to
improve their products, automate tasks,
and make smarter decisions. Machine
learning engineers in the US earn around
$112,000 on average with plenty of room
for growth as you gain experience. So,
if you want to jump into this exciting
field, learning machine learning can
open doors to highpaying in- demand
jobs. So in this video I'll guide you
through the ultimate road map to master
machine learning in 2025 one step at a
time. So let's get started. So in the
first month start with the foundations
of programming. So programming is a
language you'll use to communicate with
your computer and bring machine learning
algorithms to life. So this month is all
about Python, the language of choice for
most machine learning practitioners. So
here's what to focus on. First, learn
Python basics. Begin with Python's
fundamentals like variables, data types,
loops and functions. So spend time
writing small programs daily to get
comfortable. After that explore the key
libraries like numpy, pandas and
scikitlearn. So numpy is for numerical
operations. It makes handling large data
sets faster and easier. And pandas is to
manipulate and analyze data. So pandas
allow you to filter, sort and reshape
data in a breeze. And then scikitlearn
is for implementing algorithms in just a
few lines of code. So now you might have
heard about R, another language used in
machine learning. But don't stress about
it now. Python will serve you well,
especially as a beginner, because it's
simpler and more flexible. So aim to
spend an hour or two each day coding. By
the end of this month, you'll have a
solid base to build on. Now, in the
second month, get organized with version
control and data structures. So this
month is about learning how to organize
and manage your code effectively and
sharpening your problem solving skills
with data structures and algorithms. So
first is version control with git. So
think of git as your project history
tracker. So imagine working on a big
project and making changes then
realizing something went wrong. You want
to go back to an earlier version, right?
So that's where git comes in. And here's
what you should practice. Number one is
committing changes. So save different
versions of your work as you progress.
And then branching which means work on
separate features without affecting your
main code. And then comes merging which
means combining changes from different
versions once they are ready. So you
have to set up an account on GitHub or
GitLab to store your projects online. So
not only will this be super useful, but
it'll also start building your
portfolio. Now next is data structures
and algorithm. So think of data
structures like tools in a toolkit. So
each one like arrays, stacks, cues, etc.
serves a specific purpose. So here's how
to approach them. Number one, arrays and
lists. Now arrays and lists are for
storing data in sequence. After that,
you can get familiar with stacks and
cues. So stacks and cues are for tasks
that need ordered data access. And then
you have sorting and searching
algorithms. So these make your programs
more efficient. And that's super
important in machine learning where data
can get massive. So the goal here is to
build up your problem solving skills
which are key to machine learning
success. So take it slow, practice daily
and you'll see progress. Now in the
third month, learn to access data with
SQL. So in machine learning, a lot of
work involves accessing and organizing
data from databases. So SQL, a
structured query language, is your
ticket to getting the data you need for
training ML models. So here's what you
should focus on. Select and where. So
these commands help you pull specific
pieces of data and then you can move on
to joins. Joins usually combine data
from different tables. So this is so
powerful that you'll use it all the
time. And then comes group by and
aggregate functions. They are great for
summarizing data to find patterns. So
spend time working with sample databases
you can find online and practice writing
queries. Being comfortable with SQL will
save you time when preparing data for
your models. Now after completing the
third month you can move on to
mathematics which is building your
analytical mind. So this month we are
tackling the math behind machine
learning. So don't worry you don't need
to be a math genius but understanding
certain concepts will make everything
feel less mysterious. So in this month
you have to focus on linear algebra. So
this is the math behind how models see
data. So you can study vectors, matrices
and operations like multiplication. Next
comes calculus. So you'll use calculus
to help your models learn. So you have
to focus on derivatives and gradients
which help minimize errors in your
model. And then you can move on to
probability and statistics. So
understanding probability helps you make
sense of data. So learn about
distributions like normal distribution,
bormal distribution and then variance
and standard deviation. So once you have
learned maths, next you'll be moving on
to data handling and visualization which
is the heart of machine learning as you
all know. So with Matt under your belt,
it's time to dig into data handling and
visualization. So data preparation is
vital because your model is only as good
as the data you feed it. So number one
comes data manipulation. So using pandas
and numpy, you'll clean and organize
your data. You might be removing missing
values like clean up messy data so it
doesn't confuse your model. And then
you'll learn transforming variables like
converting data into formats that work
for models. And then you will move on to
encoding categorical data like changing
text data like female or male into
numbers. Now once you're done with data
manipulation, next comes data
visualization. So visualization is how
you get to see your data before training
a model. So here you have to learn
mattplot lip and seabboard. So you can
create line charts, histograms, scatter
plots and heat maps. So this lets you
explore patterns and spot outliers. So
understanding these patterns in your
data is crucial for building effective
models. Now in the sixth month you'll be
moving on to the machine learning
fundamentals. So now it's time to start
building your own models. So you will
focus on two main types of machine
learning this month. Number one comes
the supervised learning. So this is when
you train a model on label data where
the outcome is already known. So you'll
work with algorithms like linear
regression which predicts a continuous
outcome. Then you'll work with decision
trees which breaks down decisions into a
tree structure. And then you have
support vector machines under supervised
learning which updates data into
classes. Now after supervised learning
comes unsupervised learning. So here
your model identifies patterns in data
without labeled outcomes. So two popular
techniques in unsupervised learning is
number one clustering like K means
clustering which means group similar
data points and then you have
dimensionality reduction. This reduces
data complexity by focusing on key
features. So you can use scikitle learn
to try out these algorithms on sample
data sets. So this will give you
hands-on experience with model training
and you will learn to fine-tune them to
get better results. Now before moving
on, if you are interested in advancing
your career in the field of AI and
machine learning, simple learns
post-graduate program delivered in
collaboration with Purdue University and
IBM is a perfect opportunity. This
highly ranked program offers a
comprehensive curriculum covering
essential topics like machine learning,
deep learning, NLP, computer vision,
reinforcement learning, generative AI,
prompt engineering, and many more. With
hands-on experience to 25 plus projects
and access to 20 plus cutting edge
tools, you will gain the skills needed
to excel in today's competitive job
market. So join now and elevate your
expertise with the backing of Produce
academic excellence and IBM's
industry-leading insight. You can find
the course link in the description box
and pin comments. Now moving on to the
seventh month, you'll be building and
training models with advanced libraries.
So by now you have experimented with
some basic models. So let's step it up
with advanced tools like TensorFlow and
PyTorch. So these libraries offer more
flexibility and power. So TensorFlow and
PyTorch. So here you can start with
simple models and work your way up. So
these libraries allow for building
neural networks which you'll be studying
more on the next month. Now once you
have become familiar with TensorFlow and
PyTorch, you can move on to model
training and evaluation. So you have to
learn to split data into training and
testing sets and evaluate models using
metrics like accuracy and precision. So
your goal this month should be to get
comfortable with these libraries and
understand how they handle data and
model training behind the scenes. So
once you are done with this, you'll be
moving on to the eighth month where
you'll be dealing with advanced machine
learning. So this month's concept will
be number one on n symbol learning which
means combining multiple models to get
better predictions. So here you'll be
learning about bagging for example
random forests here multiple decision
trees make predictions and then you have
boosting like ada boost xg boost so
models learn from each other's mistakes
over here and after ensemble learning
comes deep learning. So here you explore
neural networks which mimic the human
brain. So you'll learn about neural
network basics. So you can start with
simple fully connected networks and then
you can move on to back propagation and
gradient descent. So these helps your
model learn and improve. So you can use
TensorFlow or PyTorch to practice
building neural networks. So you can
work on projects to reinforce these
concepts. Now moving on, you have two
specialize on topics like NLP and
computer vision. So machine learning
applications are so powerful and here
you'll get a taste of two major fields
which is NLP or natural language
processing. So here they work with text
data with tasks like sentiment analysis
and text classification. So you can
start with basic pre-processing like
tokenization, stop word removal and move
to building simple NLP models. After
that you can try computer vision. So for
image data you have to learn CNN
convolutional neural networks. So these
network analyze visual patterns making
them ideal for image classification. So
you practice with open data sets like
text, documents or images and apply the
concepts you will learn to see results
in real world applications. Now in the
10th month you'll be dealing with model
deployment which is bringing your models
to life. So here you'll be using Flask
or Django. So you can use these
frameworks to create a web API so users
can interact with your model. For
example, build a web app that lets
people upload images for classification.
And then you can also try out Docker. So
package your model and its dependencies
so it can run on any machine. So this is
super helpful for deploying models
without compatibility issues. So by the
end of this month, you'll be able to
share your models with the world. So
moving on to the 11th month, you'll be
starting with cloud and production. So
this month, you'll learn how to deploy
models on the cloud and ensure they
perform well in real world environments.
So you'll be dealing with cloud
platforms like AWS, Google Cloud or
Azure. So you have to learn to deploy
models of the cloud provider
accessibility and scalability. And then
comes monitoring and maintenance. So
understand how to track your models
performance over time and update it as
needed. So these skills are essential
for maintaining models in production and
ensuring they stay reliable. And finally
you will be creating real world projects
and portfolio building. So here you have
to choose topics that interest you and
showcase your skills. So first you can
start with full projects. So complete
projects that go from data cleaning and
model building to deployment. So ideas
could be a sentiment analysis tool or an
image recognition app. And then you have
to build your portfolio. So organize and
document your projects, host them on
GitHub and create an online portfolio to
share with potential employers or
collaborators. So by following this road
map, you'll be well prepared to handle
real world machine learning challenges
and have an impressive portfolio to show
for it.
>> Welcome to machine learning tutorial
part two. My name is Richard Kersner
with the SimplyLearn team. That is
www.simplearn.com.
Get certified, get ahead. Today in our
second tutorial, we're going to cover K
means linear regression along with going
over the quiz questions we had during
our first tutorial. What's in it for
you? We're going to cover clustering.
What is clustering? K means clustering
which is one of the most common used
clustering tools out there including a
flowchart to understand K means
clustering and how it functions and then
we'll do an actual Python live demo on
clustering of cars based on brands. Then
we're going to cover logistic
regression. What is logistic regression?
Logistic regression curve and sigmoid
function. And then we'll do another
Python code demo to classify a tumor as
malignant or benign based on features.
And let's start with clustering. Suppose
we have a pile of books of different
genres. Now we divide them into
different groups like fiction, horror,
education, and as we can see from this
young lady, she definitely is into heavy
horror. You can just tell by those eyes
and the maple Canadian leaf on her
shirt. But we have fiction, horror, and
education. And we want to go ahead and
divide our books up. Well, organizing
objects into groups based on similarity
is clustering. And in this case, as
we're looking at the books, we're
talking about clustering things with
known categories. But you can also use
it to explore data. So you might not
know the categories. You just know that
you need to divide it up in some way to
conquer the data and to organize it
better. But in this case, we're going to
be looking at clustering in specific
categories. And let's just take a deeper
look at that. We're going to use K means
clustering. K means clustering is
probably the most commonly used
clustering tool in the machine learning
library. K means clustering is an
example of unsupervised learning. If you
remember from our previous thing, it is
used when you have unlabeled data. So we
don't know the answer yet. We have a
bunch of data that we want to cluster to
different groups. Define clusters in the
data based on feature similarity. So
we've introduced a couple terms here.
We've already talked about unsupervised
learning and unlabeled data. So we don't
know the answer yet. We're just going to
group stuff together and see if we can
find an unanswer
connect. We've also introduced feature
similarity. Features being different
features of the data. Now, with books,
we can easily see fiction and horror and
history books. But a lot of times with
data, some of that information isn't so
easy to see right when we first look at
it. And so, K means is one of those
tools where we can start finding things
that connect that match with each other.
Suppose we have these data points and
want to assign them into a cluster. Now
when I look at these data points, I
would probably group them into two
clusters just by looking at them. I'd
say two of these group of data kind of
come together. But in K means we pick K
clusters and assign random centrids to
clusters where the K clusters represents
two different clusters. We pick K
clusters and say random centroidids to
the clusters. Then we compute distance
from objects to the centrids. Now we
form new clusters based on minimum
distances and calculate the centrids. So
we figure out what the best distance is
for the centrid. Then we move the
centrid and recalculate those distances.
Repeat previous two steps iteratively
till the cluster centroid stop changing
their positions and become static.
Repeat previous two steps iteratively
till the cluster centroid stop changing
and the positions become static. Once
the clusters become static, then K means
clustering algorithm is said to be
converged. And there's another term we
see throughout machine learning is
converged. That means whatever math
we're using to figure out the answer has
come to a solution or it's converged on
an answer. Shall we see the flowchart to
understand make a little bit more sense
by putting it into a nice easy step by
step? So we start, we choose K. We'll
look at the elbow method in just a
moment. We assign random centrids to
clusters and sometimes you pick the
centrids because you might look at the
data in a in a graph and say ah these
are probably the central points. Then we
compute the distance from the objects to
the centrids. We take that and we form
new clusters based on minimum distance
and calculate their centrids. Then we
compute the distance from objects to the
new centrids. And then we go back and
repeat those last two steps. We
calculate the distances. So as we're
doing it, it brings into the new centrid
and then we move the centrid around and
we figure out what the best which
objects are closest to each centrid. So
the objects can switch from one centroid
to the other as the centroidids are
moved around and we continue that until
it is converged. Let's see an example of
this. Suppose we have this data set of
seven individuals and their score on two
topics A and B. Uh so here's our subject
in this case referring to the person
taking the uh test and then we have
subject A where we see what they've
scored on their first subject and we
have subject B and we can see what they
score on the second subject. Now let's
take two farthest apart points as
initial cluster centroidids. Now
remember we talked about selecting them
randomly or we can also just put them in
different points and pick the furthest
one apart so they move together. Either
one works okay depending on what kind of
data you're working on and what you know
about it. So we took the two furthest
points one and one and five and seven.
And now let's take the two farthest
apart points as initial cluster
centrids. Each point is then assigned to
the closest cluster with respect to the
distance from the centrids. So we take
each one of these points in there. We
measure that distance. And you can see
that if we measured each of those
distances and you use the the
Pythagorean theorem for a triangle in
this case because you know the x and the
y and you can figure out the diagonal
line from that or you can just take a
ruler and put it on your monitor. That'd
be kind of silly but it would work if
you're just eyeballing it. You can see
how they naturally come together in
certain areas. Now we again calculate
the centroidids of each cluster. So
cluster one and then cluster two and we
look at each individual dot. There's
one, two, three. We're in one cluster.
Uh the centrid then moves over. It
becomes 1.8 comma 2.3. So remember it
was at 1 and one. Well, the very center
of the data we're looking at would put
it at the one point roughly 22, but 1.8
and 2.3. And the second one, if we
wanted to make the overall mean vector,
the average vector of all the different
distances to that centrid, we come up
with 4, 1, and 54. So we've now moved
the centrids. We compare each
individual's distance to its own cluster
mean and to that of the opposite cluster
and we find build a nice chart on here
that the as we move that centrid around
we now have a new different kind of
clustering of groups and using uklidian
distance between the points and the mean
we get the same formula you see new
formulas coming up. So we have our
individual dots distance to the mean
centrid of the cluster and distance to
the mean centrid of the cluster. Only
individual three is nearer to the mean
of the opposite cluster cluster two than
its own cluster one. And you can see
here in the diagram where we've kind of
circled that one in the middle. So when
we've moved the clust the centroidids of
the clusters over one of the points
shifted to the other cluster because
it's closer to that group of
individuals. Thus, individual 3 is
relocated to cluster two, resulting in a
new partition. And we regenerate all
those numbers of how close they are to
the different clusters. For the new
clusters, we will find the actual
cluster centroidids. So now we move the
centrids over. And you can see that
we've now formed two very distinct
clusters on here. On comparing the
distance of each individual's distance
to its own cluster mean and to that of
the opposite cluster, we find that the
data points are stable. Hence, we have
our final clusters. Now if you remember
I brought up a concept earlier K mean on
the K means algorithm choosing the right
value of K will help in less number of
iterations and to find the appropriate
number of clusters in a data set we use
the elbow method and within sum of
squares WSS is defined as the sum of the
squared distance between each member of
the cluster and its centrid and so you
see we've done here is we have the
number of clusters and as you do the
same K means algorithm over the
different clusters and you calculate
what that centrid looks like and you
find the optimal you can actually find
the optimal number of clusters using the
elbow the graph is called as the elbow
method and on this we guessed at two
just by looking at the data but as you
can see the slope you actually just look
for right there where the elbow is in
the slope and you have a clear answer
that we want two different to start with
k means equals two a lot of times people
end up computing k means equals 2 3 four
five until they find the value which
fits on the elbow joint. Sometimes you
can just look at the data and if you're
really good with that specific domain
remember domain I mentioned that last
time you'll know that that where to pick
those numbers and where to start
guessing at what that k value is. So
let's take this and we're going to use a
use case using k means clustering to
cluster cars into brands using
parameters such as horsepower, cubic
inches, make, year, etc. So, we're going
to use the data set cars data having
information about three brands of cars,
Toyota, Honda, and Nissan. We'll go back
to my favorite tool, the Anaconda
Navigator with the Jupiter notebook. And
let's go ahead and flip over to our
Jupyter notebook. And in our Jupyter
Notebook, I'm going to go ahead and just
paste the uh basic code that we usually
start a lot of these off with. We're not
going to go too much into this code
because we've already discussed numpy.
We've already discussed mapplot library
and pandas. Numpy being the number
array, pandas being the pandas data
frame and mattplot for the graphing. And
don't forget uh since if you're using
the Jupyter notebook, you do need the
mattplot library in line so that it
plots everything on the screen. If
you're using a different Python editor,
then you probably don't need that
because it'll have a popup window on
your computer. And we'll go ahead and
run this just to load our libraries and
our setup into here. The next step is of
course to look at our data which I've
already opened up in a spreadsheet. And
you can see here we have the miles per
gallon, cylinders, cubic inches,
horsepower, weight pounds, how you know
how heavy it is, time it takes to get to
60. My card is probably on this one at
about 80 or 90. What year it is? So this
is you can actually see this is kind of
older cars and then the brand Toyota,
Honda, Nissan. So the different cars are
coming from all the way from 1971 if we
scroll down to uh the 80s. We have
between the 70s and 80s a number of cars
that they've put out. And let's uh we
come back here. We're going to do
importing the data. So we'll go ahead
and do data set equals and we'll use
pandas to read this in. And it's uh from
a CSV file. Remember, you can always
post this in the comments and request
the data files for these either in the
comments here on the YouTube video or go
to simplylearn.com and request that. The
car CSV, I put it in the same folder as
the code that I've stored. So, my Python
code is stored in the same folder, so I
don't have to put the full path. If you
store them in different folders, you do
have to change this and double check
your name variables. And we'll go ahead
and run this. And uh we've chosen data
set arbitrarily because, you know, it's
a data set we're importing. And we've
now imported our car CSV into the data
set. As you know, you have to prep the
data. So, we're going to create the X
data. This is the one that we're going
to try to figure out what's going on
with. And then there is a number of ways
to do this, but we'll do it in a simple
loop so you can actually see what's
going on. So, we'll do for i and x.c
columns. So, we're going to go through
each of the columns. And a lot of times
it's important I I'll make lists of the
columns and do this because I might
remove certain columns or there might be
columns that I want to be processed
differently. But for this we can go
ahead and take x of i and we want to go
fill na and that's a pandas command. But
the question is what are we going to
fill the missing data with? We
definitely don't want to just put in a
number that doesn't actually mean
something. And so one of the tricks you
can do with this is we can take x of i.
And in addition to that, we want to go
ahead and turn this into an integer
because a lot of these are integers. So
we'll go ahead and keep it integers. And
me add the bracket here. And a lot of
editors will do this. They'll think that
you're closing one bracket. Make sure
you get that second bracket in there if
it's a double bracket. That's always
something that happens regularly. So
once we have our integer of x of yi,
this is going to fill in any missing
data with the average. And I was so busy
closing one set of brackets, I forgot
that the mean is also has brackets in
there for the pandas. So we can see
here, we're going to fill in all the
data with the average value for that
column. So if there's missing data is in
the average of the data it does have.
Then once we've done that, we'll go
ahead and loop through it again
and just check and see to make sure
everything is filled in correctly. And
we'll print and then we take x is null.
And this returns a set of the null value
or the how many lines are null. And
we'll just sum that up to see what that
looks like. And so when I run this and
so with the X, what we want to do is we
want to remove the last column because
that had the models. That's what we're
trying to see if we can cluster these
things and figure out the models. There
is so many different ways to sort the X
out. For one, we could take the X and we
could go data set, our variable we're
using, and use the eyelocation, one of
the features that's in pandas, and we
could take that and then take all the
rows and all but the last column of the
data set. And at this time, we could do
values. We just convert it to values.
So, that's one way to do this. And if I
let me just put this down here and print
X, it's a capital X we chose. and I run
this, you can see it's just the values.
We could also take out the values and
it's not going to return anything
because there's no values connected to
it. What I like to do with this is
instead of doing the location which does
integers more common is to come in here
and we have our data set and we're going
to do data set dot or data set columns.
And remember that lists all the columns.
So if I come in here, let me just mark
that as red and I print data set.c
columns.
You can see that I have my index here. I
have my MPG cylinders everything
including the brand which we don't want.
So the way to get rid of the brand would
be to do data columns of everything but
the last one minus one. So now if I
print this, you'll see the brand
disappears. And so I can actually just
take data set columns minus one and I'll
put it right in here for the columns
we're going to look at.
And let's unmark this.
And unmark this.
And now if I do an x.ad
I now have a new data frame. And you can
see right here we have all the different
columns except for the brand at the end
of the year. And it turns out when you
start playing with the data set, you're
going to get an error later on and it'll
say cannot convert string to float
value. And that's because it for some
reason these things the way they
recorded them must have been recorded as
strings. So we have a neat feature in
here on pandas to convert. And it is
simply convert objects.
And for this we're going to do convert
oops convert underscore
numeric numeric equals true. And yes, I
did have to go look that up. I don't
have it memorized the convert numeric in
there. If I'm working with a lot of
these things, I remember them, but um
depending on where I'm at, what I'm
doing, I usually have to look it up. And
we run that. Oops, I must have missed
something in here. Let me double check
my spelling. And when I double check my
spilling, you'll see I missed the first
underscore in the convert objects. And
when I run this, it now has everything
converted into a numeric value because
that's what we're going to be working
with is numeric values down here.
And the next part is that we need to go
through the data and eliminate null
values. Most people when they're doing
small amounts, you working with small
data pools discover afterwards that they
have a null value and they have to go
back and do this. So, you know, be aware
whenever we're formatting this data,
things are going to pop up and sometimes
you go backwards to fix it. And that's
fine. That's just part of exploring the
data and understanding what you have.
And I should have done this earlier, but
let me go ahead and increase the size of
my window one notch.
There we go. Easier to see.
So, we'll do 4 I in working with X dot
columns. will page through all the
columns. And we want to take X of I and
we're going to change that. We're going
to alter it. And so with this, we want
to go ahead and fill in X of I. Pandas
has the fill in a. And that just fills
in any non-existent missing data. And
we'll put my brackets up. And there's a
lot of different ways to fill this data.
If you have a really large data set,
some people just void out that data
because if and then look at it later in
a separate exploration of data. One of
the tricks we can do is we can take our
column and we can find the means
and the means is in there or quotation
marks. So we take the columns, we're
going to fill in the non-existing one
with the means. The problem is that
returns a decimal float. So some of
these aren't decimals. Certainly, you
may need to be a little careful of doing
this, but for this example, we're just
going to fill it in with the integer
version of this. Keeps it on par with
the other data that isn't a decimal
point.
And then what we also want to do is we
want to double check. A lot of times you
do this first part first to double
check, then you do the fill, and then
you do it again just to make sure you
did it right. So, we're going to go
through and test for missing data. And
one of the re ways you can do that is
simply go in here and take our X of I
column. So it's going to go through the
X of I column. It says is null. So it's
going to return any any place there's a
null value. It actually goes through all
the rows of each column is null. And
then we want to go ahead and sum that.
So we take that, we add the sum value.
And these are all pandas. So is null is
a panda command and so is sum. And if we
go through that and we go ahead and run
it
and we go ahead and take and run that,
you'll see that all the columns have
zero null values. So we've now tested
and double checked and our data is nice
and clean. We have no null values.
Everything is now a number value. We
turned it into numeric and we've removed
the last column in our data. And at this
point, we're actually going to start
using the elbow method to find the
optimal number of clusters. So, we're
now actually getting into the sklearn
part. Uh, the K means clustering on
here. I guess we'll go ahead and zoom it
up one more notch so you can see what
I'm typing in here.
And then from sklearn going to or
sklearn
cluster, we're going to import K means.
I always forget to capitalize the K and
the M when I do this. So it's capital K,
capital M K means.
And we'll go and create a um array WCSS
equals we'll make it an empty array. If
you remember from the elbow method from
our slide
within the sums of squares, WSS is
defined as the sum of squared distance
between each member of the cluster and
it centrid. So we're looking at that
change in differences as far as a
squared distance. And we're going to run
this over a number of K mean values.
In fact, let's go for I in range. We'll
do 11 of them.
Range zero of 11.
And the first thing we're going to do is
we're going to create the actual we'll
do it all lowercase.
And so we're going to create this object
from the K means that we just imported.
And the variable that we want to put
into this is in clusters. We're going to
set that equals to I. That's the most
important one because we're looking at
how increasing the number of clusters
changes our answer. There are a lot of
settings to the K means. Our guys in the
back did a great job just kind of
playing with some of them. The most
common ones that you see in a lot of
stuff is how you enit your K means. So
we have K means plus plus. This is just
a tool to let the model itself be smart
how it picks it centrids to start with
its initial centroidids. We only want to
iterate no more than 300 times. We have
a max iteration we put in there. We have
the infinite the random state equals
zero. You really don't need to worry too
much about these when you're first
learning this. As you start digging in
deeper, you start finding that these are
shortcuts that will speed up the process
as far as a setup. But the big one that
we're working with is the inclusters
equals I. So, we're going to literally
train our K means 11 times. We're going
to do this process 11 times. And if
you're working with big data, you know,
the first thing you do is you run a
small sample of the data so you can test
all your stuff on it. And you can
already see the problem that if I'm
going to iterate through a terabyte of
data 11 times and then the K means
itself is iterating through the data
multiple times. That's a heck of a
process. So you got to be a little
careful with this. A lot of times though
you can find your elbow using the elbow
method. Find your optimal number on a
sample of data especially if you're
working with larger data sources. So we
want to go ahead and take our K means
and we're just going to fit it. If
you're looking at any of the sklearn,
very common that you fit your model. And
if you remember correctly, our variable
we're using is the capital X. And once
we fit this value, we go back to the um
array we made. And we want to go and
just append that value on the end.
And it's not the actual fit we're
pinning in there. It's when it generates
it, it generates the value you're
looking for is inertia. So k
means.inertia will pull that specific
value out that we need.
And let's get a visual on this. We'll do
our PLT plot. And what we're plotting
here
is first the x axis, which is range 0
11. So that will generate a nice little
plot there. And the wcss for our y axis.
It's always nice to give our uh plot a
title.
And let's see, we'll just give it the
elbow method for the title. And let's
get some labels. So let's go ahead and
do PLT X label.
And what we'll do, we'll do number of
clusters for that. And PLT Y label. And
for that, we can do oops, there we go.
WCSS since that's what we're doing on
the plot on there. And finally, we want
to go ahead and display our graph, which
is simply plt. Oops.
Show. There we go. And because we have
it set to inline, it'll appear inline.
Hopefully I didn't make a type error on
there.
And you can see we get a very nice
graph. You can see a very nice elbow
joint there at uh two and again right
around three and four. And then after
that there's not very much. Now as a
data scientist, if I was looking at
this, I would do either three or four.
And I'd actually try both of them to see
what the u output look like. And they've
already tried this in the back. So,
we're just going to use three as a setup
on here. And let's go ahead and see what
that looks like when we actually use
this to show the different kinds of
cars.
And so, let's go ahead and apply the K
means to the cars data set. And
basically, we're going to copy the code
that we loop through up above where K
means equals K means number of clusters.
And we're just going to set the number
of clusters to three since that's what
we're going to look for. And you could
do three and four on this and graph them
just to see how they come up
differently. It'd be kind of curious to
look at that. But for this, we're just
going to set it to three. Go ahead and
create our own variable Y k means for
our answers. And we're going to set that
equal to Whoops, my double equal there
to K means. But we're not going to do a
fit. We're going to do a fit predict is
the setup you want to use. And when
you're using untrained models, you'll
see um a slightly different because
usually you see fit and then you see
just the predict. But we want to both
fit and predict the k means on this. And
that's fit underscore predict. And then
our capital x is the data we're working
with.
And before we plot this data, we're
going to do a little pandas trick. We're
going to take our x value and we're
going to set x as matrix. So we're
converting this into a nice rows and
columns kind of setup. But we want the
we're going to have columns equals none.
So it's just going to be a matrix of
data in here. And let's go ahead and run
that.
A little warning. You'll see this
warnings pop up because things are
always being updated. So there's like
minor changes in the versions and future
versions. Let's set a matrix. Now that
it's more common to set it values
instead of doing as matrix, but mass
matrix works just fine for right now and
you'll want to update that later on. But
let's go ahead and dive in and plot this
and see what that looks like. And before
we dive into plotting this data, I
always like to take a look and see what
I am plotting. So let's take a look at
why K means. I'm just going to print
that out down here. And we see we have
an array of answers. We have 2 1 0 2 1
2. So it's clustering these different
rows of data based on the three
different spaces it thinks it's going to
be.
And then let's go ahead and print X and
see what we have for X. And we'll see
that X is an array. It's a matrix. So we
have our different values in the array.
And what we're going to do, it's very
hard to plot all the different values in
the array. So we're only going to be
looking at the first two or positions
zero and one. And if you were doing a
full presentation in front of the board
meeting, you might actually do a little
different and and dig a little deeper
into the different aspects because this
is all the different columns we looked
at. But we'll only look at columns one
and two for this to make it easy. So
let's go ahead and clear this data out
of here and let's bring up our plot. And
we're going to do a scatter plot here.
So pl scatter.
And
this looks a little complicated. So
let's explain what's going on with this.
We're going to take the x values
and we're only interested in y of k
means equals 0, the first cluster. Okay?
And then we're going to take value zero
for the x-axis. And then we're going to
do the same thing here. We're only
interested in k means equals 0, but
we're going to take the second column.
So we're only looking at the first two
columns in our answer or in the data.
And then the guys in the back played
with this a little bit to make it
pretty.
And they discovered that it looks good
with a size equals 100. That's the size
of the dots. We're going to use red for
this one. And when they were looking at
the data and what came out, it was
definitely the Toyota on this. We're
just going to go ahead and label it
Toyota. Again, that's something you
really have to explore in here as far as
playing with those numbers and see what
looks good. We'll go ahead and hit enter
in there. And I'm just going to paste in
the next two lines, which is the next
two cars. And this is our Nissa and
Honda. And you'll see with our scatter
plot, we're now looking at where Y_K
means equals 1. And we want the zero
column and YK means equals 2. Again,
we're looking at just the first two
columns, zero and one. And each of these
rows then corresponds to Nissan and
Honda.
And I'll go ahead and hit enter on
there. And uh finally, let's take a look
and put the centrids on there. Again,
we're going to do a scatter plot.
And on the centrids, you can just pull
that from our K means, the uh model we
created cluster centers. And we're going
to just do um
all of them in the first number and all
of them in the second number, which is
01 because you always start with zero
and one.
And then they were playing with the size
and everything to make it look good.
We'll do a size of 300. We're going to
make the color yellow. And we'll label
them. It's always good to have some good
labels. Centroidids.
And then we do want to do a title. PLT
title.
And pop up there. PLT title. So you
always make want to make your graphs
look pretty. And we'll call it clusters
of car make. And one of the features of
the plot library is you can add a
legend. It'll automatically bring in it
since we've already labeled the
different aspects of the legend with
Toyota, Nissan, and Honda.
And finally, we want to go ahead and
show so we can actually see it. And
remember, it's in line. Uh so if you're
using a different editor that's not the
Jupyter notebook, you'll get a popup of
this. And you should have a nice set of
clusters here. So we can look at this
and we have a clusters of Honda in
green, Toyota in red, Nissan in purple.
And you can see where they put the
centroidids to separate them.
Now when we're looking at this, we can
also plot a lot of other different data
on here as far because we only looked at
the first two columns. This is just
column one and two or 01 as as you label
them in computer scripting. But you can
see here we have a nice clusters of car
making. and we were able to pull out the
data and you can see how just these two
columns form very distinct clusters of
data. So if you were exploring new data
you might take a look and say well what
makes these different almost going in
reverse you start looking at the data
and pulling apart the columns to find
out why is the first group set up the
way it is. Maybe you're doing loans and
you want to go, well, why is this group
not defaulting on their loans and why is
the last group defaulting on their
loans? And why is the middle group 50%
defaulting on their bank loans? And you
start finding ways to manipulate the
data and pull out the answers you want.
So now that you've seen how to use K
mean for clustering, let's move on to
the next topic. Now let's look into
logistic regression. The logistic
regression algorithm is the simplest
classification algorithm used for binary
or multiclassification problems. And we
can see we have our little girl from
Canada who's into horror books is back.
That's actually really scary when you
think about that with those big eyes. In
the previous tutorial, we learned about
linear regression, dependent and
independent variables. So to brush up,
y= mx + c. Very basic algebraic function
of uh y and x. The dependent variable is
the target class variable we are going
to predict. The independent variables X1
all the way up to XN are the features or
attributes we're going to use to predict
the target class. We know what a linear
regression looks like. But using the
graph, we cannot divide the outcome into
categories. It's really hard to
categorize 1.5, 3.6, 9.8. Uh for
example, a linear regression graph can
tell us that with increase in number of
hours studied, the marks of a student
will increase, but it will not tell us
whether the student will pass or not. In
such cases where we need the output as
categorical value, we will use logistic
regression. And for that, we're going to
use the sigmoid function. So you can see
here we have our marks 0 to 100, number
of hours studied. That's going to be
what they're comparing it to in this
example. And we usually form a line that
says y = mx + c. And when we use the
sigmoid function, we have p = 1 / 1 + e
the minus y, it generates a sigmoid
curve. And so you can see right here
when you take the ln, which is the
natural logarithm. I always thought it
should be nl, not ln. That's just the
inverse of uh e your e to the minus y.
And so we do this, we get ln of p 1 - p
= m * x + c. That's the sigmoid curve
function we're looking for. And we can
zoom in on the function and you'll see
that the function as it deres goes to
one or to zero depending on what your x
value is. And the probability if it's
greater than 0.5, the value is
automatically rounded off to one
indicating that the student will pass.
So if they're doing a certain amount of
studying, they will probably pass. Then
you have a threshold value at the 0.5.
It automatically puts that right in the
middle usually. And your probability if
it's less than 0.5, the value run it off
to zero indicating the student will
fail. So if they're not studying very
hard, they're probably going to fail.
This, of course, is ignoring the
outliers of that one student who's just
a natural genius and doesn't need any
studying to memorize everything. That's
not me, unfortunately. Have to study
hard to learn new stuff. problem
statement to classify whether a tumor is
malignant or B9. And this is actually
one of my favorite data sets to play
with because it has so many features and
when you look at them, you really are
hard to understand. You can't just look
at them and know the answer. So it gives
you a chance to kind of dive into what
data looks like when you aren't able to
understand the specific domain of the
data. But I also want you to remind you
that in the domain of medicine, if I
told you that my probability was really
good at classified things that say 90%
or 95% and I'm classifying whether
you're going to have a malignant or a B9
tumor, I'm guessing that you're going to
go get it tested anyways. So you got to
remember the domain we're working with.
So why would you want to do that if you
know you're just going to go get a
biopsy? Because you know it's that
serious. This is like an all or nothing.
just referencing the domain. It's
important. It might help the doctor know
where to look just by understanding what
kind of tumor it is. So it might help
them or aid them on something they
missed from before. So let's go ahead
and dive into the code and I'll come
back to the domain part of it in just a
minute. So use case and we're going to
do our normal imports here where we're
importing numpy, pandas, seabour, the
mattplot library and we're going to do
mattplot library in line since I'm going
to switch over to Anaconda. So, let's go
ahead and flip over there and get this
started. So, I've opened up a new window
in my Anaconda Jupyter Notebook. And by
the way, Jupyter Notebook, uh, you don't
have to use Anaconda for the Jupyter
Notebook. I just love the interface and
all the tools that Anaconda brings. So,
we got our import numpy aspy
number array. We have our pandas pd.
We're going to bring in Seabor to help
us with our graphs as SNS. So many
really nice tools in both Seabour and
Mattplot library. And we'll do our
mapplot library.pipplot as plt. And then
of course we want to let it know to do
it in line. And let's go and just run
that. So it's all set up. And we're just
going to call our data data. Not
creative today. Uh equals pd. And this
happens to be in a CSV file. So we'll
use a pdread_csv.
And I happen to name the file. renamed
it data forp2.csv.
You can of course um write in the
comments below the YouTube and request
for the data set itself or go to the
SimplyLearn website and we'll be happy
to supply that for you. And let's just
um open up the data before we go any
further and let's just see what it looks
like in a spreadsheet.
So when I pop it open in a local
spreadsheet, this is just a CSV file,
comma separated variables. We have an
ID. So I guess the U categorizes for
reference or what ID which test was
done. The diagnosis M for malignant, B
for B9. So there's two different options
on there. And that's what we're going to
try to predict is the M and B and test
it. And then we have like the radius
mean or average the texture average,
perimeter mean, area mean, smoothness. I
don't know about you, but unless you're
a doctor in the field, most of the
stuff, I mean, you can guess what
concave means just by the term concave,
but I really wouldn't know what that
means in the measurements they're
taking. So, they have all kinds of stuff
like how smooth it is, uh, the symmetry,
and these are all float values. You just
page through them real quick, and you'll
see there's, I believe, 36, if I
remember correctly, in this one.
So there's a lot of different values
they take and all these measurements
they take when they go in there and they
take a look at the different growth, the
tumorous growth. So back in our data and
I put this in the same folder as a code.
So I saved this code in that folder.
Obviously if you have it in a different
location, you want to put the full path
in there and we'll just do uh pandas
first five lines of data with the data
head. And we run that. We can see that
we have pretty much what we just looked
at. We have an ID. We have a diagnosis.
If we go all the way across, you'll see
all the different columns coming across
displayed nicely for our data.
And while we're exploring the data, our
uh Seabor, which we referenced as SNS,
makes it very easy to go in here and do
a joint plot. You'll notice the very
similar to because it is sitting on top
of the U plot library. So, the joint
plot does a lot of work for us. And
we're just going to look at the first
two columns that we're interested in,
the radius mean and the texture mean.
We'll just look at those two columns and
data equals data. So that tells it which
two columns we're plotting and that
we're going to use the data that we
pulled in. Let's just run that. And it
generates a really nice graph on here.
And there's all kinds of cool things on
this graph to look at. I mean, we have
the texture mean and the radius mean
obviously the axes. You can also see
and uh one of the cool things on here is
you can also see the histogram. They
show that for the radius mean where is
the most common radius mean come up and
where the most common texture is. So
we're looking at the tech the on each
growth it's average texture and on each
radius it's average uh radius on there
gets a little confusing because we're
talking about the individual objects
average. And then we can also look over
here and see the the histogram showing
us the median or how common each
measurement is. And that's only two
columns. So let's dig a little deeper
into Seabor. They also have a heat map.
And if you're not familiar with heat
maps, a heat map just means it's in
color. That's all that means. Heat map.
I guess the original ones were plotting
heat density on something. And so ever
since then it's just called a heat map.
And we're going to take our data and get
our corresponding numbers to put that
into the heat map. And that's simply
data.coR
for that. That's a pandas expression.
Let's remember we're working in a pandas
data frame. So that's one of the cool
tools in pandas for our data. And let's
just pull that information into a heat
map and see what that looks like. And
you'll see that we're now looking at all
the different features. We have our ID.
We have our texture. We have our area,
our compactness, concave points. And if
you look down the middle of this chart
diagonal going from the upper left to
bottom right, it's all white. That's
because when you compare texture to
texture, they're identical. So they're
100% or in this case perfect one in
their correspondence.
And you'll see that when you look at say
area or right below it, it has almost a
black on there. when you compare it to
texture. So these have almost no
corresponding data. They don't really
form a linear graph or something that
you can look at and say how connected
they are. They're very scattered data.
This is really just a really nice graph
to get a quick look at your data.
Doesn't so much change what you do, but
it changes verifying. So when you get an
answer or something like that or you
start looking at some of these
individual pieces, you might go, "Hey,
that doesn't match. according to showing
our heat map, this should not correlate
with each other. And if it is, you're
going to have to start asking, well,
why? What's going on? What else is
coming in there? But it does show some
really cool information on here. I mean,
we can see from the ID, there's no real
one feature that just says if you go
across the top line that lights up.
There's no one feature that says, hey,
if the area is a certain size, then it's
going to be B9 or malignant. It says
there's some that sort of add up and
that's a big hint in the data that we're
trying to ID this whether it's malignant
or B9. That's a big hint to us as data
scientists to go okay we can't solve
this with any one feature. It's going to
be something that includes all the
features or many of the different
features to come up with a solution for
it. And while we're exploring the data
let's explore one more area and let's
look at data isnull. We want to check
for null values in our data. If you
remember from earlier in this tutorial,
we did it a little differently where we
added stuff up and sum them up. You can
actually with pandas do it really
quickly. Data.isnull and summit. And
it's going to go across all the columns.
So when I run this,
you're going to see all the columns come
up with no null data.
So we've just just to rehash these last
few steps. We've done a lot of
exploration. We have looked at the first
two columns and seen how they plot with
the seabour with a joint plot which
shows both the histogram and the data
plotted on the XY coordinates. And
obviously you can do that more in detail
with different columns and see how they
plot together. And then we took and did
the Seabor heat map the SNS
heat mapap of the data. And you can see
right here where it did a nice job
showing us some bright spots where stuff
correlates with each other and forms a
very nice combination or points of
scattering points. And you can also see
areas that don't.
And then finally, we went ahead and
checked the data. Is the data null
value? Do we have any missing data in
there? Very important step because it'll
crash later on. If you forget to do this
step, it will remind you when you get
that nice error code that says null
values. Okay. So, not a big deal if you
miss it, but it it's no fun having to go
back when you're when you're in a huge
process and you've missed this step and
now you're 10 steps later and you got to
go remember where you were pulling the
data in.
So, we need to go ahead and pull out our
X and our Y. So, we just put that down
here and we'll set the X equal to. And
there's a lot of different options here.
Certainly we could do X equals all the
columns except for the first two because
if you remember the first two is the ID
and the diagnosis. So that certainly
would be an option. But what we're going
to do is we're actually going to focus
on the worst. The worst radius, the
worst texture, parameter area,
smoothness, compactness, and so on. One
of the reasons to start dividing your
data up when you're looking at this
information is sometimes the data will
be the same data coming in. So if I have
two measurements coming into my model,
it might overweigh them. It might
overpower the other measurements because
it's measuring it's basically taking
that information in twice. That's a
little bit past the scope of this
tutorial. I want you to take away from
this though is that we are dividing the
data up into pieces and our team in the
back went ahead and said hey let's just
look at the worst. So I'm going to
create a an array and you'll see this
array radius worst texture worst
perimeter worst. We've just taken the
worst of the worst and I'm just going to
put that in my X. So this X is still a
pandas data frame but it's just those
columns. And our Y, if you remember
correctly, is going to be Oops, hold on
one second. It's not X. is data. There
we go. So, x equals data and then it's a
list of the different columns, the worst
of the worst. And if we're going to take
that, then we have to have our answer
for our y for the stuff we know. And if
you remember correctly, we're just going
to be looking at
the diagnosis. That's all we care about
is what is it diagnosed? Is it B9 or
malignant? And since it's a single
column, we can just do diagnosis. Oh, I
forgot to put the brackets. There we go.
Okay. So, it's just diagnosis on there.
And we can also real quickly do like an
X do. If you want to see what that looks
like and Y head
and run this and you'll see um it only
does the last one. I forgot about that.
If you don't do print, you can see that
the the Y.D is just mm because the first
ones are all malignant. And if I run
this, the X do head is just the first
five values of radius worst, texture
worst, parameter worst, area worst, and
so on. I'll go ahead and take that out.
So, moving down to the next step, we've
built our two data sets, our answer and
then the features we want to look at.
In data science, it's very important to
test your model. So we do that by
splitting the data
and from sklearn model selection we're
going to import train test split. So
we're going to split it into two groups.
There are so many ways to do this. I
noticed in one of the more modern ways
they actually split it into three groups
and then you model each group and test
it against the other groups. So you have
all kinds and there's reasons for that
which is past the scope of this and for
this particular example isn't necessary
for this. We're just going to split it
into two groups. one to train our data
and one to test our data. And the
sklearn uh.mmodel selection we have
train tests split. You could write your
own quick code to do this where you just
randomly divide the data up into two
groups but they do it for us nicely
and we actually can almost we can
actually do it in one statement with
this where we're going to generate four
variables capital X train capital X
test. So we have our training data we're
going to use to fit the model and then
we need something to test it and then we
have our y train. So we're going to
train the answer and then we have our
test. So this is the stuff we want to
see how good it did on our model. And
we'll go ahead and take our train test
split that we just imported.
And we're going to do X and our Y, our
two different data that's going in for
our split. And then the guys in the back
came up and wanted us to go ahead and
use a test size equals.3.
That's test size. Random state. It's
always nice to kind of switch a random
state around, but not that important.
What this means is that the test size is
we're going to take 30% of the data and
we're going to put that into our test
variables, our Y test and our X test.
And we're going to do 70% into the X
train and the Y train. So, we're going
to use 70% of the data to train our
model and 30% to test it. Let's go ahead
and run that and load those up. So now
we have all our stuff split up and all
our data ready to go. And now we get to
the actual logistics part. We're
actually going to do our create our
model. So let's go ahead and bring that
in from sklearn. We're going to bring in
our linear model and we're going to
import logistic regression. That's the
actual model we're using. And let's
we'll call it log model.
Oops, there we go. Model. And let's just
set this equal to our logistic
regression that we just imported. So now
we have a variable log model set to that
class for us to use. And with most the
uh models in the sklearn, we just need
to go ahead and fix it. Fit do a fit on
there. And we use our x train that we
separated out with our y train. And
let's go ahead and run this. So once
we've run this, we'll have a model that
fits this data that 70% of our training
data.
Uh, and of course it prints this out
that tells us all the different
variables that you can set on there.
There's a lot of different choices you
can make, but for Word do, we're just
going to let all the defaults set. We
don't really need to mess with those on
this particular example. And there's
nothing in here that really stands out
as super important until you start
fine-tuning it. But for what we're
doing, the basics will work just fine.
And then let's we need to go ahead and
test out our model. Is it working? So
let's create a variable Y predict. And
this is going to be equal to our log
model. And we want to do a predict.
Again, very standard format for the
sklearn library is taking your model and
doing a predict on it. And we're going
to test y predict against the y test. So
we want to know what the model thinks
it's going to be. That's what our y
predict is. And with that, we want the
capital xx test. So we have our train
set and our test set. And now we're
going to do our y predict. And let's go
ahead and run that.
And if we uh print
y predict, let me go ahead and run that.
You'll see it comes up and it predents a
prints a nice array of uh B and M for B9
and malignant
for all the different test data we put
in there. So, it does pretty good. We're
not sure exactly how good it does, but
we can see that it actually works and is
functional. Was very easy to create.
You'll always discover with our data
science that as you explore this, you
spend a significant amount of time
prepping your data and making sure your
data coming in is good. Uh there's a
saying, good data in, good answers out.
Bad data in, bad answers out. That's
only half the thing. That's only half of
it. Selecting your models becomes the
next part as far as how good your models
are. and then of course fine-tuning it
depending on what model you're using. So
we come in here, we want to know how
good this came out. So we have our Y
predict here, log model.predict X test.
So for deciding how good our model is,
we're going to go from the
sklearn.metrics,
we're going to import classification
report. And that just reports how good
our model is doing. And then we're going
to feed it the model data. And let's
just print this out. and we'll take our
uh classification report
and we're going to put into there
our test our actual data. So this is
what we actually know is true and our
prediction what our model predicted for
that data on the test side. And let's
run that and see what that does.
So we pull that up. You'll see that we
have um a precision for B9 and malignant
B and M. And we have a precision of 93
and 91, a total of 92. So it's kind of
the average between these two of 92.
There's all kinds of different
information on here. Your F1 score,
your recall, your support coming through
on this. And for this, I'll go ahead and
just flip back to our slides that they
put together for describing it. And so
here we're going to look at the
precision using the classification
report. And you see this is the same
print out I had up above. Some of the
numbers might be different because it
does randomly pick out which data we're
using. So this model is able to predict
the type of tumor with 91% accuracy. So
we look back here that's you will see
where we have uh B9 and malignant. It
actually has 92 coming up here. We're
looking about a 92 91% precision. And
remember I reminded you about domain.
So, when we're talking about the domain
of a medical domain with a very
catastrophic outcome, you know, at 91 or
92% precision, you're still going to go
in there and have somebody do a biopsy
on it. Very different than if you're
investing money and there's a 92% chance
you're going to earn 10% and 8% chance
you're going to lose 8%, you're probably
going to bet the money because at that
odds, it's pretty good that you'll make
some money. And in the long run, you do
that enough, you definitely will make
money. And also with this domain, I've
actually seen them use this to identify
different forms of cancer. That's one of
the things that they're starting to use
these models for because then it helps a
doctor know what to investigate. So that
wraps up this section. We're finally
we're going to go in there and let's
discuss the answers to the quiz asked in
machine learning tutorial part one. Can
you tell what's happening in the
following cases? Grouping documents into
different categories based on the topic
and content of each document. This is an
example of clustering where K means
clustering can be used to group the
documents by topics using bag of words
approach. So if you gotten in there that
you're looking for clustering and
hopefully you had at least one or two
examples like K means that are used for
clustering different things then give
yourself a two thumbs up. B identifying
handwritten digits in images correctly.
This is an example of classification.
The traditional approach to solving this
would be to extract digit dependent
features like curvature of different
digits etc. and then use a classifier
like SVM to distinguish between images.
Again, if you got the fact that it's a
classification example, give yourself a
thumb up. And if you're able to go, hey,
let's use SVM or another model for this,
give yourself those two thumbs up on it.
C. Behavior of a website indicating that
the site is not working as designed.
This is an example of anomaly detection.
In this case, the algorithm learns what
is normal and what is not normal,
usually by observing the logs of the
website. Give yourself a thumbs up if
you got that one. And just for a bonus,
can you think of another example of
anomaly detection? One of the ones I use
it for in my own business is detecting
anomalies in stock markets. Stock
markets are very fickled and they behave
very erratic. So finding those erratic
areas and then finding ways to track
down why they're erratic. Was something
released in social media? Was something
released you can see where knowing where
that anomaly is can help you to figure
out what the answer is to it in another
area. D predicting salary of an
individual based on his or her years of
experience. This is an example of
regression. This problem can be
mathematically defined as a function
between independent years of experience
and dependent variables salary of an
individual. And if you guess that this
was a regression model, give yourself a
thumbs up. And if you were able to
remember that it was between independent
and dependent variables and that terms,
give yourself two thumbs up. Summary. So
to wrap it up, we went over what is K
means and we went through also the chart
of choosing your elbow method and
assigning a random centrid to the
clusters, computing the distance and
then going in there and figuring out
what the minimum centroidids is and
computing the distance and going through
that loop until it gets the perfect
centrid. And we looked into the elbow
method to choose K based on running our
clusters across a number of variables
and finding the best location for that.
We did a nice example of clustering cars
with K means even though we only looked
at the first two columns to make it
simple and easy to graph. You can easily
extrapolate that and look at all the
different columns and see how they all
fit together. And we looked at what is
logistic regression. We discussed the
sigmoid function. What is logistic
regression? And then we went into an
example of classifying tumors with
logistics. I hope you enjoyed part two
of machine learning. So in today's
session we will discuss what RNN model
is. Moving ahead we will see why should
we use RNN. After that we will see how
does RNN work recurrent neural network.
After covering these topics we will move
forward and see types of RNN recurrent
neural network and applications of RNN.
At the end we will do a hands-off lab
demo of sentiment analysis using RNN. So
before starting let us have a simple
question to brush our knowledge. So
question is what are the application of
RNN? Okay, NLP,
time series, image captioning and all of
the above. Please answer in the comment
section below and we will update the
correct answer in the pin comments or
you can pause this video, give it a
thought and answer in the comment
section. Before we move on to the
programming part, let's discuss what RNN
is and proceed further for the same. So
what is RNN? Recurrent neural network.
So RNN work on the principle of saving
output on a particular layer and feeding
this back to the input in order to
predict the output of the layer. This is
how can convert a feed neural network
into a recurrent neural network RN. The
node in different layers of neural
network are compressed to form a single
layer of recurrent neural network. A B
and C are the parameters of neural
network. Now that you understand what
RNN is, let's look at the way why RNN.
Okay. So why RNN? RNN were created
because there are few issues in the feed
forward neural network cannot handle the
sequential data considers only the
current input cannot memorize previous
input. Okay. So the solution of these
issues is RNN and RNN can handle
sequential data accepting the current
input data and previously received input
data. So RNN can memorize previous input
due to their internal memory. So moving
forward let's see how does RNN networks
work. Okay. So the input layer X takes
an input to the neural network and
process it and the passes it into the
middle layer. The middle layer edge can
consist of multiple hidden layers each
with its own activation function and
weight and biases. If you have a neural
network where the various parameters of
different hidden layers are not affected
by the previous layer that is the neural
network does not have the memory then
you can use RNN. So the RNN will
standardize the different activation
function and weights and biases so that
each hidden layer has the same
parameter. Then instead of creating
multiple hidden layers, it will create
one end loop over it as many time it has
required. So moving forward let's see
types of RNN. So there are four types of
RNN
one to one,
one to many, many to many and many to
one.
So let's see one to one RNN. So this
type of neural network is known as the
vanilla neural network. It is used for
general machine learning problem which
has a single input and a single output.
Now see
one to many RNN. This type of neural
network has a single input and multiple
outputs. An example of this is a image
captioning. Now let's see many to one
RNN. This RNN take a sequence of input
and generates a single output. Sentiment
analysis is a good example of this kind
of neural network where a given sentence
can be classified as expressing positive
or negative sentiment. And the last one
is many to many RNN. This RNN takes a
sequence of inputs and generates a
sequence of output. Machine translation
is the one of the example. So moving
forward, let's see application of
recurrent neural network. First one is
image captioning. RNNs are used to
caption an image by analyzing the
activities present. The second one is
time series prediction. Any time series
problem like predicting the prices of
stocks in a particular month can be
solved using RNN. And the third one is
natural language processing. Text mining
and sentiment analysis can be carried
out using RNN or NLP. Natural language
processing. The fourth one is machine
translation. Given an input in one
language, RNNs can be used to translate
the input into different language as
output. So now let's move to the
programming part. First we will import
some libraries major libraries for the
first we will import for the data frame.
So I will write import
pd.
The second one is import numpy
as np.
So pandas is a software library written
for the python programming language for
data manipulation and analysis. In
particular, it offers a data structure
and operations for manipulating
numerical tables and the time series.
And this numpy numpy is a library for
the Python programming language adding
support to four large multi-dimensional
array and matrices along with a large
collection of highle mathematical
function to operate on these arrays.
Okay. So for plotting we will import
some libraries like seabon
as
SNS. This is nothing just a short form
of we don't have to write again and
again CON c we can write SNS. So then
another one is from
wordcloud
port
mattplot lib
dot
pip plot
s plt. library.
Okay, so Seabone is a library that uses
Matt plot lib underneath to plot graphs.
It will be used to visualize zandom
distribution and the word cloud is a
visual representations
of words. Cloud creators are used to
highlight popular words and phrases
based on frequency and relevance. They
provide you with quick and simple visual
insights that can lead to more in-depth
analysis. And this mattplot lil mattplot
lib is a plotting library for the python
programming language and its numerical
mathematic
ext extension numpy. It provides an
object- oriented API for embedding plots
into application using general purpose
UI.
Okay. Like tinker wxython QT or gtk.
So let's import some
NLTK
natural language toolkit.
So
I will write import
NLTK.
Okay. from
NLTK
dot stem
importizer
then from analytic dot corpus
imports
and
from
NL ticket dot tokenize
port
tokenize
NLTK the natural language toolkit or
more commonly NLTK is a suit of
libraries and programs for symbolic and
statical natural language processing for
English written in Python programming
language and this is stop words. Stop
words are words that are so common they
are basically ignored by typical
tokenizers and this word tokenize is a
function in Python that splits a given
sentence into words using the analytical
library. Okay. So let's import some
scikitlearn
library. So for that I will write from
skarn
dot model
collection
import
train
test.
Okay. Then from skarn
dot feature
extraction
dot text import
vectorizer.
And then from
skarn dot matrices
matrix
import
confusion metric
classification.
Okay.
So, scikitlarn is a free source software
machine learning library for Python
programming language. It features
various classification, regression and
clustering algorithms including support
vector machine learning, logistic
regression and many others like random
forest classifier. And this train test
split method is used to split our data
into train and test set. First, we need
to divide our data into features like X
and Y labels. And this TF ID vectorzer
converts a collection of raw documents
into a matrix of TF features. The fast
text or what to vectorizer what
embedding Python implementation and this
confusion matrix. A confusion matrix is
a table that is used to define the
performance of a classification
algorithm. Okay.
Then we'll import some libraries like
prom skarn
do linear model
port
logistic
regression.
So then from
colonm
port
and from
import
random
forests classifier.
Okay. Then from
skarn dot name base
portoli
base.
Okay.
So everything is correct. You will see
while running. So logistic regression
estimate the probability of an event
occurring such as voted or didn't vote
based on a given data set of the
independent variable.
L SVC logistic regression estimate
sorry linear support vector machine SVC
is an algorithm that attempts to find a
hyper plane to maximize the distance
between classified samples and this
random forest classifier creates a set
of decision trees from a randomly
selected subset of the training set
and this Bernoli NBoli
name base is a part of the name base
family it is based on Bernoli
distribution ution and accept only
binary values that is zero or one.
So let's import some tensorflow. So
import
tensorflow
dot
compad dot v2 and
then import
tensorflow
data sets
as tfds.
So, TensorFlow is a free and open-source
library for machine learning and
artificial intelligence across a range
of task but has a particular focus on
training and inference of deep neural
networks. Okay, let's import warnings.
Nothing. Warning.
The warnings
import
string
import.
So everything is basic. Just let's see
the pickle. Typically is a Python is
primarily used in serializing and
deserializing a Python object structure.
Okay, let's run it. Let's see how many
error
after that we will load the data set and
uh we will go through data
visualization. Okay. Word cloud cannot
import name word cloud. Okay. C C will
be capital here.
random forest.
Okay, it's still loading here. Let's
see. Okay, so loading is done. So now
let's load the data set. So we'll write
data equals to PD dot
read
CSV
name
test.
So you can find this data set on the
description box below.
According to question
polarity
ID,
comma date,
comma
query
per you forget comma
polarity. Okay.
Seems fine.
Let me change this first
is using RNN.
Okay.
So here I will write data plus data dot
sample.
Let's
do one.
Okay. So, let me like brief uh tell you
that what we are going. Okay.
Let me brief you like what we will do in
this sentiment analysis using RNA. So in
this demo like you will see uh text
processing on Twitter data set and after
that we will perform different machine
learning algorithms on the data such as
logistic regression random forest
classifier SVC nas to classify positive
and negative dudes. After that I will
also build RNN recurrent neural network
which is the best fit for such textual
sentiment analysis. Okay. Since it's a
sequential data set which is requirement
for the RNN network. So let's dive into.
So now
we will see the data data visualization
data set details target like the
polarity of the tweets zero negative.
Okay. then the date like date of the
tweet and the polarity and the user that
what tweeted then the text okay so I
will write print
data set
data
shape
okay
let me first do like this. Yeah.
So there are 20
or you can say two like rows and six
number of columns. Okay. So it is a huge
data. I will you can find this data set
from the description box below. So here
let's see the data
and why I use head. Head is used for
like
for showing
top 10 rows of the data set. If you will
use tail instead of head, it will show
the last 10 rows of the data set. Okay.
Here polarity zero. Zero means negative
and four means positive. Okay. Like you
can consider 01.
This is ID, date, then query. then user
then the text.
Okay. So
here I will do data
clarity.
Okay. These are the 04. Okay.
Uniqueness. Zero means negative and the
four means positive. replacing the value
four as one for the ease of
understanding what I said to you you can
consider as 01. So data
polarity
to data
polarity
to one
and then data.
So now you can see 0 1 0 1 0 1 1 0.
Okay.
So if you will write only head it will
show the top five rows only. Okay.
So now let's use one Python function
describe
data dotribe.
So as you can see here count is two lakh
and the mean of the particular row is
this and the ID is this standard
deviation minimum value the 25% the 50%
and the 75% and the maximum
okay let's see the number of positive
versus negative tagged sentence okay
so here I I will write positives
to data
polarity
data dot polarity
= 1.
Then it is
data
polarity
data dot polarity
is equals to zero.
total
length of the data is
dot format
data
dot shape. Yep.
Now I will print
the total length, the negative and the
positive. Okay. So number of positive
Okay.
Format
positives.
So I will copy
this and paste it here.
And here I will do the changes for the
negatives.
Okay. Now let's see.
So here polarity is not defined.
So as you can see the total length of
the data is two lakh and the number of
positive sentences is like one lakh 46
and number of negatives okay spelling
this
the number of negative text sentences
99,954.
Okay. So now we have a brief data.
So now let's get a word count p of text.
So for this I will write
count
words
done
length of
start split.
Okay.
And now let's plot a word count
distribution for both positive and
negative. So I will create a bar plot.
So for that I will write it
word
count.
data
text
dot apply
but count.
Okay, then I will write P positive= data
then
count
data dot polarity
is equals to 1
and
let me copy this Here
I will write zero
and okay then
plt dot figure
and figure size
equ= to
12 Thanks.
Okay. Then plt
LT dota
45
then plt dot x label
word count
plt dot y label
and frequency
we'll write uh g
dot
comma n
Uh
alpha also 0.5 Five
positive.
Okay. Then let's make a legend also.
Location should be
Right.
False.
Data word count equals to
Okay, my bad.
So as you can see the positive and the
negatives.
Okay.
So these are the like word count
distribution for both positive and
negative. Okay.
Now let's uh what we can do we can do
the get like get the common words in
training data set for the training data
set. So for that I will do
from
collections
import
counter
or
words
to
for
test
data
text
line
dot split
forward. Word
and words.
If length of
word
than two
all
dot
one dot lower
here I can write counter
all words
dot most
common then I need 20.
So as you can see these are the most
common word used like in every sentence
the and you for have that I am but just
like this out over all.
So these are the most common words like
it used the is used like 64,000 times
and like this UR is used for 8,000 times
something like that. So now we will do
some data pro data processing. Okay. Now
let's do the data processing.
So
div
and SNS dot current plot
data
polarity.
Okay,
these are the uh negatives and this
positives.
There is a slight change I guess that is
why it's not looking
so much of different like there's a
slight
46 different so that is why it's looking
almost same. Okay.
So now removing the unnecessary columns
like query, user, word count, data dot
drop,
date
query
and word count.
X = 1
comma
place= to true.
Okay.
Uh A will be true.
So here I will write data
what is this? No. Okay my bad.
So here I will write data dot drop
id
comma
one
then data dot head
the data see we have only the to the
polarity and the text. Okay.
So
now uh let's see the null values.
So data
dot
um
data
print. Okay.
So there is no null values. So now
converting pandas's object to a string
type.
For that we have to write
text
to data
text.
Yeah.
Get as type.
Yeah. So now download the stop words
NLTK.
Download
words.
words
as you said
stop words
it's in English
stop
This
These are some, you know, stop words.
So moving forward, let's download
NLTK dot download.net.
So the pre-processing steps taken are
like lower casting each text is
converted to lower case then remover of
URLs will do this we will do okay links
starting with http or https or ww are
replaced by like commas and removing
usernames removing short words removing
stop words like limitization is the
process will do of for the converting a
word to its base Okay. So for that
what I will do
we'll just copy the whole code for you.
We'll explain you one by one what I've
done.
Okay.
So this is a course for the URL pattern
for removing all the WW, HTTPS and HTTP
type of thing and removing
them. Then I have used pattern for the
lower casting removing all the URLs.
Okay.
Then removing all the usernames like at
the red and removing punctuations
and stop words.
Okay. Like this.
So now what we have to do data
processed
weights
then data
Next
dot apply
lambda
x
process
then tweets.
Okay.
Then print
next.
reprocessing.
It is taking time
It will be completed. It will return
here the text prep-processing is done.
Okay.
As you can see the text prep-processing
is done. So now let's check
data dot add
10.
As you can see see the at the rate and
this slices are gone.
Okay.
So now the text is pre-processed.
So now what we will do? We will analyze
the data. So now we are going to analyze
the pre-processed data to get an
understanding of it. We will plot word
clouds for positive and negative dudes
from our data set and see which words
occurs the most. Okay. First we will uh
create for the negative words or
negative tweets you can say. So I will
write pl dot figure
then figure size
15.
Okay. Then word cloud also
word cloud
x words
2,00 comma
width = to 1,600
comma
height = to 800
rate
dot join the data dot polarity
and I will write here polarity
okay equals equals to zero
again
then
processed tweets.
Okay.
Then here
I have to write plt dot show
me show
wcolation
linear.
Perhaps you forget the comma here.
So 2000
then comma width
dot generate
here. what I can do.
Let me run now. Let's see. Hope this
time it will work.
Guess is still loading.
As you can see this is
okay like today I am and work don't wish
they need much. These are the most
negative tweets. Okay, words from
negative tweets you can say,
right? So, let's see the positive
tweets. Okay,
so the thing will be same.
Let me copy
paste it here. So for this I will do one
it will take a little bit of time
to come loading like as you can see hit
can't. Okay sorry
these are the negative words.
Okay still loading. So let's wait for
like few seconds.
Now you can see the positive words like
love, okay, good, lol
and awesome something like that. Okay,
so these are some
positive words. So now let's do the
vectorzation and splitting the data like
storing into input variable process to X
and output variable polarity to Y. Okay,
we'll do that.
So x = to data
possessed
with
values
and pi= to data
entity
dot
values.
Okay.
Now I will write here print
dot shape
print y dot.shape.
Okay cool. So now what we will do we
will convert text to word frequency
vectors. Okay. TF to IDF. So this is an
acronym that stand for term frequency to
inverse document frequency which are the
components of the resulting scores
assigned to each word. Okay. So term
frequency this summarize how often a
given word appears within a document and
inverse document frequency this
downscales word that appear a lot across
documents. Okay. So now here we will
convert a collection of raw documents to
a matrix of TF to IDF features. Okay.
And then I will write
enter
kazut
riser
and sublinear
x = to
dot with
transform
printed.
Okay.
number of feature
comma length
vector
do get
their names.
So number of feature words are like 1703
to 1.
Okay. Now we will do like
now let's print the shape.
So now we will do the split uh spread to
train and test. So the pre-provised data
is divided into two sets of data
training data and the testing data. So
data set upon which the model would be
trained on contains 80% data and the
test data is the data set upon which
model would be tested again contains 20%
of data. So for that I will write extra
test
comma
Test
test size
= to 0.20 2
random
state
101.
Okay. Random state.
So what I will do? I will do the you
know print the shape of X train, Y
train, X test, Y test like how many
columns are there? Rows not column
exactly the rows are there. Okay.
So we'll paste there. So see
extra train like this is a total was
like two lakh.
Okay. So 1 lakh 60,000 in training as we
discussed earlier like 80% in training
and 20% in testing. Okay.
So now let's do the model building.
Okay. Model evaluating functions. So now
let's make a model.
Okay.
And first I will do I will write and
then I will explain you the whole. Okay.
So here what I did uh this will tell you
the accuracy of the model of training
data and the testing data. Okay. Then we
will predict the values for test data
set and the evaluation for the data set.
Then we will compute and plot the
confusion matrix.
Okay, the both the categories negative
positives. Okay, group name will be true
negative and the false positive. Okay,
so there's nothing that's let's run it.
So now what we will do? We will do first
for the logistic regression. So here I
will write LG equals to
logistic
regression.
Okay. Then history
equals to LG do fit
X train,
Y train
with model
evaluate
LG. Now let's see
this is for the logistic regression.
Okay,
as you can see the accuracy of the
training data is 83% the testing data is
77%.
Okay.
So this is the confidence matrix the
predictive value like these are the
categories.
Now let's see for the linear SPM. For
that I will write SPM
equals to
SVC
then SVM
dot fit
train.
Then model
evaluate
of SVM.
Okay.
And after that we will do for random
forest and the N base. Okay. Then we
will start with the RNN.
So as you can see the accuracy of
training data is very pretty good 93%
and logation is 83% and the testing is
less than
regression model. Let's see for the
random forest. So I will write here RF
equals to
random forest
fire
m=
to 20
criterion = to
tropy Okay.
Then max
depth equals to 50.
Then RF dot fit
X train,
Y train
and model
evaluate.
Okay,
loading. Let's see the accuracy how it
will come.
After this we will do for the name base
and after that we will move on to the
our main model RNN recurrent neural
networks.
It's still loading.
Guess it will take little bit of time.
So as you can see the confusion matrix.
Okay. So training data accuracy is 75%
very less. So now let's see the last
model name base. Okay. So, NB equals to
NB
NB dot fit
SP,
wide train.
Okay. Then model
evaluate.
Oh, NAB base training 867.
So, as for
linear SEC has the best
test training uh accuracy you can say
and the best testing accuracy is 7670
76.45 4 five
see logistic regression. So now let's
move to the our main model RNN. So what
is RNN recurrent neural network at the
start
are the state-ofthe-art algorithm for
sequential data and are used by Apple CD
and Google search voice. It is the first
algorithm that remembers its input due
to an internal memory which make it
perfectly suited for machine learning
problem that involve sequential data.
And there is one more thing embedding
layer. Embedding layer is one of the
available layers in KAS. This is mainly
used in natural language processing
related applications such as language
modeling but it can also be used with
other tasks that involve neural networks
while dealing with NLP problems. We can
use pre-trained word embedding such as
glow.
Alternately we can also train our own
embeddings using kas emitting layer.
LSTM layer long short-term memory
networks usually called LSTMs I have
made already many videos you can check
it out were introduced by Skyer these
have widely been used for speech
recognition language processing
sentiment analysis and text prediction
before going deep into LSTM we should
first understand the need of LSTM which
can be explained by the drawback of
practical use of RNN so let's start with
RNA
Okay.
So here I will importing some libraries.
Okay. So after that I will write import
kas
version
2.110. Okay fine.
So now let's
paint
X test, comma,
white train.
Let's do train
test
weights, comma, data dot polarity
dot values.
Then test
size equals to 0.2. Test size 0.2 means
like 80 and 20%
thing 80 to training and then 20% to
testing.
Okay.
and let's
the model evaluation. Okay.
So I will these are relu sigmoid all the
you know the layers.
So now this epoch it will run till 5,000
like count will go till 5,000. Okay see
the 5,000 and it will go to 1 to 10. So
it will take time. So I will get back to
you after this completing this. Okay.
Now as you can see uh the box
ran successfully. Okay. So what should I
do? But I will give some space here. So
now we will see the positive and
negative outcome. Okay. This is
something like testing. Okay. We will
test. We will predict. we will give one
uh a sentence and then we will predict
it is coming right or wrong. The
accuracy is giving a right or wrong.
Okay. So here I will write
sequence
equals to tokenizer
dot text
to
sequences.
Okay, then I'll write this
data science
article.
This was
okay.
So here I will write test equals to P
sequences
and here I will write sequence
Comma max length
to
max length.
Then I will write here prediction equals
to model.
We write model
then we'll write model 12
dot predict
then test.
Okay.
If diction
is greater than 0.5 means 50%.
Then
it should print
positive.
Okay.
Else
negative.
Okay. Let me run this.
Okay. Sequential
object has no okay spellic
see the negative because here is the
word worst it is showing correct. Now
check from the RNN model. So model
equals to kas dot models dot load
models. Here we will load RNN model. RNN
model
SG file. It is pre-trained model. Okay.
Pretend RNN model. So sequence
tokenizer
dot text
to sequences.
Then
I will write here this this
ML
course
is best.
Okay. S equals to P sequences
sequence
X
okay then prediction equals to model dot
predict
Then test
if prediction is greater than 0.5
in
positive 0.5 means 50% more than 50%.
Else
print negative
attribute load models
positive because this ML course is best.
So there is no negative word.
Okay.
So what we will do now we will do model
saving loading and prediction. Okay. So
for that uh I will write import pickle
file = to open
vectorzer
then
here I will Pickle
dot dump
dump
vector file vector.
Okay.
So like this I have to write for name
base logitation SVM and random forest.
So
what I will do
right here.
Okay. Let's run this.
Okay.
Now what we have to do? We have to
predict using saved model. Okay.
What we will do here? We will load model
first and we will predict. Okay. So
first I will write the function name
load
models
and we will load the vectorzer. So file
equals to
open
vectorzer
dot pickle
IBizer
file.
file dot close.
Now I'm loading the logistic regression
model. So for that we have to write open
B
LG to pick
code
file
then file dot close
then
riser
LG.
Okay.
Yeah. So now we will predict the
sentiment. So for that I will write here
predict
riser
text.
Okay. Uh so here we will predict the
sentiment. So for that
text equals to
process
then demands for
sentiment in
text.
Then text
data
dot transform.
This is
okay. Then sentiment
model dot predict.
So here I will make a list of text with
sentiment. So for that I will write data
equals to empty array. Then for text
prediction
and zip
text x
sentiment
dt
append
text prediction.
Okay.
Then we will convert the list into pas
data frames. So for that I will write df
= to ad dot data plane
comma columns
person
next
comma
sentiment
then df equals to df dot
Replace
comma 1
positive
and here.
Okay.
So at last I will write here if
to
then here we will loading the model
vectorzer
comma ng plus load.
Here we text to classify like what
should be in the list. So like text
here I will like I love machine
name.
So
John
be so
So here df equals to date
the command text
then print
df.
I love machine learning. Positive. B is
so active. Positive. J I feel so good.
Negative. Okay. There is
and
yeah. See now it's coming. Okay.
This is how you can do the sentiment
analysis using uh RNN model. Here we
have loaded RNN model. So it is showing
right. So let's do them as poses
right here.
Add
one.
negative. Okay,
RNN model is working. So right, today we
are going to explore K nearest neighbors
or KN&N which is one of the most popular
algorithms in data science. Python is a
powerful tool for data science and KNN
is great for classifying data by
predicting the category of sample base
on its closest neighbor. This algorithms
is used in many field like healthcare,
finance and agriculture helping us make
decision based on data. The best part it
is really easy to use. You just need to
pick a number for K and choose a
distance function to compare data
points. However, KN&N has it downsides.
It doesn't work well with the large data
set and it require proper scaling of the
data to get accurate result. In this
video, we will show you how KNN work
with real data set, the Iris data set.
We'll walk you through simple Python
code and demonstrate how to find the
best K value to maximize your model's
accuracy. So stay tuned to seekn in
action. So welcome to the demo part. So
I'm here using Google Collab. So you can
use any of your favorite ID like Jupyter
notebook, Intelligi, Visual Code Studio,
anything. Okay. So let me rename this
file as KNN classification.
Okay, cool. So let me tell you that KN&N
can be used for the classification
regression predictive problems. So KN
falls in the supervised learning family
of algorithms. Okay, so we will measure
the distance between the K neighbors and
the first step will be we will choose
the number of K of neighbors. Then uh uh
we'll take the k nearest neighbors of
the new data point according to your
distance metric. And the step three will
be our among these case neighbors count
the number of data points of each
category. Okay. Step four we will assign
the new data points to the category
where you counted the most neighbors.
Okay. So let's start. First let's import
some library. import
numpy
as np and let's import
pandas sp. So everyone knows what is
numpy and the pandas. Okay, so numpy is
a library for the python programming
language adding support for uh you know
large multi- dimensional arrays. Okay.
Along with the large collection of uh
what to say uh highlevel mathematical
functions and various pandas is a
software library written for the Python
programming language for the data
manipulation and analysis uh all the
data frames and uh data structures it
offers for the manipulating numerical
tables. Okay. And the time series you
can see. So moving forward uh we'll
import our data set. So you can download
the data set from the description box
below. Okay. data set
equals to pb dot read
csv. The data name is iris dot csv.
Okay. This is how you read uh your data
in python. Okay. Yeah. Data set is
loaded. So data set dot shape
shape. Okay. Yeah. So data uh set dot
shape is used for how many numbers of
rows and columns present in your data
set. Okay, 150 rows and six columns. So
let me tell you brief about data set
this data set. So this data set include
three Iris species with uh 50 samples
each as well as some properties about
each flower. So one flower species is
linearly separable from the other two
you can say but the other two are not
linearly separable from each other.
Okay. And this shape I told you we can
get a quick idea of how many instances
of rows and columns are present in our
data set. So let's see our data set.
Data set dot
head. So head is used for uh you know by
you can see top five rows of your data
set using head and if you will use tail
you instead of head you can see the last
five rows of your data set. Cool. Okay.
So columns are ID sample length sample
width petal length petal width and the
species. Cool. Then moving forward let's
describe our data sets. So these are the
basics uh basic function. Okay.
of Python you can say.
So data set.escribed what describes do
is it will give you count of all the
rows mean value standard deviation value
minimum value what is the 25% okay of
all the values in the particular row.
What is the 50%? What is the 75%? What
is the maximum? Maximum is 150 you can
say. Okay. and 25% of 150 is 38.25 25
this okay of all the columns if it is
normal uh you know character so it won't
give you any data okay cool yeah so
moving forward uh let's now take a look
at the number of instances row belong to
each classes okay so we will write data
set dot
group by
species
dot size. Okay. Uh spec S is capital
that's why it's showing the error. Yeah.
So you can see Iris Satossa are 50, Iris
verical are 50 and virginica is 50.
Okay. And the data type type is integer.
Cool. So as you can see data set
contains six columns like ID, sample
length, sample width and petal length,
petal width and spacing. The actual
features are described by columns 1 to
four. the last columns labels or
samples. Okay. So firstly we need to
split data into two arrays like X
features and Y labels. So how we will do
this? By writing code like feature
columns equals to
sample length sample width petal length
petal width. Just remember you are
writing correct name. Okay. Then x = to
data set
feature
columns.
Okay. Dot
values.
Then y = to
data set
species dot values. Cool. Then let me
run it. So this is uh how we can split
the data set okay into two arrays X and
the Y. Okay. In X there are feature
columns. These four columns are there
and in Y species column is there. Okay.
And what is the species column this
Satossa venica and all this. Cool.
Then now we will do label encoding. So
as you can see labels are categorical
Kverse classifier does not accept string
labels. So we need to use label encoder
to transform them into numbers. Okay.
Then iris satossa correspond to zero.
Iris vericy color correspond to one and
I is virginica correspond to two. Okay.
012. Cool. So how we can write from
skarn dot pre-processing
import
label
encoder okay so what we'll do label
encoder transform them into numbers okay
so I will write
l equals to
label encoder
okay my bad label encoder. Okay. Then y
= to ele alate transform y. Okay. So now
I will run it. Yeah. Correct. So
splitting data set into training set and
the test set now. Okay. So now we'll
split data set into training set and
test set to check later on whether or
not a classifier work correctly or not.
Okay. So here I will write from skarn
dot not cross. I will write it here.
Skarn domodel selection
import
train test split. So what is train test
split? Train test split is a model
validation procedure that reveals how
your model performs on your new data.
Okay. And what is X train access Y train
Y test in Python. Okay. Let me first
write it and then I will let you know.
Okay. Then I will write here
x train comma x test comma y train comma
y test. Okay
train test
split
then x comma y
comma test size
0.2 and the random is this. Okay. Okay.
Some error came. Okay. Underscore model
selection.
Yeah. Cool. So what is X train X test? Y
train Y test. Okay. So X train and Y
train sets are used for training and
fitting the model. Okay. So the X test
and the Y test are the set used for
testing the model and it's predicting
the right outputs level. Okay. So here
you can see test size is 0.2. into means
80% is for testing or sorry 80% is for
training and 20% is for testing for the
new data. Cool. Yeah. So now we will see
some uh let's do some data
visualization. Okay. So here I will
write import
mattplot lib
dotpipplot
as plt
then import
cb
as sns
then here I will write person mattplot
lib
in line. So what is mattplot lib?
Mattplot lib is a plotting library for
the python programming language and it's
numerical mathematic extension. Okay. So
numpy it provides an object API for
embedding plots into application using
generating purpose GUI toolkits like
kintter and python or gtk. Okay. Whereas
seabon seabon is a library for making
statical graph in python. It builds on
the top of mattplot lip and integrates
closely with pandas data structure.
Okay, seborn helps us to explore and
understand the data. Okay, so I will run
it here. I will write from
pandas dot plotting
import
parallel
coordinates.
Then plt dot figure
size should be
15, 10. Okay.
Then parallel coordinates.
Okay. Then data set dot drop. I don't
need id,
x is one.
Okay. then comma spaces
then plt dot title
and let parallel
coordinates
plot okay and
you can give some font size
equals to 20
then font
weight
equals to
okay Let's add bold only bold then plt
dot x label
then
features
comma
font size
to 15 then plt
dot y label then I'll write here
features
values
comma font size
equals to 15. Okay. then plt dot legend
then locals to 1 comma
I will write have frame on
equals to true comma shadow
equals to true comma face color
equals to White
T should be capital
and comma edge color
equals to
okay plt dot show okay some error is
there plig
size okay spelling mistake Take.
Okay. One more error. PLT. Legend. Okay.
Face color. Okay. Some spelling stick.
So yeah, let me make it output in full
screen. Yeah. So parallel coordinates is
a plotting technique for plotting you
know multivariate data. So it allows one
to see clusters in the data and to
estimate other stat visually. So using
parallel coordinate points uh you know
are represented as the connected line
segments as you can see. Okay. And each
vertical line represent one attribute
and one set of connected line segments
represent one data point. Okay. And
points that tend to cluster will appear
closer together. Okay. So this uh this
color is iris satossa and this is irisy
color and this ba one is iris virginica
as you can see in the legend. Okay,
cool. So moving forward let's create
another graph and curves. Okay so here I
will write from
p and dotplotting import scatter. Oh,
this uh plotting import
Andreo curves.
Okay, then plt
dot figure. It's AI based. So it's
giving me suggestions. Suggestions.
Suggestions. Okay. So sometimes
suggestions are good but not always.
Yeah. Let's carry on. Figure size is 15,
10.
Then Andrew
curves
then I will add data set dot drop id
access spaces and plt and curves plot.
Okay fine. Okay, let me add we don't
need X label and all. Let me add legend
plot
legend.
Then same LOC equals to 1. Then
proposition size.
Then I will write here size
is 15
of frame on equals to two comma shadow
equals to true. So this is uh truly
based upon you if you want to add legend
or not or you can skip. If you want to
skip you can skip. Okay. Face color
equals to white. Then
edge color equals to black. Okay. Then
plt dot show.
Yeah. So no error is there. So let me
first view output in full screen. Okay.
So Andrew curves these are the endoc
curves. Okay. You can see the graph in
the curve. So and curve allow one to
plot multivariate data as a large number
of curves. So that are created using the
attributes of samples okay as
coefficient for 4year series. Okay. So
by coloring these curves differently for
each class. It is possible to visualize
data clustering. So curves belongs
belonging to samples of the same class
will usually be closer together and they
form large structure. Okay. As you can
see here and this is the legend why we
are setting the four color is should be
white and yeah and the edge color should
be black okay and frame on shadow should
be there you can see the shadow okay
like this if you want to skip you can
skip this part legend part legend part
okay but like it's good to have it's
good practice to have this cool so let's
create one small pair plot okay so I
will write here plt dot
figure
then I will write SNS dot pair plot
then I will write data set
dot drop
we'll drop again id we don't need comma
xis is one then I will write u equals to
species
size equals to three.
Then markers equals to
O SD. Yeah. Cool. Then plt dot show.
It's running. Yeah. So first let me make
it to full screen. Yeah. So here pair
wise is useful when you want to
visualize the distribution of the
variable or the relationship between the
multiple variable separately within
subsets of your data set. Okay. So here
this blue one is stoa verol and the
virginica. Okay. So these are some uh
graph you can say or here you can see
sample length. Okay. So green one is
virginica and versol are almost having
same saple length. Okay. And here are
the different types of graphs. Okay. If
you don't want to uh let's say if you
don't know how to read this graph you
can use this graph or you can use this
graph. Either you can use this graph.
Okay. That's the power of pair pair wise
plots you can say. And the sample width
cm. Okay. Satossa. No. Okay, virginica
and verticola again having almost same
sample width. Okay, and here as well.
Okay, this stoa is in different form.
Okay. Yeah. So, we'll uh create one
small uh pair of u one small graph then
we'll move forward. Okay. Uh I will
write here plt dot. So now we are
creating box plot figure.
Okay. Then data set dot drop.
Then again ID commais should be one
dot box plot. Okay. Then figure size I
chose this. Yeah. Let's run it. Yeah.
Again see this is the box plot. Petal
length same petal width sap length sle
width okay according to width it is
showing like from this to this width we
have the veric color and from this to
this width this length sorry sample
length we have uh that virginica lies
here Iris virginica and from here to
here iristosa lies okay the sample cool
and like you can use 3D models you can
use different types of charts. Okay, I
did four and four are enough to read the
data set. Okay, so now we will do uh
some cannon classification. Okay, we
will make prediction and we will see the
accuracy of our data set. Okay, so now
what I will do? I will write here from
skarn
dot
neighbors
import k neighbors. Okay.
Write from skarn dot
neighbors
import
k neighbor
classifier.
Okay.
Then I will write here from
skarn
dot matrix
import
confusion
matrix
comma accuracy
accuracy score. Okay, then I will write
from
skarn dot model selection
import
cross
value score. Okay, so here what we I did
uh we are fitting the classifier to the
training set and loading the libraries
basically. Okay, then we'll initiate a
learning model K equals to three. Okay.
So here I will write classifier.
Let me give one space. Classifier equals
to
neighbors.
Classifier
and
neighbors
to three. Okay. Then fitting the model.
I will write here classifier
dot fit
x train,
y train.
Okay. Then uh now we will predicting the
test uh test set results. Okay. So y
prediction equals to
classifier dot predict
x test. Cool. Okay. From Okay. Spelling
mistake.
Okay.
Okay.
Yeah. I guess it's fine now. So now uh
let's evaluate the prediction. So here I
will uh build the confusion matrix.
Okay. So here I will write cm confusion
matrix equals to confusion
matrix
uh y test
y prediction. Okay. Then here I will
write cm. Okay. So what is confusion
matrix basically? So I confusion matrix
uh you know is a table that shows how
well a model performs by comparing its
prediction to the actual values. Okay.
So what it show is a confusion uh matric
display the number of correct and
incorrect prediction of each class in
models whose either it can give you true
positive true negative false positive
false negative either it can give you
zero or one. Okay. So yeah moving
forward let's calculate the model
accuracy. This is the main part. If the
model accuracy is low it means uh your
analysis of you know classification is
not good. Okay. So accuracy it should be
more than 80 at least. accuracy
equals to
accuracy
score
y test
comma y prediction into 100 otherwise it
will give me in points so I will write
model
canon model accuracy
is
oh I will write plus STR I will round
it. Okay. Accuracy
I will add the person. So why I wrote
this accuracy to round? So I don't want
after points I need only two numbers.
Okay. I don't want like 6 7 8 9 10 11 12
like this. Okay. I need only like 80.20
like this. Cool. Let's run this. see
KN&N model accuracy is 96.67 67 that's
why I wrote two here and the percent
should be there so 96 which is very very
very good okay so this is how you can
find uh the accuracy so now let's find
the optimal number of neighbors in K
okay basically finding the best K so we
will use uh using cross validation
parameter okay tuning so first I will
create the list of K for KN okay so here
I will write K
list
equals to list
range
1 comma 50A 2.
Here I'm I will create the list of CV
score. Okay. So here I will write CV
scores.
Okay. Okay. I have to give brackets.
Yeah.
So we'll here perform the 10fold cross
validation. Okay, I will explain you
what is cross validation. Don't worry.
So first let me write for K N K listN
equals to
K neighbor classifier
and here I will add N neighbor
neighbors equals to K. Then scores
equals to cross
value score
KN&N then X train
Y train
okay
I will write here cross validation
equals to 10
comma scoring equals to I will write
here accuracy
Okay. And then here cv
course dotappend
to course dome.
Okay. So now yeah let me run it. Okay.
Comma some error came. Okay. The error
is the scoring parameter.
Yeah. Why? Because here accuracy you can
see and I did the spelling mistake.
Okay. So now what is cross validation?
Of course cross validation uh you know
determine the accuracy of your machine
learning model by partitioning the data
into two different groups. Okay called
training set and testing set. You can
see train test and the testing set. Okay
X and Y. So the data is randomly
separated into a certain number of
groups or subsets called folds. Okay you
can see the 10 folds we have wrote. Each
fold contains about the same amount of
okay and there is one more thing
validation. So validation is a technique
for assessing the accuracy of the model
on data set. Okay. And this cross
validation we did on new data set. Cool.
So now let's find the best K. Okay. So
here I will write best K equals to K
list. Okay. Before that I will write one
thing. I will write here MSE was
changing to mclassification error. Okay.
So equals to one
that's X for X in for X in
CV scores.
Okay. Okay, list I will write MSE dot
index
minimum
MSE.
I will use this bracket square bracket.
Okay. So then I will write print
the best
optimal
number of
neighbors
is person
best K. Okay, let's run it. See the best
optimal number of neighbors. Okay. N
neighbors is 9. Okay. So in the K
nearest neighbor KN algorithms. Okay. So
K represent the number of neighbors that
are considered when classifying a query
point. Okay. See the if we'll classify
this particular point you will get the
six. Okay. 1 2 3 4 5 6. Okay. And uh let
me show you. If you will classify this
portion only, so you will get 1 2 3 4 5
6 points like this. Okay. So the best
optimum number of neighbor is nine.
>> What is Python? Python is a high-level
object-oriented programming language
developed by Guido Van Roum in 1989 and
was first released in 1991.
Python is often called a batteries
included language due to its
comprehensive standard library. A fun
fact about Python is that the name
Python was actually taken from the
popular BBC comedy show of that time
Montipython's Flying Circus. Now let's
look at the top features of Python
first. So Python has a simple structure
and a clearly defined syntax. This
allows the learners to pick up the
language quickly. So it is easy to learn
and use.
Python can run on different operating
systems such as Windows, Linux, and Mac,
making it a portable language. It
enables programmers to develop the
software for several competing platforms
by writing a program only once.
Third, Python is freely available at the
official website since it is open
source. This means that source code is
also available to the public.
Now, Python uses an object-oriented
approach that encapsulates code within
objects.
Python provides a collection of
libraries for various tasks such as
machine learning, web development, and
data analysis. And finally, in Python,
you don't need to assign the data type
of the variable. When you assign some
value to the variable, it automatically
allocates the memory to the variable at
runtime.
Now, with that, let's move on to the
uses of Python programming.
So, Python programming language is used
to develop desktop applications and
build web applications too. It is
popularly used in the field of data
science, machine learning and artificial
intelligence to analyze data, build
predictive models and make business
decisions. Python is also widely used in
game development. Now, let's see some of
the popular Python frameworks and
libraries.
Python can be used for web development
using frameworks like Zango, Flask,
Pyramid and Churi.
Now you can build graphical user
interfaces using libraries and
frameworks such as Tkinter or just KER.
You can also use PI GTK, PIQT or PYJS or
Python JavaScript.
Now, Python is also used to perform
machine learning tasks using libraries
such as TensorFlow, PyTorch,
Scikitlearn, Mattplot Lib, and Scypi.
You can also perform mathematical
computations using numpy and pandas.
Now, let's look at the best ids that you
can use to write programs in Python and
perform specific tasks. So, we have
Jupyter notebook, which is part of the
Anaconda distribution that is widely
used these days. Even for our demo in
this video, we'll be using Jupyter
Notebook. I'll show you in a while. Then
we have the visual code editor from
Microsoft. This is also one of the
preferred IDEs by learners and
companies. Then we also have the popular
text editor called Sublime Text editor.
Then we also have PyCharm followed by
Python and Spider as our top idees. Now
let's look at the top companies that are
using Python in our day-to-day work.
So we have Google, Kora, Facebook, even
Netflix, Spotify, and Instagram. Now
there are other top product- based,
service- based and startups that also
use Python programming. So what really
is Python programming language?
Python is an object-oriented highle
programming language that supports
built-in data structures and dynamic
semantics.
It supports multiple programming
paradigms such as structured,
object-oriented and functional
programming.
Python is often described as batteries
included language because it has a
comprehensive collection of standard
libraries. Python supports different
modules and packages which allows
program modularity and code reuse.
Python was developed by Guido Van Rosum
and its implementation started in
December 1989.
Python 1.0 version was released in the
year 1994. Python 2.0 came out in
October 2000 while Python 3.0 was
released in December 2008.
Now that you have got an understanding
of the Python programming language,
let's now look at the top 10 reasons why
you should learn Python.
So at number 10, we have ease of use.
One of the most common reasons to like
Python is that it is quite easy to learn
and code. It provides a simple syntax
that improves readability and makes it
easier to understand. So developers can
create any desktop or machine based
application using this language. Python
is very versatile and is instrumental in
artificial intelligence and machine
learning. We will talk about this later
in the session. Compared to Java or C++,
it has fewer lines of codes.
In the example here, we are printing a
hello world program in Java. As you can
see,
if you have to write a program in Java,
you first have to declare the class name
along with its scope.
Next, using curly braces, you need to
pass the main method along with its
arguments. And then using
system.out.print print len method you
can print hello world that's quite a
tedious task isn't it
the same task of printing hello world
can be done using just one line of code
in python as shown here you can write
the print function and pass whatever you
want to display inside the brackets and
that will print the output it is so
simple
that is why Python is considered as a
highle language and it's open source
You can just download it from the
website and start using it.
At nine, we have active community.
You need a community to learn new
technology and friends are your best
asset when it comes to learning a
programming language. Python has large
community support.
It has an extensive and active community
to assist engineers, developers,
analysts, and data scientists with
expert support in case of programming
errors or issues with the software. You
can just go ahead and put your queries
in the community forum. The community
members will address your queries in
real quick time. Communities like Stack
Overflow also brings many Python experts
together to help learners.
Python enhancement proposals or PEP is
where the proposals and the improvements
are announced. Also, there are a set of
recommendations or core values called
the Zen of Python written by Tim Peters
that represents the guiding principles
for Python development.
Up next at 8, we have portable and
extensible.
Multiple cross- language operations can
be performed effectively because Python
is portable and extensible in nature.
For example, if the users have a Python
code written on Windows and they want to
execute on a Mac operating system or
Linux operating system or Solaris, they
can easily do it without any amendment.
They can also run this code on any
platform flawlessly and without any
interrupt.
Due to its extensibility feature, you
can integrate other programming
languages such as Java,Net, C and C++
codes with Python. The components of
other programming languages can be used
with Python and thus it can be used to
make a crossplatform suitable
application too. So it is a really good
feature that Python provides.
The next reason to learn Python is
testing frameworks.
Python supports several built-in
flawless testing tools and frameworks
that help in debugging and speeding of
workflows.
Some of the tools and frameworks
supported by Python are Piest, Selenium
and Splinter. This is the reason for
which every tester tries to use Python
based tools and frameworks to test any
application or code or to validate it in
an easier manner. Piest is the most
recommended testing framework for
functional, integrational, and unit
testing. You can run Selenium test
scripts using Python programming
language to automate various tasks. And
Splinter is an open-source tool for
testing web applications using Python.
It lets you automate browser actions
such as visiting URLs and interacting
with their items.
At number six, we have libraries and
packages.
Another reason why Python has become so
popular in the industry these days is
that it has a massive collection of
libraries and packages that make your
task simple and easy. It has a range of
libraries, packages, frameworks, and
modules for data manipulation,
statistical calculation, web
development, machine learning, and data
science.
Python programmers have developed tons
of free and open-source libraries that
you can use. You can find many of them
via Python package index, the repository
of Python software. Python provides the
default package called pip. Anaconda is
a third party Python ecosystem. Other
examples include numpy, sci and zango.
Then we have scripting and automation.
Python is not just a programming
language. It can also be used for
writing scripts for automating tasks and
workflows without human intervention.
The code can be written in the form of
scripts and executed later. Further, it
is interpreted by the machine and
checked for errors at runtime. The
machine is used to read and interpret
the code. Once the developer checks the
code, it can further run or be used
several times without any interruption.
This allows you to automate a set of
certain tasks within a program or the
same code can be used with other
applications as well.
At number four, we have web development.
Another reason to learn Python is that
it makes the web development process so
much easier.
It provides a wide collection of
frameworks that make it easier for
developers to develop web applications.
Some of the examples are Zango, Flask,
Pyramid, Turbo Gears, CherryPie, etc.
These frameworks are written in Python
which makes the code a lot faster and
stable.
The task which used to take hours in PHP
can be finished in minutes using Python.
Python is also used for web scraping.
Django offers many elements of intricate
programs such as template design,
management panel, signing in, signing
up, signing out, URL routing, etc.
Once the user establishes the framework,
all these features become ready to use.
Flask is a microwave framework written
in Python.
of all the components that are part of
this module, they are all ready to
execute in the server context.
Pinterest and LinkedIn use Flask.
Pyramid offers more attributes than
Flask. It will assist users with URL
routing and authentication support.
Turbo Gears is a highly recommended and
scalable framework that supports
features such as authentication,
caching, identification, management of
sessions, and pluggable applications.
Up next at number three, we have machine
learning.
The growth of machine learning has been
phenomenal in the last 5 years and it's
rapidly changing the world around us.
Python is one of the most preferred
programming languages for machine
learning because of its simple syntax
and support for several machine learning
libraries.
Using different libraries and functions
in Python, the system can learn and
train itself from past data.
Once the system is trained, it can then
learn to adjust itself to new inputs.
Finally, it can make predictions and
perform humanlike tasks automatically.
At number two, we have data science.
Machine learning and data science go
hand in hand. Python is robust, scalable
and provides extensible visualization
and graphics options. Hence, it is
widely used in data science.
Python has libraries such as numpy for
numerical computation of data, pandas
for operations to manipulate data on
numerical tables and time series. It
also provides simply for symbolic
computation and sci for technical and
scientific computations.
It has another library called pyrain
which is sought for python based
reinforcement learning, artificial
intelligence and neural network library.
Scikitlearn is the machine learning
library for creating classification,
regression and clustering algorithms.
And finally, it provides PyTorch and
TensorFlow for deep learning.
Finally coming to the most important and
the top reason to learn Python which is
career opportunities and salary.
Python language provides a variety of
job opportunities and promises a high
growth graph with huge salary prospects.
It is been used by most of the tech
giants.
Industry leaders using Python are
Amazon, Google, Facebook, IBM, NASA,
Netflix and YouTube.
Next, you can see the Google trends
but I have considered three programming
languages Python, Java and C++. I have
compared them for the past 12 months.
You can see it clearly on your screens
that Python has become a frontr runner
in terms of popularity and web search
volume. It means people are interested
in Python. They want to learn it and use
it in their work. You can also check for
the YouTube search.
There also you will find that Python
programming language is the most
searched language on YouTube.
Now on your screens you can see the
report of PPL which is popularity of
programming language index. It is
created by analyzing how often language
tutorials are searched on Google. It is
a leading indicator. The raw data comes
from Google trends. The bar graph
depicts that Python is the most popular
and widely used programming language
across the globe followed by Java then
JavaScript and C.
The popularity of programming language
index can help you decide which language
to study or which one to use in a new
software project.
The next graph shows the popularity of
Python and Java over the years starting
from 2004 till the current period which
is 2020. Worldwide, Python is the most
popular language. Python grew the most
in the last 5 years by 19.4%. 4% and
Java lost the most by minus 7.2%.
Now let's talk about the different
career opportunities and the job roles
that you can get into if you learn
Python language.
First, you can become a Python developer
where you will be asked to write and
test codes, debug programs, and
integrate applications with third party
web services.
Second, you can become a web developer.
Here you will be responsible for writing
serverside web application logic. Python
web developers usually develop back-end
components, connect the application with
third party services and support the
front-end developers by integrating
their work with the Python application.
You can also become a data analyst if
you know Python. As a data analyst, you
have to gather data from multiple
sources using scripts. analyze that
data, develop and implement databases
and data collection systems.
You can become a data scientist. As a
data scientist, you need to understand
the challenges in business and come up
with the best solutions using modern
tools and techniques to analyze,
visualize, and build prediction models
to make business decisions.
Lastly, you can be a machine learning
engineer where you can develop
intelligent machines that can learn from
vast volumes of data and apply knowledge
without human intervention.
So there's a lot of scopes if you learn
Python. But before we move on, let's
understand first what is Jupyter
Notebook. So guys, as you can see all
over here that Jupyter Notebook is a
popular open-source tool that basically
allows you to create and share documents
which contains codes, equations, you can
have visualizations also. Basically, it
is used for data analysis, machine
learning and scientific research which
makes it a very essential tools for
developers like data scientists and
researchers alike. Now before installing
Jupyter notebook I request you that you
have Python installed in your system. So
the requirement should be Python 3.6 or
greater. So now let us officially
navigate to the Python's website. So
guys as you can see all over here. So on
python.org if I click on download
Python. So we're going to see that all
over here download Python 3.125. So as I
already told you that the requirement of
Python should be greater than 3.6. So
just you can click all over here and you
can see the download has started.
So guys as you can see all over here
that we have installed the Python. Now
let us open the file. So you can see the
given software is going to installed on
this directory. Okay. So just click all
over here. So guys as you can see all
over here the Python installation of
3.125 is in progress. Let's wait for
some time till it gets installed.
So as you can see guys all over here
that we have successfully installed our
Python. Now let us open our terminal and
let us check whether Python is correctly
installed. So we are going to type
python
/ version.
So as you can see all over here we have
successfully installed our Python. So
guys that was our prerequisite. Now
there are two ways to install Jupyter
notebook. The first one can be pip.
Okay, pip is a package manager or using
Anocanda distribution. So let us see
with pip first. So guys, pip is a
package manager which is used to install
and manage software packages libraries
written in Python. So you can see all
over here that the Python with version
greater than 3.6 have default pip
installed in them. Okay. So we can use
pip command to install our Jupyter
notebook. So guys as you can see all
over here we have come to the official
documentation of jupitter.org and it is
saying that installing Jupyter lab with
pip command. So what you can do guys you
can just copy all over here. You can go
right all over here and click on this.
Now as you can see all over here it has
started downloading the Jupyter lab.
So guys, we are going to install our
Jupyter lab with the pip command. So
this is the official documentation of
Jupyter notebook. Okay? And just all you
have to do is copy this and type on your
terminal. So as you can see all over
here it has started downloading the
packages which is required to download
the Jupyter notebook. Let us wait for
some time.
Okay guys, so we have successfully
completed this step. Now let us move on
to our next step. So as you can see all
over here. So we have installed. Okay.
Then what we have to do then you can
type this. We can launch the Jupyter lab
with this command on the terminal. Now
let us wait. So as you can see all over
here guys, we have successfully
installed our Jupyter notebook. So you
can go all over here and just create a
new notebook and you can also choose
your kernel and you can start working on
your Jupyter notebook. Suppose I'll show
you one snippet. So 3 + 5. Let us try to
run this notebook. So as you can see it
is giving us the eight as answer. So it
is following the Python syntax and in
this way we have successfully installed
our Jupyter notebook using the pip
command. So now as you can also see all
over here you can also install Jupyter
notebook with this command pip install
notebook and then you can just open it.
This is also an another alternative.
Similarly, you can install with VA also
same command and just open the VA. Now,
if you are using any other operating
system like Mac OS or Linux, then you
can install by brew install Jupyter Lab.
So, home will be the package manager for
Mac OS and Linux. So, I hope so you are
pretty clear with how to install Jupyter
notebook with the pep command. Now, I
have downloaded Anacondas from this
official website. So as you can see all
over here this is the official website
of Anaconda. Okay. Now just type your
email and you can just download it. So
similarly as you can see after
installing I'm going to launch my
installer and let us click next. Okay.
Let us click agree. Okay. And let us
install this on the given directory.
Let us wait for some time till the
installation gets complete.
So guys as you can see all over here we
have completed our installation of
Anoconda. So just click on finish and
you can say we have successfully
installed our Anocanda. Now let us open
our Anaconda navigator. So just click
on.
So as you can see all over here just
right click on this and our Anocanda
navigator will be opened. So as you can
see all over here this is our Anocanda
navigator and it is loading the packages
and for us to install the Jupyter
notebook. So as you can see all over
here just click on launch. So guys if
you click on launch it is going to open
our Jupyter notebook. So as you can see
all over here it is saying launching the
Jupyter notebook and it is hosted on
localhost 8889. So this is our hosted
Jupyter notebook and in similarly you
can create a new notebook all over here
and in this way you can start working
>> LLMs. If you ever wondered how machine
learning can now understand and generate
humanlike text, you are in the right
place. From chatboards like Chat GPT to
AI assistant that powers search engines,
LLMs are transforming how we interact
with technology. One of the most
exciting advancement in this space is
Google's Gemini or OpenAI Charging large
language model designed to push the
boundaries of what AI can achieve. In
this video, we will explore what LLMs
are, how they work, and why models like
Geminy are critical for the future of
AI. Google Gemini is part of a new wave
of AI models that are smarter, faster,
and more efficient. It is designed to
understand context better, offer more
accurate responses and integrate deeply
into service like Google search and
Google Assistant, providing more
humanlike interactions. So we will break
down the science behind LLMs including
their massive training data set,
transformer architecture and how models
like Gemini use deep learning innovation
to change industries. Plus we will
compare Google Gemini to other popular
LMS such as OpenAI Chity models showing
how each of these technologies is used
to power chat bots, virtual assistants
and other AIdriven application. By end
of this video, you will have a clear
understanding of how large language
models like Gemini work, their key
features, and what they mean for their
future AI. Don't forget to like,
subscribe, and hit the bell icon to
never miss any update from Simply Learn.
So, what are the large language models?
Large language models like Chargen
pre-trained transformer 4 o and Google
Gemini are sophisticated AI system
designed to comprehend and generate
humanlike text. These models are built
using deep learning techniques and are
trained on vast data set collected from
the internet. They leverage self
attention mechanism to analyze
relationship between words or tokens
allowing them to capture context and
produce coherent relevant responses.
LLMs have significant application
including powering virtual assistant
chatboards, content creation, language
translation and supporting research and
decision making. Their ability to
generate fluent and contextually
appropriate text has advanced natural
language processing and improved human
computer interaction. So now let's see
what are large language model used for.
Large language models are utilized in
scenarios with limited or no domain
specific data available for training.
These scenarios include both few short
and zero short training approaches which
rely on the model's strong inductive
bias and its capability to derive
meaningful representation from a small
amount of data or even no data at all.
So now let's see how are large language
models trained. Large language models
typically undergo pre-training on a
board. All encompassing data set that
shares statical similarities with the
data set specific to the target task.
The objective of pre-training is to
enable the model to require highlevel
feature that can later be applied during
the finetuning phase for specific task.
So there are some training processes of
LLM which involves several steps. The
first one is text prep-processing. The
textual data is transformed into a
numerical representation that the LLM
model can effectively process. This
conversion may be involve techniques
like tokenization encoding and creating
input sequences. The second one is
random parameter initialization. The
model's parameter are initialized
randomly before the training process
begins. The third one is input numerical
data. The numerical representation of
the text data is fed into the model of
processing. The model's architecture
typically based on transformers allows
it to capture the conceptual
relationship between the words or tokens
in the next. The fourth one is loss
function calculation. A loss function
calculation measures the discrepancy
between the model's prediction and the
actual next word or token in a syntax.
The LLM model aims to minimize this loss
during training. The fifth one is
parameter optimization. The model's
parameter are registered through
optimization technique. This involves
calculating gradient and updating the
parameters accordingly gradually
improving the model's performance. The
last one is iterative training. The
training process is repeated over
multiple iteration or epox until the
model's output achieve a satisfactory
level of accuracy on that given task or
data set. By following this training
process, large language model learn to
capture linguistic patterns, understand
context and generate coherent responses
enabling them to excel at various
language related tasks. The next topic
is how do large language models work. So
large language models leverage deep
neural network to generate output based
on patterns learned from the training
data. Typically a large language model
adopts a transformer architecture which
enables the model to identify
relationship between words in a sentence
irrespective of their position in the
sequence. In contrast to RNAs that rely
on recurrence to capture token
relationship transformer neural network
employ self attention as their primary
mechanism. Self attention calculates
attention scores that determine the
importance of each token with respect to
the other token in the text sequence
facilitating the modeling of intricate
relationship within the data. Next,
let's see application of large language
models. Large language models have a
wide range of application across various
domains. So here are some notable
application. The first one is natural
language processing NLP. Large language
models are used to improve natural
language understanding tasks such as
sentiment analysis, named entity
recognition, text classification, and
language modeling. The second one is
chatbot and virtual assistant. Large
language models power conversational
agents, chatbots, and virtual assistant
providing more interactive and humanlike
user interaction. The third one is
machine translation. Large language
models have been used for automatic
language translation enabling text
translation between different languages
with improved accuracy. The fourth one
is sentiment analysis. LLMs can analyze
and classify the sentiment or emotion
expressed in a piece of text which is
valuable for market research, brand
monitoring and social media analysis.
The fifth one is content recommendation.
These models can be employed to provide
personalized content recommendations
enhancing user experience and engagement
on platforms such as news website or the
streaming services. So these application
highlight the potential impact of large
language models in various domains for
improving language understanding
automation. So hello guys welcome to
this demo part of this video. So here
what I will do I will go to new then
Python 3 file
then here
I will give it the name called
exploratory
data
analysis. Basically we will so we have
one data set file of
roller coaster basically. So we will be
using that and you can download that
file from the description box below from
the below link driving link. Okay. So we
will be doing some small basic functions
using Python and later on we will uh
make some good charts. Okay. We will
remove duplicates and all we will do all
that
thing. We'll do data preparation. We'll
do feature engineering. Okay. And uh
we'll remove the duplicates. We'll check
for the duplicates. We'll make charts
like histogram, KD blocks, box plot and
like many more similar to that like heat
map. We will make scatter plot group by
comparison. Okay, we all do that, right?
So just stick with me and you can write
side along with me here this code. Okay.
Okay. So let's start with importing
pandas first.
I guess everyone know what is pandas and
numpy
fine
as np then I will import
numpy
as np why this is np and pd sorry my bad
pd is here because I don't want to write
again and again this pandas this numpy
so basically I can write this small
version okay Yeah. So if you guys don't
know what is panda. So panda is very
popular library for working with data.
Okay. It's goal is to be the most
powerful and flexible opensource tool.
Okay. So it has reached that goal. So
data frames are the center of pandas.
And what is data frame? A data frame is
structured like a table or a
spreadsheet. Okay. The rows and the
columns. Okay. Whereas numpy numpy is an
open-source Python library again that
facilitates uh you know efficient
numerical operation on large quantities
of data. Okay. So there are many
functions in numpy as well those we can
use in the pandas data frame. Got it. So
we'll import one more
dot piplot
dotp. Okay. as
plt. Okay, there is one more library
mattplot lip for plotting the graph and
and there is one more import
seabbon
as SNS. So what is se? Sebon is again
the Python data visualization library
based on Matt plot lip. Okay, it
provides a highle interface for drawing
attractive and informative statical you
know graphs. Okay. Then I will write
here plt dot style
dot use.
We'll write ggplot.
Got it. ggplot.
Fine.
So yeah,
let me run it. So what is ggplot? So
ggplot is an again this is an opensource
data visualization package for
historical programming. Okay. So yeah uh
you can say or a general scheme for data
visualization which breaks up graph into
you know semantic components such as
scales and layers. Okay. ML lip py
pyab
p. Okay. Yeah. So now
we will import our data set df. DF means
data frame. You can write any word of
your choice. Then pd again pandas dot
read
csv used for readings CSV file. Okay.
Excel files. Got it? Then here I will
write my
this coaster
dot CSV. Okay. I'm not writing any path
because my this uh data set is here
itself. Coaster Coaster CO this is
coaster.csv CSV. Okay. If you have your
data set in another location as in like
in C drive, D drive or whatever. Okay.
You can give that path.
Okay. Let me run it. Yeah. So now we
will do some data understanding. Okay.
We'll see data frame shape head and tail
data types and describe like small small
function. We will use DF dot shape.
Right? So we have 1087 rows and 56
columns in our data set. Okay. Then we
will see df.hat
five.
So df do.head means it will give me top
five rows of my data set. Okay. You can
see 1 2 3 4 5. Okay. Five rows and 56
columns. Here 56 column but we have 1087
rows. Okay. So we have coaster name,
length, speed, location, status, opening
date, type, this, this, this, this.
Okay, we will do
uh you know we'll make some graphs using
this these columns. Okay, and there is
one more df.tail
again last five rows. So df.tail gives
you the last last five rows. See 1086
1085. Okay. And if you want to see
full data, it is here. Okay. 0 to 108.
Fine. Yeah. So we have one more df doc
columns to check
all the columns.
See coaster name, land, the speed,
location, status, opening date, type and
these all are my columns names. Fine. So
we have one more data types. Actually we
have like many data types but let me
show you some important ones or you can
say some the basics on one. Okay. So DF
dot D types. Dypes means data types.
Coaster name is object type length
object and inversion is float. Okay.
We'll check inversion.
Where is inversion? Yeah float type.
Okay,
then everything is object and eer
introduces int numeric one and latitude
float. Okay, so basically we have three
types float, int and object.
Okay, then what we will do? Let's just
quickly check this describe.
Okay. So what is the count of this
inversion 932?
Okay. It won't include any you know what
is it empty cells. Okay. It won't count
empty cells. Right. Mean of this
particular table then standard deviation
then mean minimum value 25% 70% 75% and
max. It will describe you all this.
Okay. And there is one more df.info info
to get the info. See coaster name 1087
value non null object then length 953
null. Okay. Like this.
Yeah. So now moving forward what we will
do? We will do data preparation. Okay.
What comes in data preparation like
dropping irrelevant columns and rows
which we don't want. Okay. Then second
thing is identifying duplicates columns.
Then third is renaming columns.
Then we'll do some feature creations,
right?
So if you want to drop a column, okay,
how you can drop?
So you have to write just df dot drop.
First I will give here stag data
repage.
Okay. Yeah. So how you can drop a
column? Okay. DF.
Then here I will write
opening
date=
to 1.
Okay.
Opening. Okay. X is wrong.
Instead of this
maybe what is the opening date? Okay. O
is capital here.
Instead of this you can use double
equals to
okay axis is not defined.
My bad. So what you have to do? You have
to give this here and yeah.
Okay. Next is not defined.
Yeah. Fine. Okay. So as you can see here
first I will show you this question name
length speed location status and opening
date is there fine.
So if you will go here question name
length speed location status nothing
opening date is there right so this is
how you can drop a table fine for a
while I'm making this as a comment maybe
in future
upcoming you know making graph I'll
leave this okay so yeah so I will write
here df equals to
df
poster name,
comma, location
then comma
status.
Then here I will write manufacturer.
Fine. Then again comma.
Then I will give here.
Okay. Year
introduced
then
comma
latitude.
Wait I will tell you why I'm doing this.
Then longitude
latitude longitude
then
type main.
Okay. Then
opening
date
clean.
Fine.
Then speed into
m/ hour r
speed
and comma what else then
height
foot
inversion
clean
geforce
clean dot copy. Yeah.
So here I will write okay
type in
okay one more mistake is here
fine.
So now what I will do
I will write here opening
date
clean equals to PD2
dot date time
DF
opening
date
Clean fine.
What happened?
So what does this pd do to date time do?
Okay. So it converts argument to data
time date time. Okay. So this function
you know you can say converts a scalar
array like or series or data frame
dictionary like to a pandas datetime
object. Okay.
So now let's rename the columns.
Okay.
Then to you know for better things DF
equals to DF dot rename
columns equals to
poster
name. Then
poster
name,
year
introduced then
here
introduced.
Okay.
Opening date clean to
open
date.
Okay. Then I will write speed
I then write caps speed.
Got it? Then
height
into foot
I will write it as
catch 50.
So why I'm writing this because this is
a very good practice as a professional
way or as a data analyst or as a
business analyst whatever you are
working on machine learning projects or
whatever this is a good practice.
Okay. So inverions
in versions
clean
then should be like
invers
fine.
Let me run it.
Okay. Now let's check the column.
Yeah. So now you can see our column name
is changed.
Fine.
So I will write this na
dot sum.
So what does this is na dot sum do? So
is na function returns a boolean value
of you know true if the value is n and
the false otherwise and the sum function
returns the sum of the true values which
equals to the number of n values in the
column. So here we have zero and n here
same the status 213
here 217 5
height foot is 916 okay so now what we
will do we will write here df dot
location
df dot duplicate
I told you in starting
we'll
duplicate
Okay.
So, question name, location, status,
man. Okay. It's duplicated now. Okay.
So, now what we will do? I will check
duplicates for the coaster name. So, you
can write df. Loc
then df
df dot
duplicated and subset equals to
poster
name then dot head. Okay. Head of five.
Now everyone know right what do
what this head does.
Okay. Yeah. So here you can see so these
are duplicate okay of the question name
right.
Why? because you can just check the
typing
and all. Fine. So now checking with an
example duplicate. Let's check
dot query
question name
crystal.
Now you will get each equals cyclone.
Okay. I took this crystal B cipher.
Fine.
Then run it. So now you can see 39 and
43 are the same.
Okay. Everything is same. So this one is
duplicate. So now what I will do? I will
write here. Just let me give some space.
Yeah. DF dot columns.
Then I will write here DF
dot location.
Then TF dot
duplicated
subset
question
name
then
location
then
opening
date.
Okay.
dot
d set
index
then drop that.
Okay.
So now what we will do we will do some
feature understanding.
Okay.
So now we will do some feature
understanding.
Okay.
So in this we will plot some feature
distribution like histogram KD box plot.
Okay. For that I will write here DF
year introduced.
Okay. Then I will write value
count.
So now what I will do? Let's create bar
chart. Okay.
Ax goals to der
introduced dot
value
counts. Okay. Then dot
add 10 max. I need
then dot plot kind equals to I need bar
then title is top
I write what what what should we give
top 10
years
coasters introduced.
Okay, then
I will write here ex dot set
X label. Then I will write here
introduced.
Okay. Then I will write here ex dot
set Y label.
than
count. Okay, let me run it. Okay,
unexpected character after line
continuation.
Okay,
we can remove this
unexpected intended.
Now let me run it. Yeah, so this is our
bar plot. Okay. So this is how you can
create bar plot using your data. Okay.
So a bar plot is you know one of the
most common types of graphics in this
data visualization or EDA. It shows the
relationship between a numerical and the
categorical variable. Okay. So each
entity of this categoric variable is
represented as a bar.
Got it? So now let's do some more. So
what uh let's make a stogram class.
Okay,
histogram.
So for that I will write
a ex = to df
speed
r then plot
kind equals to
Then comma I will write bins here. Okay.
Bins equals to 20
then
title
coaster
speed.
Okay. meter per hour then AX
dot set
X level
speed
okay forgot to give this
okay yeah so histogram displays
numerical data by grouping data into
bins of equal width so here each bin is
plotted as a bar whose height correspond
to how How many data points are in that
bin? So bins you can say are also
sometimes called the intervals or
classes or buckets. Basically
it's the same.
So here for this this much is the bin.
For this this much is a bin. This is
like that. Okay.
So now let's create KD plot. So for that
it's very simple. Ax = to DF.
Then I can write as speed in me of R
then plot
kinda
then title
I can write the coaster
speed
then ex dot set
X
speed
can.
Yeah. So, KD, what does KD means? A
kernel density estimate. This plot is a
method of, you know, visualization the
distribution of observation in a data
set analog to a stola. So, KD represent
the data using a continuous probability
density curve in one or more dimension.
Okay. In one or more dimension that
curve fine. So now let's do some feature
relationship. So in this we will make
scatter plot, heat map correlation and
pair plot. Okay. Or we can also do some
group by comparison. Fine. So now let's
first make
scatter plot. Okay. So, df dot plot
kind equals to
scatter
comma
x = to
speed
m/ hour
comma y = to
height in foot.
Fine.
Then comma let's give the title
equals to coaster speed
versus
height.
Fine then pl do
okay some error is there
speed meter per hour. Okay. Okay. S
should be capital and H should be
capital.
Yeah.
So this is scatter plot speed versus
height. This is a speed versus height.
Okay. If the I can see here height is
directly proportional to speed somewhat
because if you can see the 350 is the
speed less than
120 km/h. No it's not like that. Okay
fine.
So a scatter plot identifies the
possible relationship between change
observed in two different sets of
variable. Here the variables are height
and the speed. Okay. It provides a
visual and aesthetical uh you know means
to test the strength of relationship
between two variable. Fine. Now let's
okay now let's make one more to give you
better idea. Okay with legend.
Now I will wait SNS dot scatter
plot then x = to
speed
r
y = to
height into foot
whatever you can say then hue will be
there the year
introduced.
Okay. Then data equals to DF.
Then ex= to set title.
Set title. Then
coaster
speed versus
height.
then plt dot show.
Yeah.
So now you can see
so this color you know the light color I
will let me zoom it. Okay. So this color
we have some here which are introduced
in '90s. Okay. And this color which are
introduced in 1925 and these color which
are introduced in 2000. Okay. So you can
see like this as well.
So now let's make pair plot. Okay, it's
look amazing.
So let me write SNS dot
dot pair plot then TF comma
variable is equals to I need
introduce
then speed
meter per R
comma S should be capital height
foot. Just remember the spelling, okay?
It's case sensitive.
Inversion
inversions, comma,
inversions, comma, geforce.
Okay. Comma. U I will put what? What?
What? What? Okay. I will put type
main.
Fine then pl do
okay let me run it okay some error is
there
year introduced
y capital
let's say here so that's why I'm saying
just remember the proper spelling again
some
error in versions is version.
Now let me run it.
Okay. Again some
what inversions?
Okay. Let me check here while renaming
inversions only.
Okay. Fine.
in versions.
Let me paste in but nothing changes. Let
me check again.
Okay, some error is there.
Wait. So yeah, you can see it's running
fine.
So after this what I will do?
So what is the pair plot? So pair plot
function allows the user to get you know
then an axis grid via which each
numerical variable is stored in the data
is shared across the x and the y okay in
the structured column.
So this is how see you know type mean
wood other and steel this red one is
wood other or the blue one and the
purple are the steel one okay the
different different format okay now
let's create the last graph which is
heat map okay so let me write df
correlation equals to df
here introduced,
comma,
speed m/a
height
into foot
canions,
comma
G force
then
drop then coalition okay DFO
introduced I
capital yeah
so So this is correlation values of all
the things right. So now let's write SNS
dot
heat map
heat map TF C
not will be true.
Okay.
Now let me run it. Yeah. So to create
heat map in Python so you can use this.
Okay. C bond library for this heat map.
So this function takes a data frame you
know as a input and generates a heat map
type of things as the output. Okay. So
this is how you can perform EDA using
any data set or you can show your data
or insights with a beautiful
representation using graphs and all like
this. Fine. Web scraping is a powerful
technique that allows you to
automatically extract data from website.
Turning the vast amount of information
available online into something you can
easily analyze and use. Whether you are
gathering data for research, building a
data set for machine learning project,
or just curious about how websites work
behind the scenes. Web scraping is an
essential skills to have in your
toolkit. On the other hand, Python is
one of the most popular programming
languages for web scraping thanks to its
simplicity and the wealth of libraries
available. In this video, we will
explore how to use Python to scrape data
from website. And we will dive into
practical examples using Python
libraries like request and beautiful
soup to fetch and parse web content. But
it's not just about the code. Web
scripping comes with its own set of
challenges and ethical constitution. We
will talk about how to scrape
responsibly, respecting the rules set by
websites and ensuring that your scraping
activity don't negatively impact the
sites you are collecting data from. So
by the end of this video, you will have
a solid understanding of how to start
scraping data from the web using Python.
Whether you are new to programming or
looking to add web scraping to your
skill set, this video will give you the
knowledge and tools you need to get
started. So let's jump in and see how
Python can help you unlock the full
potential of the web. So without any
further ado, let's get started. So here
I am using this Google Collab for the
web scraping. Okay. You can use your own
like Jupyter notebook, Visual Code
Studio, any thing. Okay. So here I'll
write scraping
using Python.
Okay. Then here first you have to
install some libraries like
you know uh request
and you have to install beautiful. So
you have to install that you know pandas
because we will create one data frame
and we will save it then we will check
our data. Okay. And you can install some
basic Python library like numpy and all
that. Okay. So here first I will import
request.
Okay. Then I will write from PS4
import
beautiful
soap.
Okay.
So
then I will write import
pandas
as pay.
Fine. Then now what I will do? Okay
let's see what is this request and all
the so request is an HTTP client library
for the Python programming language. So
request is one of the most you know
downloaded Python libraries. Okay. It's
like more like over 2,000 not exactly
2,000 sorry 200 or 300 million monthly
download. Okay. So what it does it maps
the HTTP protocol onto Python subject
oriented semantics. And here beautiful
soap. So
SOAP is a Python package for you know
parsing the HTML and XML documents
including those with you know malphone
markup. It creates a parse tree for
documents that can be you know used to
extract data HTML from HTML which is
useful for web scraping. Let me run it
here. Our second step will be
define the URL and the headers. Okay,
URL and the headers.
So URL okay from where you want to
extract your data. So I will write here
simply learn
then this okay let's open this PMP
certification
it yeah so here I will write this
and headers
equals to
so these are the headers okay so like in
which you know device or you're working
on which browser you are working on. So
these are for the headers. So now what
what I will do I will send a get request
to the simple page. Okay to this URL for
that you know right let me write this
sending
get request.
Okay, here write response
equals to
request
dot get
then URL
comma headers
equals to headers.
Okay, let me run it. Okay, working fine.
So now uh let's check if the request is
successful or not. For that if
response
dot status code plus equals to = 200
then
so equals to
false.
Okay. Then response
dot content
do
HTML
dot parser. Okay.
Here.
So here I will write course
titles. Why? because here I'm you know
initializing the list to store our
titles or whatever the things are okay
basically the data okay why I'm writing
course here because this is again course
page that's why nothing else okay here I
will give the empty list
so now we will find all course titles
based on the actual HTML structure what
is HTML structure
you have to go to the page right click
Click then inspect.
Okay. According to this HTML structure
means what is the class name? What is
the you know this PMP certification is
which heading? H1 heading, H2 heading,
which heading it is. Okay. So let's see
which heading it is. Okay. Okay. This
PMP certification training H1 heading.
Okay. Just remember H1 heading.
So here I will write for
quotes in soap
dot find
or
h1.
Okay.
Then title
custom course text
dot strip.
Okay.
Then here I'm write course
titles
dot append
title. Okay.
Now what I will do? I will check
if any data was extracted before it not.
Okay, I will write if course
titles.
Here I will create a data frame. DF
equals to PD dot data frame.
In this I will write course title. it
will be our you know uh that column
name. So for I will write it course
data. Okay.
Then
course
titles.
Okay.
Then here I will write df
dot to
csv.
Then uh just create simply learn dot
CSV. Okay. So our data will be saved in
this simply learn dot csv. Okay. CSV.
Fine. Here what I will do I will write
index equals to false.
data.
Print data is saved in CSV.
Simply learn dot CSV.
Fine.
So here I will write
else
web page.
Okay.
write field
to
retrieve the web page. Okay, that's it.
Let me run it.
Okay, here if response storage to this
group
spawn object has no attribute
status code. Okay.
Okay. Title.
It is text.
Some minor spelling mistakes are there.
Okay. Dear.
Sorry. Sorry. My bad
again.
Okay. Sorry.
D will be capital, F will be capital.
Yeah. So now you can see data is saved
in simply dot CSV. So what I will do? I
will write data equals to everyone know
how to read CSV in Python. Read
CSV.
What was our file name? Simply learn
CSV.
Okay. Let me copy paste. Run it. Link
fine data.
Okay.
PMP certification training data is fed.
Okay. Here you can see H1 we gave and
that's why PMP certification training
came. Let's take what's in H2.
Okay. In H2 there is leading premier PMI
partner something is there. Okay. I will
change here
H1 to H2 then I will run it. Okay
fine.
See leading premier P P P P P P P P P P
P P P P P P P P P P PMI. Okay, let me
make it bigger. Yeah. So now you can see
course data we mentioned. So
if you will see this H1 didn't come why
because we have mentioned H2 that's why.
So these all are the H2.
Okay. So what is this? I don't know.
Okay. This is a time series graph. This
is Google collaping. Okay. So this is
how you can literally retrieve your data
from any website. Okay. Just remember
that some websites don't give you access
to their you know for their scrapping
just like Amazon don't give so you have
to use at that time API and this is not
ethical also to use someone's data okay
without asking or whatever you can
>> welcome to the neural network tutorial
my name is Richard Kersner I'm with the
simply learn team what's in it for you
well today we're going to cover what is
a neural network what can neural neural
networks do, how does a neural network
work, types of neural networks, and then
we're going to jump into a use case to
classify between the photos of dogs and
cats, and we'll do that on the KAS with
the TensorFlow in the back, but it's a
Python script. So, that's always my
favorite part is when we dive into the
actual script. So, what is a neural
network? So, hi guys. I heard you want
to know what a neural network is. Here
we have uh looks like he just went
shopping at a red tag sale. My robots's
back. So as a matter of fact, you have
been using neural network on a daily
basis. In today's world, it's just
amazing how much we use our new
technology. We're not even aware of it.
When you ask your mobile assistant to
perform a search for you, you know, like
saying you're Google or Siri or whoever
you use, Amazon Web, self-driving cars.
So that's the newest thing coming out.
They're just now trying to make those
legal in different states in the US and
around the world. Even in the UK, they
now have self-driving cars going up and
down the street. It's pretty amazing.
These are all neural network driven.
Computer games use it. A lot of computer
games are driven by neural networks in
the back end as part of the game system
and how it adjusts to the players. And
it's also used in processing the map
images on your phone. So every time you
do a navigation someplace and it opens
it up, they now use neural networks to
help you find the quickest way to get
there. Neural network. A neural network
is a system or hardware that is designed
to operate like a human brain. In
today's development, this is so
important to understand because we don't
have anything else to compare it to. I'm
sure someday in the future, the computer
will redefine or the neural network or
the AI artificial intelligence will
redefine what these mean. But as far as
we can today's world, in today's
commercial development, we have to
compare it to what humans do. So, it's
we want to compare and how it operates
to a human brain and how it solves
problems like a human does. What can a
neural network do? And really, we're
just going to dive in deeper to we just
covered and look at other examples. So,
what can a neural network do? Well,
let's list out the things neural
networks can do for you. Translate text.
Boy, we got Google Translate and
Microsoft has their own translate. They
have some really cool. They actually
have an earpiece. It's supposed to start
translating as you talk. What a cool
technology. What a cool time to live.
Identify faces. Can you imagine all the
uses for facial identification? In the
case of our uh sample or our code that
we're going to look at later, we'll be
identifying dogs and cats. So, not quite
as detailed as uh understanding whose
face belongs to who. I I'm waiting for
the Google glasses to come out so I can
see who's who and the identify faces as
I'm walking around. Have a little name
tag over them. Not out there yet, but
boy, we are close. We can identify the
faces and they have all kinds of
technologies to bring that information
back to us. Recognize speech goes along
with the translate text. So now as
you're talking into your assistant, it
can use that to do commands, turn lights
on, all kinds of things you can do with
recognizing speech. Read handwritten
text. They're starting to translate all
these old text documents that they've
had in storage instead of doing it
individually where somebody's going
through each text by themsel in a room.
Picture like an old Raiders of the Lost
Arc theme where he's in the back, you
know, archaeologist studying the text.
Now it's fed into a computer. They take
a picture. They even use neural networks
to take a scroll that is so messed up
that they can't undo the scroll and they
x-ray it and then they use that x-ray to
translate the text off of it without
ever opening the scroll. I mean just way
cool stuff they're starting to do with
all this. And of course control robots.
What would be a neural network without
bringing in the robots? And we have our
own favorite robot in the middle who
goes to our red tag cell and goes
shopping for us. So, you know, these are
just a few of the wonderful things that
neural networks are being applied to.
It's such an infant stage technology.
What a wonderful time to jump in. And
there are a lot of other things it goes
into. I mean, we could spend just
forever talking about all the different
applications from business to whatever
you can even imagine. They're now
applying neural networks to help us
understand. So, now we talked a little
bit about all the cool things you can do
with a neural network. Let's dive in and
say, how does a neural network work? So
now we've come far enough to understand
how neural network works. Let's go ahead
and walk through this in a nice
graphical representation. They usually
describe a neural network as having
different layers. And you'll see that
we've identified a green layer, an
orange layer, and a red layer. The green
layer is the input. So you have your
data coming in. It picks up the input
signals and passes them to the next
layer. The next layer does all kinds of
calculations and feature extraction.
It's called the hidden layer. A lot of
times there's more than one hidden
layer. We're only showing one in this uh
picture, but we'll show you how it looks
like in a more detail in a little bit.
And then finally, we have an output
layer. This layer delivers the final
result. So the only two things we see is
the input layer and the output layer.
Now let's make use of this neural
network and see how it works. Wonder how
traffic cameras identify vehicles
registration plate on the road to detect
speeding vehicles and those breaking the
law? They got me going through a red
light the other day. Well, last month.
That's like the horrible thing. They
send you this picture of you and all
your information because they pulled it
up off of your license plate and your
picture. I shouldn't have gone through
the red light. So, here we are and we
have an image of a car and you can see
the license plate on there. So, let's
consider the image of this vehicle and
find out what's on the number plate. The
picture itself is 28x 28 pixels and the
image is fed as an input to identify the
registration plate. Each neuron has a
number called activation that represents
the grayscale value of the corresponding
pixel range. And we range it from zero
to one. One for a white pixel and zero
for a black pixel. And you can see down
here we have an example where one of the
pixels is registered as like 082.
Meaning it's probably pretty dark. Each
neuron is lit up when its activation is
close to one. So as we get closer to
black on white, we can really start
seeing the details in there. And you can
see again the pixel shows us one up
there. It's like part of the car and so
it lights up. So pixels in the form of
arrays are fed to the input layer. And
so we see here the pixels of a car image
fed as an input. And you're going to see
that the input layer which is green is
one dimension while our image is
two-dimension. Now when we look at our
setup that we're programming in Python,
it has a cool feature that automatically
does the work for us. If you're working
with an older neural network pattern
package, you then convert each one of
those rows so it's all one array. So
you'd have like row one and then just
tack row two onto the end. You can
almost feed the image directly into some
of these neural networks. The key is
though is that if you're using a 28x 28
and you get a picture of this 30x30,
shrink the 30x30 down to fit the 28x 28.
So you can't increase the number of
input in this case green dots. It's very
important to remember when you work on
neural networks. And let's name the
inputs x1 x2 x3 respectively. So each
one of those represents one of the
pixels coming in. And the input layer
passes it to the hidden layer. And you
can see here we now have two hidden
layers in this image in the orange. And
each one of those pixels connects to
each one of those hidden layers. And the
interconnections are assigned weights at
random. So they get these random weights
that come through. If x1 lights up, then
it's going to be x1 times this weight
going into the hidden layer. And we sum
those weights. The weights are
multiplied with the input signal and a
bias is added to all of them. So as you
can see here we have X1 comes in and it
actually goes to all the different
hidden layer nodes or in this case uh
whatever you want to call them network
setup the orange dots and so you take
the value of X1 you multiply it by the
weight for the next hidden layer. So X1
goes to hidden layer 1 X1 goes to hidden
layer two X1 goes hidden layer 1 node
two hidden layer one node three and so
on. And the bias a lot of times they
just put the bias in as like another
green dot or another orange dot and they
give the bias a value one and then all
the weights go in from the bias into the
next node. So the bias can change. We
always just remember that you need to
have that bias in there. There's things
that can be done with it. Generally most
of packages out there control that for
you so you don't have to worry about
figuring out what the bias is. But if
you ever dive deep into neural networks,
you got to remember there's a bias or
the answer won't come out correctly. The
weighted sum of the input is fed as an
input to the activation function to
decide which nodes to fire. And for
feature extraction, as a signal flows
within the hidden layers, the weighted
sum of inputs is calculated and is fed
to the activation function in each layer
to decide which nodes to fire. So here's
our feature extraction of the number
plate. And you can see these are still
hidden nodes in the middle. And this
becomes important. We're going to take a
little detour here and look at the
activation function. So, we're going to
dive just a little bit into the math so
you can start to understand where some
of the games go on when you're playing
with neural networks in your
programming. So, let's look at the
different activation functions before we
move ahead. Here's our friendly red tag
shopping robot. And so, one is a sigmoid
function. And the sigmoid function which
is 1 over 1 + e to the minus x takes the
x value and you can see where it
generates almost a zero and almost a one
with a very small area in the middle
where it crosses over and we can use
that value to feed into another
function. So if it's really uncertain it
might have a 0.1 or 2 or 3 but for the
most part it's going to be really close
to one and really close to this case
zero zero to one the threshold function.
So if you don't want to worry about the
uncertainty in the middle, you just say,
"Oh, if x is greater than or equal to
zero, if not, then uh x is zero." So
it's either zero or one. Really
straightforward. There's no in between
in the middle. And then you have the
what they call the reel relu function.
And you can see here where it puts out
the value, but then it says, well, if
it's over one, it's going to be one. And
if it's uh less than zero, it's zero. So
it kind of just deadends it on those two
ends, but allows all the values in the
middle. And again, this like the sigmoid
function allows that information to go
to the next level. So it might be
important to know if it's a 0.1 or a
minus.1. The next hidden layer might
pick that up and say, "Oh, this piece of
information is uncertain or this value
has a very low certainty to it." And
then the hyperbolic tangent function.
And you can see here it's a 1 - e to the
-2x over 1 + e - 2x. And it's very much
along the same theme, a little bit
different in here in that it goes
between minus one and one. So you'll see
some of these it goes 0ero to one, but
this one goes minus one to one. And if
it's less than zero, it's, you know, it
doesn't fire and if it's over zero, it
fires. And it also still puts out a
value. So you still have a value you can
get off of that just like you can with
the sigmoid function and the relu
function. Very similar in use. And I
believe the originally used to be
everything was done in the sigmoid
function. That was the most uh commonly
used. And now they just kind of use more
the reloo function. The reason is one,
it processes faster because you already
have the value and you don't have to add
another compute the 1 / 1 + e to the
minus x for each hidden node and the
data coming off works pretty good as far
as putting it into the next level. If
you want to know just how close it is to
zero, how close is it not to
functioning, you know, is it minus.1
minus.2 usually they're float values.
You get like minus point minus.00138
or something. So, you know, important
information, but the Reu is most
commonly used these days as far as the
setup we're using. But you'll also see
the sigmoid function very commonly used
also. Now that you know what an
activation function is, let's get back
to the neural network. So, finally, the
model would predict the outcome of
applying a suitable activation function
to the output layer. So, we go in here,
we look at this, and we have the optical
character recognition OCR is used on the
images to convert it into a text in
order to identify what's written on the
plate. And as it comes out, you'll see
the red node. And the red node might
actually represent just the letter A. So
there's usually a lot of outputs when
you're doing text identification. We're
not going to show that on here, but you
might have it even in the order. It
might be what order the license plates
in. So you might have ABCDE E FG, you
know, all the alphabet plus the numbers.
And you might have the 1 2 3 4 5 6 7 8 9
10 places. So it's a very large array
that comes out. It's not a small amount
of uh, you know, we show three dots
coming in, eight hidden layer nodes, you
know, two sets of four. We just show one
red coming out. A lot of times this is
uh, you know, 28 * 28. If you did 30 *
30, that's, you know, 900 nodes. So 28
is a little bit less than that uh, just
on the input. And so you can imagine the
hidden layer is just as big. Each hidden
layer is just as big if not bigger. Then
the output is going to be there's so
many digits. You know, it's a lot.
There's it's a huge amount of input and
output. But we're only showing you just,
you know, it' be hard to show in one
picture. And so it comes up and this is
what it finally gets out in the output
as it identifies a number on the plate.
And in this case, we have 08-d3858.
Error in the output is back propagated
through the network and weights are
adjusted to minimize the error rate.
This is calculated by a cost function.
When we're training our data, this is
what's used and we'll look at that in
the code when we do the data training.
So, we have stuff we know the answer to
and then we put the information through
and it says yes, that was correct or no,
cuz remember we randomly set all the
weights to begin with. And if it's
wrong, we take that error. How far off
are you? You know, are you off by is it
if it was like minus one, you're just a
little bit off. If it's like minus 300
was your output, remember when we're
looking at those different options, you
know, hyperbolic or whatever, and we're
looking at the could doesn't have an
limit on top or bottom. it actually just
generates a number. So if it's way off,
you have to adjust those weights a lot.
But if it's pretty close, you might
adjust the weights just a little bit.
And you keep adjusting the weights until
they fit all the different training
models you put in. So you might have 500
training models and those weights will
adjust using the back propagation. It
sends the error backward. The output is
compared with the original result and
multiple iterations are done to get the
maximum accuracy. So, not only does it
look at each one, but it goes through it
and just keeps cycling through these the
data making small changes in the network
until it gets the right answers. With
every iteration, the weights at every
interconnection are adjusted based on
the error. We're not going to dive into
that math because it is a differential
equation and it gets a little
complicated, but I will talk a little
bit about some of the different options
they have when we look at the code. So,
we've explored a neural network. Let's
look at the different types of
artificial neural networks. And this is
like the biggest area growing is how
these all come together. Let's see the
different types of neural network. And
again, we're comparing this to human
learning. So here's a human brain. I
feel sorry for that poor guy. So we have
a feed for forward neural network.
Simplest form of a they call it a ann a
neural network. Data travels only in one
direction input to output. This is what
we just looked at. So as the data comes
in, all the weights are added, it goes
to the hidden layer, all the weights are
added, it goes to the next hidden layer,
all the weights are added, and it goes
to the output. The only time you use the
reverse propagation is to train it. So
when you actually use it, it's very
fast. When you're training it, it takes
a while because it has to iterate
through all your training data. And you
start getting into big data because you
can train these with a huge amount of
data. The more data you put in, the
better trained they get. The
applications vision and speech
recognition actually they're pretty much
everything we talked about a lot of
almost all of them use this form of
neural network at some level radio basis
function neural network this model
classifies a data point based on its
distance from a center point. What that
means is that you might not have
training data. So you want to group
things together and you create central
points and it looks for all the things
you know some of these things are just
like the other. If you've ever watched
the Sesame Street as a kid, that dates
me. So, it brings things together and
this is a great way if you don't have
the right training model, you can start
finding things that are connected you
might not have noticed before.
Applications power restoration systems.
They try to figure out what's connected
and then based on that they can fix the
problem if you have a huge power system.
and self-organizing neural network
vectors of random dimensions are input
to discrete map comprised of neurons. So
they basically find a way to draw they
call them they say dimensions or vectors
or planes because they actually chop the
data in one dimension, two dimension,
three dimension, four, five, six. They
keep adding dimensions and finding ways
to separate the data and connect
different data pieces together.
Applications used to recognize patterns
in data like in medical analysis. The
hidden layer saves its output to be used
for future prediction. Recurrent neural
networks. So the hidden layers remember
its output from last time and that
becomes part of its new input. Uh you
might use that especially in robotics or
flying a drone. You want to know what
your last change was and how fast it was
going to help predict what your next
change you need to make is to get to
where the drone wants to go.
Applications text to speech conversation
model. So, you know, I talked about
drones, but you know, just identifying
on Lexus or Google Assistant or any of
these, they're starting to add in I'd
like to play a song on my Pandora, and
I'd like it to be at volume 90%. So, you
now can add different things in there,
and it connects them together. The input
features are taken in batches like a
filter. This allows a network to
remember an image in parts. Convolution
neural network. today's world in photo
identification and taking apart photos
and trying to you know have you ever
seen that on Google where you have five
people together this is the kind of
thing separates all those people so then
it can do a face recognition on each
person applications used in signal and
image processing in this case I use
facial images or Google picture images
as one of the options modular neural
network it has a collection of different
neural networks working together to get
the output so wow we just went through
all these different types of neural
networks. And the final one is to put
multiple neural networks together. I
mentioned that a little bit when we
separated people in a larger photo and
individuals in the photo and then do the
facial recognition on each person. So
one network is used to separate them and
the next network is then used to figure
out who they are and do the facial
recognition. Applications still
undergoing research. This is a cutting
edge. you hear the term pipeline and
there's actual in Python code and in
almost all the different neural network
setups out there they now have a
pipeline feature usually and it just
means you take the data from one neural
network and maybe another neural network
or you put it into the next neural
network and then you take three or four
other neural networks and feed them into
another one. So how we connect the
neural networks is really just cutting
edge and it's so experimental. I mean
it's almost creative in its nature.
There's not really a science to it
because each specific domain has
different things it's looking at. So if
you're in the banking domain, it's going
to be different than the medical domain
than the automatic car domain. And
suddenly figuring out how those all fit
together is just a lot of fun and really
cool. So we have our types of artificial
neural network. We have our feed forward
neural network. We have a radial basis
function neural network. We have our
Cohen self-organizing neural network,
recurrent neural network, convolution
neural network, and modular neural
network where it brings them all
together. And u no the colors on the
brain do not match what your brain
actually does, but they do bring it out
that most of these were developed by
understanding how humans learn. And as
we understand more and more of how
humans learn, we can build something in
the computer industry to mimic that, to
reflect that. And that's how these were
developed. So exciting part, use case
problem statement. So this is where we
jump in. This is my favorite part. Let's
use the system to identify between a cat
and a dog. If you remember correctly, I
said we're going to do some Python code.
And you can see over here, my hair is
kind of sticking up over the computer,
cup of coffee on one side, and a little
bit of old school. A pencil and a pen on
the other side. Yeah, most people now
take notes. I love the stickies on the
computer. That's great. That's that is
my computer. I have sticky notes on my
computer in different colors. So, not
too far from uh today's programmer. So,
the problem is is we want to classify
photos of cats and dogs using a neural
network. And you can see over here we
have quite a variety of dogs in the
pictures and cats and you know just
sorting out it is a cat is pretty
amazing. And why would anybody want to
even know the difference between a cat
and a dog? Okay, you know why? Well, I
have a cat door. It'd be kind of fun
that instead of it identifying, instead
of having like a little collar with a
magnet on it, which is what my cat has,
the door would be able to see, oh,
that's the cat. That's our cat coming
in. Oh, that's the dog. We have a dog,
too. That's a dog I want to let in.
Maybe I don't want to let this other
animal in cuz it's a raccoon. So, you
can see where you could take this one
step further and actually apply this.
You could actually start a little
startup company idea, self-identifying
door. So, this use case will be
implemented on Python. I am actually in
Python 3.6. It's always nice to tell
people the version of Python because
that does affect sometimes which modules
you load and everything. And we're going
to start by importing the required
packages. I told you we're going to do
this in Kass. So we're going to import
from KAS models sequential from the Kass
layers conversion 2D or COV2D max
pooling 2D flatten and dense. And we'll
talk about what each one of these do in
just a second. But before we do that,
let's talk a little bit about the
environment we're going to work in. And
uh you know, in fact, let me go ahead
and open a uh the website, KASS's
website, so we can learn a little bit
more about KASS. So here we are on the
Kurass website, and it's uh ke.io.
That's the official website for Kurass.
And the first thing you'll notice is
that Kurass runs on top of either
TensorFlow, CNTK, and I think it's
pronounced Thano or Theo. What's
important on here is that TensorFlow and
the same is true for all these, but
TensorFlow is probably one of the most
widely used currently packages out there
with the KAS. And of course, you know,
tomorrow this is all going to change.
It's all going to disappear and they'll
have something new out there. So, make
sure when you're learning this code that
you understand what's going on and also
know the code. I mean, look, when you
look at the code, it's not as
complicated once you understand what's
going on. The code itself is pretty
straightforward. And the reason we like
KAS and the reason that people are
jumping on it right now, it's such a big
deal is if we come down here, let me
just scroll down a little bit. They talk
about user friendliness, modularity,
easy extensibility, work with Python.
Python's a big one because a lot of
people in data science now use Python,
although you can actually access Kass
other ways. Is if we continue down here
is layers. And this is where it gets
really cool. When we're working with
KASS, you just add layers on. Remember
those hidden layers we were talking
about? And we talked about the reelu
activation. You can see right here. Let
me just up that a little bit in size.
There we go. That's big. I can add in an
eelu layer. And then I can add in a
softmax layer in the next instance. We
didn't talk about softmax. So you can do
each layer separate. Now if I'm working
in some of the other kits I use, I take
that and I have one setup and then I
feed the output into the next one. This
one I can just add hidden layer after
hidden layer with the different
information in it which makes it very
powerful and very fast to spin up and
try different setups and see how they
work with the data you're working on.
And we'll dig a little bit deeper in
here. And a lot of this is very much the
same. So when we get to that part, I'll
point that out to you also. Now just a
quick side note, I'm using Anaconda with
Python in it. And I went ahead and
created my own package and I called it
the Kass Python 36 because I'm in Python
36. Anaconda is cool that You can create
different environments really easily. If
you're doing a lot of different
experimenting with these different
packages, probably want to create your
own environment in there. And the first
thing, as you can see right here,
there's a lot of dependencies. A lot of
these you should recognize by now if
you've done any of these videos. If not,
kudos for you for jumping in today. PIP,
install, numpy, sci, the scikitlearn,
pillow, and h5py
are both needed for the tensorflow and
then putting the kass on there. And then
you'll see here uh and pip is just a
standard installer that you use with
Python. You'll see here that we did pip
install TensorFlow since we're going to
do KAS on top of TensorFlow. And then
pip install and I went ahead and used
the GitHub. So git plusgit and you'll
see here github.com. This is one of
their releases, one of the most current
release on there that goes on top of
TensorFlow. And you can look up these
instructions pretty much anywhere. This
is for doing it on Anaconda. Certainly
you'd want to install these if you're
doing it in Iuntu server setup. you
you'd want to get I don't think you need
the H5 py and aru but you do need the
rest in there because they are
dependencies in there and it's pretty
straightforward and that's actually in
some of the instructions they have on
their website so you don't have to
necessarily go through this just
remember their website on there and then
when I'm under my uh Anaconda navigator
which I like you'll see where I have
environments and on the bottom I created
a new environment and I called it KAS
Python 36 just to separate everything
you can say I have Python 3.5 and Python
36 I used to have a bunch of other ones,
but it kind of cleaned house recently.
And of course, once I go in here, I can
launch my Jupyter Notebook, making sure
I'm using the right environment that I
just set up. This, of course, opens up
my um in this case, I'm using uh Google
Chrome. And in here, I could go and just
create a new document in here. And this
is all in your um browser window when
you use the Anaconda. Do you have to use
Anaconda and Jupyter Notebook? No. You
can use any kind of Python editor,
whatever setup you're comfortable with
and whatever you're doing in there. So,
let's go ahead and go in here and paste
the code in. And we're importing a
number of different settings in here. We
have import sequential. That's under the
models because that's the model we're
going to use as far as our neural
network. And then we have layers and we
have conversion 2D, max pooling 2D,
flatten dense. And you can actually just
kind of guess at what these do. We're
talking we're working in a 2D
photograph. And if you remember
correctly, I talked about how the actual
input layer is a single array. It's not
in two dimensions. It's one dimension.
All these do is these are tools to help
flatten the image. So, it takes a
two-dimensional image and then it
creates its own proper setup. You don't
have to worry about any of that. You
don't have to do anything special with
the photograph. You let the carass do
it. And we're going to run this. And
you'll see right here they have some
stuff that is going to be depreciated
and changed because that's what it does.
Everything's being changed as we go. You
don't have to worry about that too much.
If you have warnings, if you run it a
second time, the warning will disappear.
And this has just imported these
packages for us to use. Jupiter's nice
about this that you can do each thing
step by step. And I'll go ahead and also
zoom in there. A little control plus.
That's one of the nice things about
being in a browser environment. So, here
we are back. Another sip of coffee. If
you're familiar with my other videos,
you notice I'm always sipping coffee. I
always have a in my case latte next to
me, an espresso. So the next step is to
go ahead and initialize. We're going to
call it the CNN or classifier neural
network. And the reason we call it a
classifier is because it's going to
classify it between two things. It's
going to be cat or dog. So when you're
doing classification, you're picking
specific objects. You're specific. It's
a true or false. Yes, no. It is
something or it's not. So first thing
we're going to create our classifier and
it's going to equal sequential. So their
sequential setup is the classifier.
That's the actual model we're using.
That's the neural network. So we call it
a classifier. And uh the next step is to
add in our convolution. And let me just
do a uh let me shrink that down in size
so you can see the whole line. And let's
talk a little bit about what's going on
here. I have my classifier and I add
something. What am I adding? Well, I'm
adding my first layer. This first layer
we're adding in is probably the one that
takes the most work to make sure you
have it set correct. And the reason I
say that is this is your actual input.
And we're going to jump here to the part
that says input shape equals 64x 64x3.
What does that mean? Well, that means
that our pictures coming in. And there's
these pictures. Remember we had like the
picture of the car was 128x 128 pixels.
Well, this one is 64x 64 pixels. And
each pixel has three values. That's
where these numbers come from. And it is
so important that this matches. I
mentioned a little bit that if you have
like a larger picture, you have to
reformat it to fit this shape. If it
comes in as something larger, there's no
input notes. There's no input neural
network there that will handle that
extra space. So, you have to reshape
your data to fit in here. Now, the first
layer is the most important because
after that, KAS knows what your shape is
coming in here and it knows what's
coming out and so that really sets the
stage. Most important thing is that
input shape matches your data coming in.
And you'll get a lot of errors if it
doesn't. You'll go through there and
picture number 55 doesn't match it
correctly. And guess what it does? It
usually gives you an error. And then the
activation, if you remember, we talked
about the different activations on here.
We're using the reelu model. Like I
said, that is the most commonly used now
because one, it's fast. Doesn't have the
added calculations in it. It just says
here's the value coming out based on the
weights and the value going in. And um
from there, you know, it's uh if it's
over one, then it's good or over zero,
it's good. If it's under zero, then it's
considered not active. And then we have
this conversion 2D. What the heck is
conversion 2D? I'm not going to go into
too much detail in this because this has
a couple of things it's doing in here, a
little bit more in-depth than we're
ready to cover in this tutorial. But
this is used to convert from the photo
cuz we have 64x 64x3 and we're just
converting it to two-dimensional kind of
setup. So it's very aware that this is a
photograph and that different pieces are
next to each other. And then we're going
to add in uh a second convolutional
layer. That's what the cov stands for
2D. So it's these are hidden layers. So
we have our input layer and our two
hidden layers and they are
two-dimensional because we're dealing
with a two-dimensional photograph. And
you'll see down here that on the last
one, we add a max pooling 2D and we put
a pool size equals 22. And so what this
is is that as you get to the end of
these layers, one of the things you
always want to think of is what they
call mapping and then reducing.
Wonderful terminology from the big data.
We're mapping this data through all
these layers. And now we want to reduce
it to only two sets. In this case, it's
already in two sets because it's a 2D
photograph. But we had, you know, two
dimensions by we actually have 64x 64
by3. So now we're just getting it down
to a 2x two. Just the two dimension
two-dimensional instead of having the
third dimension of colors. And we'll go
ahead and run these. We're not really
seeing anything in our run script
because we're just setting up. This is
all set up. And this is where you start
playing because maybe you'll add a
different layer in here to do something
else to see how it works and see what
your output is. That's what makes KAS so
nice is I can with just a couple flips
of code put in a whole new layer that
does a whole new processing and see
whether that improves my run or makes it
worse. And finally, we're going to do
the final setup, which is to flatten
classifier, add a flatten setup. And
then we're going to also add a layer, a
dense layer, and then we're going to add
in another dense layer. And then we're
going to build it. We're going to
compile this whole thing together. So,
let's flip over and see what that looks
like. And we've even numbered them for
you. So, we're going to do the
flattening. And flatten is exactly what
it sounds like. We've been working in a
two-dimensional array of picture, which
actually is in three dimensions because
of the pixels. The pixels have a whole
another dimension to it of three
different values. And we've kind of
resized those down to 2x two. But now
we're just going to flatten it. I don't
want to have multiple dimensions being
worked on by tensor and by kas. I want
just a single array. So, it's flattened
out. And then step four, full
connection. So we add in our final two
layers. And you could actually do all
kinds of things with this. You could
actually leave out this some of these
layers and play with them. You do need
to flatten it. That's very important.
Then we want to use the dents again.
We're taking this and we're taking
whatever came into it. So once we take
all those different the two dimensions
or three dimensions as they are and we
flatten it to one dimension. We want to
take that and we're going to pull it
into units of 128. They got that. You're
say where did they get 128 from? You
could actually play with that number and
get all kinds of weird results. But in
this case we took the 64 + 64 is 128.
You could probably even do this with 64
or 32. Usually you want to keep it in
the same multiple whatever the data
shape you're already using is in. And
we're using the activation the re lu
just like we did before. And then we
finally filter all that into a single
output. And it has how many units? One.
Why? Because we want to know whether
true or false. It's either a dog or a
cat. You could say one is dog, zero is
cat. Or maybe you're a cat lover and
it's one is cat and zero is dog. And if
you love both dogs and cats, you're
going to have to choose. And then we use
the sigmoid activation. If you remember
from before, we had the reel and there's
also the sigmoid. The sigmoid just makes
it clear it's yes or no. We don't want a
any kind of in between number coming
out. And we'll go ahead and run this.
And you'll see it's still all in setup.
And then finally, we want to go ahead
and compile. And let's put the compiling
our um classifier neural network. And
we're going to use the optimizer atom.
And I hinted at this just a little bit
before. Where does atom come in? Where
does an optimizer come in? Well, the
optimizer is the reverse propagation.
When we're training it, it goes all the
way through and says error and then how
does it readjust those weights. There
are a number of them. Atom is the most
commonly used and it works best on large
data. Most people stick with the atom
because when they're testing on smaller
data, see if their model is going to go
through and get all their errors out
before they run it on larger data sets.
They're going to run it on atom anyway,
so they just leave it on atom most
commonly used. But there are some other
ones out there. You should be aware of
that that you might try them if you're
stuck in a bind or you might blur that
in the future, but usually atom is just
fine on there. And then you have two
more settings. You have loss and
metrics. We're not going to dig too much
into loss or metrics. These are things
you really have to explore KAS because
there are so many choices. This is how
it computes the error. There's so many
different ways to on your back
propagation and your training. So we're
using the atom model, but you can
compute the error by um standard
deviation, standard deviation squared.
They use binary cross entropy. I'd have
to look that up to even know what that
is. There's so many of these. A lot of
times you just start with the ones that
look correct that are most commonly used
and then you have to go read the KAS
site and actually see what these
different losses and metrics and what
different options they have. So, we're
not going to get too much into them
other than to reference you over to the
KAS website to explore them deeper, but
we are going to go ahead and run them.
And now we've set up our classifier. So,
we have an object classifier. And if you
go back up here, you'll see that we've
added in step one. We added in our layer
for the input. We added a layer that
comes in there and uses the reelu for
activation. And then it pulls the data.
So this is even though these are two
layers, the actual neural network layer
is up here. And then it uses this to
pull the data into a 2x two. So into a
two-dimensional array from a
three-dimensional array with the colors.
Then we flatten it. So there's our adder
flatten. And then we add another dense
what they call dense layer. this dense
layer goes in there and it it downsizes
it to 128. It reduces it. So you can
look at this as uh we're mapping all
this data down the two-dimensional setup
and then we flatten it. So we map it to
a flatten map and then we take it and
reduce it down to 128 and we use the
reel again. And then finally we reduce
that down to just a single output and we
use a sigmoid to do that to figure out
whether it's yes, no, true, false, in
this case cat or dog. And then finally
once we put all these layers together we
compile them. That's what we've done
here and we've compiled them as far as
how it trains to use these settings for
the training back propagation. So if you
remember we talked about training our
setup and when we go into this you'll
see that we have two data sets. We have
one called the training set and the
testing set. And that's very standard in
any data processing is you need to have
that's pretty common in any data
processing is you need to have a certain
amount of data to train it and then you
got to know whether it works or not. Is
it any good and that's why you have a
separate set of data for testing it
where you already know the answer but
you don't want to use that as part of
the training set. So in here we jump
into part two fitting the classifier
neuron network to the images and then
from KAS let me just zoom in there. I
always love that about working with
Jupyter Notebooks. You can really see.
We're going to come in here. We do the
cross pre-processing an image. And we
import image data generator. It's so
nice of KAS. It's such a high-end
product right now going out. And since
images are so common, they already have
all this stuff to help us process the
data, which is great. And so, we come in
here, we do train data gen, and we're
going to create our object for helping
us train for reshaping the data so that
it's going to work with our setup. and
we use an image data generator and we're
going to rescale it. And you'll see here
we have one point which tells us it's a
float value on the rescale over 255.
Where does 255 come from? Well, that's
the scale in the colors of the pictures
we're using. They're value from 0 to
255. So, we want to divide it by 255 and
it'll generate a number between 0 and 1.
They have sheer range and zoom range.
Horizontal flip equals true. And this,
of course, has to do with if the photos
are different shapes and sizes. Like I
said, it's a wonderful package. You
really need to dig in deep to see all
the different options you have for
setting up your images. For right now
though, we're going to just stick with
some basic stuff here. And let me go
ahead and run this code. And again, it
doesn't really do anything because we're
still setting up the pre-processing.
Let's take a look at this next set of
code. And this one is just huge. We're
creating the training set. So the
training set is going to go in here and
it's going to use our train data gen we
just created flow from directory. It's
going to access in this case the path
data set training set. That's a folder.
So it's going to pull all the images out
of that folder. Now I'm actually running
this in the folder that the data sets
in. So if you're doing the same setup
and you load your data in there and
you're doing this, make sure wherever
your Jupyter notebook is saving things
to that you create this path or you can
do the complete path if you need to, you
know, C colon slash etc. And the target
size, the batch size and class mode is
binary. So the classes, we're switching
everything to a binary value. Batch
size. What the heck is batch size? Well,
that's how many pictures we're going to
batch through the training each time.
And the target size 64x 64. A little
confusing, but you can see right here
that this is just a general training and
you can go in there and look at all the
different settings for your training
set. And of course with different data,
we're doing pictures. There's all kinds
of different settings depending on what
you're working with. Let's go ahead and
run that and see what happens. And
you'll see that it found 800 images
belonging to one classes. So we have 800
images in the training set. And if we're
going to do this with uh the training
set, we also have to format the pictures
in the test set. Now, we're not actually
doing any predictions. We're not
actually programming the model yet. All
we're doing is preparing the data. So,
we're going to prepare a training set
and the test set. So, any changes we
make to the training set at this point
also have to be made to the test set.
So, we've done this thing. We've done a
train data generator. We've done our
training set. And then we also have
remember our test set of data. So I'm
going to do the same thing with that.
I'm going to create a test data gen and
we're going to do this image data
generator. We're going to rescale one
over 255. We don't need the other
settings, just the single setting for
the test data gen. And we're going to
create our test set. We're going to do
the same thing we did with the test set
except that we're pulling it from the
test set folder. And we'll run that. And
you'll see in our test set we found
2,000 images. That's about right. We're
using 20% of the images as test and 80%
to train it. And then finally, we've set
up all our data. We've set up all our
layers, which is where all the work is
is cleaning up that data, making sure
it's going in there correctly. And we're
actually going to fit it. We're going to
train our data set. And let's see what
that looks like. And here we go. Let's
put the information in here. And let's
just take a quick look at what we're
looking at with our fit generator. We
have our classifier.fit
generator. That's our back propagation.
So the information goes through forward
with a picture and it says, "Oh, you're
either right or you're wrong." And then
the error goes backward and reprograms
all those weights. So we're training our
neural network. And of course, we're
using the training set. Remember, we
created the training set up here. And
then we're going steps per epic. So it's
8,000 steps. Epic means that that's how
many times we go through all the
pictures. So we're going to rerun each
of the pictures. and we're going to go
through the whole data set 25 times, but
we're going to look at each picture
during each epic 8,000 times. So, we're
really programming the heck out of this
and going back over it. And then they
have validation data equals test set.
So, we have our training set and then
we're going to have our test set to
validate it. So, we're going to do this
all in one shot and we're going to look
at that and they're going to do 200
steps for each validation and we'll see
what that looks like in just a minute.
Let's go ahead and run our training
here. And we're going to fit our data.
And as it goes, it says epic one of 25.
You start realizing that this is going
to take a while. On my older computer,
it takes about 45 minutes. I have a dual
processor. You know, we're processing uh
10,000 photos. That's not a small amount
of photographs to process. So, if you're
on your laptop, you know, which I am,
it's going to take a while. So, let's go
ahead and uh go get our cup of coffee
and a sip and come back and see what
this looks like. So, I'm back. You
didn't know I was gone. That was
actually a lengthy pause there. I made a
couple changes. Let's discuss those
changes real quick and why I made them.
So, the first thing I'm going to do is
I'm going to go up here and insert a
cell above and let's paste the original
code back in there. And you'll see that
the original thing was steps per epic
8,000, 25 epics, and validation steps
2,000. And I changed these to 4,000
epics or 4,000 steps per epic, 10 epics,
and just 10 validation steps. And this
will cause problems if you're doing this
as a commercial release. But for demo
purposes, this should work. And if you
remember our steps per epic, that's how
many photos we're going to process. In
fact, let me go ahead and get my drawing
pen out. And uh let's just highlight
that right here. We have 8,000 pictures
we're going through. So for each epic,
I'm going to change this to 4,000. I'm
going to cut that in half. So, it's
going to randomly pick 4,000 pictures
each time it goes through an epic. And
the epic is how many processes. So, this
is 25. And I'm just going to cut that to
10. So, instead of doing 25 runs through
8,000 photos each, which you can do the
math of 25 * 8,000, I'm only going to do
10 through 4,000. So, I'm going to run
this 40,000 times through the processes.
And the next thing I not you'll you'll
want to notice is that I also changed
the validation step. And this would
cause some major problems in releasing
cuz I dropped it all the way down to 10.
What the validation step does is it says
we have 2,000 photos in our training or
in our testing set and we're going to
use that for validation. Well, I'm only
going to use a random 10 of those to
validate. So, not really the best
settings, but let me show you why we did
that. Let's scroll down here just a
little bit and let's look at the output
here and see what that what's going on
there. So, I've got my drawing tool back
on, and you'll see here it lists a run.
So, each time it goes through an epic,
it's going to do 4,000 steps. And this
is where the 4,000 comes in. So, that's
where we have. We have epic one of 10,
4,000 steps. So, it's randomly picking
half the pictures in the file and going
through them. And then we're going to
look at this number right here. That is
for the whole epic, and that's 24, 411
seconds. And if you remember correctly,
you divide that by 60, you get minutes.
If you divide that by 60, you get hours.
Or you can just divide the whole thing
by 60 * 60 which is 3600. If 3600 is an
hour, this is roughly 45 minutes right
here. And that's 45 minutes to process
half the pictures. So if I was doing all
the pictures, we're talking an hour and
a half per epic times 36 or no 25. They
had 25 up above 25. So that's roughly a
couple days. A couple days of
processing. Well, for this demo, we
don't want to do that. I don't want to
come back the next day. Plus, my
computer did a reboot in the middle of
the night. So, we look at this and we
say, "Okay, let's we're just testing
this out. My computer that I'm running
this on is a dual core processor. Uh,
runs 0.9 gigahertz per second. For a
laptop, you know, it's good about 4
years ago, but for running something
like this, it's probably a little slow.
So, we cut the times down. And the last
one was validation. We're only
validating it on a random 10 photos. And
this comes into effect because you're
going to see down here where we have
accuracy, value loss, value accuracy,
and loss. Those are very important
numbers to look at. So the 10 means I'm
only validating across 10 pictures. That
is where here we have value. This is ACC
is for accuracy. Value loss. We're not
going to worry about that too much. And
accuracy. Now accuracy is while it's
running, it's putting these two numbers
together. That's what accuracy is. And
value accuracy is at the end of the
epic. What's our accuracy into the epic?
What is it looking at? In this tutorial,
we're not going to go so deep, but these
numbers are really important when you
start talking about these two numbers
reflect bias. That is really important.
We just put that up there. And bias is a
little bit beyond this tutorial, but the
short of it is is if this accuracy,
which is being our validation per step
is going down and the value accuracy
continues to go up, that means there's a
bias. That means I'm memorizing the
photos I'm looking at. I'm not actually
looking for what makes a dog a dog, what
makes a cat a cat. I'm just memorizing
them. And so the more this discrepancy
grows, the bigger the bias is. And that
is really the beauty of the KAS neural
network. It has a lot of built-in
features like this that make that really
easy to track. So let's go ahead and
take a look at the next set of code. So
here we are into part three. We're going
to make a new prediction. And so we're
going to bring in a couple tools for
that. And then we have to process the
image coming in and find out whether
it's an actual dog or cat if we can
actually use this to identify it. And of
course the final step of part three is
to print prediction. We'll go ahead and
combine these. And of course you can see
me there adding more sticky notes to my
computer screen hidden behind the
screen. And you know last one was don't
forget to feed the cat and the dog.
So let's go and take a look at that and
see what that looks like in code and put
that in our Jupyter notebook. All right.
And let's paste that in here. And we'll
start by importing numpy as np. Numpy is
a very common package. I pretty much
import it on any Python project I'm
working on. Another one I use regularly
is pandas. They're just ways of
organizing the data. And then np is
usually the standard in most machine
learning tools as the return for the
data array. Although you know you use a
standard data array from Python. And we
have cross pre-processing import image.
This should all look familiar because
we're going to take a test image and
we're going to set that equal to in this
case cat or dog one as you can see over
here. And you know let me get my drawing
tool back on. So let's take a look at
this. We have our test image we're
loading and in here we have test image
one. And this one hasn't data hasn't
seen this one at all. So this is all
new. Oh, let me shrink the screen down.
Let me start that over. So here we have
my test image and we went ahead and the
cross processing has this nice image
setup. So we're going to load the image
and we're going to alter it to a 64x 64
print. So right off the bat, we're going
to cross is nice that way. It
automatically sets it up for us so we
don't have to redo all our images and
find a way to reset those. And then we
use also to set the image to an array.
So again, we're all in pre-processing
the data just like we pre-processed
before with our test information and our
training data. And then we use the
numpy. Here's our numpy that's uh from
our um right up here. Import numpy as in
p expand the dimensions test image axis
equal zero. So it puts it into a single
array. And then finally all that work
all that pre-processing and all we do is
we run the result. We click on here we
go result equals classifier predict test
image. And then we find out, well, what
is the test image? And let's just take a
quick look and just see what that is.
And you can see when I ran it, it comes
up dog. And if we look at those images,
there it is. Cat or dog. Image number
one. That looks like a nice floppy eared
lab. Friendly with his tongue hanging
out. It's either that or a very floppy
eared cat. I'm not sure which. But
according to our software, it says it's
a dog. And uh we have a second picture
over here. Let's just see what happens
when we run the second picture. We can
go up here and change this uh from dog
image one to two. We'll run that. and it
comes down here and says cat. You can
see me highlighting it down there as
cat. So, our process works. You're able
to label a dog a dog and a cat a cat
just from the pictures. There we go.
Cleared my drawing tool. And the last
thing I want you to notice when we come
back up here to when I ran it, you'll
see it has an accuracy of one and the
value accuracy of one. Well, the value
accuracy is the important one because
the value accuracy is what it actually
runs on the test data. Remember, I'm
only testing it on. and I'm only
validating it on a random 10 photos and
those 10 photos just happened to come up
one. Now, when they ran this on the
server, it actually came up about 86%.
This is why cutting these numbers down
so far for a commercial release is bad.
So, you want to make sure you're a
little careful of that when you're
testing your stuff that you change these
numbers back when you run it on a more
enterprise computer other than your old
laptop that you're just practicing on or
messing with. And we come down here and
again, you know, we had the validation
of cat. And so we have successfully
built a neural network that could
distinguish between photos of a cat and
a dog. Imagine all the other things you
could distinguish. Imagine all the
different industries you could dive into
with that. Just being able to understand
those two difference of pictures. What
about mosquitoes? Could you find the
mosquitoes that bite versus the
mosquitoes that are friendly? It turns
out the mosquitoes that bite us are only
4% of the mosquito population, if even
that, maybe 2%. There's all kinds of
industries that use this and there's so
many industries that are just now
realizing how powerful these tools are.
Just in the photos alone, there is a
myriad of industries sprouting up. And I
said it before, I'll say it again. What
an exciting time to live in with these
tools and that we get to play with. So
key takeaways. Well, we covered what is
a neural network. We use all kinds of
processing the map images on your phone.
We talked about things that a neural
network can do. translate text, identify
faces all the way to control robots, you
know, lots of exciting things. How does
a neural network work? So, we discussed
that with the different layers going
from the picture to the input layer to
the hidden layers and their weights to
the final output layer. We also talked
about how it does the math and computing
the output as yes or no, categorically
true false. We discussed types of
artificial neural networks. A lot of
vocabulary there from the feed forward
neural network which is the most
commonly used. That's the one the neural
network we used is a feed forward neural
network that does backward propagation
to train. And there's a lot of other
ones out there. There's the radial
biases, the cohen self-organizing
recurrent neural network, convolution
neural network, modular neural network.
The big one was modular because it
incorporates pieces of all the other
ones. So that whatever you're working on
now is a huge conglomerate of multiple
networks. Just all cutting edge. All of
it's new. People even working on it
don't even know where it's going. Again,
very exciting times. And finally, we dug
through my favorite part. You can see
with my uh latte on one side, my old
school pens and pencil, and all my
sticky notes working away. That's not
actually me, by the way. You probably
guessed that. And we walked through and
actually did a cat and dog photo, a
simple cat and dog photo. And you could
see where some of the problems are in
processing large amounts of photographs
and data where that starts to become
going from a single machine on my laptop
with this, you know, lower amount of
resources all the way to big data. How
if you're processing hundreds and
thousands of these photos, this now
needs to be set up on an enterprise
machine or even on a cluster of
computers. Again, significantly past the
scope of this. The neat part about it
though is once you write this code, most
of this code, they now have tools that
you can almost take the same ideas, if
not the actual code, and push it right
onto a cluster computation. So really
cool times for this Python. My name is
Richard Kersner with the SimplyLearn
team. That's www.simplearn.com.
Get certified, get ahead. Although deep
learning is uh been around for a while,
it is just in its infant stages of
development as far as exploding on the
market. I mean it is right now they're
building robots with it. Deep learning
is used to train robots to perform human
tasks. Music composition. Deep neural
nets can be used to produce music by
making computers learn the patterns
involved in composing music. Image
colorization. Neural network recognizes
objects and uses information from the
images to color them. Machine
translation. Given a word, phrase or a
sentence in one language, neural
networks automatically translate them
into another language. Google Translate
is one such popular machine translator
you may have come across. And you'll
notice in here we didn't show any
examples of straight numbers like uh
projective cells in a business tracking
your favorite stock. You can certainly
do those with machine languages, but
this is the next level. Uh save that for
your regression models, your linear
regression where you're actually
processing and crunching just straight
numbers. With machine learning and deep
learning, we're going to a whole new
level as far as what we can figure out
on the computer. What's in it for you?
We're going to cover what is deep
learning. We're going to take a look at
the biological versus artificial
intelligence. What is neural network
activation function in your neural
network and the cost function and how do
neural networks work. How do neural
networks learn? So there's a little
you'll see a switch right there. We just
went from how are they working in the
math in the background to exactly how
are they learning. We'll be implementing
the neural network. We'll do a gradient
descent deep learning platforms and
we'll give an introduction to TensorFlow
and implementation in TensorFlow. That's
Google's platform that they open sourced
recently and it's probably one of the
most cutting edges in deep learning and
even it is still in the infant stage
which is one of the reasons they
released it to open source. What is deep
learning? Deep learning is a sub field
of machine learning that deals with
algorithms inspired by the structure and
function of the brain. And you can see
we have a nice picture here. We have
artificial intelligence which is kind of
the big bubble that encompasses all
these different things we're talking
about. This is ability of machine to
imitate intelligent human behavior. And
in there we have machine learning
application of AI that allows a system
to automatically learn and improve from
experience. And if you looked at any of
our other videos, you'll know that
machine learning covers a lot. So deep
learning is a subcategory of that. But
don't forget machine learning has all
kinds of other tools that people use to
do very basic uh descriptive and
predictive and postcriptive uh
analytics. And then you have deep
learning application of machine learning
that uses complex algorithms and deep
neural nets to train a model. Let's take
a look at the biological neuron versus
the artificial neuron. Now remember in
the human brain and and this is true for
most animals there are a lot of
different neurons going on. So this is
the very basic one. I mean there's
hundreds of different cells involved. So
when we talk about neural networks and
this is why I say it's in a very infant
stage. They're really basing it on uh
just the most basic thing that we're
able to figure out going on in the
neural networks. And you can see right
here we have dendrites fetch information
from an adjacent neurons and pass them
on as inputs. So you have your data
coming in and your data going out. Any
computer model should be looking at that
what's coming in what's going out. The
data is fed as an input to the neuron.
So we look at the artificial neuron. You
can see we have our inputs. They come in
each one is specially weighted into the
neuron and then the neuron has an
output. The cell nucleus processes the
information received from the dendrites
and the neuron processes the information
provided as inputs. Axons are the cables
over which the information is
transmitted and the information is
transferred over weighted channels. So
you can look at that uh I mentioned
weights briefly but you alter the data
coming in. So those weights are what
causes different information coming in
to be weighted differently and processed
differently. And the synapses receive
the information from the axons and
transmit it to the adjacent neurons.
That's in your biological model. And
then when we look at the artificial
neuron, the output is a final value
predicted by the artificial neuron. So
as we dig deeper into looking at the
theory behind the neural network and we
kind of flip back and forth between
these because there's two huge aspects
of it. One is from the outside. What are
you seeing and what's going on from the
inside so you can find to do what you
need to do and give the best results you
can. And we start off with what do we
feed? We feed an unlabeled image to a
machine which identifies it without any
human intervention. And so you can see
here we have a circle that comes in at
784 pixels and it comes in by 28x 28.
And you can see how it colors in the um
the circle on there. And we put a
triangle in. The triangle in also comes
in as 28x 28 and it has 784 pixels. So
you'll see between these two both of
them are 784 pixels. This machine is
intelligent enough to differentiate
between the various shapes. So that's
what we want to use our neural network
to do is to say hey this is a circle.
This is a triangle. That's more of a
categorical. You can also do a
regression model where you're actually
putting out float value or a numerical
value. We'll be looking at the true
false or the categorical model mostly
because that's where you usually start
at the different there is no real
difference when you as far as the way
the internal functioning goes when you
start flipping between them other than
well we'll talk about that in just a
minute. So you can actually go between
the two quite easily and the neural
network provides this capability. So
we're going to use this capability to
look between those two. One of the
things I want you to note in here is
that we're looking at 784 pixels. We're
looking at 784 inputs. That's very
different than stock with a high low or
last year's sales based on date or we're
looking at just a couple of numbers and
they're very clear. They're numbers.
They're very clear what they are, which
is something you'd put into a machine
learning linear regression model. This
is a step up from that in that we're
looking at complex patterns and how do
you figure those complex patterns out.
So, a neural network is a system modeled
on the human brain. And we looked at
that comparing the two. Let's go ahead
and look deeper into the neural network
itself. We have our inputs coming in. So
the inputs are fed to a neuron that
processes a data and gives us an output.
Input and output. This is the most basic
structure of a neural network known as a
perceptron. So if you see the term
perceptron, that's what we're talking
about. We're talking about this single
node that has inputs and an output.
However, neural networks are usually
much more complex. Let's start with
visualizing a neural network as a black
box. And I always love that symbol. It's
a black box. It's kind of magical. We
have our inputs coming in and we want
certain outputs. The box takes inputs,
processes them, and gives an output.
Let's have a look at what happens within
this box. And you can see me there in my
uh secret agent getup and I got my
hidden hood and everything. I guess I'm
part of the uh black skull or something
like that group. Uh so let's take a look
at what happens within this magic box.
And remember, we're skipping back and
forth between the theory of what's going
on in the box, which you have to know
how to fine-tune and how to build,
versus looking at it from the outside.
We're programming this box, and we have
an input and an output to the box as a
whole. Within the box exists a network
that is a core of deep learning. And you
can see here we're showing one layer and
we have our grid coming in. The network
consists of layers of neurons. Each
neuron is associated with a number
called the bias. And you can think of
the bias uh if you overly simplify this
and we're doing a linear regression
model. This is your y intercept in your
uklidian geometry. You have to have
something that offsets it. And so you
always have a bias in these cells.
Neurons of each layer transmit
information to neurons of the next layer
over channels. And so you can see each
of our layers going through from left to
right. These channels are associated
with numbers called weights. These
weights along with the biases determine
the information that is passed over from
the neuron to neuron. So just like the
bias is your y intercept in uklidian
geometry. You could look at the an one
weight. Remember this is very
complicated. So we're not looking at
just one weight. You could look at the
weight as your slope of the line. Or if
you're doing x= uh my y + c, it would be
the m value. Neurons of each layer
transmit information to neurons of the
next layer. And you can see here as they
light up going across into the final
layer. and then to the output. And in
this case, the output is going to be
either uh a square in this one or it
might light up the other one which is a
circle. The output layer emits a
predicted output. So in this case, we're
looking at a classification uh true
false. Is it a circle? Is it a triangle?
Is it a square? Let's now go deeper.
What happens within the neuron? So we're
going to dig deeper and start getting a
little bit closer to some of the math.
Don't worry, you don't have to be a
calculus expert and know your
differential equations. Even though this
is one giant differential equation, you
don't need to understand those to
understand what's going on. Within each
neuron, the following operations are
performed. The product of each input and
the weight of the channel it's passed
over is found. This is simply addition.
We're going to sum up the weight times
the output from the previous channel and
plus the bias. Sum of the weighted
products is computed. This is called the
weighted sum. Bias unique to the neuron
is added to the weighted sum. The final
sum is then subjected to the particular
function and we'll discuss those that
particular function. That part is really
important because those functions uh
have a huge impact on how well your
model performs under different
conditions. The final sum is then
subject to a particular function. This
is the activation function. So if you
ever hear the term activation function,
that's what we're talking about. What
activates this cell and what doesn't. As
we dig deeper into activation function,
an activation function takes the
weighted sum of the input as its input
adds a bias and provides an output. And
a lot of times you'll actually see one
formula for the sum of the weight the
weighted sum and the bias. You'll just
see that as a single line of everything
added together. And here we've broken it
apart because it makes it clear that
this bias is not computed the same as
the weighted sums. Here are the most
popular types of activation function.
And I always find these interesting
because at one point I was sitting at a
table with a gentleman who was finishing
his PhD. He was in his last year and he
said he went through all this stuff and
he ended up just trying the four
different activation functions on this
particular problem he was working on. So
knowing the math behind it doesn't
necessarily mean you're going to know it
right away. Uh so even somebody who
might have a PhD and be doing the
calculations on this comes back out of
it and ends up just trying the different
uh um activation functions to see what's
going to make a difference. And a lot of
times that's a final step. That's the
kind of thing where you built your whole
model. You've come back and you're like
wait a minute can I do a better deal
with a sigmoid function or the threshold
or the rectifier. Knowing what they're
doing is important so you can explain it
to somebody else. And again you probably
do this on a small set of data. If
you're working with big data, uh you
don't want to take down the full server
farm just to test out your three
different series. You take a small
portion of that data, test it, and then
you put it through to the big data. So
let's take a look at this. We have the
sigmoid function, and it's used for
models where we have to predict the
probability as an output. It exists
between zero and one. And you'll see
that's true of all of our activation
functions we're working with. Either the
cells on or off, it's true or false. And
there might be a little variation in
there which as an output could be used
to compute uncertainty in your solution.
So if you're getting a 7 with this
activation function, it might be well
I'm not sure if that's really a square
or I'm not sure that's really a
triangle. And that might be a flag for
it to be looked at by a human observer
at least in today's models where we're
at right now. And you can see here we
have the formula is simply equals 1 over
1 + e the minus x where x is your value
coming in. and it's going to give you a
result that looks very similar to the
graph on there which is somewhere
between zero and one. Um, and right in
the middle you can see that there's a
huge uh kind of you can go through all
the different values and uncertainties
involved. So the sigmoid function is
probably the default on most of them. Uh
the next one is the threshold function.
It is a threshold-based activation
function. If x value is greater than a
certain value, the function is activated
and fired. Else not. Pretty
straightforward. Yes, no, true, false.
um I don't want to test for
improbabilities. I just want a straight
answer. I don't want to know if there's
a partial value on there. It either is
true or it's false. And the rectifier
function, it is the most widely used
activation function. I would debate
that. Um rectifier is pretty common one,
although I see that the sigmoid function
is used to be the basic one, but it's up
there. The rectifier function is very
commonly used. You get the output of X
if X is positive and zero otherwise. And
you can see here again just like um uh
it's either you kind of get a value
going up there. So max of x of zero. So
it's it's again it's like the threshold
function. Yes, no, true, false. Uh it's
either zero or it's uh some kind of
progressive value. And then we have the
rectifier function. I would argue with
this because the sigmoid function used
to be the most common one. But with the
rectifier function, it now says it is
the most commonly used or widely used
activation function and gives an output
of X if X is positive and zero
otherwise. This is kind of nice because
it now says absolutely not or it gives
you a value of probability. Now, when I
say a value of probability, be very
careful there. I'm not saying that it's
going to tell you this is 75% chance of
being a circle. I'm going to tell you
that it says, hey, if this says 0.1, it
probably needs to be looked at or 2 or
3. It's going to depend on your data as
to what that value means. In general,
that just means it's flagging it that if
it's not a one, then chances are it
needs to be looked at by a person and
re-evaluated. And there's a hyperbolic
tangent function. This function is
similar to sigmoid function is bound to
a range of minus1 to 1. So you can see
there's our 1 - eus 2x and 1 plus over 1
+ eus 2x. Again, it's very similar to
the sigmoid function. The bonus of the
hyperbolic function is you have that
variable coming through the middle. So
again, you can look at it and you have a
little bit more weight as far as you can
process that down the line. That's a
little bit more advanced than than what
we're looking at right now. And a lot of
times it's not even necessary in a lot
of our different uh uses for these
activation functions. Now, we looked at
activation functions and I kind of said
those are a little bit like a black box
because even if you know all the math, a
lot of times you end up just playing
with them to find out what works. And it
also depends on what model you're
working with, whether you need a flat
yes, no, true, false, or you need to
have something in the middle that says,
hey, this isn't quite a one. You might
need to process this with the human
intervention. And you could look at
that. Uh, one example would be
self-driving cars. You don't want a car
to be yes, no, I'm going to go through
the the light. You want it to be like,
okay, if it's uh almost yes, maybe we
stop and have human intervention so we
don't get an accident. Cost function is
something you can really see and measure
and is very important. The cost value is
the difference between the neural net's
predicted output and the actual output
from a set of labeled training data. So
we have our group of data that's a
square circle and since we're looking at
geometrical shapes, we've had somebody
already labeled that data. They've
already said this is a triangle, this is
a square. And so if this is coming up
and it's giving us and it's saying a
square is a triangle and it's saying a
triangle is a circle, the output is
wrong. And so that output can then be
measured in the versus the actual output
and that's the cost. Uh you might also
hear this as error because that's the
error value being returned. How far off
is it? And what we're looking for is the
least cost or the least error value. And
it's obtained by making adjustments to
the weights and biases iteratively
throughout the training process. And
this this is called back propagation.
And we're going to look in that a little
deeper as we look into an example. It's
really hard to see when you're just
looking at arrows without actual numbers
and where that flow is coming from. But
you can look at this is here's our
inputs. They put out a prediction. The
prediction comes out and says, "Hey,
we've already labeled this data cuz
we're in training mode and the training
data is off. This is the cost. Can we
send that error or that cost back and
adjust those weights?" And we do it in
very small increments across large
amounts of data so that those weights
minimize that cost or that error. But
what happens within these neurons? So
let's look at a little example of this.
Kind of helps if you have some kind of
visual. Let's build a neural network
that predict bike prices based on a few
of its features. And we'll see here we
have our CC, our mileage, and our ABS.
And these are our three input layers.
And then we have the bike price and the
output layer. Now, it doesn't do us very
good to just uh pump it in from the
beginning and pump it out. And to be
honest, I would use a machine learning
linear regression model on this since
these are just straight numbers. But
because we want a simple example, we're
going to put this through and show you
as a neural network what that looks
like. And we got to put a hidden layer
in there. The hidden layer helps in
improving the output accuracy. And you
could look at this as a bunch of ores.
So it might say, hey, when we compare
these three values on the first hidden
layer neuron, we're looking at one set
of features and then we might weight
them in the second one. So these are a
bunch of different ores kind of how the
math comes out in behind the scenes. And
then they go out of course to the bike
trace or the output layer. And each of
the connections have a weight assigned
with it. And you'll see here we have a
mileage CC with the weight one and
weight two going into our first neuron.
And you'd also have your ABS going in
there. And so X1 * weight 1 + X2 *
weight 2 plus the bias of one. And step
two is our activation. The activation
function coming in there. When does this
fire? And the neuron takes a subset of
the inputs and processes it. And then we
go through and we do that with the um
second hidden layer neuron and the third
one and so on. So you process each layer
in order going forward. Now when I told
you this is in its infant stage, they
now have neurons that fire into the same
layer or back a layer so that you now
have a time series and there's all kinds
of wild things that they're
experimenting with on these layers. This
basic setup has been around since the
mid90s. It's only now because of our
technology that it's open to almost
everybody to play with it. And that's
why I say this is in an infant stage in
development is this basic math is here,
but what we can do with it is amazing.
And what they're actually doing with all
these different things is amazing. And
so we're just at the beginning of how to
use all these different tools and our
deep learning and our neural networks.
Uh and so once we have our hidden layer
computed, the information reaching the
neurons in the hidden layer is subjected
to the respective activation function.
And so each one of these fires an
activation output uh and then those are
each weighted to the final output layer.
So the processed information is now sent
to the output layer once again over
weighted channels. And you could look at
this as each one of these is um I always
look at this as like a group of people.
They're all looking at the bulletin
board and the first person says this is
what I project sales for the company and
the second person and the third and so
on. And then their perspectives are
weighted based on their expertise. So
your accountant might have a very high
weight where the um maybe your janitor
has a very low weight because their
expertise is not in accounting and then
that goes into the output layer and once
in the output layer it goes uh the
output which is the predicted value is
compared against the original value. So
now we have our output layer and since
we have like already a list of uh bikes
with their the different setups and what
their value is we can now generate an
error from this. The cost function
determines the error in prediction and
reports it back to the neural network.
So this is the cost. This is how far off
it is. This is your error coming back.
And as you can see, this is back
propagation going on. So now our error
is going in reverse because we know
we're not completely correct on this
particular channel. The weights are
adjusted in order to reduce the error.
So each time we go back, we are changing
those weights to reduce that error. and
we change them in small increments. You
don't want to fit one input. Remember,
you might have a data pool with a
terabyte of data. You don't want to
solve for the first set of data that
comes in and that be the main solution
because everything else will be off.
This is going to confuse you. That's
also called a bias. So, we have the bias
in the cell where we're adding a value,
the kind of like the y intercept, and we
have a bias of the whole neural network,
which means that it's weighted towards
one set of answers. So we want to make
small changes in these weights so we
don't create a bias and the weights are
adjusted in order to reduce the error or
the cost. The network is now trained
using the new weights. Once again the
cost is determined and back propagation
is continued until the cost cannot be
reduced any further. So let's go ahead
and plug in values and see how our
neural network works. So here we come in
here and initially our channels are
assigned with random weights. This is
important because if you assign them all
with the same weight, you might be able
to reproduce it. But it turns out that
if I put all my weights as one or all my
weights as zero, it takes longer to
train where if you have random weights,
they already have like a little bit of
adjustment and ores built in and that
will give us a better answer and train
faster. Our first neuron takes a value
of mileage and CC as inputs. So here
comes our computation whatever those
inputs are. And we do that again with
the second neuron with those values
coming in. You can see here we have
weight three and so on and then our
third neuron coming down and of course
our fourth neuron. So we're adding all
these different values coming in here in
our hidden layer. The process value from
each neuron is sent to the output layer
over weighted channels. So again here's
our weights coming in and we have N1,
N2, N3 and N4. Once again the values are
subjected to the activation function and
a single value is emitted as the output.
On comparing the predicted value to the
actual value, we clearly see that our
network requires training. So, here we
have it that our bike price uh we put
out, we thought it was worth 2,000 on
our random weights and the bike actually
was $4,000 on there. Guessing that's not
US dollars cuz that'd be a very
expensive bike. But maybe it is. There's
some $2,000 $4,000 bikes out there. The
cost function is calculated and back
propagation takes place. And this is
pretty simple. You can look at that as
our um we're subtracting one value from
the other. We square it and then we take
half of that and that is propagated back
up. And each layer generates its own
errors. Let's go back one because you
have your predicted Y and your actual Y.
That goes back to the first layer. And
then based on the value of the cost
function, certain weights are changed.
So when we look at the next layer, that
error is not the original 4,000 - 2,000
squar / 2. This error is based on the
error of each cell generated. How far
off is that cell as far as its weights.
We're not going to show you. It's
actually a very complicated differential
equation. And you can probably write it
out if you wanted to. You just write out
each formula that goes into the next
level and you add them all together and
you can write it out all the way
through. Computers make it so you don't
have to. And our neural network is
considered trained when the value for
the cost function is minimum. So when we
get our error way down as low as we can,
that's when our neural network is
trained. And there I mean just recently
they've come up with all kinds of
different means for measuring that
particular value. a little bit beyond
the scope of today's neural network, but
you can actually you can actually see
it. You know, how far do you do this
until the neural network doesn't need to
be trained anymore and you can overtrain
a neural network. Now, the tools that
we're looking at automatically let you
know when to stop, which is really nice.
And that is just like I said, we're at
the beginning stages in neural networks
and it's just really cool what they can
do now and how much of it's automated
and how much of it is experimental.
Right now, let's take a look at gradient
descent. But what approach do we take to
minimize the cost function? So here we
have nice error thing coming in. This is
our cost or our error. Uh let's start
with plotting the cost function against
the predicted value. And so you can see
they fed in multiple y's and these are
the errors coming in and the cost of
each of these inputs and changes going
on. Note we start at a random point on
the curve. So usually you put in you
know you pick up your data and you
randomly pick where to start in your
data. A lot of times you just run it
from the beginning because you're going
through so much data, it's not that big
of a deal. But you start with one point
going in. So your forward propagation
goes through. You're going to go ahead
and find your cost or your error. It
points that on the curve. And you can
see how we're plotting it right here.
Since the gradient at this point is
positive, we may move right. So we're
going to move a little bit to the right
on here. And this time the gradient is
negative. We move a little bit to the
left. Eventually we try out the point
where the gradient is zero. This is a
least value of cost function. You have
to be a little careful with this because
this particular I mean they make it look
nice and simple in this graph. Sometimes
these curves look like stair steps and
so there is global minimums and then
there is local there might be a local
point where the gradient is zero but
it's not the global one. Uh so it might
be way off to the left where it just
happens to step down a little bit and
you think you're in the right gradient.
And with that we have all the right
weights and we can say our network is
trained. So here we have um just some
major these are some of the big names
out there right now in development for
deep learning platforms. TensorFlow
which we'll actually do an example in in
a minute. Deep learning for J which is
in the Java platform. Uh so if you're a
Java programmer uh by the way is
TensorFlow is accessed most people are
using Python to access it but it is a
system that's kind of separate from a
lot of the programming languages which
makes it a lot more um flexible as far
as use. Deep learning forj is Java based
and then cross is just exploding right
now. And this is interesting. Cross is
uh working with TensorFlow. It actually
can sit on top of TensorFlow. And it can
also do its own thing. Uh so if you're
studying deep learning, you're getting
into it, you want to know the basics of
TensorFlow, but you also are going to
want to know the upper level of KAS
sitting on top of TensorFlow. We're just
looking at TensorFlow today though in
our example. And there's also Torch on
there. There's a bunch more that we
didn't list on here. Um even sklearn or
the uh side package in Python has a
neural network you can program a very
basic one and it is the same basic one
that you could do in TensorFlow if you
stripped everything out of it and then
TensorFlow has a lot of tools they've
added in and so has KAS but we're going
to be looking specifically at TensorFlow
in our example and TensorFlow is an
open- source tool used to define and run
computations on what they call tensors
very common language now so you more and
more we see the term tensor as being a
standard in the uh deep learning
language and this was originally
developed by Google. So let's dig a
little bit big in there. What are
tensors? Tensors are just another name
for arrays. So a tensor of dimension
five. You can see here we have ab kmq
whatever. So it's an array coming in.
And the tensor of dimension 54 more like
a picture. Very common to see that in a
picture. You can also see a tensor even
more detailed than a picture as we go to
the next one. Tensor of dimension 333.
This is 3D space. You might have a
picture that also has colors. That might
be the third dimension. You might have
four dimensions because you have both
your grid and your different color
channels and your zplot. You can see
where you can now process a very
highlevel set of data coming in whether
as an image or features. They could be
features that have nothing to do with
images. So there's a lot of stuff you
can do now with the tensors coming in.
Thus this where the term tensorflow
comes from. So we have um right now the
TensorFlow is the most popular library
in deep learning and I did mention KAS
now works with TensorFlow. So there's a
lot of stuff you can do between the two.
Uh it's an open-source software library
developed by Google. Uh so they hit a
roadblock and they realized hey this is
an infant stage technology. You know we
thought it was going to be the next
greatest thing and we were going to have
a hold on it but it's really infant as
far as how it's applied and what we can
do with it. Let's open source it so
everybody can work on it. uh let's take
it to the next level. And that's really
what open source does to a lot of these
uh packages when they release them. And
you can run on either a CPU or a GPU. So
when we look at the details, if you have
your graphic processing units, um what's
nice about those is they run a lot
faster. The downside is you have to play
with them a little bit to get them up
and running. And it's a hardware
upgrade. When we run it, I'll be running
it in the CPU mode. I have played with
it in my GPU on my personal computer.
you know, it does increase the
processing. Uh, but I did run into some
version problems with my Python and
stuff like that. And when I did finally
work it out, I went back to the CPU
because it didn't increase my speed
enough for what I was working on. But in
a larger group, you might be able put
that on. If you're working with a larger
stack of computers, you might want to
run it in the GPU. You can create a data
flow graphs that have nodes and edges.
So there's our edges coming in. We
didn't talk about edges, but that's very
up and cominging way of looking at your
analytical data is how do different
nodes connect? What do those edges look
like in between them? And it's used for
machine learning applications such as
neural networks. It is mostly a neural
network, but they have all kinds of
tools which sit on top of our basic
neural network. They have new stuff
evolving into the TensorFlow library.
So, it's very much uh just exploding.
great time to jump into TensorFlow
because there's all kinds of cool things
we're doing with it and all kinds of
cool applications you can now use uh
TensorFlow for. So let's take a look at
implementation in TensorFlow and we're
going to build a neural network to
identify handwritten digits using the uh
Mnest database or the MNIST database and
that stands for modified National
Institute of Standards and Technology
database. It is a collection of 70,000
handwritten digits and the digit labels
identify each of the digits from 0ero to
nine. This is a cool example because
it's simple enough that you could
actually run this through some basic
machine learning categorizing algorithms
and train them and you'll get about the
same answer because again it's it's
simple grid. The digits on the grid
don't have a huge amount of variation
like you would say an automated driving
car looking at the environment. So you
can still do this with a lot of your um
different linear models and stuff like
that. You can solve this and you'll get
about the same answer. When I ran a
comparison between TensorFlow and
between some basic uh regression models
or category models uh in machine
learning, they came up pretty even as
far as their output. Uh so this is kind
of where we start to see the complexity
of something coming in this case a
tensor you know or a grid of uh
information where the deep learning
model does as good as the regular models
and when you get past this kind of
complexity and features suddenly the
neural networks come up with better
answers better solutions and a better
build and that's why there's such a move
into neural networks is we live in a
complicated world and it's just really
cool we can do with this. So the
handwritten digits from the um NIST
database, they come in, the data set is
used to train the machine, a new image
of a digit is fed and the digit is
identified. Um and if you've looked at
any of our other machine learning tools
where we're doing training, uh where we
train our uh model to fit and then you
test it out, this should look pretty
familiar. Uh and there is some tools out
there for say untrained categorizing uh
where it's just looking for features
that fit together. So there are tools
that don't need that training. But this
is where uh when we talk about neural
networks, we do need to train them. And
this is what we're looking at.
So for this I'm going to use the
Anaconda Navigator just because it's a
very nice visual tool. You might be in
PyCharm or one of your other IDEs for
editing Python because we are looking at
Python TensorFlow. And under Anaconda,
we have the notebook, which is something
we use pretty regularly. And they have
the Jupyter Lab. The Jupyter Lab is the
Jupyter notebook, but with tabs and a
few new features. So, we'll be using the
Jupyter Lab today. And under the
environment, you'll want to go ahead and
and uh if you haven't yet, uh you'll see
that I have a number of different setups
in here. Right now I have the Python
version 36 and the TensorFlow. In this
case I have TensorFlow 1.12. If we
scroll down you can see that uh here we
go. TensorFlow and it's version 1.12.
And in here if you haven't yet you'll
need to install those and go in and just
open our terminal. And u if you've never
used the Anaconda or if you're in your
other thing you might have something
simple like pip. Is what I use for my
install. And you can simply do install
TensorFlow. And that should bring in the
most current version. Now, when I
installed this a few months ago, Python
version, I'm not going to run this
because I already have installed on
here. Python version 3.7, the newest one
out, still had a couple glitches with
the TensorFlow. I believe they've fixed
it as of writing of this, but um I'm
going to stick with 3.6 just so I don't
get any surprises on there. So, this is
Python version 36 with TensorFlow 1.12
on here. And if you haven't installed it
yet, you also want to install Numpy for
this example. That's Numbers Python or
uh nu py. You can just simply run an
install on there. Keep in mind if you're
in Anaconda uh and you've created one of
these environments specific to this,
keep withd.
If you're going to use pip, keep with
pip. Don't install one package with pip
and one under because that's how they
track those version numbers and how they
fit together and you can end up with a
problem. they don't pip doesn't see cond
and vice versa. Uh so just keep that in
mind when you're running your installs.
We'll go ahead and open up Jupyter Lab
and we're going to launch that. So
here's my Jupyter Lab. One of the really
cool features of Jupyter Lab is you have
tabs now. So you can open up multiple uh
notebooks. And this is nice cuz I have
my notes I'm working on and then our
actual window we're looking in. And
we'll go ahead and zoom in a little bit
here. There we go. So you have a nice u
hopefully easy to see fonts. And then
we'll go ahead and do a simple or get
our imports out of the way. Um, and so
we're going to import our TensorFlow as
TF. Uh, that's pretty much a standard
for TensorFlow, numpy, our numbers
Python as py, and we'll import our matt
plot library as plt. Again, these are
very common. So if you see TF or py or
plt, this is a standard that most people
use. Do you have to? No, you could just
do import numpy instead of doing as py.
And then from
tensorflow.acamples.tutorial
tutorials. This is always nice because
they actually include data set we're
going to play with. So, we're going to
import input data. So, there's our data
coming in. That's all we're doing is
telling it this is where it's coming
from. And if we're going to tell where
it's coming from, we need to go ahead
and create a variable with that
information in it. And we'll just call
this uh mnist
or minced. You know, I don't really know
how they pronounce that. I should
probably look that up. It's a very
common data set to use. And there's our
input data. And we're going to read data
sets. And this is um if you look at
this, we imported input data from our
TensorFlow. And so this is a TensorFlow
read statement for their tutorials. So
this isn't like some special Python
setup. This is just their setup. Makes
it easy to pull it in. So once we get
into their data sets, we need to go
ahead and tell it what kind of data set.
And again, this is what we brought in,
but it's going to be the nint data. And
this part is very important. one hot
equals true. This means that instead of
importing a value from 0 to 9, we
evaluate the data set. It's going to
bring it in as one hot. Whenever you see
one hot encoder, we're flattening that
out. And we have true false for zero,
true false for one, true false for two.
So our output, if you remember from our
output uh from the slide we did earlier,
uh in this case, I grabbed the one for
bike price. Doesn't really matter which
one we use. This has one output. So we
have our bike price on this. We're going
to have instead of one output, we're
going to have 10 outputs representing
each of the digits in there. And this
code really isn't going to show us
anything. It's good to see what we're
actually looking at. So um let's go
ahead and do a figure ax equals go into
our plot library subplots 10, 10. And
that is if you remember we talked about
tensor. Tensor being data coming in.
This is a 10 by 10 grid or 100 pixels on
there. And if we're going to display it,
uh let's go do K0
for I and range 10. Just a simple loop
through on the data. Let's do what is
it? Uh for J and range 10. And I
actually misqued that 10 uh 10 x 10 is
not the actual size of the pixels. Uh
the actual pixels are going to be um we
look at the shapes and we'll get into
that in just a second here. We'll take a
quick look at shape on there. Uh turns
out they're uh what are they? are, I
believe, 28x 28. Uh, so let's take a
look at that. And we're just going to
plot these. What are we looking at? What
are we working with? As a data
scientist, you should always be looking
back at your data and seeing what it
looks like and get that human
perspective because you just never know.
You know, the the computer may put
something out that looks makes no sense.
And at that point, you want to go back
and reevaluate what you did. Uh, so
we're going to go ahead and plot. We're
going to plot 10 digit, you know, 10 of
the digits by 10 of the digits. And
here's our ax. We'll create the J on our
subplots and we're going to do an image
show. We're going to look at the
training image for images of K. And then
we want to reshape this. We're going to
reshape this. And we're going to reshape
this 28x 28. That's how I knew I had it
wrong is cuz I looked down my notes. I
was like, oh no, that says 28. It's not
10 x 10. And I should know that already
cuz I've done enough messing with this
data set that I should have remembered.
Uh, but it's 28 x 28. And the aspect
we're going to do is auto. And this is
all, if you look at this, here's our
variable NIST. the NIST is coming from
data set. Uh so this is all part of the
TF TensorFlow learning or examples
tutorial in there. And then we'll go
ahead and do K plus equals 1. So we just
keep paging through our different um
images. And let's see what that looks
like. Let's go ahead and do a plot show.
Uh and we'll go ahead and run this so we
can take a look and see what we have
here. And so we have a nice plot here.
And you can just see that we have uh
some random numbers showing up in each
one of these little subplots. If you're
wanting a copy of this code, put a note
down in the YouTube video and let us
know or come visit us at
www.simplearn.com
and we'll send you out a copy of what
we're working on and get a copy of that
for your own setup. Uh so now we've
taken a look and we can just see we have
here's our pictures that are coming on.
We plotted them so we have an idea of
what we're looking at. Let's go ahead
and uh print. Let's look at the shape of
the features. Uh so when we have this we
have our nest train images and we'll do
the shape on there. Let's take a look
and just see what we're looking at uh as
far as uh our count and everything. And
so you can see here we have 55,000.
That's basically how many images we have
and this by 784. And in this data set
there's also our labels. So let's take a
look at that. We have our net train
labels shape. Let's take a look and see
what that looks like. Uh and there we
have 10 because there's 10 digits. So we
brought in that's our output we're
looking at. And so we have there we go
55,000. They match. They should match
because you should have equal numbers in
both of those. You know, here's our data
in and here's our answer. If you
remember, this is a bunch of zeros and
with one each each one will be 0001
would be what letter four or something
like that. So, let's take a look at what
our one hot encoding did for the first
observation. And this is when we're
exploring data, you really want to dig
in there and just see what the heck am I
looking at. So, we're going to look at
the labels. And this would be the first
label that comes up. And we'll go ahead
and run this. And we look at that. You
can see this is what I'm talking about.
0 0 or 1 is 0 2 is 0 3 is 0 four is 0 5
is 0 6 is 0. 7 equals 1. So our very
first label is a seven. But our very
first label comes up that it's a seven.
And so we don't have like 0 through 9.
We have a bunch of zeros and just the
one to mark it as a seven on here. So
now we've kind of looked a quick look at
the data. And in here you might ask some
questions like what is 784? 24 * 24.
Remember that's the size of our grid on
there or our tensor coming in. So 784 is
a setup on there. And we've gone through
all this viewing the data. We'll go
ahead and start looking at our
tensorflow. So let's take our X
variable. This is going to be our
training set. We'll do a placeholder and
then we're going to have these come in
as float. Now if I remember correctly,
they're actually, you know, zero or one
for the values because they're either
but we have them coming in as a float
value. And we have a little bit of a
shape coming in here. And there's our
784. Uh so we let it know that this is
what's what our input is for our
TensorFlow. And this is our training
set. So we'll just put a label on there
to help us uh track that train set. And
then W. And with W, we'll go ahead and
do TF variables. And we'll do this as uh
zeros variables TF zeros. And we'll set
this as as 784 by 10. 10 being the
output. 784 being our number of
variables in and this is our weights.
Remember we have a bias in there too.
And I'll go back over this in just a
second as we see how that fits together
in our tensorflow. And we'll do this one
um with our variables again. We have 10.
So we're going to do the bias. We're
going do it the same kind of format and
setup on here. And so we'll do that as
as TF zeros of 10. So we'll just create
an array of 10 there. And this is our
bias. So with these three lines um and
there's actually they're coming out with
the eager execution which would bypass
some of what we're doing. But this is
important to understand is the first
thing you have to do with TensorFlow is
we have to allocate a space for the
variables and our TF placeholder and our
TF variable with our weights and our
biases. This actually hasn't done
anything yet. So all it is is
placeholders. That's why it's okay to
use zeros. Um you could have just as
easily used ones or anything else and it
wouldn't matter. The next stage is to go
ahead and set up some of the functions
going on. But before we do that, just
note that this hasn't done anything.
Even if I execute it, all it's done is
created placeholders until we do the
final initialization. And so we need to
go ahead and set up. We'll do y=
tf.n.oftmax.
And the code for this is tf.mmoxw.
And this is our uh sum. Let's just put a
note here so we can keep track of what's
going on. We're finding weighted sum of
inputs plus the bias. Uh so there's our
plus b the bias and then we need to go
ahead keep um let's do y underscore and
again another placeholder and this one
we'll set um it actually we'll put in as
tf placeholder on here tf placeholder
float none 10. There's our one hot
encoder going on there. So our 10 values
coming out and we'll do a cross entropy
on here and this is going to be minus tf
reduce sum and we'll do y here's our y
underscore which is remember we have
your y output and your actual output. Uh
so this will be our y underscore time
the tf log of y. And then finally um
before we do the actual initialization
of all our variables we'll set up our
train step. This equals our gradient
descent optimizer. Very important.
Remember we looked at that chart on our
um uh slides and so we've set up all
these formulas and here's our gradient
descent optimizer and as it keeps
looking it keeps looking for that zero
value. That's what we're doing with that
particular formula. So let's take a look
and see what we're doing here. We just
put together all of our pieces for
TensorFlow. And you know the devil's in
the details. We have here our training
set coming in. We have to put a
placeholder on there. We have our uh
variables with their weights. We have
our biases coming out and then we put in
our uh the weighted sum. So here's
summizing our summation here. Then we
have our y variable output. So there's
our y um how it works and then of course
the actual output on there. And then we
have our cross entropy coming in and
that's our minus tf.reduce sum the y *
the tf log of y. And then the training
step gradient descent optimizer and
we're using a 0.01 in this and we're
going to minimize cross entropy. So,
we're going to let it do all the work.
So, once we've set up all of these
different layers, we've allocated for
them, we need to go ahead and initialize
them. So, we're going to do an init tf
initialize, and it's going to be all
variables. Uh, one of the cool things is
they're in the process of doing away
with this. So, all these steps would be
bundled into one instead of having to
have placeholders. You initialize them
in the same process going on. And then
finally, everything in TensorFlow is
based on your session. Now, this is
changing that there's other options to
be able to run this, but we want to go
ahead and do uh session. There's our TF.
There's a TF session. And then we want
to go ahead and do session run. And what
are we going to run? Well, we did
initialization of all our variables. Uh
so, this is what we're running. And this
is we're actually once we do this, we
actually create our TensorFlow object.
So, this whole piece of code right here
is our TensorFlow object. We have our
input coming in with our weighted
variables coming in. Our soft max for
our metal going out. How does it add it
together for our y value? Uh and then we
have the actual uh float value coming
out. Checking on our all the way down.
So you can see all the different stages
going through that we're setting up. Um
and this is one of the reasons that a
lot of people like TensorFlow is because
you can designate all these different
pieces one step at a time. This is also
one of the reasons people don't like
TensorFlow is because you have to
designate all the different layers
coming down and there's a lot of steps
being made right now to minimize this to
make it either easier to automate it or
to allow you to do more complicated
things and all those steps are still at
play. So it's worth looking into the
more advanced version what's going on
with KAS on top of TensorFlow. It's also
important to understand what's going on
in these individual levels if you're
going to play with them. It's important
to understand, hey, what's going on with
the uh finding the weighted sum of the
inputs plus the bias because there's
other ways to do that. There's all kinds
of other tools in there now, but this is
the basic setup that you want to do on a
TensorFlow coming in. And we want to go
ahead and just run and admit our
TensorFlow. So, let's go ahead and do
that. Let's run this. We do get a
warning here because uh there's a move
to use global variables. This is one of
the changes they're making, but it as
far as this example, it's not going to
make a difference because we're doing
once we initialize it. This is
initializing our variables. And again,
these are only placeholders up here
until we initialize them. And I would
highly suggest put a note down there or
or go over to simplylearn.com and let
them know and have them email you a copy
of the code. So, you can actually play
with this code right here because this
is the body of what's going on in
TensorFlow. This is the build in neural
networks. And then once we've done that,
now comes kind of the fun part is we
need to go ahead and train it. Uh so
we've created our TensorFlow, we've
created our uh network and now we need
to go ahead and train it. Uh so let's
put together that training code and
let's just do uh for I in range u 0 to
1,000. So we're just going to look at uh
the first 10,000 in our training. And
the way we pull that data from our mints
train next batch of 100. Uh so you look
at this. We're going to be doing groups
of 100 and then there's going to go
through a thousand of them. This is very
important that TensorFlow builds this
in. This is one of the downsides of
doing sklearn or one of the older
packages is they don't let you batch
groups in. Uh they wanted to have it all
up front and then you have to build your
own batch programs right now. Uh this
lets us go ahead and do that. And you
can see here we have batch x of s, batch
y of s. So there's our x and our y. You
could look at this as our training of X
and our train of Y or the data N and the
answer in. Uh and then we simply do our
session run. Uh so here's our session
that we've created. We're going to run
it and we want to do the train step. We
initialized our train step up here and
our TF. And so there's our train step
feed. It's a dictionary. Dictionary
coming in which we're going to create
right here. uh is x is our batch x of
our sample comma and our y underscore is
going to be our batch of our y sample.
Uh and so this goes through and we've
now hit the run button and we've trained
our session. We've trained this setup on
here. And once we've trained it, then we
need to go ahead and find out how good
our accuracy was and actually start
running some predictions through there.
Uh so we'll go ahead and create a a
correct prediction. And this is where
our tf.equal equal. We'll use our argmax
y of one and tf argmax of y of
underscore of one to help us get the
correct predictions on there. And then
we want to use that to feed into an
accuracy. And so our accuracy is going
to be tf reduce uh mean and we'll take
that and we'll do um a cast and this is
the correct prediction that we're
sending in there. And it is a uh float
value. Keep it simple. And let's go
ahead and print this out so we can see
what we're looking at. Uh so what are we
printing out? Uh we need to do a session
run. This session run is going to be on
the accuracy. Where did accuracy comes
from? This is our we're casting our TF
on there with the correct predictions on
that. So here's our accuracy feed in. So
it needs a dictionary for the data
coming in. We're going to create our
dictionary and x is going to be our nest
test images and y there is going to be
our nest.est
labels. Let me just double check and
make sure I have that typed in there
correctly. There we go. Oh, and let's go
ahead and run that and see what comes
up. And we end up with a N165
for our accuracy, which means our
trained neural network does a pretty
good job letting us know what these
different symbols are in guessing that a
seven and a three and a four, uh,
something that as humans we kind of take
for granted. I even have trouble reading
this. So, I don't know if I would be
able like that first one, I would sit
there for a long time figuring out
that's a seven versus a two. That could
have easily been a two to me. Sing seem
to do a pretty good job analyzing this
data. And this is used to analyze
something very complicated on these
images, very different than uh just a
straight value of uh cost of sales and
here's our return and our marketing. Uh
we can now create this nice neural
network that does all kinds of cool
things. Do you know friends that
according to the lending statistics the
demand for AI and ML specialist is
projected to surge by 40% between 2023
to 2027.
And on an average, an ML engineer is
expected to earn around 133 and $336 per
year. So if you are an aspiring ML
engineer and thinking about what
innovative projects you can show in your
portfolio, then your wait is over cuz in
this video I'll be covering eight
amazing ML projects that you can
showcase in your resume. So guys, let's
start first with a beginner level
project and the first project that we
are going to encounter that is home
value prediction. So guys, this project
aims to develop a predictive model to
estimate the value of residential
properties. The model will analyze
various features such as location,
square, footage, number of bedrooms and
bathrooms, age of the property and other
relevant factors. By leveraging
historical property data, the model will
be able to provide accurate home value
predictions which can be useful for real
estate agents, buyers and sellers. So
guys, the programming language that we
are going to use all over here will be
Python and machine learning libraries
that we will be using will be
scikitlearn, tensorflow, kas and for
data handling libraries we have pandas,
numpy and for visualization we have to
use mattplot and seabon. Now what will
be the approach for this one guys? So
guys the first one that we have a data
collection. So here what is going to
happen guys? So first you have to
collect the historical property data
from the sources like Zillow
retailer.com. You can also get database
from the public real estate databases
like Kaggle data sets where you have
Zillow home value prediction. Ensure
that the data set include features like
location where you have latitude,
longitude, square footage, number of
rooms, year built, property type and
previous sales. The next step that comes
is data cleaning. You have to handle the
missing values by using imputation
techniques or removing incomplete
records. Removing outliers that may skew
the model's prediction, normalize or
standardize the data to ensure
consistency. The third one that we have
is feature engineering. You have to
create new features such as proximity to
schools, crime rates and access to the
public transportation. Encode categorial
variables, example property type,
location using techniques like one hot
encoding. Generate interaction features
that capture relationship between
existing features. The fourth one that
we have is model selection. Use
regression models like linear
regression, random forest, gradient
boosting, neural networks. Experiment
with different models to identify the
best performing one. Now in the next
phase all you have to do guys is model
training and evaluation. Split the data
set into training and test sets. Train
the model on a training set and evaluate
their performance on the testing set
using metrics like RSM which means root
mean squared error. You can use cross
validation to ensure the model's
robustness and avoid overfitting. The
sixth one that we have all over here is
hyperparameter tuning. You can optimize
the model's hyperparameter using
techniques such as grid search or random
search to improve accuracy. And if
you're looking forward to deploy your
model, then you can develop a web
interface using flask or Django to allow
users to input property features and get
predictions. You can deploy the model on
the cloud platform like AWS for
scalability.
Now if we talk about the complexity
level of this, we all know that it is a
beginner level project. Now let us move
on to the one more set that is music
genre classification and generation. So
guys this is also one of the most
beginner level project. This project
aims to develop a system that can
classify music tracks into different
genres and generate new music
composition within specified genre. The
goal is to build a model that analyzes
audio features to categorize music and
uses deep learning techniques to create
new music. This project introduces
advanced concept of audio processing,
deep learning and generative models. So
guys, what will be used in this? So
we'll have programming language that
will be Python. For audio processing,
we'll be using librosa. For machine
learning libraries, we'll be using
tensorflow, kas, pytorch. For data
handling libraries, we'll be using
pandas, numpy. For visualization, we'll
be using mattplot, seabon. For the data
set guys, you can use gtzan music genre
data set or you can get it from free
music archive. So guys in the first
phase we are going to have data
collection. You can obtain data sets
containing music tracks and their
corresponding genre from the sources
from GTN music genre data set and the
free music archive. Ensure that a data
set includes diverse genre and
substantial number of tracks per genre.
Next we'll go for data prep-processing.
Use library librosa to load and
pre-process audio files including
feature extraction such as mil
frequency, septal coefficients, chroma
features and spectral contrast. You can
normalize the extracted features to
ensure consistent input for the given
models. Now if you talk about feature
engineering guys, you can extract
additional features from the audio files
such as tempo, beat, zero crossing rate
etc. Create a feature matrix that
represents the extracted audio features.
Then go for the model selection. Use
conventional neural network or recurrent
neural networks for the music genre
classification. Split the data set into
training and testing data set. Now if we
talk about model training and evaluation
guys then you can train the selected
classification model on the training
set. Evaluate the model's performance on
the testing set using metrics like
accuracy, precision, recall and F1
score. Use confusion matrices to
understand the classification
performance across different genres. Now
if we talk about model selection and
training for the music generation, what
you will do guys? You can use the
generative adversial networks or
recurrent neural networks such as LSTM,
long short-term memory for music
generation. Train the generative model
on the data set to create new music
sequences. Next, we have model training
and evaluation. You can train the
generative model on sequences of audio
features. You can evaluate the generated
music by listening tests and by
objective metrics like inception score
or fche audio distance. If I talk about
hyperparameter tuning guides, we can
optimize the model. You can use the
hyperparameters using techniques like
grid search or random search to improve
performance. If I talk about deployment
guys, you can deploy these models on
cloud platforms like AWS. Now let us
move on to our next project. So guys,
the complexity of this project is at the
beginner level. Now let us move to the
intermediate level projects. Next
project that we have all over here is
sentiment analysis of Twitter data. This
project aims to develop a sentiment
analysis model that can classify to its
side positive, negative or neutral. The
goal is to analyze public sentiment on
various topics or events using natural
language techniques. So guys, what will
be used all over here? So in this we
will have programming language like
Python. Okay. NLP libraries, NLTK spacy.
For machine learning libraries you can
use scikitlearn, tensorflow, kas. For
data handling libraries we have pandas,
numpy. For visualization we have
mattplot lilip seon and we can use the
API Twitter for data collection. Now how
you going to work on it guys? So guys if
I talk about the data collection use a
Twitter API to collect tweets based on
specific hashtags like keywords or
topics. Extract relevant fields like
tweet text, user information, timestamp
etc. Then if I talk about data
prep-processing guys, clean the tweet
text by removing special characters,
links, mentions, hashtags and stop
words. Tokenize the text and perform
limization or stemming to reduce the
words to their base form. Next, if you
talk about feature engineering guys, you
convert the clean text data into
numerical representation using TF, back
of words or word embedding. Now if I
talk about model selection guys, you can
choose a classification algorithm such
as logistic regression, n bias or LSDM.
Split the data set into training and
testing data set. Now if I talk about
model training and evaluation, then you
can train the selected model on the
training set. Evaluate the model's
performance on the testing set using
metrics like accuracy, precision,
recall, and fn score. Use cross
validation to ensure the model's
robustness. If I talk about
hyperparameter tuning guys, you can
optimize the model's hyperparameters
using grid search or random search to
improve the performance for deployment
which can be optional. You can deploy
your model on AWS for real-time
sentiment analysis. So if I talk about
the complexity level guys, its
complexity is intermediate. So guys, our
next project is customer segmentation
using K means clustering. This project
aims to segment customers into distinct
groups based on their purchasing
behavior and demographic information.
The objective is to understand customer
segments and tailor marketing strategies
accordingly. So guys, what programming
languages we'll be using? So basically
we'll be using Python. For machine
learning libraries, we will have
scikitlearn. For data handling
libraries, we'll have pandas, numpy. For
visualization libraries, we'll have
mattplot lab, seabon. And the data set
source will be e-commerce transaction
data. how we are going to work on this
one. For data collection, we can obtain
a data set of e-commerce transactions
that include customer demographics,
purchase history, and product
information. Next, we'll have data
prep-processing. For data
prep-processing, we are going to do the
cleaning of the data by handling the
missing values and outliers. Then, for
feature engineering, we are going to
create features like total purchase
amount, purchase frequency, and recency
of the purchases.
Then, we are going to proceed for the
model selection. You can use K means
clustering to segment the customers into
distinct groups. You can determine the
optimal number of clusters using methods
like album methods or silhou.
Now if I talk about model training and
evaluation, you can train the K means
model on the process data set. You can
evaluate the quality of clusters by
analyzing intracluster and intercluster
distances. Next we have the evaluation.
You can visualize the clusters using
techniques like PCA, principal component
analysis, TSN etc. Next we have the
hyperparameter tuning. Now now you can
tune this model and interpret the
characteristic of each segment. You can
develop a target marketing strategies
for each segment based on unique
behavior and preferences. Now deployment
is optional. You can develop a dashboard
using flask or Django to visualize
customer segments and track marketing
campaigns. So guys if I talk about the
complexity of this project. So this is
an intermediate level project. So guys
for data set you can use the Kaggle's
customer segmentation data set which is
available at the Kaggle's platform. Now
the third intermediate level project
that we have all over here is building a
chatbot with Rasa. This project aims to
build an intelligent chatbot using Rasa
framework. The chatbot will be capable
of understanding user queries and
providing appropriate responses making
it useful for customer support, personal
assistance or information retrieval.
What languages we are going to use? So
it will be Python based. We'll have the
NLP libraries like Rasa, NLTK, Spacey.
For machine learning libraries, we are
going to have scikitlearn, tensorflow,
kas, etc. For data handling libraries,
we are going to use pandas, numpy. So
guys, this was what we are going to do
it and how you can work on this one by
collecting the data, collect the
conversation data and FAQs from the
target domain, annotate the data to
create training examples for the
chatbot. Next comes is data
prep-processing. Clean the text data by
removing special characters and
normalizing the text. You can tokenize
and limitize the text to prepare for a
training. The third one we have the
models training. You can use Rasa's
NLU's component to train a model for
intent recognition and entity
extraction. You can define a dialog
management policies to handle different
conversation flows. Next guys, you can
perform the feature engineering and
integration. You can integrate the Ras
NLU and core components to build
complete chatbot. You can connect the
chatbot to messaging platform like
Facebook Messenger etc. For model
selection and testing what you can do
guys you can test this chatbot with
various inputs to ensure that it handles
the scenarios appropriately and you can
also select the right model using this.
Now if I talk about model training guys
what you have to do you have to collect
the user feedback and conversational
logs to continuously work on training
the model. Next, similarly you have to
retrain the model periodically with the
new data to see how it is working. So
that will be your evaluation. Now for
the hyperparameter tuning, what are you
going to do guys? You have to check in
those scenarios where it is able to tune
up with those scenarios where it can
handle the input appropriately. And next
is deployment. So guys, for deploying
it, you can use AWS. So guys, for a data
set, you can use Ras open source. So
that's a very good data set for you to
proceed. So guys, if you talk about
difficulty of this project, this is an
intermediate level project. Now let us
move on to the advanced level projects.
For advanced level projects, the first
one that comes up to my mind is movie
similarity from plot summaries. Now this
project aims to develop a system that
can find out recommended movies similar
to a given movie based on their plot
summaries. By analyzing the textual
content of the movie plot summaries, the
model will identify similarities and
suggests movies with similar themes,
story lines or genres. This project
introduces beginners to natural language
processing text similarly measures and
recommendation systems. What languages
we are going to use guys? We'll be using
Python NLP libraries like NLTK spacy
machine learning libraries like
scikitlearn data handling libraries like
pandas numpy visualization we can use
numpy and data set source will be IMDb
or kegel so guys this process is also
involving the data collection then you
have to go for data cleaning then
feature engineering next model selection
so similar process as I have discussed
in other projects so you have to also go
through the same one next what you have
to do device. Similarly, what you have
to do, you have to train the model, then
evaluate the model, then hypertune it
and finally proceed for the deployment.
So, this is overall process of this
project. Try to research on the website
a lot like how you can extract it. So,
guys, you can use kegel or towards data
science to research more about this
project. Now guys, if I talk about the
difficulty of this project and this is
an advanced level project. Now let us
move on to the next one that we have all
over here that is image segmentation
project for brain tumor prognosis. This
is a very very amazing project and
definitely you can put up on your
portfolio. Basically guys this project
aims to develop an image segmentation
model to identify and delinate brain
tumors from MRI scans. The goal is to
accurately segment the tumor regions
which can aid in prognosis treatment
planning surgical interventions. This
project introduces intermediate level
concepts of computer visions, deep
learning and medical image analysis. So
guys, what we'll be using all over here
for programming languages we can use
python for deep learning libraries we
can use tensorflow, kas, pytor. For
image processing libraries, we can use
opencv, scikit image. For data handling
libraries, we can use pandas, numpy. For
visualization, we can use mattplot,
cbond. Now if I talk about what is the
process of developing this project the
first step will be same data collection.
So next step you have to go for data
prep-processing. Third step you have to
do the model selection where you can use
CNN models for image segmentation task.
Then you go for model selection. Moving
ahead you're going to have the model
training and evaluation. You have to
split the data set into training and
validation and testing data sets. Next
proceed for the evaluation phase. Okay,
evaluate the model with certain metrics.
So here I can give you certain idea like
you can use dice coefficient,
intersection over union or accuracy.
Then go for hyperparameter tuning where
you have to optimize the model's
hyperparameters.
You can use grid search or random search
as we have discussed. And finally you
can deploy this model on AWS. Now guys
we have come to the final project. This
is also a very amazing project guys. So
guys the complexity level of this
project is advanced level.
Now let us move on to our final project
that is the impact of climate change on
birds. This is a very very amazing
project and definitely you can add it on
your resume. This project aims to
analyze the impact of climate change on
the bird population and migration
patterns by examining various climatic
factors and their correlation with bird
species data. The project seeks to
predict how climate change might affect
bird behavior and distribution. This
project will introduce you some advanced
level concepts like time series
analysis, environmental data modeling,
etc. So guys, what programming languages
we'll be using for data analysis? You
can see we'll have pandas, numpy. For
machine learning libraries, we are going
to have scikitlearn, tensorflow. For
visualization, we're going to have
mattplot lab, plotly. For geospatial, we
are going to have geopandas, folium. For
data source, we're going to have public
data sets on bird observation and
climate data sources from eird. Now what
will the process flow for this one guys?
First you have to proceed for data
collection. Gather bird observation from
the data like EIRD which provides
extensive record of bird sightings.
Okay. And for climate data you can
collect it from NOA including
temperature, precipitation and other
relevant climatic factors over the time.
And similar next process will be the
data prep-processing. Then you have to
proceed for feature engineering. Then
you have to go for model selection.
Okay. Moving ahead you have to go for
model training. then evaluation, then
hyperparameter tuning and finally you
have to deploy the model. So research
about this project, see what models you
are going to use. Suppose I can give you
a hint about this. You can use time
series analysis models like ARMA or ML
models. You can also use random forest
or gradient boosting for predicting
impact on the bird population. So guys,
use Google exhaustively to research
about this project. This is also a very
amazing project and it's going to give
you a lot of idea. Now if I talk about
the complexity level of this, it is an
advanced level project.
>> Welcome to deep learning interview
questions. My name is Richard Kersner
with the SimplyLearn team. That's
www.simplearn.com.
Get certified, get ahead. Today we're
going to help you prepare for interview
questions dealing with deep learning.
And we're going to go from the very
basics of neural networks and deep
learning into some of the more commonly
used models so you can have an
understanding of what kind of questions
are going to come up and what you need
to know in interview questions. We'll
start with a very general concept of
what is deep learning. This is where we
take large volumes of data in this case
on cats and dogs or whatever. A lot of
times you use um a training setup to
train your model. Remember it's kind of
like a magic black box going on there.
And then we use that to extract features
or extract information and in this case
classify the image of a cat and a dog.
So the primary takeaway we're talking
about deep learning is it learns from
large volumes of structured and even
unstructured data and uses complex
algorithms to train neural network. It
also performs complex operations to
extract hidden patterns and features.
And if we're going to discuss deep
learning in this very uh simplified
overview and we also have to go over
what is a neural network. This is a
common image you'll see of a drawing of
a forward propagation neural network and
it's it's a human brain inspired system
which replicate the way humans learn. So
this has inspired how our own neurons
and our brain fire but at a much
simplified level. Obviously it's not
ready to take over the human uh
population and and be our leader yet.
Not for many years. It's very much in
its infant stage. But it's inspired by
how our brains work. Um and they use a
lot of other inspirations. You can study
brains of moths and other animals that
they've used to figure out how to
improve these neural networks. The most
common one consists of three layers of
network and this is generally how you
view these networks is you have an
input, you have a hidden layer and an
output. And the neural network is uh
broken up into many pieces. But when we
focus just on the neural network, it's
always on the hidden layers that we're
making all the adjustments and figuring
out how to best set up those hidden
layers for their functions to both train
faster and to function better. When we
look at this, of course, we have our
input, hidden, and output. Each layer
contains neurons called as nodes perform
various operations. And you can see here
we have the list of the nodes. We have
both our input nodes and our output
nodes and then our hidden layer nodes.
And it's used in deep learning algorithm
like CNN, RNN, GN, etc. We'll address
some of these models a little closer, at
least the most common models as we go
down the list and we study the deep
learning and the neural network
framework. Let's start with what is a
multi-layer perceptron or MLP a lot of
time as they're referred to. And you'll
see these abbreviations. I'll be honest,
I have to write them down on a piece of
paper and go through them because I
never remember what they all mean even
though I play with them all the time.
What is a multi-layer perceptron? Well,
if you look at the image on the right,
it's very similar to what we just looked
at. You have your input layer, your
hidden layer, and your output layer. And
that's exactly what this is. It has the
same structure of a single layer
perceptron with one or more hidden
layers except the input layer, each node
in the other layers uses a nonlinear
activation function. What that means is
your input layer is your data coming in
and then your activation function is
based upon all those nodes and weights
being added together and then it has the
output. MLP uses supervised learning
method called back propagation for
training the model. Very key word there
is back propagation. Single layer
perceptron can classify only linear
separable classes with binary output 01.
But the MLP can classify nonlinear
classes. So let's break this down just a
little bit. The multi-layer perceptron
with an input layer and a hidden layer
and an output layer. As you see that it
comes in there, it has adds up all the
numbers and weights depending on how
your setup is. That then goes to the
next layer. That then goes to the next
hidden layer if you have multiple hidden
layers. And finally to the output layer.
The back propagation takes the error
that it sees. So whatever the output is,
it says, hey, this has an error to it.
It's wrong. And then sends that error
backwards from where it came from. And
there's a lot of different functions
used to uh train this based on that
error and how that error goes backwards
in the notes. Uh so forward is you get
your answers. Backward is for training.
You see this every day. Even my uh
Google Pixel phone has this. It they
train the neural network which takes a
lot more data to train than it does to
use. And then they load up that neural
network into in this case I have a Pixel
2 which actually has a built-in neural
network for processing pictures. And so
it's just the forward propagation I use
when it processes my photos, but when
they were training it, you use the back
propagation to train it with the errors
they had. We'll be coming back to
different models that are used. For
right now though, multi-layer
perceptron, MLP, put that down as your
vocabulary word and of course back
propagation. What is data normalization
and why do we need it? This is so
important. We spend so much time in
normalizing our data and getting our
data clean and setting it up. Uh so we
talk about data there's a pre-processing
step to standardize the data. So
whatever we have coming in we don't want
it to be a uh you know one gigabyte file
here a 2 GBTE picture here and a 3
kilobyte text there. Even as a human I
can't process those all in the same
group. I have to reformat them in some
way that loops them together so they're
a standardized format. We use this uh
data normalization and and
pre-processing to reduce and eliminate
data redundancy. A lot of times the data
comes in and you end up with two of the
same images or uh uh the same
information in different formats. Then
we want to rescale values to fit into a
particular range for achieving better
convergence. What this means is with
most neural networks they form a bias.
We've seen this in recently in attacks
on neural networks where they light up
one pixel or one piece of the view and
it skews the whole answer. So suddenly u
because one pixel is really bright uh it
doesn't know what to do. Well when we
start rescaling it we put all the values
between say minus one and one and we
change them and refit them to those
values. It helps get rid of that bias
helps fix for some of those problems.
And then finally we restructure the data
and improve the integrity. We want to
make sure that we're not missing values
um or we don't have partial data coming
in. One way to look at this is uh bad
data in bad data out. And so you want
clean data in and you want good answers
coming out. One of the most basic models
used is a Boltzman machine. So let's
address what is a Boltzman machine. And
if you know we just did the MLP
multi-layer perceptron. So now we're
going to come into almost a simplified
version of that. And in this we have our
visible input layer and we have our
hidden layer. The Boltzman machines are
almost always shallow. They're usually
just two-layer neural nets that make
stochastic decisions whether a neuron
should be on or off. True or false? Yes.
No. First layer is a visible layer and
second layer is the hidden layer. Nodes
are connected to each other across
layers, but no two nodes of the same
layer are connected. Hence, it is also
known as restricted Boltzman machine.
Now that we've covered a basic MLP or
multi-layer perceptron, and we've gone
over the Boltzman machine, also known as
the restricted Boltzman machine, let's
talk a little bit about activation
formulas. And this is a huge topic that
can get really complicated but it also
is automated. So it's very simple. So
you have both a complicated and a simple
at the same time. So what is the role of
activation functions in a neural
network? Activation function decides
whether a neuron should be fired or not.
That's the most basic one and that
actually changes a little bit because
it's either whether fired or not in this
case activation function or what value
should come out when it's fired. But in
these models, we're looking at just the
boltsman restricted layers. So this is
what causes them to fire. Either they
don't or they do. It's a yes or no,
true, false, all or nothing. It accepts
the weighted sum of the inputs, the bias
as input to any activation function. So
whatever activation function is, it
needs to have the sum of the weights
times the input. So each input, if you
remember on that model, and let's just
go back to that model real quick. And
then you always have to add a bias. And
you can look at the bias if you remember
from your uklitian geometry. You draw a
straight line. Formula for that line has
a y-coordinate at the end. It's always
um cx plus m or something like that
where m is where it crosses the
ycoordinates. If you're doing a straight
line with these weights, it's very
similar, but a lot of times we just add
it in as its own weight. We take it as a
node of a one value coming in and then
we compute its new weight. And that's
how we compute that bias just like we
compute all the other weights coming in.
The node which gets fired depends on the
y value. And then we have a step
function. And the step function this is
where remember I said it's going to get
complicated and simple all at the same
time. We have a lot of different step
functions. We have the sigmoid function.
We have just a standard step function.
We have the ru is pronounced like ray
the ray of from the sun and lu like a
name. So ru function. And we have the
tangent h function. And if you look at
these, they all have something similar.
They all either force it to be um one
value or the other. They force it to be
in the case of the first three a zero or
one. And in the last one, it's either a
minus one or one. And you can easily
convert that to a 0, one, yes, no, true,
false. And on this, one of the most
common ones is the step function itself
because there is no middle value. There
is no um uh discrepancy that says, well,
I'm not quite sure. But as you get into
different models, probably the most
commonly used used to be the sigmoid was
most commonly used, but I see the relu
used more often. Really, depending on
what you're doing, you just have to play
with these and find out which one works
best depending on the data in your
output. The reason to have a non01
answer or something kind of in the
middle is when you're looking at this
and it's coming out, you can actually
process that middle ground as part of
the answer into another neural network.
So it might be that the relu function
says hey this is only a 6 not a one and
uh even though the one is what's going
into the next neural network or the next
hidden layer as an input the 6 value
might also be going in there to let you
know hey this is not a straight up one
or straight up zero it's someplace in
the middle this is a little uncertain
what's coming out here so it's a very
powerful tool in the basic neural
network you usually just use the step
function it's yes or no let's take a um
a big step back and take a kind of an
overview. The next function is what is a
cost function that we're going to cover.
This is so important because this is
your end result that you're going to do
over and over again and use to decide
whether the model is working or not,
whether you need to try a different step
function, whether you need to try a
different activation, whether you need
to try a fully different model used. Uh
so what is the cost function? Cost
function is a measure to evaluate how
good your model's performance is. It is
also referred as loss or error used to
compute the error of the output layer
during back propagation. There's our
back propagation where we're training
our model. That's one of our key words.
Mean squared error is an example of a
popular cost function. And so here we
have the cost function C = half of Y - Y
predicted. Um and then you square that.
So the first thing is um you know real
quick if you haven't done statistics
this is not a percentage. It's not a
percentage of how accurate it is. is
just a measurement of the error and we
take that error if we're training it and
we push that error backwards through the
neural network and we use that through
the different training functions
depending on what model you're using to
train the neural network. So when you
deploy the network you're usually done
training it because it takes a lot of
computational force to train it. Um this
is a very simple model and so you deploy
the train one. Uh but we want to know
how your error is and so how do we do
that? Well you split your data. part of
your data is for trading and part of
your data is for testing. And then we
can also test the error on there. So
it's very important. And then we're
going to go one more step on this. We
got to look at both the local and the
global setup. It might work great to
test your data on what you have on your
computer, but that's different than in
the field. So, when we're talking about
all these different tests and the error
test as far as your loss, you don't you
want to make sure that you're in a
closed environment when you do initial
testing, but you also want to open that
up and make sure you follow up with the
testing on the larger scale of data
because it will change. It might not fit
the larger scale. There might be
something in there in the way you
brought the data in specifically or the
data group you used or um any of those
could cause an error. So, it's very
important to remember that we're looking
at both the local and the global context
of our error. And just one other side
note on a lot of the newer models of
neural networks by comparing the error
we get on the data our training data
with a portion of the test data we can
actually figure out how good the model
is whether it's overfitted or not. We'll
go into that a little bit more as we go
into some of the different models. So we
have our output. We're able to um figure
out the error on it based on the square
means usually although there's other uh
functions used. So we want to talk about
what is gradient descent? Another
vocabulary word gradient descent is an
optimation algorithm to minimize the
cost function or to minimize the error.
Aim is to find the local or global
minima of a function. Determine the
direction the model should take to
reduce the error. So as we're looking at
this, we have our uh squared error that
we just figured out the co based on the
cost function. It says how bad is my
model fitting the data I just put
through it. And then we want to reduce
that error. So how do you figure out
what direction to do that in? Well, it
could be that you're looking at just
that line of that line of data coming
in. So that would be a local minima. We
want to know the error of that
particular setup coming in. And then you
have your global your global minima. We
want to minimize it based on the overall
data we're putting through it. And with
this we can figure out the global
minimum cost. We want to take all those
local minimum costs of each piece of
data coming in and figure out the global
one. How are we going to adjust this
model to fit all the data? We don't want
it to be biased just on three or four
lines of data coming in. We want it to
kind of extrapolate a general answer for
all the data coming in. But this of
course uh we mentioned it briefly about
back propagation. This is where really
comes in handy is training our model.
Neural network technique to minimize the
cost function helps to improve the
performance of the network. Back
propagates the error and updates the
weights to reduce the error. So as you
can see here is a very nice depiction of
a back propagation. We have our
predicted y coming out and then we have
since it's a training set we already
know the answer and the answer comes
back and based on case of the square
means was one of the functions we looked
at uh one of the activation functions
based on cost function that cost
function then depending on what you
choose for your back propagation method
and there's a number of them will change
the weights it will change the weight
going to each of one of those nodes in
the hidden layer and then based upon the
error that's still being carried back
it'll change the weights going to the
next hidden layer and then it computes
an error level on that and sends that
back up. And you're going to say, well,
if it computes the error into the first
hidden layer and fixes it, why would it
stop there? Well, remember, we don't
want to create a biased neural network.
So, we only make small adjustments on
these weights. We don't make a big
adjustment that changes everything right
off the bat. So, no matter how far back
you go, you're always going to have a
small amount of error, and that's still
going to continue to go all the way back
up the hidden layers. For right now,
focus on the back propagation is taking
that error and moving it backwards on
the neural network to change the weights
and help program it so that it'll have
the correct answers. So far, we've been
talking about forward propagation neural
networks. Everything goes forwards, goes
left to right. Uh but let's let's take a
little detour and let's see what is the
difference between a feed forward neural
network and a recurrent neural network.
Now, this is in the function, not when
we're training it using the back
propagation. So, you've got new
information coming in and you want to
get the answer and there's a couple
different networks out there and we want
to know we have a feed forward neural
network and we have a new uh vocabulary
term recurrent neural network. A feed
forward neural network signals travel in
one direction from input to output. No
feedback loops considers only the
current input cannot memorize previous
inputs. One example of one of these feed
forward neural networks. And we've
covered a number of them, but one of the
ones that has a big highlight nowadays
is the CNN, a convolutional neural
network. TensorFlow, the one put out by
Google is probably most known for their
CNN, where the information goes forward.
It uh first takes a picture, splits it
apart, goes through the individual
pixels on the picture, so it picks up a
different reading, then calculates based
on that, goes into a regular feed
forward neural network, and then gives
you a categorization on there. Now,
we're not covering the CNN today, but we
do have a video out that you can look up
on YouTube put out by SimplyLearn, the
convolutional neural network. wonderful
tutorial. Check that out and learn a lot
more about the convolutional neural
network. But you do need to know that
the CNN is a forward propagation neural
network only. So it's only moving in one
direction. So we want to look at a
recurrent neural network. Signals travel
in both directions making it a looped
network. Considers the current input
along with the previous received inputs
for generating the output of a layer.
Has the ability to memorize past data
due to its internal memory. And you can
see they have a nice uh image here. We
have our um input and for some reason
they always do the recurrent neural
network um in reverse from bottom up in
the images. It's kind of a standard
although I'm not sure why. Your X goes
into your hidden layer and your hidden
layer the answer for part of the answer
from that it generates feeds back into
the hidden layer. So now you have an
input of both X and part of the hidden
layer and then that feeds into your
output. Now if we go back to the forward
let me just go back a slide and we're
looking at uh our forward propagation
network. One of the tricks you can do to
use just a forward propagation network
is if you're in a what they call a time
sequence, that's a good uh term to
remember or a time series meaning that
it's sequential data. Each term comes
after the other. You can trick this by
creating your input nodes as with the
history. So if you know that uh you have
values one, five and seven going in and
you know what the output is from one
what those outputs are, you can expand
the input to include the history input.
That's one of the ways to trick a
forward propagation network into looking
at that. But when you do with a
recurrent neural network, you let the
hidden layer do that for you. It sends
that data and reprocesses it back into
itself. What are some of the
applications of recurrent neural
network? The RNN can be used for
sentiment analysis and text mining.
Getting up early in the morning is good
for health and it's a positive
sentiment. One of the catches you really
want to look at this when you're looking
at the language is that I could switch
this around and totally negate the
meaning of what I'm doing. So, it no
longer be positive. So, when you're
looking at a sentence, knowing the order
of the words is as important as the
meaning of the words. You can't just
count how many good words there are
versus bad words to get positive
sentiment. You know, have to know what
they're addressing. And there's lots of
other different uses. Uh, kids are
playing football or soccer as we call it
in the US. RN can help you caption an
image. So based on previous information
coming in, it refeeds that back in and
you have a image setter. And then time
series problems like predicting the
prices of stocks in a month or quarter
or sell of product can be solved using
an RNN. And this is a really good
example. You have whatever your stocks
were doing earlier this month will have
a huge effect of what they're doing
today if you're investing. So having an
RNN model, a recurrent neural network
feeding into itself what was happening
previously allows it to take that model
and program in that whole series without
having to put in the whole a month at a
time of data. You can only put in one
day at a time. But if you keep them in
order, it will look back and say, "Oh,
this because of what happened yesterday,
I need some information from that and
I'm going to use that to help predict
today's." And so on and so on. We're
going to go back to our activation
functions. Remember I told you uh ReLU
was one of the most common functions
used. Uh so let's talk a little bit more
about ReLU and also softmax. Softmax is
an activation function that generates
the output between zero and one. It
divides each output such that the total
sum of the outputs is equal to one. It
is often used in the output layers.
Softmax L of the N equals E to L the N
over the absolute value of E to the L.
So what does this function mean? I mean
what is actually going on here? So we
have our uh output nodes and our output
nodes are giving us uh let's say they
gave us 1.2.9 and point4. As a human
being I look at that and I say well the
greatest value is 1.2. So whatever
category that is if you have three
different categories maybe you're not
just doing if it's a cat or it's a dog
or u oh let's say it's a cow. We had
cats and dogs earlier. Why the cats and
dogs are hanging out with a cow. I don't
know. But we have a value and it might
say 1.2 2 is a cat, 0.9 is the dog, and
point4 is a cow. Uh, for some reason, it
thinks that there's a chance of it being
any one of these three items, and that's
how it comes out of the output layer.
Well, as a human, I can look at 1.2 and
say this is definitely what it is. It's
definitely a cat or whatever it is. Uh,
maybe it's looking at different kinds of
cars might be a better whether it's a
car, truck, or a motorcycle. Maybe
that'd be a better example. Well, from a
computer standpoint, that might be a
little confusing because they're just
numbers waving at us. And so with the
soft max, we want all those numbers to
always add up to one. So when I add
three numbers together, I want the final
output to be one on there. And so it
goes through this formula changes each
of these numbers. In this case, it
changes them to 46.34
and 2. They all add up to one. And
that's a lot easier to register because
it's very set. It's a set output. It's
never going to be more than one. It's
never going to be less than zero. And so
you can see here that there's probably a
pretty high chance that it's the first
one. So you're as a human being, we have
no problem knowing that. But this output
can then also go into say another input.
So it might be an automated car that's
picking up images and it says that image
in front of us is probably a big truck.
We should deal with it like it's a big
truck. It's probably not a motorcycle.
Um or whatever those categories are.
That's the softmax part of it. But now
we have the ru. Well, what where's the
ru coming from? Well, the ru is what's
generating the 1.2 and the 0.9 and the
point4. And so if you remember our relu
stands for rectified linear unit and is
the most widely used activation
function. We looked at a number of
different activation functions including
tangent h the step function. Remember I
said the step function is really used if
that's what your actual output is
because then you know it's a zero or
one. But the relu if you have that as
your output you now have a discrepancy
in there. And if that's going into
another neural network or another
process having that discrepancy is
really important. and it gives an output
of x if x is positive and zero
otherwise. So it says my x value is
going to be somewhere between zero or
one and then the uh usually unless it's
really uncertain the output's usually a
one or zero and then you have that
little piece of uncertainty there that
you can send forward to another network
or you can look at to know that there's
uncertainty involved and is often used
in the hidden layers. This is what's
coming out of the hidden layers into the
output layer usually or as we reference
the uh convolution neural network the
CNN you'd have to go to another video to
review the RLU is the most common used
for convolutional part of that network
has a bunch of little pieces that are
very simplified looking at all the
different images or different sections
of the map and the RLU works really good
for that like I said there's other
formulas used but that this is the most
common one and you'll see that in the
hidden layers going maybe between one
layer and the next layer. So just a
quick recap, we have our soft max, which
means that if you have uh numerous
categories, only one of them is going to
be picked, but you also want to have
some value attached to it, how well it
picked it, and you put that between 01.
So it's very uh standardized. So we have
our soft max. We looked at that. Let's
go back one. We looked at that here
where it transforms the numbers. And
then we have our ReLU function which
takes the information in the summation
and puts it between a zero and a one
where it's either clearly a zero or
depending on how confident our model is,
it'll go between the zero and one value.
What are hyperparameters? Oh, this is a
great interview question.
Hyperparameters. When you are doing
neural networks, this is what you're
playing with most of the time once
you've gotten the data formatted
correctly. A hyperparameter is a
parameter whose value is set before the
learning process begins. Determines how
a network is trained and the structure
of the network. This includes things
like the number of hidden units, how
many hidden layers are you going to have
and how many nodes in each layer.
Learning rate. Learning rate is usually
multiplied once you figured out the
error and how much you want to change
the weights. We talked about or I
mentioned it earlier just briefly. You
don't want to just make a huge change
otherwise you're going to have a biased
model. So you only take little
incremental changes and that's what the
learning rate is is those small
incremental changes. Epics, how many
times are you going to go through all
the data in your training set? So one
epic is one trip through all the data.
And there's a lot of other things
depending on which model you're working
with and which programming script you're
working with. Like the Python sklearn
package will have it slightly different
than say Google's TensorFlow package
which will be a little bit different
than the Spark machine learning package.
So these are just some examples of the
hyperparameters. And so you see in here
we have a nice image of our data coming
in and we train our model. Then we do a
comparison to see how good our model is.
And then we go back and we say, "Hey,
this this model's pretty good, but it's
biased." So then we send it back and we
change our hyperparameters to see if we
can get an unbiased model or we can have
a better prediction on it that matches
our data closer. What will happen if
learning rate is set too low or too
high? We have a nice couple graphs here.
We have one over here. It says a
learning rate set too low. And you can
see that it slowly works its way down
the curve. And on the right you can see
a learning rate set too high. It's just
bouncing back and forth. When your
learning rate is too low, that's what we
studied two slides ago. That's what the
learning rate was. Training of the model
will progress very slowly as we are
making very tiny updates to the weights.
We'll take many updates before reaching
the minimum point. So I just mentioned
epic going through all the data. You
might have to go through all the data a
thousand times instead of 500 times for
it to train. Learning rate too high
causes undesirable divergent behavior to
the loss function due to drastic updates
and weights. At times it may fail to
converge or even diverge. So if you have
your learning rate set too high and it's
training too quickly, maybe you'll get
lucky and it trains after one epic run,
but a lot of times it might never be
able to train because the weights are
changing too fast. They they flip back
and forth too easy. And you see down
here we've introduced uh two new terms
converge and diverge. Converge means
that our model has reached a point where
it's able to give a fairly good answer
for all the data we put in. All those
weights have adjusted and it's minimized
the error. Diverge means that the data
is so chaotic that it can never manage
to to train to that data. The data is
just too chaotic for it to train. So we
have two new words there. Converge and
diverge are important to know. Also what
is dropout and batch normalization?
Dropout is a technique of dropping out
hidden and visible units of a network
randomly to prevent overfitting of data.
It doubles the number of iterations
needed to converge the network. So here
we have our standard neural network and
then after applying dropout. Now it
doesn't mean we actually delete the
node. The node is still there and we're
still going to use that node. What it
means is that we're only going to work
with a few of the nodes. Um, a lot of
times I think the most common one right
now used is 20%. Uh, so you'll drop out
20% of the nodes when you do your
training, you reverse propagate your
data and then you'll randomly pick
another 20 nodes the next time you go
through an epic data training. So each
time you go through one epic, you will
randomly pick 20 of those nodes not to
not to mess with. And this allows for
less overfitting of the data. So by
randomly doing this you create some I
guess it just kind of pulls some nodes
off to the side and says we're going to
handle the data later on so we don't
overfit. Batch normalization is a
technique to improve the performance and
stability of neural network. The idea is
to normalize the inputs in every layer
so that they have mean output and
activation of zero and standard
deviation of one. This question covers a
lot of different things which is great.
It's a great uh interview question
because it pulls in that you have to
understand what the mean value is. So a
mean output activation of zero that
means our average activation is zero. So
when you normalize it remember usually
we're going between minus1 and one on a
lot of these. It's a very standard
setup. So you have to be very aware that
this is your mean output activation of
zero. And then we have our standard
deviation of one. So we want to keep our
error down to just a one value. The
benefits of this doing a batch
normalization is it provides
regularization. It trains faster, higher
learning rates and weights are easier to
initialize. What is the difference
between batch gradient descent and
stochastic gradient descent? Batch
gradient descent. Batch gradient
computes the gradient using the entire
data set. It takes time to converge
because the volume of data is huge and
weights update slowly. So you can look
at the batches. A lot of times if you're
using big data, batch the data in, but
you still go through a full epic. You
still go through all the data on there.
So bash gradient descent means you're
going to use it to fit all the data and
look for a convergence there. Stochastic
gradient descent. Stochastic gradient
computes the gradient using a single
sample. It converges much faster than
batch gradient because it updates weight
more frequently. Explain overfitting and
underfitting and how to combat them.
Overfitting happens when a model learns
the details and noise in the training
data to the degree that it adversely
impacts the execution of the model on
the new information. It is more likely
to occur with nonlinear models that have
more flexibility when learning a target
function. An example of this would be um
if you're looking at say cars and trucks
and motorcycles, it might only recognize
trucks that have a certain box-like
shape. It might not be able to notice a
flatbed truck unless it's only a
specific kind of flatbed truck or only
Ford trucks because that's what it saw
on the training set. This means that
your model performs great on your train
data and great on maybe a small test
amount of data, but when you go to use
it in the real world, it leaves out a
lot and start and is not very functional
outside of your small area, your slow
laboratory data coming in. Underfitting,
doing the opposite when you underfit
your data. Underfitting alludes to a
model that is neither well-trained on
training data nor can generalize to new
information. Usually happens when there
is less and improper data to train a
model. has a bur performance and
accuracy. So if you're using underfitted
data and you generate a model and you
distribute that in a commercial zone,
you'll have a lot of people unhappy with
you because it's not going to give them
very good answers. So we've explained
overfitting and underfitting. So now we
want to ask how to combat them.
Combating overfitting and underfitting,
resampling the data to estimate the
model accuracy, k-fold cross validation,
having a validation data set to evaluate
the model. So when we do the reampling,
we're randomly going to be picking out
data and we'll run it a few times to see
how that works depending on our random
data and how we sample the data to
generate our model and then we want to
go ahead and validate the data set by
having our training data and then
keeping some data on the side uh testing
data to validate it. How are weights
initialized in a network? Initializing
all weights to zero. All the weights are
set to zero. This makes your model
similar to a linear model. So if you
have linear data coming in, doing a
basic setup like that might work. All
the neurons in every layer perform the
same operation given the same output and
making the deep net useless. Right?
There's a key word. It's going to be
useless if you initialize everything to
zero. At that point be looking into some
other uh machine learning tools.
Initializing all weights randomly. Here
the weights are assigned randomly by
initializing them very close to zero. It
gives better accuracy to the model since
every neuron performs different
computations. And here we have the
weights are set randomly. We have our
input layer, the hidden layers and the
output layer. And W equals NP random
random N layer size L, layer size L
minus one. This is the most commonly
used is to randomly generate your
weights. What are the different layers
in CNN? Convolutional neural network.
First is the convolutional layer that
performs a convolutional operation. We
have our other video out if you want to
explore that more. and go into detail
exactly how the C the convolutional
layer works in the CNN as far as
creating a number of smaller uh picture
windows that go over the data. Uh the
second step is as a relu layer relu
brings nonlinearity to the network and
converts all the negative pixels to
zero. Output is rectified feature map.
So it goes into a mapping feature there.
Pooling layer pooling is a down sampling
operation that reduces the
dimensionality of the feature map. So we
have all our relu layer which is pulling
all these little maps out of our
convolutional layer. It's taking that
picture and little creating little tiny
neural networks to look at different
parts of the picture. Uh then we need to
pull it together and then finally the
fully connected layer. So we flatten our
pooling layer out and we have a fully
connected layer recognizes and
classifies the objects in the image. And
that's actually your forward propagation
reverse propagation training model
usually. I mean there's a number of
different models out there of course.
What is pooling in CNN and how does it
work? Pooling used to reduce the spatial
dimensions of a CNN performs down
sampling operation to reduce the
dimensionality. Creates a pulled feature
map by sliding a filter matrix over the
input matrix. I mentioned that briefly
on the previous slide. Um it's important
to know that you have if you see here
they have a rectified feature map. And
so each one of those colors like the
yellow color that might be one of the a
smaller little neural network using the
ReLU. You'll look at it'll just kind of
um go over the main picture and look at
all the different areas on the main
picture. So you might step one 2 3 four
spaces. Um and then you have another one
that's also looking at features and it
has a 2785. Each one of those is a map.
So it might be the first one might be a
map looking for cat ears and the second
one looking for human eyes. When it does
this, you then have this rectified
feature map looking at these different
features and the max pooling with a 2x
two filters and a stride of two. Stride
means instead of skipping every pixel,
you're going to go every two pixels. You
take the maximum values and you can see
over here we look at a pulled feature
map. One of the features says, hey, I
had a max value of eight. So somewhere
in here we saw a human eye labeled as
eight. Pretty high label. And maybe
seven was a human hand and maybe four
was cat whiskers or something that we
thought might be cat whiskers. Four is
kind of a low number in this particular
case compared to the other ones. So you
have your full pool feature map. You can
see the process here is we have our
stepping, we look for the max value and
then we create a poolled feature map of
the maxed values. How does a LSTM
network work? That's long shortterm
memory. So the first thing to know is
that an LSTMs are a special kind of
recurrent neural network capable of
learning long-term dependencies.
remembering information for long periods
of time is their default behavior. We
did look at the RNN briefly talked about
how the hidden layer feeds back into
itself. With the LSTM has a much more
complicated feedback and you can see
here we have the hidden layer of T minus
one and the hidden layer that's what the
H stands for hidden layer of T and the
formulas going in. As we can see here we
have the hidden layers we have T minus
one and then H of T where T stands for
time. So this is a series remember
working with series and we want to
remember the past and you can see you
have your ex your input of t and that
might be a frame in a video as a frame
comes in they usually use in this one
the tangent h activation formula but you
also see that it goes through a couple
other formulas the omega formula and so
when it combines these that then goes
into the next layer your next hidden
layer that then goes into the data
that's submitted to the next input so
you have your x of t + one. So when you
have that coming in, then you have your
H value that's coming forward from the
last process. And depending on how many
of these omega structures you put in
there depends on how long-term the
memory gets. So it's important to
remember this is more for your long-term
recurrent neural networks. The three
steps in an LSTM, step one decides what
to forget and what to remember. Step
two, selectively update cell state
values. So based on what we want to
remember and forget, we want to update
those cell values and then decides what
part of the current state make it to the
output. So now we have to also have an
output on there. What are vanishing and
exploding gradients? This is a great
question that affects all our neural
networks. While training an RNN, your
slope can become either too small or too
large and this makes the training
difficult. When the slope is too small,
the problem is known as vanishing
gradient. So our slope, we have our
change in x and our change in y. When
the slope decreases gradually to a very
small value, sometimes negative, and
makes training difficult. When the slope
tends to grow exponentially instead of
decaying, this problem is called
exploding gradient. The slope grows
exponentially. You can see a nice graph
of that here. Issues in gradient
problem, long training time, poor
performance, and low accuracy. What is
the difference between epic, batch, and
iteration in deep learning? Epic. An
epic represents one iteration over the
entire data set. So that's everything
you're going to go ahead and put into
that training model. Batch. We cannot
pass the entire data set into the neural
network at once. So we divide the data
set into a number of batches. And then
iteration. If we have 10,000 images as
data and a batch size of 200, then the
epic should run 10,000 times over 200.
So that means we have our total number
over the 200 equals 50 iterations. So in
each epic we're running over all the
data set, we're going to have 50
iterations. And each of those iterations
includes a batch of 200 images in this
case. Why TensorFlow is the most
preferred library in deep learning? Uh
well, first TensorFlow provides both C++
and Python APIs that makes it easier to
work on. Has a faster compilation time
than other deep learning libraries like
KAS and torch. TensorFlow supports both
CPUs and GPUs computing devices. So
right now TensorFlow is at the top of
the market because it's so easy to use
for both programmer side and for
hardware side and for the speed of
getting something up and running. What
do you mean by tensor in TensorFlow?
Tensor is a mathematical object
represented as arrays of higher
dimensions. These arrays of data with
different dimensions and ranks that are
fed as input to the neural network are
called tensors. And you can see here we
have a tensor of dimensions five, four.
So it's a two-dimensional tensor coming
in. Um, you can look at an image like
this that each one of those pixels is a
different value if it's a black and
white. So, it might be zero and ones and
then each one represents a black and
white image. In a color photo, you might
um either find a different value system
or you might have a tensor value that
has the xy coordinates as we see here
plus the colors. So, you might have
three more different dimensions for the
three different images, the red, the
blue, and the yellow coming in. And even
as you go from one layer or one tensor
to the next, these layers might change.
We might flatten them, might bring in
numerous. In the case of the convergence
neural network, we have all those
smaller different mappings of features
that come in. So each one of those
layers coming through is a tensor. If it
has multiple dimensions coming in and
weights attached to it, what are the
programming elements in TensorFlow?
Well, we have our constants. Constants
are parameters whose value does not
change. To define a constant, we use
tf.constant command. Example, A equals
TF.Constant 2.0 TF float 32. So it's a
tensor float value of 32. B equals TF
constant 3.0. Print AB. If we did a
print of AB, we'd have um TF.stant and
then of course uh B is that instance of
it. Variables. Variables allow us to add
new trainable parameters to graph. To
define a variable, we use TF.variable
command and initialize them before
running the graph in session. Example W
equals TF variable.3 DT type TF float 32
or B equals a TF variable minus 3, D
type float 32. Placeholders.
Placeholders allow us to feed data to a
TensorFlow model from outside a model.
It permits a value to be assigned later.
To define a placeholder, we use TF
placeholder command. Example A equals TF
placeholder b= a * 2 with the TF session
as SESS result equals session run B,
feed dictionary equals A3.0. 0 print
result. Uh so we have a nice example
there of a placeholder session. A
session is run to evaluate the nodes.
This is called as the tensorflow
runtime. So for example, you have a= tf
constant 2.0 b= tf constant 4.0 c= a
plus b. And at this point you'd go ahead
and create a session equals tf session.
And then you could evaluate the tensor C
print session run C. That would input C
as an input into your session. What do
you understand by a computational graph?
Everything in TensorFlow is based on
creating a computational graph. It has a
network of nodes where each node
performs an operation. Nodes represent
mathematical operation and edges
represent tensors. Since data flows in a
form of a graph, it is also called a
data flow graph. And we have a nice
visual of this graph or graphic image of
a computational graph. And you can see
here we have our input nodes, our add
multiply nodes and our multiply node at
the end. And then we have the edges
where the data flows. So we have from A
going to C, A going to D. You can see we
have a two flowing, a four flowing.
Explain generative adversarial network
along with an example. Suppose there is
a wine shop that purchases wine from
dealers which they will resell later. So
we have our dealer going to the wine,
our shop owner that then sells it for a
profit. But there are some malfactor
dealers who sell fake wine. In this
case, the shop owner should be able to
distinguish between fake and authentic
wine. The forger will try to different
techniques to sell fake wine and make
sure certain techniques go past the shop
owner's check. So, here's our forger
fake wine shop owner. The shop owner
would probably get some feedback from
the wine experts that some of the wine
is not original. The owner would have to
improve how he determines whether a wine
is fake or authentic. Goal of forger to
create wines that are indistinguishable
from the authentic ones. Goal of shop
owner to accurately tell if the wine is
real or not. There are two main
components of generative adversarial
network. And we refer to as a noise
vector coming in where we have our
forger who's going to generate fake wine
and then we have our real authentic wine
and of course our shop owner who has to
figure out whether it's real or fake.
The generator is a CNN that keeps
producing images that are closer in
appearance to the real images while the
discriminator tries to determine the
difference between real and fake images.
The ultimate aim is to make the
discriminator learn to identify real and
fake images. What is an autoenccoder?
The network is trained to reconstruct
its inputs. It is a neural network that
has three layers. Here the input neurons
are equal to the output neuron. The
network's target outside is same as the
input. It uses dimensionality reduction
to restructure the input. Input image
comes in. We have our Latin space
representation and then it goes back out
reconstructing the image. It works by
compressing the input to a Latin space
representation and then reconstructing
the output from this representation.
What is bagging and boosting? Bagging
and boosting are ensemble techniques
where the idea is to train multiple
models using the same learning algorithm
and then take a call. So we have in here
where we're bagging. We take a data set
and we split it. We're going to have our
training data and our test data. Very
standard thing to do. Then we're going
to randomly select data into the bags
and train your model separately. So we
might have bag one, model one, bag two,
model two, bag three, model 3, and so
on. In boosting the emphasis is to
select the data points which give wrong
output in order to improve the accuracy.
So in boosting we have our data set
again we split it to test data and train
data and we'll take a bag one and we'll
train the model. Data points with wrong
predictions then go into bag two and we
then train that model and repeat.
machine learning is which is a subset of
artificial intelligence, right? That's
uh basically
um machines learning from data
in order to uh make decisions
essentially. Um so this was a big
departure from the rules-based systems
at the time, right? That were explicitly
programmed to make decisions. So just
think of an example like a really big
kind of if this then that then that then
that and and else if this this this
right so bunch of rules that had to be
pre-programmed in order to um come out
with some final answer. Uh with machine
learning it's the exact opposite of
that. we're actually training something
from examples from existing data um in
order to predict something or um
make some type of decision. Uh and so
we're going to learn about the various
ways we can do machine learning. But if
you guys remember we um talked about
some of this like the differences and
the uh basically rules-based approaches
to learning from data approach. Um and
in included in that is going to be uh
complex unstructured data. So things
like images, text, audio. What handles
those really well is uh deep learning
which we will get to in the course after
this. But uh those are certainly in
there as learning from data even complex
data.
So we had this picture uh and I think
this is kind of around where we left off
last time was uh just distinguishing
between those three terms. We see
artificial intelligence, deep learning
and machine learning kind of used
interchangeably, but this is really how
they fit in. Artificial intelligence is
kind of a broad anything mimicking human
intelligence. Um which doesn't have to
be learning from data, but uh machine
learning is part of that. And then um
one way to accomplish machine learning
is to use neural nets which is the focus
of uh deep learning. Um and so deep
learning has been has found a lot of
success especially recently with uh
those complex data types like images,
speech, text, right? So deep learning
used all over the place. Even in um
modern like generative AI, we see deep
learning used quite a bit. Um it really
anything that's using neural nets is uh
going to be deep learning.
Um and again we'll focus on that later
but we're going to be mainly focused on
machine learning for this course.
Primarily machine learning that does not
use neural networks. Okay. So just
models that are not necessarily neural
networks
be our focus.
So in machine learning we had an example
of a game uh essentially um learning
what decisions to make uh based on the
uh kind of current um state of the
board. This could be a um you know
machine learning example that uh learns
from many previous examples. So a lot of
data around these games are used to
train these um kind of robots that can
play these games and play them at a very
high level. Um so there's been a lot of
successes actually in machine learning
and deep learning um around
uh playing games like chess or go
um using machine learning algorithms. So
pretty cool.
All right, so I think this is where we
ended. Last time we said there's a bunch
of different use cases for machine
learning. So um recommendation system is
going to be a big one and we will
actually study that uh in one of our
final lessons of this course. Um chat
bots like generative AI doing sentiment
analysis chat bots we'll study later but
those are certainly an application of
learning from data in order to uh
generate responses to text prompts
right. Um spam filtering that's a good
example like classifying an email as
spam or not spam. Um that that gets
trained from examples and uh learning
from data such as previous emails. Um
social media posts analysis is another
kind of text data um use case but you uh
can do a lot with that text like you can
predict the sentiment um you can predict
uh the category of what what the post is
talking about um those kind of things
all can be done with machine learning
>> and many other use cases not on this
list that we will uh cover
>> you know as we as we go further.
Okay, so this is where we kind of left
off. Um, so what's doing all the hard
work here is
>> uh machine learning algorithms. So these
are things that will um these are things
that will learn from the data. So they
are uh they they are basically um
algorithms or sets of rules that uh or
mathematical rules I should say not
formal rules like in the in the sense of
a rule system but mathematical um
formulas and mathematical uh rules
essentially that help us learn from the
data. So they correlate the data to some
type of outcome. So some type of
prediction uh whether that's going to be
as we will see whether that could be
like a number like we're predicting a
price or demand or sales
um or it could be a category like is
this transaction fraud or not fraud or
what's the probability that this is
fraud um so we have different kinds of
predictions we can make with machine
learning
um but uh we will study the kind of the
differences of those coming up. Um, but
machine learning algorithms are really
what power they're kind of the models,
right? They're the models that help
power uh machine learning to actually
learn from data.
So, we're going to spend a lot of time
in this course studying those algorithms
like the different models that we can
build and what their differences are,
what their strengths are, what their
weaknesses are. We'll we'll learn a lot
about those.
Okay. So I guess you can imagine like
everything is so data dependent, right?
Um we're learning from data. So uh it
makes sense that the quality of data
really really matters here in
determining how strong the model can be.
Um so you see this graph here charting
kind of the um high quality data um
versus just uh any old data but a decent
enough quantity of it. Um you can see
that performance and the performance is
measured by some evaluation metric. Um,
so think of it as uh something like an
accuracy. Like if we were predicting
fraud or not fraud, how accurate can our
model get at actually detecting fraud,
um, it gets better and better and better
the graph shows that the higher quality
of data that we have. So there's kind of
that there's a there's a saying in
machine learning um, called garbage in
garbage out. What that means is if you
have poor data, even the best model in
the world, poor data is not going to
result in having a good model that can
be accurate and perform well. Um, so it
needs to be high quality, meaning um
there needs to be a decent amount of it
and it needs to be labeled appropriately
as we will will talk about
um and it needs to not have any, you
know, significant outliers. it needs to
be clean, not have those missing values,
all of those things. Um, you can you
have a good chance at deriving good
predictions from higher quality data
as this kind of shows.
Okay.
So, one thing we're going to learn um as
we go along is
quantity matters as well. So, not only
quality, but a decent amount of it. And
um we're going to learn those kind of
rules of thumb like how much data do I
need for certain algorithms. Um one
thing that we will see is that uh the
the basic machine learning models that
we'll study don't need as much as a
neural network would. It you know neural
networks are going to require a lot more
um than a basic machine learning model
learning model. So uh that's something
we will see as we go along. But uh this
is something we'll talk about and
discuss with each model that we study is
kind of how much data do we actually
need to produce a high quality model.
Okay, any questions uh so far?
Okay, let's talk about the different
types of machine learning that we're
going to discuss. Pime, there's going to
be two primary ones that we will study
in this course and then a couple others
that'll be a little bit more advanced
that we won't get to but worth knowing
about. Um, so there's going to be four
total that we'll study or talk about and
they'll be on this list here, which is
um supervised learning and unsupervised
learning. Now, I'd say the majority of
our focus will probably be on supervised
learning, and we'll talk about what that
means, but we'll also cover unsupervised
learning as well. And so, we'll look at
the most popular techniques in each of
these types of machine learning.
Um,
and then we'll talk about these two, but
not really study them because they're
more advanced topics um that that will
be beyond the scope of what we'll do.
But uh these are going to be um
different styles of machine learning
that are going to be characterized by um
what kinds of predictions they make,
what kind of data they need and require.
Um and uh what kind of outcomes they're
actually producing. Um, so let's let's
get into each of these, but uh the the
one that we'll probably spend the
majority of our time on is going to be
supervised learning, but we will study
unsupervised learning as well. We'll
study both and we're going to talk about
we're going to define both of those um
coming up. And again, these will be a
little bit more advanced topics that we
won't spend too much time on.
Um, but but we'll discuss their
relevancy in machine learning um and
give a good definition to it.
Okay.
All right. Let's start with supervised
learning. Now this is going to be uh a
term that really refers to
using examples. So using labeled
examples. So here we say labeled data to
help our model train. In other words,
help our model be able to predict guided
by specific input output pairs. So
supervised really refers to the fact
that we have answers. We have examples,
we have answers with those and we use
that collection of data to build our
model off of so that we can predict
um those kinds of things like a price,
like a category, like a spam not spam.
in this in this slide like we would be
predicting if this shape is a square, a
triangle or a circle.
Um but but when we build a model for
that, we have data that has an answer
attached to it. Right? We've talked
about this before a little bit with
labels. So there's a guide there that
can guide us towards building our model.
there's an actual every every example
has an answer and that answer is really
critical to help build our model off of.
So, um that's it's almost like you have
um a you have a bunch of exercises
in let's say like a math textbook. You
have a bunch of exercises and you have
the answers and that way you can kind of
check your work. You think about model
training um that is the really a lot of
that process of model training as we are
going to discover is um basically
checking our work against these answers
in our data in our training data.
Okay. So supervised learning is any type
of machine learning that involves
learning from labeled data in order to
predict outcomes. Okay, predict outcomes
like now the the outcomes can be
numerical. They can be like a price,
temperature, demand, sales, revenue.
They can be numerical, but they can also
be categorical. So they can be like
spam, not spam, fraud, not fraud,
cancer, not cancer. Um, dog, cat,
giraffe, those kind of categories. Um,
we could predict those. It's some type
of outcome. Okay, some type of outcome.
The key is we're using labeled examples
to guide our model building. That's why
it's called supervised learning.
So we know in our data we know what the
inputs are. Of course, those are going
to be think of the inputs as like all of
our columns and then we have a special
label column that represents the output
we're trying to predict. So if you think
about that housing price data, the label
could be the price. And that's something
we would build a model to predict, but
we have answers for all of our examples
in our rows. We have answers to help
guide our model building.
They help tweak our model because we
know the answer ahead of time. So
they're they're really good examples to
build our model off of.
Okay. So that's that's supervised
learning.
Uh in this example is circle not in the
prediction because it's not part of the
test data even though it's in the
labeled data.
Um no it just not necessarily. It just
means that like we learn against all of
these examples that have these answers
and then when we observe new examples um
we can try to predict what those would
be based on what we've seen before. So I
if there was a you know it's just a
coincidence we only have two two
examples in our test data like we could
have a circle here in which case we
would predict circle
that's fine or at least we would hope
our model would predict circle right
that's what we're hoping may or may not
get it right
um but it's it's only not there because
we only like we're just assuming that we
only have two examples we're testing
against but in reality we would probably
do a lot more than two.
It's just it's just a coincidence
really.
In reality, we would test against a lot
more data. And we're actually going to
see why we would do that. Like why would
we train our model and then kind of use
additional data um to to evaluate it?
It's actually really important that we
do that step to get a sense of how good
our model is before we take it out in
the real world. So if we apply our model
that we build on our label data to
um this kind of set of test data that we
haven't been exposed to before. It helps
give us a sense of how good is our
model. So it's tested is usually used
for evaluation.
So that's something that's something
we'll study.
How do we train? Uh it depends on the
model. Um so training will be a sense uh
will be an algorithm that will um
basically update the model according to
the data. These labeled examples. Um
every model is going to be different in
exactly how it trains. So we're going to
we're going to talk about that when we
get to the individual models that we'll
study.
But uh loosely speaking, they're going
to use the data to adjust itself. Like
imagine adjust like tuning a bunch of
knobs. Um, like the best example I can
give you is we I think I did this one
last week where you have kind of a
function
that predicts the price and let's say it
has
um weights like weight one with feature
one, weight two with feature two,
weight three with feature three. So
imagine we had three input features and
we we built an answer according to that.
Essentially what we would do to train
the model is adjust these
um in order to get this correct based on
our our labeled examples.
Okay.
So that's something we're going to learn
about coming up shortly when we when we
actually dive into model. Every model is
going to be slightly different in how it
trains, but at a high level it's going
to use the training data with those
examples, right? the labeled examples to
help guide the formula essentially to
adjust to generate the proper kind of
model here.
The these things are going to be
adjusted according to the data
in order to produce the correct output.
So think about these as knobs that will
turn.
Okay.
Uh which type of machine learning is
used? Uh probably supervised um which is
what we're talking about now. So
probably supervised because most people
want to
um build some type of model to predict
something.
Uh so yeah, I'd say I'd say supervise.
Yes, we're are we are definitely going
to learn how to train. Yeah, we'll see.
We'll do the code. Um, I'll tell you
about how it's done. Yeah, we're
definitely going to learn it. But what I
was saying is it's kind of on a model
bymodel basis.
So, I want to wait till we get into the
individual models, then we'll talk about
how they're trained.
But yeah, we'll we'll learn how to do
that.
But yeah, supervisor is used all over
the place. Even even for uh generative
models, they use supervised learning
because um like an LLM
is going to use labeled examples in
order to train, right? In order to train
how to generate responses according to
prompts. Um it needs to learn against a
lot of text examples.
So that supervised learning is what um
results in that model,
right? Learning from those labeled
examples.
Okay.
It is yeah image image uh a lot of um
yeah a lot of image processing is
supervised like object detection. So the
YOLO model is an object detection model.
Yes. Um because it has to be trained
right. It has to be trained on uh it has
to be trained on images
with labels such as this is what object
is in this image. This is the box around
the object.
Um yes. So if if it's if it ever uses
label data to train and build the model,
it is supervised. So YOLO is definitely
supervised and we actually we will we
will cover the YOLO model later on in
our deep learning course. We talk about
object detection.
So we'll we'll study that.
But yeah, it's supervised
Okay. So on the slide we have some
common supervised learning algorithms
that are we will study. So all of these
we will study and understand what they
do and how they work but just giving you
some to name them. linear regression is
kind of the one I just drew out which is
the um this is the prototypical like
easiest to understand model that is kind
of the um exactly like this where we
have a weight times a feature
um a weight times a feature and then a
weight times a feature
and on and on and on. You can have as
many as you want.
um that is a linear regression. And so
that is um that's a supervised model
because we need this value here and we
need all of our inputs in order to um
actually train this model and generate
all those weights
um that that is uh that uses um labeled
examples to help tune all those knobs.
Um same with all these other models. So,
we're going to talk about decision
trees. We're going to talk about
logistic regression and and SVMs, which
are support vector machines. We'll talk
about all of those, but they're all
examples of supervised uh supervised
learning.
Okay, we'll talk about all of these.
They're all supervised because they all
require labeled examples in order to
train them and and then subsequently use
them. Okay.
Okay. So what are some use case
examples? So for for instance in uh
supervised learning we may be predicting
temperature based on yearly temperature
trends. So we would have that yearly
data as our um as our labeled examples
and those would supervise the learning
of a model that predicts temperature.
Um, same thing with predicting crop
yield based on um, seasonal crop quality
changes. So maybe we have a bunch of
features relating to crop quality. We
could predict crop yield. Um, we would
just need historical examples with those
labels, right? What the crop yield is
for each time period. Let's say we would
just need those uh, supervised examples
and we could easily build a model off of
it.
Um
uh this this last one sorting waste
based on known waste items and their
corresponding waste types. Um that's
kind of like spam. It's like filtering
basically like a spam filtering. Um so
think of it like the the shapes example.
We sorting things into squares, circles,
triangles. Um, same kind of idea here
where we have a bunch of examples on
what those um what those waste items
should uh should belong to, like what
wastist bins they would go to, for
example. Um, and those could be labeled
and therefore then we could um
understand what category of waste they
belong to.
Um, same thing with spam. Something is
fraud or not fraud. spam or not spam,
cancer or not cancer. All of those are
going to be supervised learning examples
because they're going to require in
order to train them, they're going to
require data that has those labels.
Okay? So, anything that has labels is
going to be supervised learning. So
again, this is where we will spend
probably the the majority of our time is
doing supervised learning problems, ones
that we have labeled data. We're
building a model and we're going to
predict those those uh labels
essentially.
Okay, before we go to unsupervised, any
questions about uh supervised.
Okay.
All right. So supervised requires labels
in order to have an example to go off of
to build your model. And that's because
you're predicting those kind of outcomes
like spam or not spam, cancer or not
cancer. Now unsupervised learning is
completely different. It's the opposite.
So unsupervised learning is where we do
not use labels whatsoever. So we're not
using any labels at all. So it's it it
can be completely unlabeled or even if
it's labeled, we're not using labels in
any way. But um we primarily would say
it's unlabeled data. We have no guidance
because we're not using the labels in
any way. We have no guidance to um
predict anything, but that's because
we're not really predicting anything in
unsupervised learning. Generally, what
we're doing is looking for some
structure or pattern.
Okay, with unsupervised learning, we're
looking for some structure or pattern.
So, um, one type of example that's very
very popular is going to be this second
one, which is, um, identification
identification of user groups based on
similarities or commonalities. Now, this
is going to be a problem basically known
as clustering
and it's a problem we will study quite a
bit. there's going to turn out to be
lots of different algorithms that can
accomplish clustering. So what
clustering attempts to do is basically
say um we have data that's like this and
then data over here and then data over
here. Let's just group these together.
So like this should be one group, this
should be one group and this should be
one group. And we can find those
structures and say okay this is group
one, this is group two and this is group
three.
one, two, three. And we can basically
build what we would call clusters of
data um based on how close together the
points are kind of located in these kind
of cluster zones like these boxes I've
drawn.
Okay. Now, that doesn't require any
label to do which is really fascinating.
So, unsupervised, you don't need any
label at all to accomplish the
algorithm. Um so, clustering is one good
example.
um finding outliers or anomalies is
another. So we don't necessarily have
any label of what is an outlier or what
is an anomaly. We are deriving that from
the features alone. There's no guidance.
There's no label um to doing like
outlier detection or anomaly detection.
Okay, so that's another good example.
One that's not listed on here um but is
also really important that we will study
is something known as dimensionality
reduction.
So dim reduction and what that what this
focuses on is basically compressing the
data set a bit. So we take our data and
basically compress it um so that but we
do it in such a way that we retain as
much information as we can. This is a
very like smart compression and what it
does is it lowers the dimension. Um
dimension think of the dimension as like
number of columns.
Number of columns.
So imagine we had 100 columns in a data
frame. What we could do is actually
reduce that down to 10. So like 10% of
that. So we reduce it down to 10. And um
but those 10 are it's not like we
chopped out um 90 other columns. We um
smartly kind of compressed all that
information into these 10 new columns um
that are compressed versions of the
hundred that we used to have. Um so
dimensionality reduction is is another
unsupervised technique. It requires no
guidance, no label to do, but is um a
really useful technique to reduce the
size of your data if you're doing things
with it. Um so this is another one that
we will we'll study how to do it and
basically more details behind it, what
the algorithms are.
Um we'll so probably those two in
unsupervised will spend the most amount
of time on clustering and dimensionality
reduction.
uh and supervised if some data is
present but we didn't label it means in
example we had circle triangle square in
the training data we add pentagon
but we didn't label that in that case
uh yeah so every um in supervised
learning, every row, think about it as
like every row in our data frame needs
to have a label
uh associated to it. It needs to have a
a column that represents the label.
So if we've never seen Pentagon before,
I can't use that as a label.
So, it has to the Pentagon has to exist
in the data if I'm going to be able to
predict it,
right? So, I can't predict, right? If
we've never seen it before, we have no
examples to go off. We have no guidance.
So, how could we predict that?
Right? We can't predict it.
if it's if it's in there. So if if we
have labels of Pentagon, let's say, then
yeah, we could predict Pentagon. We
could
remove. Remove what?
We wouldn't if it was talking about the
Pentagon, we wouldn't remove that. No,
let me go back to that page. We wouldn't
remove it. Um, it's just if it's not in
our labels, we're not going to be able
to predict it. So, Pentagon's a good
example here. Uh, Pentagon is not one of
our labels. So, it currently is not in
our data set as one of the labels. We
only have data that's either a triangle,
circle, or a square. We don't have
pentagon. So, I would never be able to
predict pentagon. I'll never be able to
do that if I haven't seen examples of it
before.
Okay. But let's say we had that in
there.
So we had Pentagon.
So if we had Pentagon, um we could have
an example of it in our labels
and then yeah, we it could be then we
could predict it.
Yeah. Yeah. The the don't get worried.
Don't worry about the test data. So the
test data is just saying here's a new
here's a shape. What is it? Okay, that's
a square. Here's a shape. What is it?
Okay, that's a triangle. And we could
have as many of those examples as we
want in our test data. So we could have
a circle and say, okay, what's this?
Should be circle,
right? The test data can be whatever it
whatever it wants. But yeah, if if we've
never seen Pentagon before, we're never
going to be able to predict it.
These are the the label data and labels
are basically the talking about the same
thing. The labels just mean what are the
categories that are present in our data.
So in this data we only have three
labels that are present.
So the labels is are relative to our
label data, right? It's saying
what labels,
excuse me, what labels uh do we have
in our data and we only have those three
circle, triangle, square. So so Pentagon
would not be part of those labels. We
couldn't predict it.
No. So unsupervised is not going to make
a prediction. That's the big difference
with unsupervised. They're not going to
make a prediction like this. Um so
unsupervised is not going to make a
prediction. It's going to do something
different like um basically say like
these guys are similar, these are
similar, these are similar, this is a
cluster, this is a cluster, this is a
cluster. It's not going to make a
prediction. That's what supervised
learning does.
Clustering, yes, which is unsupervised,
yes, clustering does not require any
labels. Unsupervised just means we don't
have any labels. We don't require any
labels.
So the other thing unsupervised might do
is it might say
and again without the labels it might
say that this is an outlier.
it might say that this guy is an outlier
because there's only there's only one of
those and they're not like the other. So
that that's something that um that's
something that uh unsupervised could do.
Um it it yeah and no. It kind of labels
a cluster in the sense that um it would
basically assign a number to it like
this is cluster one, this is cluster
two, this is cluster three.
It'll assign a number to it, but it's
not a very meaningful it doesn't assign
like a prediction label in the in the
traditional sense of a label.
It does provide like a numerical index
for the cluster to because what we want
to know is like okay this guy has the
cluster of one. This guy belongs to
cluster one. This guy belongs to cluster
one. This guy belongs to cluster two.
This guy belongs to cluster two. Does
that make sense? So there needs to be
some like index of what cluster you
belong to.
So it's kind of like a label but not in
the traditional like prediction sense.
Very
good. So again, unsupervised, no labels.
You're doing things like
identifying clusters,
um identifying outliers, doing
dimensionality reduction. These are all
like structure and pattern oriented
things. They're not predictions of a
label. Okay? They're not which is what
we would see in supervised learning.
Okay. So an example would be that we
take we put in the data um we can group
together uh data such as images into
categories based on similarities um
which would be like those clusters. So
there's no these would be groups that we
don't have any label on ahead of time
like we don't have we don't say that
this image should belong to this this
image should belong to this we derive
that from the characteristics of the
data. Um so think like a good example is
um customer groups. So we would identify
customers based on like okay do they
have similar spending levels? How many
days do they go shopping in a week? How
much money do they spend? And we can
kind of group together customers based
on similar qualities.
Clustering will find those groups that
should exist.
um it will discover those groups based
on um the similarities in the data, but
there's no labels that that say like
this person should be in this group,
this person should be in this ahead of
time. There's no labels of that. It gets
derived during the algorithm. It's
unsupervised,
right? There's no unsupervised really
literally means no guidance. There's no
guidance to doing it. We just derive
that from the structure of the data
which is the similarities.
Okay.
All right. So,
a couple more for you. So we had um
supervised which uses the labels. We
have unsupervised which uses no labels
looking for structure. And then we have
something that's kind of in between
which is um what is known as
semiupervised learning. And this is
where you use a combination of a little
bit of label data, but most of your data
is actually unlabeled data. Um, and you
try to get some use out of that label
data in order to um build a model out of
it. And so uh it uses the um it uses
that label data to um generally provide
some guidance on usually what happens
with semi-supervised learning is you use
your label data to kind of predict what
the label should be for the unlabelled
data and then you can go from there. So
you can create artificial labels on this
unlabeled data and then you can use all
of it once it's all been labeled kind of
like a supervised learning uh approach.
So but but this is semi-supervised
basically refers to the fact that you
start out with most of your data not
being labeled but you do have some
labeled examples and what you can do is
basically extrapolate those labels into
the unlabeled data set and then provide
some artificial labels and then now
everything has a label you can do
supervised learning.
Okay. So, it falls kind of between um
supervised and and unsupervised.
Uh and there so this is this is kind of
rare. Most of the time you're not going
to do that. You're actually just going
to um prefer to just start with all
label data. That's usually the preferred
approach. Most of the time you'll
actually just be doing supervised
learning, not really semi-supervised
learning. So, it's pretty rare, but um
it it could like if Yeah, it could if
the if we had a lot of examples of
Pentagon and we wanted and so they were
unlabeled and then we tried to guess
what kind of shape they were um and
provide an artificial label uh and then
um then use that whole data set to build
a model off of then yeah it could it
could fall into this category. Okay.
They Oh, going back to the question,
they still use some kind of label data
like age, gender. They use uh that's
those aren't those aren't really labels.
That's the features. So, yeah, they
still use the core features of the data.
They just don't have any like labels in
the traditional sense of a label. Like
you should think of a label as something
we are trying to predict.
So whether that's a price, whether
that's like a category like spam, not
spam, cancer, not cancer, it's something
we'd be interested in kind of
predicting. And so um in our data, we
would have an answer for every row. We'd
have one of our columns would be like
the the result like the outcome answer
that we're trying to predict. That's the
label.
So in unsupervised, we don't have any of
the labels.
We do have just the regular features
like gender, age, income, square
footage,
bedrooms, bathrooms, all those things.
Okay.
So, we have semi-supervised that falls
in between supervised. Now, the reason
it falls between is be is because
there's a decent amount of data that's
unlabeled. In fact, a majority of it
unlabeled. But what we can do is try to
label it. We can try to take what we
know from our existing labels and
predict an artificial label and then use
all that data together in kind of a
supervised fashion for a model down the
road.
So that's kind of what this picture uh
says is we can try to take um you know
maybe we try to infer some labels based
on we have some some labelled data here.
We have most of our data is unlabeled
and we try to supply some labels to it.
Um like maybe we have a babies category
of teens, a tween, uh you know youth and
um adults. Um and then we try so we we
take our our labels and we try to
extrapolate those into artificial labels
for this unlabelled data so that we can
use it now because then everything has a
label at this point and then we can just
go ahead and do supervised learning from
there.
So we can do supervised from there. What
we would prefer to do and what we'll do
in this course
um is just start with supervised. We'll
just start with the labels. We won't try
to derive artificial labels usually.
We'll just start with labels.
So one example in the real world is
something like Google photos which um
whenever you take a picture it can
provide uh uh labels based on previous
uh images in your library. So it can it
can produce tags or um labels on those.
Uh generally when you take that picture
it's kind of unlabeled unless you go in
and specifically provide some tags and
some labels. But um if you don't do that
it can still it can still uh make it can
artificially create one of those based
on the other label data that you already
have.
So that's um
that's an example.
Okay.
All right. Last one in terms of machine
learning. So we have supervised, we have
unsupervised.
Uh then we had semi-supervised which is
somewhere in between a mixture of having
some unlabelled data and label data. Um
now we're going to talk about
reinforcement learning which is
completely different. Um it's it's
completely different than the other
three. It's a type of machine learning
where we uh basically learn from
interaction with the environment. And
you might ask what are we learning? We
are learning what actions to take in the
environment. Um and the way we do that
is by reinforcing
positive actions that lead to a a
reward. Um, so that's where the word
reinforcement comes from is we we
basically uh imagine like a child
that's, you know, learning from trial
and error. Like they're trying to crawl,
they're trying to walk and they keep
falling down. um eventually they learn
how to do it through trial and error and
they might get a reward or they might um
reinforce some of those positive
movements that lead them to walk or
crawl um or they might learn from the
penalties, right? They might learn from
uh some type of feedback. So they might
learn from falling down like, "Oh, that
hurts. I should uh support myself a
little bit better, right?" Or be a
little more coordinated. Um
and so they they learn from those
actions and their interaction with the
environment. Um
uh so this is a complex um algorithm
essentially uh it's it deals a lot with
um again taking actions. Usually when
you take an action something changes in
the environment um then you kind of
observe some type of feedback. So, think
about like a a board game where you're
trying to figure out what move you
should make or another good example is
like with a robot um trying to navigate
a maze. So, like what route should it
take? Should it move forward? Should it
move backward? Should it move left or
right? Those are different actions it
can take. Also, like a self-driving car,
should it should it turn? Should it
speed up? Should it slow down? Those are
all good examples of things that have
been trained from reinforcement
learning.
Uh yeah. So real world examples would be
like in a board game uh a a reward would
be like if you win the game. Um or if
you like capture a piece like in
checkers or chess, that's a reward. A
penalty would be like if you lose the
game or lose one of your pieces, that
could be a a penalty.
um in a board game or sorry in like a a
robot navigation task, it could get
rewards for um moving in the right
direction
um towards the exit or like when it like
let's say you wanted to train a robot on
how to open the door and navigate a
room. Um you would penalize it for
bumping into the wall.
Um you would give it a reward for moving
usually oh like oh the algorithms
themselves usually it's like a a step
function um it's usually it's like a
discrete function that kind of is based
on the state so the reward it could be
like um like depending on the let's
let's go back to the board game example
like the reward could be like or even
the maze let's say like a navigating the
maze like getting to this let's say this
was the exit
And this was the entrance.
Then if they make it to here, they get a
numerical like if they make it to the
exit, they get a numerical reward of
like plus 100, let's say. So it's just a
number. And then if they uh like if they
bump if they go into here, like let's
say this is kind of like a death trap or
like a pit, this this would be like a
minus 100. So it could be like discrete
numerical values could be the reward if
they're moving in the right direction.
Like let's say we want to encourage
going this way then we could give
smaller intermediate rewards like this
should be a plus like if you move
forward this is a plus five this is a
plus 10 this is a plus 15 if you're
moving in the wrong direction away from
the exit. Um that would be like a minus5
or a minus 10. Does that make sense? So
they're they're numerical in nature and
what you're trying to do is collect the
most reward. You're trying to get the
largest reward you can through trial and
error. So you you try this out many many
many times. You basically simulate
running through this maze many many
times. And what dictates it what
dictates like where I should go is based
on what I've observed in the past. It's
almost like you're a child remembering
like, okay, what move should I make from
this space? Like if I'm here, if I'm
here, which way should I go? Should I go
down? Should I go right? Should I go
left? You kind of know that from
experience.
Does that make sense? Based on the
reward that I've seen in the past, like
when I've moved down, I've gotten a
higher reward than moving left or right.
Does that make sense? So, yeah, it's
it's a numerical value
as a reward.
Yeah, that's a great question. Um, how
does it differentiate rewards based on
gain and loss? I chess. So it's it's a
very comp complicated uh answer but
essentially every so in the chess board
you can think of the board as like every
every um
space is a state.
So I could be in this state I could be
in this state and then it's not not only
is every every uh space but where all
the other pieces are. So there's lots of
states that are possible.
Um, so
the way there's a way to quantify
essentially what's the value of taking a
certain action like moving my piece
left, moving it right, moving it up or
down um given the rest of the state. So
you're you're right, it may be
beneficial to sacrifice. Um, but we
would learn that through experience that
okay, the best move in this situation is
to sacrifice.
We would we would have to learn that
through trial and error many many many
times which is to say like okay if I'm
in this current state of the world right
all these pieces are distributed in this
way the best move for me right now in
the long run
to get the most reward in the long run
is to actually sacrifice my piece and
move it right move it into like a bad
position theoretically but we know from
experience that's actually the most
long-term reward is from that position
like moving it right may be the best
for me. So what you learn is how to take
actions
and actions are usually like move right,
move left, move up, move down. You think
about like a self-driving car though,
that's going to be like slow down, speed
up, turn your wheel 10 degrees. Um those
kind of actions.
So the the short answer is it's there's
a calculation there that you learn what
the long-term value of every state is
every unique state
and then you're trying to basically say
what action should I take from that
state
given that current state of the
Okay.
And I really I really like reinforcement
learning. It's actually probably my
favorite field of machine learning.
Unfortunately, we won't be covering it
um in our main uh course. We have
offered uh electives around
reinforcement learning in the past. So,
um stay tuned. Maybe when we get to the
end of this program, uh we'll offer an
elective on it and if enough people sign
up for it, we'll we'll run it. But, um
we it's not part of our we don't really
cover reinforcement learning as part of
our main topics. It's it is an advanced
uh more advanced topic than than what
we'll cover, but um I I really enjoy it.
I find it very fascinating.
Okay. So, all of this is kind of um
illustrating what I was saying, which is
um you think of like uh the thing that's
interacting in the environment like the
robot or the car or the human moving a
chest piece is known as the agent. It's
interacting with the environment by
taking actions which updates the state
um of of the environment. So that's
that's why you see this word state here.
This gets updated constantly every time
you take an action. Um ultimately what
reinforcement learning is trying to do
is learn the best action like what would
be the best action to take. Um
and the best action is is the one that
leads to the most long-term reward.
That's the best action. Um, so you have
to uh you have to learn what you know
what leads to a good reward by kind of
experiencing this over and over and over
through trial and error. So there's a
lot of um kind of simulation or letting
the robot try something a lot um in
order to kind of learn what's rewarding
and what's not. Think about it again
like I think a good example is like with
children, right? you kind of have to let
them try things until they learn on
their own what's what can they do and
what can they not do
what's the best actions right
so reinforcement learning has made its
way into other places so I I said like a
good example is self-driving cars or ro
robotics a lot of reinforce
reinforcement learning is used there one
place it's found its way into recently
is recommendation systems have kind of
merged with reinforcement learning
learning. Um, and this is because you
you can imagine there's kind of a
built-in reward for you clicking on a
video and kind of watching it.
Um, so that kind of reinforces that
recommendation and then uh that's where
um you can then kind of recommend a
similar thing and see if that's
rewarding and generates a click or
generates some view time or watch time
or whatever. Um so reinforcement
learning has found its way into a lot of
areas. Um recommendations being one of
them because it's just natural for the
idea of like what um should I recommend
next to generate the most reward. In
this case, the reward is kind of
correlated to did they click on it or
not or did they how long did they watch
watch for longer it's more rewarding
um those kind of things but uh place
places where reinforcement learning have
been used I said self-driving cars um
games so uh one of the most famous
examples if you want to look it up is
the um Alph Go this was in 2016 um the
Alph Go uh algorithm was a reinforcement
learning bot that beat um some of the
world's best Go players, which if you're
not familiar, Go is a um board game
that is a little bit more uh complex
than chess. It has more more uh it's a
larger board um more pieces to it. Um
but they there was a reinforcement
learning powered bot that actually um
learned how to play the game so
effective it could beat um world kind of
masters at the games was pretty amazing.
Um that's the alpha go and that was by
deep mind Google and deep mind in 2016.
That was pretty that was only 10 years
ago not that long.
Um so certain uh we said recommendation
uh even autocorrect um learning to
predict like what is the best correction
uh to generate a reward which would be
like you accept that correction or you
reject it would be a penalty. Um so
reinforced learning has been adapted to
these kind of problems very
successfully. Let's take a look at the
packages that we will use throughout. So
um of course we will rely on these three
which we've already relied on to do a
lot of things like numpy to do numerical
manipulations and calculations.
Uh mapplot lib to do any plotting and
not only map lib but maybe seabour as
well both of those to do plotting. Um,
pandas is a big one because
that's where all of our data is going to
be manipulated and prepped before it
goes into modeling.
So all of that stuff we learn from
pandis is definitely going to be applied
here in this course uh as we actually
build models. Um so of course like these
old ones that we've been working with
quite a bit um still going to be useful
here in the modeling stage. Um mainly
for different reasons though mostly to
get our data prepared to do some type of
modeling or maybe to visualize it before
we do modeling to get a sense of what it
looks like those kind of things.
Um,
sci-fi is sometimes useful for certain
uh um processing like in unsupervised
learning. We'll actually use scyp a
little bit to do dimensionality
reduction or help us do that. Um so
scypi will be used here and there and
we've seen it before with hypothesis
testing. We use scypi like the t test
and z test came from there. Um, some of
the unsupervised learning stuff will
come out of there, but the package we
will use by far the most in this course
is going to be Scikitlearn,
which is here. Um, and we've already
seen a little bit about scikitlearn in
terms of its pre-processing capability.
So, we use the uh minmax scaler and the
standard scaler from there from the
pre-processing module in scikitlearn.
but it has um many different models
built into it that we can use to help uh
do our training and predictions. Um so
it's a incredibly useful machine
learning library. It is the industry
standard machine learning library. Um if
you're going to do anything in machine
learning, it would be expected that you
know how to use scikitlearn.
Now what's really lucky about that is
that scikitlearn is a really easy
package to get used to. Nearly
everything we do in scikitlearn will
mostly follow the same pattern and so um
the code will be extremely simple. They
did a great job with that package of
making things really user friendly,
really simple. Um it's a really
fantastic package and we're going to get
a lot of practice with it uh as we go
along. Every model we build will
essentially be from scikitlearn
and not only like the models but um
doing the training doing the predictions
and then doing the evaluation will all
come from different uh scikitlearn u
modules. So that'll be really nice and
we'll get um good exposure to that
package throughout the course. So if
anything will come away from this course
as um psychit learn uh uh experts
that'll be very nice. So this is this
will be the new one for us psychitlearn
but we'll get a lot of practice with it.
Okay.
All right. So just to recap that lesson
before we move on to lesson three. Um we
talked about machine learning as
learning from data. um which is included
underneath the AI umbrella. But deep
learning is also included under machine
learning because it's still learning
from data but it's learning using neural
networks.
Um we talked about the four different
types of machine learning. We had
supervised, unsupervised,
semi-supervised and reinforcement. So
those are the the different types of
machine learning that are out there. Um
and then we talked about some of the pi
python packages uh that we will use the
main one being scikitlearn and of course
we'll use our older like pandas to
manipulate our data and get it uh pass
it into our model training etc.
But scikitlearn will be uh our go-to for
anything machine learning.
All right. So, some questions for you
guys, some checks.
So, let me know in the chat. What do you
guys think? Uh, which of the following
best describes machine learning?
Which choice do you think makes the best
is the best for this?
Very good. Very good. I see I see a lot
of choices for A and A would be the
correct choice. So machine learning is
definitely um a a subset of AI. that's
underneath that AI umbrella, but of
course we're learning from experience
and of course that experience is
recorded in the data um without being
explicitly programmed. Uh so it's the
exact opposite of BNC. We're definitely
not learning from rules and it's
definitely not just used for image and
speech recognition. It can be used for
many other things beyond those. So yeah,
A is the best choice there.
What do we say here?
Okay. What do you guys think about this?
Which example illustrates the use of
machine learning to enhance customer
experience in an ecommerce company?
In other words, what would be some what
would be some uh typical use cases of
machine learning?
Good. So I think uh C is going to be the
best answer here. Definitely C. So it's
using machine learning to do uh fraud
transactions. So so that would be a
prediction probably a supervised
learning right if if this is fraud or
not fraud. Um and then maybe some
customer behavior uh that might be
unsupervised. So maybe grouping together
customers uh clustering them based on
their data like their shopping behavior
and characteristics. Um that that might
be unsupervised but either way it's
machine learning.
Okay.
Okay. Final one. What distinguishes deep
learning from machine learning and
artificial intelligence? So what's
unique about deep learning?
Oh, very good. Yep. So, deep learning
uses neural networks as so you guys are
right on top of that. Neural deep
learning uses neural nets. That's what
makes it unique. So, machine learning
would be part A. Machine learning is
focused on learning from data.
underneath of that is learning from data
using neural networks which is what uh
deep learning is.
Very good.
All right, let's go to lesson three.
And lesson three has two notebooks.
We're going to be starting with 3.1.
So, you'll want to open up that
notebook. I'm going to go over to it
now. Give you a moment to open that up.
So, we're going to open the 3.1
notebook. Um, there's two of them. We'll
see how far if we can get into the
second one today. Probably will.
Um, but we're going to do the uh we're
going to start with 3.1 notebook. Do you
guys have this notebook? Should be in
your materials for for this course.
Let me give you a moment to open that
one.
Do you guys have it?
All right. So, we're going to start by
talking about uh supervised learning
um in our machine learning journey. So
remember, we're going to talk about uh
supervised and unsupervised after we do
supervised. Um and there's going to be a
lot to cover with supervised mainly
because um there are uh two different
types of problems we can tackle uh which
will be uh we'll talk about in a moment
predicting different kinds of values. Um
but let's talk about the kind of what
we're hoping to learn here which is um
talk about the different kinds of
problems that we'll study which are
these these categories of supervised
learning. Um those two categories are
going to be called classification and
regression. We'll talk about those and
their differences and then talk about
some applications and some uh example
algorithms
and that's just within this notebook. Um
3.2 two we'll get into uh regression in
particular
um which will be uh very very
interesting. Okay. So that'll be our
first models that we'll build will be
over there in 3.2.
Okay. So if you guys remember um
supervised learning is where we learn
from labeled data. So we have input and
outputs in our in our data set. Um and
you so you train a model on this data
that includes input features and
corresponding outputs that are that are
the labels. Right? So um the goal is to
learn a relationship between the input
and the output. Of course that's what
any model is trying to do. Um, and what
this allows us to do is then take that
model and use it to make predictions on
never-beforeseen
uh data. Right? So then we have a
predictive model out of that that we can
use um going forward on new examples.
Um so
remember we will have in our data a
bunch of features which are columns and
then generally one of those columns will
be the label that we're trying to
predict.
And our model is going to try to learn
some type of relationship between those
inputs and the output label. So the
output label could be like fraud not
fraud, cancer not cancer, uh a price, a
temperature, those kind of things.
So let's talk about that. inside of um
supervised learning there are two
different types of learning that we can
do and they're really based on the label
or sometimes that label is known as the
target that we're trying to predict. Um
and depending on that type we get these
two different categories of learning or
two different types of learning. One is
known as regression. So that's generally
when we are predicting something that is
continuous or something that is a
numerical.
So numerical
numerical value. So think of price,
think of temperature, think of revenue.
We're trying to predict something like
that. Um versus something that is
categorical. So that the predicting
something categorical would be like
fraud, not fraud, spam, not spam. um
those are discrete categories and the
problem of predicting categories is is
known as classification because we're
trying to classify examples as belonging
to one category or another.
So we have these two main types of
supervised learning problems. we have
regression and we have classification
and they're going to be handled slightly
differently
um for many reasons that we're going to
uncover. Um one of the primary reasons
is that of course we're predicting
something that's continuous in the
regression case versus something
discrete. So the models have to be
slightly different to account for that.
Um but then a step beyond that is the
evaluation has to be different too. Um I
kind of alluded to this last week, but
when you're predicting a regression,
it's very very difficult to to get the
exact numerical answer. So um generally
we don't care about that. Um generally
we don't care about getting exactly uh
we don't care about getting it exactly
right.
um we just care about getting it um
we're just we care about getting it
nearby, getting it close enough. Um
whereas classification, we do care about
getting exactly right because it's a
discrete category. So we're going to be
able to evaluate that a little bit
differently to say did we get the answer
right or wrong. Regression is going to
be did we get close? Um because it's we
assume it's going to be nearly
impossible to predict a a continuous
number. Um, that's very hard to do.
Okay.
So, any questions on
uh that?
Any questions on those two differences?
Let me give you some examples. Maybe
it'll it'll help too.
So, again, the classification is going
to be predicting uh something that's
categorical. regression is going to be
predicting something that is continuous.
So think about trying to predict the
price of a house based on those other
features we talked about before like
square footage, bedrooms, bathrooms, all
those things we predict the price. That
would be a regression problem because
the price is a continuous value.
Let's take a look at an example here.
Um, imagine we were trying to uh predict
the temperature tomorrow. That's going
to be a regression problem, a a
supervised learning kind of regression
problem because we're trying to predict
a numerical temperature.
Okay? And versus a category like a
discrete category would be this would be
a classification. So this is a
regression on the left. This is a
classification
on the right. Classification
um because we are um predicting one of
two categories. Is it just hot or cold?
Now, we're not saying exactly where that
threshold is on what's hot or cold. That
would be a decision on on what we want
to what our discrete categories actually
mean.
But, um we only have two choices, hot or
cold.
versus predicting the entire temperature
which would be um a numerical prediction
of some exact number. Right? So that'd
be a regression and then on the right
would be a classification. Um now again
why is this so different? You can see
the types of predictions we're making
are completely different. One's a
number, one's a category. But again with
evaluation it's like if the if the true
answer in our labels was 84
and we predicted 83 that's a pretty good
result. That's still pretty close.
That's pretty close to this. So from an
evaluation perspective that's pretty
good. Um whereas like if I predicted
cold and it's actually hot that's that's
a wrong answer. So they're evaluated
slightly different.
Um, and that's something we're going to
see as we talk about evaluation of our
models once we build them is depending
on if it's classification regression,
there's going to be different ways of
evaluating them.
You can kind of see why it's very
difficult to say, okay, we got exactly
84 when it could be any number. Our
model is going to be predicting a
number. That's really hard to pin down
an exact floatingoint number. So, the
best we can do is kind of say, how close
did I get? Like, this would be a worse
answer. If I got something all the way
down here, that's a really long distance
to here. That's bad. That's a bad
prediction. But if I get something
really close, that's better, right?
That's a decent prediction because it's
pretty close,
right?
Of course, being perfect would be
getting exactly right, but that would be
nearly impossible to do.
Okay.
All right. Any questions on this?
Does it make sense on regression versus
classification? We're going to use those
words quite a bit as we go along. So
regression predicting that continuous
value classification predicting a
category
and they're going to be um different
models that do that
different models being used for
regression versus different models being
used for classification.
All right, let's talk about supervised
learning. uh applications here. So just
to name a few, we have HR operations.
Imagine your recruiter tasked with
finding the best candidates. Um so
supervised learning can help by um
rejecting or accepting candidates. Now
this is something that happens quite a
bit even today. Um and that it's kind of
like uh how recommendations happen like
this this resume should be um
recommended this should not um from a
whole pool of applications. Um so
there's those kind of use cases of of um
predicting a category that would be like
a classification. Should we should we
accept or reject the the candidate?
um finance. You see this all the time
with things like risk and loan
approvals.
Um you can uh predict the the the
category of like if the if the loan if
we should accept or reject the loan
application. Um you know that would be a
classification.
Um what's interesting about
classifications by the way so it says
here like we can predict the likelihood
of a of a loan being repaid.
um is a lot of classifications um we we
say that they predict a category but
under the hood they can actually predict
a probability and we turn that
probability into a category. So um you
know like we could say what's we could
say the likelihood of her loan being
repaid is very low. Let's say it's less
than 50% probability. Um then we could
label this as reject,
right? Right? We could label that as a
rejection. Um if it's greater than 50%.
Then we could label this as accept. So
we can set a threshold there
and say okay truly we're predicting a
prob like our model spits out a
probability but we turn that into a
category by saying should we accept if
it's less than 50% we should reject if
it's greater than we should accept.
Okay. So that's something we will see
with some of our classification models
is that they actually produce a
probability and we turn that probability
into a category label
um by by doing something simple like
this putting a threshold on it um for
the for the category.
So finances is used all over the place.
Not only just loans like fraud, we
talked about fraud, not fraud. That
would be a classification.
Um predicting sales revenue, that would
be a regression, right? What is the
revenue going to be in the next two
quarters? That's going to be a
regression problem.
Uh emails like spam, not spam, that's
going to be a classification.
um that's going to operate on the that's
going to take the text input and predict
if this email is a spam or a not spam.
That's going to be a uh supervised
learning problem, but it's going to be a
classification problem,
right? Uh manufacturing supervised
learning is used to inspect and uh
quality and classify products in
different grades. For example, a factory
might use a model to check for defects.
So this is actually something that
happens is you look at images of
products as they go through the assembly
line and you can take a look at those
images and predict if it's a high
quality, low quality, medium quality. Um
so they can be this is a classification,
right? They're going into different
categories of quality. Um so it's much
much like a manual kind of intervention
by some uh QA or quality control uh
specialist.
Okay. But that's a classification.
So in the maritime industry, supervised
learning can be used to predict current.
So current level
um and that can be used to forecast uh
supply and demand. Um so those would be
like regression models that are used to
predict um kind of like temperature but
in this case like title levels.
We talked about fraud already, so that's
there. Um, that would be a
classification.
Okay,
any questions on these uh examples?
Of course, there's many more. Um
recommendation is kind of like a
supervised learning problem uh where you
are
taking examples of things that people
have viewed in the past or or reviewed
in the past and using that to predict
what they would want to watch in the
future. Um so recommendation is
supervised learning. Um and it's like a
classification, you know, trying to
predict um uh certain number of
categories of of uh shows or movies that
you would want to watch. Um
and that's something that we will study
in the future. Recommend we'll we'll
have a whole lesson dedicated to
recommendation as well.
All right.
So when it comes down to the uh actual
models themselves, so there's going to
be lots of different models that we are
going to cover. Um and they are um going
to be different in their purpose and
kind of their uh what kinds of problems
they're used for. Um and uh their their
how they actually train is going to be
different. Um, but at a high level,
they're all trying to do the same thing,
which is learn some sort of relationship
between the input data and the and the
label, right? That's really what they're
trying to do because they're all
supervised. They're they have those
labels, trying to build some
relationship there. Um, they just do it
differently.
And what we're going to study is the
pros and cons of a lot of these models,
like when would I use one of them, when
would I use another. Um, so we'll try to
talk about that as we go along. Um, but
they're all trying to learn some
relationship between the input features
and the output, right? So you have to
keep that in mind. They're trying to
model that relationship. They just do it
in different ways. Okay? So as we go
along and learn about new models, um, we
will learn the details. will learn the
ins and outs um and those pros and cons,
but they're no matter what, they're all
trying to uh learn that relationship,
right? And be able to make predictions
on new data.
Okay,
so here's a list of models that we will
cover and work on throughout the uh the
sessions that we have. um we're not
going to do them all in one one sitting,
but um the first one that we're going to
start with and that we'll cover today is
going to be linear regression.
So we will cover linear regression and
then we'll cover the rest of these guys
mostly in the context of uh
classification.
So, um, what's interesting is some of
these guys can actually be used for both
regression and classification as long as
you make, um, certain adjustments to
them. They have variations that can be
used to do classification and regression
is very interesting. Um but we're going
to start with linear regression today
and then work our way through the rest
of these models when we do um we're
going to do a separate lesson four on
classification. So these all these guys
will come from lesson four.
Um and then uh we will do this guy in
lesson three in the 3.2 notebook. We'll
do all about linear regression.
Yeah, I so logistic regression is a
classification um which is kind of
strange that its name is regression but
it's doing a classification but the the
reason is that the logistic regression
um computes a probability. So it does a
regression to predict a number but that
number is actually a probability. So it
it produces a result that's between it
produces a probability that's between um
obviously uh zero and one.
So it uh and then we take that
probability and we turn it into a
category
like a spam not spam fraud not fraud.
Um but so so logistic regression is kind
of special. It's sort of like a
regression but it's predicting a very
specific type of value which is a
probability. So for for that reason it's
a classification uh algorithm primarily.
So we'll study that one in lesson four.
Uh but yeah, that's that's why it's
under that kind of umbrella of
classification is because it's it's
producing a probability as its main
output which we can then turn into a
category as long as we interpret that
probability as um in the right way uh
like the probability of spam,
probability of not spam.
Okay.
Okay. So, let me focus on um
let me focus on linear regression. I'm
not going to go through all of these
other use cases because we haven't
learned these models yet. Um so, I don't
think they're good. Uh I don't think
it's good to read about them yet until
we've covered them. So, once we cover
them in lesson four, I'll come back and
describe these examples to you guys and
we'll see why it makes sense. But I
think for a linear regression um which
is what we'll cover next, let me talk
about that example. So a prototypical
example would be like predicting the
house prices that we've seen in that
house price data set.
So um if we wanted to uh if we wanted to
predict um if we wanted to estimate the
market value of a house so the price
um we could do that by using the
features such as number of bedrooms,
square footage, location, age of the
property. Um and you know then when a
new when a new house comes on the market
we could estimate what the price should
be based on those features. So linear
regression is a good one to predict the
price like a housing price. Um and we'll
actually practice that in the next uh
notebook.
So we'll we'll uh and then all these
other now there's descriptions of these
other models but again we haven't
covered these guys yet. So I don't want
to really go through those until we get
to those models. So we get to those I'll
come back and mention the example.
Uh can K andN be used for clustering?
No. So um the clustering model is going
to be different. It's going to be uh K
means
K means that's the primary clustering
model. Not K nearest neighbors. K
nearest neighbors is used for uh it can
be used for regression. It can be used
for classification.
So we'll we'll talk about K andN which
is the K nearest neighbors in lesson
four.
It sounds really similar. Yeah, it
sounds really similar but K means is a
clustering algorithm that's that's
slightly different
different uh there's no labels used at
all. This K nearest neighbors is a is a
supervised learning algorithm. It uses
uh labels.
Good. Any any other questions so far?
Okay.
So that being said, let's move on to the
3.2 notebook.
Let's move on to that which will be our
um first discussion around uh
regression. So going into supervised
learning and regression. Give you guys a
moment to pull up this notebook.
But yeah, you want to pull up the 3.2.
We'll do this one next. So we'll focus
in. And so our plan is to do regression
first and then we'll talk about
classification in lesson four
which we will cover all those other
models which you you could use for
classification uh on that list but then
we're going to talk about linear
regression uh first.
All right. So, we have a a big agenda.
This is a big notebook um to go through
a lot of material here surrounding
regression. So, we're we're going to
start with linear regression and see um
how we actually perform it, what that
model is doing. Um which we've kind of
seen the idea of it a little bit
already, so it should be somewhat
familiar. Um and then we'll talk about
how to adapt that linear regression idea
to um nonlinear what's called nonlinear
regression which is going to be using
like polomial
uh features. We'll talk about how to do
that. Um and then a big big big topic
for us is going to be evaluating the
model. So it'll be it'll be quite easy
to actually build it. building the model
will be really easy but evaluating and
interpreting that will be uh a lot of
interesting work there um because we
want to know what the performance of
that model is once we have it built
right we want to know how good of a
model is it is it worth using or do we
need to retrain it or get new data or
change the model up to talk about that
um how do you determine what to do based
on that performance
um and then we'll talk about here um a
couple things. We may not get to this
today, but regularization
which is used to boost the performance
uh in certain situations um whenever the
model is kind of uh performing um poorly
against test data even though it
performs pretty well on training data.
In that scenario, you can use offshoots
of linear regression that do some uh
what's called regularization. We'll talk
about that.
Um and then we'll talk about
hyperparameter tuning uh generally as a
strategy which is something you
generally do want to do when you're
training machine learning models. Um so
again these two we may not get to today
but um quite a quite a lot to get to be
prior to that mainly centered around
evaluation and building linear
regression.
Okay. So pretty cool. we'll get to our
first kind of model here. This linear
regression
to start with.
Okay,
so let's start with uh linear regression
here. Um, and really what linear
regression is attempting to do and I
want to show you this in this picture is
draw this line sometimes what is known
as the line of best fit. So this is our
model that kind of goes through the data
and it's generally a good predictor
um because if you give me um features uh
if you give me new features and let's
say they are let's say you give me a
feature that's right here.
So you say, okay, I have a feature
that's this value on the x- axis. Then I
know all I have to do is plug that into
my line equation, and I will generate a
a value that's like right here.
Okay, that's pretty that's on that line
at that input. And that's going to be my
prediction for what the output variable
should be. It's just going to be
something on that line. And what you can
see is this line is a decent estimate
for this data because it slices through
this pretty evenly. So it's a good guess
as to what the output should be given
any one of these inputs. It's a it's a
good estimator this line. And so our
goal building a linear regression is to
kind of build the equation of this line.
So we want this equation.
Equation of this line
is going to be our model.
Yes, it's going to look just like that.
MX plus B or yeah, MX plus C. It's going
to look exactly like that. uh except
that it's going to be more than just MX
because we have um generally more than
one feature. So you think of X as a
feature um it will be more than just MX.
It will generally be like uh it'll
generally look like this
and then plus maybe some bias here plus
uh an intercept. Yeah, it'll generally
look like that. So, yeah, you're exactly
right. MX plusb is the right idea.
Exactly right.
It'll generally look like that.
Nonlinear. It can be adapted to
nonlinear. Yeah. If we transform, we're
going to talk about that. If we
transform all of our features in a
nonlinear way, um we can apply linear
regression to it. Yes. And and that
would be a nonlinear regression. So yes,
we can do nonlinear things too.
We'll talk about that.
Okay. So linear regression again is the
art or science I should say not really
art but it is an exact science of
finding the equation of this line that
fits through this data. Um now why one
thing you should be thinking about is
why is this line a good predictor and
the argument is that if you take a look
at this distance from these blue points
so let's say these blue points are our
actual data points this line is going to
be found such that it minimizes this
distance
from the points to actually I should
draw it this way from the points to the
line.
So, we want this distance to be um
actually I should draw it that way, this
way. We want this distance to be kind of
at a minimum. So, it would be bad to
draw a line all the way out here because
then that's a lot of distance, right?
So, and that would be a lot of error um
contributed from not being able to
predict those points in our data set
very well. Um which is our training
data. That's why we have labels, right?
that that guide us in building this
line. Um so our goal is to build that
line especially so that this error or
this distance can be as minimum as
possible. Right? Which are all these
distances from these points to the line.
We want those to be as minimum as
possible. So our goal is to find this
equation.
So we're going to build a model that's
going to find this equation.
of the line
um such that our error
is minimal.
And what is the error? The error is the
distance
of our data points
to
to the line that we build. So
essentially what we'll do in order to
train this will be to adjust the
parameters or the or in that like I
think it's really good you brought up
the MX plus C. Basically the M and the C
will adjust. So we adjust those
accordingly to make this distance as
small as possible.
Okay to minimize that distance as much
as possible.
Okay.
So, um where is regression used? We've
already seen some examples. Here's some
more uh advertising like predicting
sales, predicting um oil and uh oil
production and demand. Those are like
forecast those are regression problems.
Um retail like demand forecasting for
inventory. Um healthc care predicting um
uh the levels of certain um uh blood
markers or you know something like that.
Um real estate predicting prices based
on those uh talked about like square
footage, bedrooms, bathrooms, those
things. So regression is used again
whenever we want to predict a number a
numerical output um that's a regression
problem.
So this kind of regression we're talking
about here is generally
um known as uh a when that equation is
linear that is known as a linear
regression. And so go back to that
picture when we have a when that
equation of the line that we find is a
linear equation meaning that it is
exactly the form I've been telling you.
So it's it's something like um weight
time feature
plus weight time feature
plus weight time feature
and then maybe some intercept um term
like some some bias term there. Um this
is a linear equation because all of the
features are to the single power. So
it's a linear power and this is a linear
combination of features with with those
different weights. So this is a linear
model
because it is uh it's what in math we
would call this a linear equation right
everything is to the first power. It
resembles mx plus b. It is a linear
equation or a linear model. Um so when
we talk about linear regression that is
a regression model so we're predicting
some continuous target that assumes we
are model our model is formed from this
kind of equation a linear equation.
So this is going to be our our model for
a linear
uh regression.
Okay.
And so when you when you train a linear
regression, your goal is to learn these
weights so that you can plug in um you
can plug in any one of your uh input
features and you um can generate a
prediction. You can which is going to be
something on that line, right? It's
going to be a value that's sitting here
on this line.
We put in all of our features and we end
up there somewhere on that line.
This output.
Okay.
Okay. Let me pause there. Any questions
on the linear model here or why it's
called linear regression?
Okay. And by the way in these notes um
this bullet point here where it says it
uses the least squares criterion to
estimate the coefficients that is
exactly what I said earlier with the
distance. So the distance is based on
the square
of this this quantity like how far away
you are from the line is based on this
square distance here and here and here
and here. So what we're trying to do is
find the least distance or least squares
which is that minimum distance. So
that's how we find all of these weights
is from minimize. We basically tune them
enough using our labels. So here's our
label which is the y. We basically plug
in our data and tune those enough to
minimize the error. It's it's a it's an
optimization problem,
right? We we're trying to find the
minimum of this quantity which is that
best fit line.
Okay.
So we have linear regression
um and we can do a simple linear
regression that only has one feature. So
if it only has one feature that's
exactly the so if there's only one input
feature sometimes that is known as um
simple regression or simple linear
regression and there's basically there's
only one feature. So one independent
variable is the feature.
There's only one feature. And so this
equation resembles the
exact equation that you guys just put in
there, which is um mx plus b,
right? It resembles exactly that. Um
we're just using different symbols for
those like beta beta 0 and beta 1. But
um basically exactly that simple line is
only one feature. So, and that's because
that line is going to um that line is
going to be generated uh according to
that equation. So, here's kind of what
it looks like.
This is the best fit line through all of
these blue dots. This is something we're
going to be able to build. We're going
to be able to build that equation um
pretty easily in scikitlearn.
So, we'll be able to find that um and it
won't be too hard. So this line will be
um y = beta 0 plus beta 1. So some
weight beta 1 times the only feature we
have x1.
Okay. So in this case um we would be
predicting sales. So sales would be the
value basically the label that we're
trying to predict and the feature that
we're putting in is uh I think it's the
number of TV expenses. Yep. TV expenses
which is on the x- axis. So there's one
feature which is um TV expense.
So um on this graph this would be this
would be our model.
Okay that would be our model. We only
have one feature and we have um these
two weights. We have an intercept B 0
and or beta 0 and then a one weight
which gets applied to that one feature
beta 1. And so our model would have
certain value for beta 0 and a certain
value for beta 1. That's what get that's
these guys get learned
learned during
model
training.
Okay. So those are what get learned
during our model training and they get
learned by a a a least what's called a
lease squares algorithm that is trying
to minimize that distance. It tries to
tweak beta 0 beta 1 to minimize this
distance of this line
um this line
to all of these points
trying to minimize this.
So imagine taking a line and kind of
moving it around and turning its its
slope, its angle um to try to find that
best fit,
which reduces that error the most.
Right? That's kind of what we're doing.
Uh can I explain? Yeah. So uh sales is
in dollars and and TV expense
um
uh
TV actually I think it's the other way
around. I think the sales is actually a
quantity. So I this is number of sales
that we have and TV expense is um I
think I think it's in dollars. So how
much money how much expense um did we
put into the into the product and then
this is how many sales did we have of
that product.
So I think it's the other way around
but what this what this graph is showing
is the blue points are our actual data
points. Okay. So so we have a collection
like we have a data frame that has so
imagine we had a data frame that has the
uh true values.
So it has the um TV expenses.
Um it has points that are like one. So
it has points that are like 120 and then
the sale sales could be like 700
700 units, let's say. And then it has um
so this is just our data set, right?
This would be like in a data frame that
we have. And then we had ones that were
um 50 and then this could be um this
could be 400 let's say and on and on and
on right so this is our data and this
data is plotted in the blue so these are
these blue points here
right so these are the blue points here
and the red points are is our model so
we built a linear regression model um
where we are putting in some values
we're putting in some fake x values
here and generating some predictions
which is this line,
this linear uh regression line, right?
And that line is derived from this data,
right? It gets learned from this
supervised uh examples.
Does that make sense?
That line is derived from the data. it's
actually um learned from like the line
of best fit is learned from that data
and the actual data is in the blue.
So you can see we're trying to build
this such that this distance is kind of
a minimum
so it's an optimal fit
to balance out these distances.
So it's just plotting. So it's just
building that relationship between the
input and output. Like when the when the
expenses are higher, um we seem to have
more sales.
Okay.
Uh what's perpendicular like the
distance? This should be this should be
perpendicular because it's a distance
here.
Is that what you mean? Like the distance
from the real points to the line? Yeah,
that should be perpendicular
because it's it's a it's a distance
formula.
Okay.
All right. So more generally now do do
we usually have one feature? No. So
generally we expand this to the more
general case where we have more than one
feature like what we see in the housing
data right where we could predict a
price but we have many different inputs
like bedrooms, bathrooms, square footage
etc.
So more broadly
instead of simple linear regression we
have what's known as multiple linear
linear regression which means we have
multiple variables or multiple features.
Um so this is exactly the equation I've
been talking about. Um so we just extend
that that one into many features. So
which is this case and then a intercept
term which is uh um there as sometimes
known as the bias. Um
but this is the intercept term to kind
of orient the line to start out in the
right place. Um and uh but this is the
um this is the equation that we would be
building the model. This is our model
essentially, right? This is the equation
we would be learning.
Intercept is like a constant. Yeah. So
if if all of the features were zero, um
this is what our our data would be. This
is what our result would be. If
basically if this was zero, this was
zero, this was zero, it would reduce to
this as the prediction. Yeah. It's like
a constant. Yes.
So in in geometry, the intercept's
actually really important because it it
orients where your line should start. So
it orients like so so these values are
kind of like the slope. They orient the
tilt of it. Like should it be tilted
like this or should it be more sloped?
But the intercept orients where it
should start like vertically like should
it start all the way up here? Should it
start more down here?
Um, that's what the intercept kind of
tells us.
Okay, so this is the situation. This is
going to be our linear regression model
that we will be building most of the
time because we will have again these
are all going to be features.
So this is some feature the X this is
some feature this is some feature
X1 etc. These are all features and what
gets learned during the training are
these coefficients. So all of these
coefficients including the beta 0ero um
will get learned. So these will get
learned
um from our data right they get learned
they will be trained from our data um in
order and and how do they get trained
it's from reducing that distance we try
to get that line of best fit by tweaking
those betas enough to uh until we reach
a minimum distance but there's there's
an algorithm behind that um that that
scikitlearn will run for us to find that
best fit Um, so we don't need to do that
manually, but that's that's the process
is basically tweaking those weights to
end up with that line of best fit. So in
higher dimensions, instead of a line,
you get more of what's called a plane
here. Um, which kind of looks like this.
So the best fit is actually this plane
where all um, it kind of dissects all
these points just like that um, in
higher dimensions. So this is uh instead
of a line you get this in in three
dimensions you get this plane like this
but it's still it's like a line of best
it's just a more general line of best
fit. It's still the same idea. Um we're
still trying to um come up with the best
coefficients to minimize that distance
from our from our points to the line.
Although in higher dimensions it's no
longer a line. It's more like a plane
like this. So you're trying to minimize
this distance from here down to the
plane
here up to the plane
in higher dimensions. So I want you to
keep in mind what we're trying to do
before we go into the code because the
code's going to make it seem really
really simple and that's because
scikitlearn is great and that's what it
does.
But we should realize that there's
something really complex going on which
is again finding the best value of these
weights
that minimizes the distance of this line
to the data points that we have. So
there's an algorithm there that will
keep trying to make adjustments to this
based on those distances. So it's going
to use those distances as a guide to
kind of tweak them to find the one that
results in the lowest amount of
distance. So we keep making tweaks, keep
making tweaks, keep making tweaks and
eventually we try to find we converge to
the set of weights that gives us that
best fitting line. Um and and there's an
algorithm there that occurs. Now luckily
that gets abstracted for us a bit behind
um scikitlearn
um finding that best fit. So there'll be
a function that we use in scikitlearn
when we build the model that will go
ahead and find the best weights for us
and that's then we now have our optimal
model right that then we can just plug
in different values of these features
and generate a prediction which is going
to be this uh result right so so that's
what we're ultimately trying to do is uh
train the model which will uh find all
those optimal weights and then uh we can
predict with it which would be plugging
in different feature values to to
generate a prediction.
Okay,
so let's see how that happens. It's
actually going to be super easy um with
scikitlearn.
So uh in this scenario we have um we're
going to import our pandas because we're
going to load our data from that. Um, so
of course we need some data to work
with. So we're going to load this uh
CSV.
Um, I
uh so I was not actually able to find
this CSV for this example, but I mean
that's okay because we'll do some we'll
do other examples where we'll work with
the data. If you happen to have it, um,
great. I didn't see it in in my files.
So just have to take the word for it
that these are the this is that TV and
sales columns here um from this data
set.
Okay. Um as an example. So um just to
see how it's fit um what we're going to
do and this is going to be a very
standard process for us for building a
model. These steps are going to be very
very standard for us which is going to
be first of all splitting the features
away from the label. That's the first
step that we always will take. So if you
take a look at this code, it's taking
all rows but only the first column.
Okay, so it's extracting all the
features from the data frame um which
happen to be which is just the first the
first column uh which is the TV uh
column right just that column there and
our target variable which is our label.
So our target variable aka the label um
is the second column, right? It's that
that sales column.
Um and so our first step here, let me
call that out here. First step is to
always split apart
features from labels.
Okay, so we put all those features into
a data frame called X and we have all of
our labels into technically a series but
uh sort of like a data frame, right? Um
called Y, which is just the um which is
just the uh uh labels. So that's just
the TV values. Um now you're going to
see why we do that. It's because we need
um our our features and labels split
apart to put them into the model
building function. It expects our
independent variables or our features to
be separated from our answers or our
labels that guide the model building.
That's the first thing you got to do is
separate those.
Okay, so this code will separate those
out into a capital X and a lowercase Y.
And that's actually pretty industry
standard notation. Whenever you split
apart all your features, usually you put
them into a data frame called capital X
and then you have a lowercase Y to
represent your labels. That's actually
pretty standard.
So it's pretty standard that um X
represents
features
and
Y represents labels
label column
whatever our label column is in this
case it is the sales because we're going
to be predicting sales
using the TV column the TV quant expense
quantity.
Yeah. So what it so the assignment is
that we are um the assignment is that we
are
uh we are um splitting apart our data.
So that when we first read in the data
um it is a data frame right that has two
columns TV and sales.
Oh perfect thank you Tim. I will I will
go ahead and so if we look at this data
it only has those two columns right it
only has those two columns. Okay. So
what we're doing with this is we are
splitting apart
our our independent variable our
features. So this this x will contain
our features
and y will contain
our label.
Does that make sense? We're splitting
this data apart. So, we're only grabbing
that first column here to be our
features. And then we're we're grabbing
the second column, which is the sales,
because we're going to predict the
sales. This is our label. We're going to
we're going to build a model to predict
the sales given the TV input, TV expense
input. So, the first thing we have to do
is split apart the features and the
label.
Okay, that's the first step we usually
will take. And the reason we have to do
that um just to reiterate, the reason we
have to do that is because our model
will expect our our data features to be
separate from the label. We will pass
those in separately.
X is TV. It's the first column
because we're using eyeling.
We are predicting the sales given the TV
expense value.
Yeah. which is why we split it into so
this is the second column right the
index one column
uh you just put in read CSV and pass in
the URL so you could so exactly the code
that was up earlier from Tim
um you just do this
and then data equals ed read CSV URL
So we split our data into X and Y here.
All right. Now, one other step that
we're going to take that's a very very
critical step and you're going to we're
going to see this step over and over and
over and over again. So splitting apart
into X and Y will become we'll do that
over and over and over and over again.
Not only that, but doing this next step,
which is what's called a train test
split. Now, let me show you what the
train test split does. It takes our data
and it's going to split apart our data
that we have, our X and our Y data. It's
going to split it apart into a
percentage that will be used to train
the data
and then a percentage that will be used
to test. Now, why would we want to do
that? It's mainly so we can do
evaluation. So, we build the model over
here and then we test it on data that
has not seen before. So, we reserve a
percentage of the data to be used for
test. Usually this this data is um
somewhere between uh 20 to 30%.
So somewhere between 20 to 30% of the
original data. So that means the
majority of it is used for training. So
the majority of the of that X and Y over
here is going to be between 70 to 80%.
will generally be used for for uh for
training. Okay. So somewhere between 20
to 30 the industry standard is some
anywhere in between there. Um a lot of
people like to use 30%, some people like
to use 20%. Um anything in that range is
acceptable. Um we will I think we
generally will favor like 30%.
um to be used for testing. But um the
the point is we don't we don't want to
mix those together. We want those to be
separated out so that we can have a fair
evaluation, right? We want to train our
data on this train our model on this
data and then see how well it performs
on this data that it has never seen
before.
Right? So in order to have data it's
never seen before, we're going to take
our x and our y and we're going to split
it using this function called train test
split that will do this kind of
splitting for us. Okay, so scikitlearn
has a function called train test split
that will go ahead and we're going to
pass our x and our y and we'll pass in a
percentage like 30% that we want to
split out into a test set and then the
remainder of that the 70% will be used
for training the model.
Okay.
So what we're going to get let me redraw
that. So what we're going to get out of
this for the train test split is we're
going to we're going to have an X and a
Y per
training and test. So we're going to get
now we're going to get an X train
and a Y train.
So we're going to get training features
and training labels. And then we're
going to get test features
to plug into our model and and test
answers or test labels
to do evaluation because what we should
be able to do is build the model over
here and then apply the model on this
data. Meaning we can take these features
and plug it into our model and then see
what answers we get and compare those
answers to this testing data. Right? We
should be able to do that to generate an
evaluation.
Okay. Now you may be wondering why do we
do any of that? What's the purpose of
that?
Evaluating it on this test data gives us
a good sense of will our model
generalize to new examples. Right? If it
performs pretty well on this data,
that's a good signal like when it's
performing pretty well on data it's
never seen before, that's a good
indicator that it's going to perform
pretty well when we use it on brand new
examples
um in the future.
Right. So that's a that's why we do this
evaluation on this data that it has not
seen before. It's going to see this
training data, right? We're going to
train the model on that data. But that
model will never be exposed to this test
data until we do the evaluation
and and generate some metrics to see how
good is this performing
and does it have a good chance of
generalizing to never before seen
examples which is what we want right
because we're going to use this model in
the real world. It's going to be being
used on new examples that it hasn't seen
before. We want it to perform well. So,
this is kind of our test, our
evaluation.
Okay. Any questions on the We're going
to do this in a moment. I'll show you
what it looks like in the code, but any
conceptually any questions on the train
test split idea. It's a very very
important idea that we um basically use
part of the data to train it and then
another part of it to evaluate. It's
very important we do that. By the way,
this has a term um this in machine
learning this is called cross
validation
because we are using one data set to
train the model and then we're cross
over we're crossing that over into
another data set to validate it which is
the uh the the testing that.
So this is called cross validation. Um
there's actually many ways to do cross
validation. That's something we'll
study. This is a very simple way of
doing cross validation. There's more
complex ways. You can take your data and
you can actually divide it into many
sections
and basically train it against most of
these and evaluate it against one at a
time and then rotate. So that's another
way to do cross validation. We're going
to study that. Um but this is the this
is the simplest way to do it here.
Okay.
So let me show you what you get when you
use train test split. So uh we're going
to import from sklearn.
We're uh from the model selection
module. Now we haven't used this before.
This is our first time using it. But
here's our model selection. We're going
to import this train test split function
and we're going to use it on our X and Y
and we're going to set a test size of
30% which is which is.3. So our test
size
is 30%.
Converted to decimal
right converted to.3 so that means we're
reserving 30% for that test set. Um you
can set a random state. Now that's
completely optional. Um the random state
is for reproducibility
because what the train test split is
going to do is it's actually going to
shuffle the data and then split it apart
into the 7030.
So um yes, the seed. Exactly. It's like
a seed. So it's it's saying like when
you do that shuffling every time I run
this notebook I'm going to get the same
result but it's going to be random the
first it's going to be random but I'm
going to be able to reproduce that
randomness with that random state. Yes,
it is like a seed.
Uh it's you can choose any number to be
your your um your random state. It 42
isn't important. You could choose zero.
You could choose one. Um, you could
choose any positive integer. Um, 42 is
kind of like the uh industry standard.
It's it's you'd have to look it up why
it is. Um, apparently 42 is a special
number. Um,
in kind of the history of development of
this stuff, there's nothing really
special about 42. You could choose a
random You could choose a random seed to
be uh zero. That's fine. It it doesn't
really it doesn't really matter.
Um you just want you can choose it to be
uh one, two, three. Um you can choose it
to be 15. You can choose it to be
anything you want it to be. It's really
so that your your shuffling is
consistent. Every time you run this
notebook, you get the same shuffle
result. So I'm always going to get the
same rows in these splits.
Hitch. There it is. I knew it was from
something.
Yeah. So 42 is kind of like a
it's it's just used ubiquitously
uh you know as kind of a um paying
tribute to the Hitchhiker's Guide to the
Galaxy, but it's no it's there's nothing
that special about 42. It doesn't it's
not going to change our result or
anything.
It's just so that this train set split
is going to shuffle our data and split
it apart into 7030.
You just want to set this to something
so that you get a cons every time we run
this notebook, we get a consistent
shuffle.
And so the data in these sets
are uh consistent. That's all.
Okay. But do you guys see how we pass in
our X and our Y and we generate four we
generate four different data uh
quantities here which is we generate
training features, test features,
training labels and test labels because
again we are generating these four
different we're generating data on these
two different sets a training set
and a test set. So we have training
features, training label,
and then test features, test label.
Okay, that's why it's so important to
split apart our data into the X and the
Y. We need those split apart in order
for this part to work.
So by the way, these two steps we will
always do for any model we build. We'll
generally do X and Y and then train test
split in order to generate the data that
we will use for building our model.
Okay. So this this data here is going to
be what we actually use to guide the
training of our model. So it's
definitely supervised, right? Linear
regression
um we we will use that
Okay, so we haven't built the model yet.
We're just getting our data split apart
and ready for the training. We haven't
actually built our model yet, right?
That'll be coming up uh in a moment. But
this is getting our data ready. We
started with our data frame. We split it
apart into uh an x and a y. And we split
that into a train test split. And um you
know then we can uh then we can go ahead
and um pass in to our model training
which we'll do in a moment.
Um you that's a good question. You could
run so what you could do is you could
run
um should we import numpy? Let's see.
We did. Okay. You could run the average
on the um you could check the MP mean on
the X train and see how it compares to
um
see how it compares to X.
So you could you could do that and see
what the average of this feature is um
compared to the average of the original.
They may not be perfect because we are
taking a reduced data set size. So I
don't think there's really any good
there's not like a one-sizefits-all
validation we can do because we're
taking a random shuffle and taking a
percent. We're taking 70% of the data
out. So we're not guaranteed to maintain
the same statistics. We can see if
they're close.
Um but does that make sense? Like we're
not guaranteed to get the same stats
because we're taking a slice of it.
We're taking 70%.
So it's not guaranteed to to to
be the same distribution really.
Delete that.
Uh is it good practice? Yes, it is.
It is. Uh 30% is the industry standard.
Anything between 20 to 30, so 0.2,
0.25.3,
any of those are acceptable. It's really
up to you. Um I mostly see 30%.
Mo I think.3 is is a good good practice
to use for sure.
Um I did explain random state. Uh random
state is so that you get consistent
shuffling. Um you can set this to any
integer that you want it to be. It it
doesn't really matter. Um you can set it
to uh 100, you can set it to 10, you can
set it to 15. Um it just ensures because
what this split will do is it will
shuffle the data first. It'll shuffle
the rows and then um split it apart into
the into the train and test sets. So you
set the random state so that the next
time you run this you get the same
consistent shuffling. That's the only
that's the only thing it it helps you
with because it is randomized but when
you set a random state um it's so that
like if you run it again you'll get the
same shuffling.
You'll get the same the shuffling
matters because it it it uh dictates
what ends up in in these sets.
Okay.
All right. So let's see let's do let's
build the model. Um and let me show you
how easy this is going to be to build
the model. And this is really how it's
going to be for every single scikitlearn
model will basically look the exact same
for training it which is what's going to
make it really really nice. So the first
thing we have to do is import our model.
So from scikitlearn we're going to be
using a linear from the linear model
package or the linear model module I
should say within sklearn we're going to
be importing the linear regression
and we're going to create an instance of
the linear regression here.
Okay, so linear regression and look how
easy this is going to be. Nearly all
nearly all sklearn models use
ffit function to train.
So every one of them, no matter which
one we use, like the decision tree, like
the um logistic regression, any of those
like we use for classification that are
going to be coming up in lesson four,
they're all going to look the same in
terms of it's going to run.fit,
which is um scikitlearn's
uh generic function for training your
model. So this will execute the training
once we run this code. And what that
again the linear regression training is
going to do that least squares distance
procedure or algorithm to try to find
the right weights. It's trying to find
those weights that minimize that squared
distance uh from our line that it's
trying to build to the data.
And what I want you to notice is what we
put into the ffit. See how we put in the
training data where we put in the
training features and we put in the
training labels. Now this is supervised.
So of course we put in the labels,
right? Of course we put in these labels
here and of course we put in our
features here. So we're putting in all
of our examples from our training split
into this ffit which is going to train
the model uh so that we can we can use
it for prediction.
Okay, it's really fast. If I run this,
it's going to be pretty much instant.
Pretty much instantly it gets trained.
And you can see here we now have a
linear regression. you can see in this
little box. Um, and it and this
information says that it has been
fitted. So, it's now ready to be used.
Right? So, we now that's it. We've
trained our model. We try that's how
easy that was. We did ffit. Now, what we
should realize is there's a lot of work
going on behind the scenes of this ffit.
Okay. There's a lot of work being done
there to do the least squares algorithm
and find those weights and and create
that line of best fit. Right? So there
there's a lot of work being going on
there that's going on there behind the
scenes, but scikitlearn is abstracting
it away for us, right? And all we have
to do is fit when we're using this code.
Really easy. Really easy. Fit. And there
we go. We've trained our linear
regression model.
And by the way, if you want to see what
the coefficients are, you can actually
extract them if you do so if you take
your lin regression and you do um
coefficients like this.
COF with a with an underscore. So this
gives us the trained
weights
coefficients
also known as the coefficients right.
Um so if you run this you can see uh
right now we have this coefficient here
um which is the only coefficient we had
on our feature. So we only had one
feature coefficient there.
And we can take a look at our intercept
which is this.
So this gives us the train weights
and so we can look at the intercept we
can look at the the the coefficient. Um
so obviously if we have multiple
features our model has many features
it's going to have more values in that
coefficient but the intercept is just
the single value 7.23
and then the coefficient
is 0.046. So that's the weight that gets
learned.
Is there a size limit? No, not really.
There's no size limit. Um,
no. You can use as much data as you
want.
There's really no size limit other than
what like what you can fit in memory.
I'd say that's the only limit is
basically what the amount of data that
can fit in memory.
Okay.
All right. Were you guys able to run
this? Were you guys able to run the
linear regression ffit?
Okay, perfect.
Perfect. Do you Okay, great. Great.
So, we have a model and we can use it to
predict. Um, and so that's actually what
we're going to do next. If we go down
here, um we're going to have a function
that's going to um build a scatter plot
of our original test data.
Um so we're going to have our test data
here.
Um,
and we're going to then take our uh
we're going to take our training data
and plot we're going to use the uh this
data versus our sales predictions. So
you can see we're going to you this is
how by the way this is how you use the
scikitlearn model to predict. You have a
fit to train it and look at the function
you use to predict. It's literally just
called predict. That's how easy it is.
and you pass in your data, all your
features into this predict and it
generates a prediction for every row. So
every row in these features in this data
frame um will end up with a prediction
using our model. So what we're going to
do is plot our training uh features
against the predicted sales to see how
good of a fit that really was.
Okay. to see to see the regression fit.
Okay. And so there's the regression fit.
We have all of our test data here
plotted in the green. We have our blue,
which is our um we have our our blue,
which is our uh um training data line
that we built our model on. So that's a
pretty decent fit. Um and then our test
data is here. We just plotted in the
green scatter. But the thing I want you
to see is this prediction, right? We we
were able to generate some predictions
on that training um by running our
predict function with our model. Now
this model has been trained. So we've
already fit it and now we're using it to
predict, right? And so we're predicting
the sales and plotting that on the y
ais. So the sales are we're using the
predicted sales there which is our blue
line. So this is our line of best fit.
So this is our model prediction.
This is our model predictions. Right?
You can see it's a pretty decent uh
line, right? Pretty decent line of best
fit.
Of course, there's some error here. Like
there, you know, it's not perfect, but
it it does a decent job of being a best
fit line.
Okay.
So look how easy that was to
just to recap this to fit our model was
a linear regression.fit and of course
we're going to do more examples. So no
worries uh on that we're going to see
this many many many times throughout
this notebook. But we have linear
regression.fit to train it and then we
have linear regression.predict
to and we pass in our features and that
generates a predicted output.
Right? So what this is actually doing is
is computing this quantity.
We could do either.
We could do either. Um, so we could do,
so one thing we could do is plot uh, so
we could swap it out. We, we could do
either one. It doesn't, it's not a big
deal to do the training set. We could
do, so we could plot X test and then we
could plot linear regression X test.
So it's it's a similar line. Um it's
just different input features, but the
line is going to be the same. Just
different inputs,
but the coefficients are the same,
right? It's the same line. It's just we
generate different outputs.
So yeah, you could do either one.
This is This is honestly this is
probably better. I see what you're
saying. This is probably better because
this is the line of best fit through
this data. So that probably makes sense
to do to do predict on the test set.
Agreed on that. Probably makes about
most sense.
But you could do either one.
Yeah, I think that would be the most I
think that makes the most sense is for
it to be on the same one just to
validate. So like we could do we could
do training here and then train and
train just to see how that data lines
up. Really, what we're trying to do is
have our scattered data and then our
line of best fit on the same plot.
That's all we're trying to do, right?
So, yeah, I think I think they should be
the same.
I think that makes sense.
These values
or which values do you want to see?
Yeah, we could uh we could generate
those if we just do um let's go down
here. So the the line values
um are going to be uh the prediction. So
um the the
uh test
predictions
equals um
test predictions equals linear
regression.predict predict x test and
then we could uh we could print out our
test predictions.
Yeah. So we can see what those actual
values are on our uh on the test set.
Yeah.
Um we will do that. Yeah. So you thought
we were checking how well our data was
trained. We will do that. Yes, we
haven't learned how to evaluate this
yet. We're going to talk about that
coming up next. Yeah, we will do that.
We just haven't learned how to do proper
evaluation
of a regression model.
But yeah, it's something we're going to
talk about for sure
and see how to do in our code.
Okay.
All right. Any other uh questions on
this example?
Again, big takeaways
fit to train it and then predict to use
it.
Predict on the features to use the model
and make predictions with it.
So here is example. We we made all the
predictions. This these are all the
values that are on that line.
These are all our predictions. And
notice they this is a truly regression,
right? These are all floating point
values. Um so this is definitely a
regression, right?
Okay.
Uh, that's a good question. Um,
I'm not sure if there is
if there's like a verbose
there's not really no there's not really
a verbose. You can I mean you can look
at the source code if you really want to
see you can view the source code to see
um how it's done. I can tell you I mean
so generally linear regression is done
in two ways. Either you use a formula um
to to solve the optimization problem of
minimizing like this this uh distance
from the points to to the line. Um
or you use something called gradient
descent which is how a lot of these
things do it is they iterate through a
bunch of different iterations where they
update these weights according to um a
certain uh basically a gradient of the
the error function. The error function
in this case is the is the squared
distance from the line to the uh to to
the points.
So uh we can compute the gradient of
that and do um gradient descent. So if
you really want to look into it, I would
do some research on like linear
regression gradient descent.
Okay, linear regression gradient descent
to see how that's uh how that's being
done. Yeah, it it's it's a pretty simple
procedure. Um, again, you have the the
notion is that you want to minimize
minimize the loss or the error. Uh, in
this case, the loss is the square
distance. So, it's like um there's like
a it's a formula. It's like a sum of a
square distance from your prediction
um or your label sorry to your model
which is the beta 0 um plus beta 1 x1
plus beta 2 x2
etc like your model and then squared. So
this squared this is the squared
distance here and you're minimizing this
guy which is like a calculus problem.
You you find you basically find the this
is this is a I'm getting so far into the
weeds of this, but this is like a
parabola and you work your way No, no,
you're good. It's it's it's a good
question. Um you work your way down to
the minimum of it. Does that make sense?
Like you're working your way down here
and you do that through a descent
process, like a descent iteration.
Um
so
that's how these are found.
Um, but you don't see that happening in
the background. But if you look at the
source code, it I guarantee you it would
be it's either going to be this or
they're going to use the they're going
to use a a a matrix formula to basically
solve an equation um that involves this.
Basically, the derivative of this set
equal to zero and you find the minimum.
Either way, you're finding the minimum
of this.
Okay. But yeah, I don't think Psycharn
has like a uh maybe there's some type of
verbose flag you can look for.
I don't think they have that though. Not
that I've seen.
All right.
So I have uh an important um concept to
talk about next which is going to be uh
called overfitting and underfitting
um which is a really important concept
that's related to the training and test
data we just split apart to do
evaluation.
And um essentially the the issue with
machine learning is that it's not
perfect and it can struggle in different
ways. And the two ways that it primarily
struggles is going to be overfitting and
underfitting.
So overfitting is a situation where the
model basically memorizes the training
data so well that it's it fails to
generalize to new examples. So what we
see with overfitting is this exact sign
here where we have really good
performance on the training data. So
when so when we do that train test split
we see a really good accuracy or really
low error on the training data but it
does not perform anywhere near that on
that test data split. So what that means
is that the model is overfitting to the
training data. It's basically memorizing
it and it's not able to generalize very
well.
Now, why does that happen? It's usually
because the model is way too complex.
And that means generally you need to do
something to reduce the complexity.
Either you need to use a simpler model
or you need to use some type of
technique to mitigate overfitting. And
we're going to we're going to study some
of those techniques coming up in this
notebook. uh we might not get to it
today, but we're going to study
particularly what can we do to prevent
overfitting because overfitting is the
more common issue with machine learning
models. They tend to do so well at
learning from data that they pick up on
small details and patterns in the
training examples that they're exposed
to. They don't do a great job at
generalizing to new examples. They can
struggle with that.
So that's overfitting is struggling to
generalize to new examples, but you do
really well on your training data. So it
appears like you have a good model, but
it it's not able to go and make
predictions on test data very well,
which means we would not want to use
that model in the real world, right?
Because it's not able to generalize
outside of what it's already seen. And
that's not a good thing if we're trying
to use it for real world examples,
right?
So overfitting is a real issue. Um you
see it all the time. I've seen it many
many times in the real world, real
industry uh work that I've done.
Overfitting is a is a challenge for a
lot of machine learning models. And so
we need some techniques to overcome
overfitting. And we're going to study
some of those uh coming up shortly.
Um, one of the things that we can do,
one of the one of the things that we can
do to detect overfitting is exactly what
we just did, which is you split apart
your data into training and testing so
that you have a chance to do an
evaluation to see if you're even
overfitting in the first place. You want
to see that performance be consistent
from train to test, right? You want to
see consistency. What you don't want to
see is performance that drops off on the
test data. It's much worse. You don't
want to see that. That means that your
model is overfit uh to your training
data and it's not going to perform well
in the real world.
Okay. So, we're going to have a couple
ways to uh overcome that. Talk about
that. Um now, the opposite can actually
happen as well, which is called
underfitting. And underfitting
refers to the fact that a model is too
simple and it actually just performs
poorly across the board. So if we see
poor performance on the training and
testing data, that's a good signal that
the model's underfit and that means it's
too simple usually and you should try
using something more complex. Um, so the
best way to combat underfitting is to
use a more complex model. And as we go
through and learn about the models,
we're going to learn about which ones
are simple and which ones are complex.
So we're going to have a scale of kind
of complexity. And if you're
underfitting, you want to bump up to the
to a more complex model. If you're if
you're overfitting, one way of combating
that is to actually go down to something
more simple. Go the opposite way to
something simpler. So we need to learn
right now we've only learned linear
regression
but we will learn other models you know
in the future and we'll we'll talk about
uh their complexity and how they're
related to each other.
Okay, but these are two issues we see
just to draw that out again is if we
have a train test split where we have
7030 split let's say and we perform
really well over here but we go to apply
that model over here and it fails its
accuracy drops off significantly more
error that's that's definitely
overfitting which is not good
right and then underfitting is just not
performing well in either case so even
on the training data itself your your
accuracy is not very good. So you're not
really learning effectively. You're
underfitting your model. So that's
that's um underfitting case.
Okay.
All right. Now the issue is that it can
be very difficult to balance these two
and get it correct. That's what makes
machine learning a little bit
challenging is getting this balance
correct of simplicity and complexity. So
you don't want to be overly complex that
you overfit, but you don't want to be
overly simple that you underfit and
you're not able to learn effectively. So
there's a bit of a tradeoff there. And
this trade-off is typically known in the
community as bias variance trade-off. Um
in which case uh it's basically like a
complexity simplicity trade-off. It's
another word for that. Um,
and so, uh, it's it's thought that, um,
if you, uh, if you have very, um, if you
have a situation where you're able to
fit the training data very well, you
risk not being able to generalize. In
other words, you risk overfitting, and
it's hard to um, it's hard to combat
that in a way. Um, and um, on the
reverse side, if you have something
really simple, um, you risk not learning
enough. Even if you're trying to combat
that overfitting, you risk not learning
enough and your model just doesn't
perform as well as it could. So, there's
a bit of a trade-off there of trying to
find the right balance between something
complex enough to learn, but something
not overly complex that it's going to
not generalize to new data. That's the
challenge. Um, like I said, we are going
to have techniques to overcome this. So
luckily there are things to basically
overcome this trade-off and um and help
us along the way so that we don't
overfit. They basically prevent
overfitting
um and allow us to use complex enough
models um that that won't be overfit.
This is in the um this was in our uh
lesson 3.2 notebook. So you want to pull
that one back up. We were working on
Monday.
Um, and just to recap this a little bit,
remember we were building a linear
regression, I wanted to recap some of
the steps we took there, um, that we
will be doing over and over again. And
really the same kind of steps, uh, that
we do here, we'll do in a lot of our
model building. Pretty much all of our
model building um, that we do, whether
it's regression or classification,
doesn't really matter. um we'll still be
doing a lot of these steps which are um
remember first we split apart our data
into kind of a features and a label
uh x and y and the reason that's
important is because um the model
training uses the features and the label
um to help train the model, right? They
use those separately. Um so we want to
split those apart whenever we can. And
so we have usually uh it's a good
practice to call your features capital X
and your labels lowercase Y. And what we
do with that is remember we immediately
split that into what we call the
training in a test set. And the picture
we had for that was something like this
where we had about 70% of the data
we used to train the model against and
then the other 30% of the data we use to
test the model against. Meaning that we
build a model over here and we apply it
to this set over here um to make
predictions. And then the that's where
the supervised learning really comes
into play, right? is on this test set.
We already have the answers. We already
have the label. And so we can apply our
model to this to the features over here.
Predict uh what the the label should be
and compare that. We can get a a metric,
right, that compares how close we are in
our prediction to the actual values. Um
and that was some of our performance
metrics. I'll recap some of those that
kind of measure that distance away from
our predictions to what the actual label
is. Um, but remember we had this train
test split function which helps us split
apart our features and our labels into
these uh four sets of data. So we have
our training features, our testing
features and then our training labels
and our testing labels. So we have all
of those and um really these two guys
are going to be used to train the model.
That's why they're called underscore
train. They're going to be used to train
that model. And then the then we're
going to predict on these set of
features and then com use those
predictions to compare to this set of
labels, right? That's on the test test
set. Um and you notice here our test
size is set to 30%. Um, that's a pretty
standard number. Anywhere between like
20 to 30% is pretty standard. Um, we'll
typically use.3, but it could be 02.
Anywhere in between is fine.
Okay, so we had that. Hopefully that uh
we remember that from Monday.
So we had a train and a test set. And
then building the model was actually
really really easy. Once you have those
train and test sets, um, we just import
our model object. So from uh scikitlearn
sklearn
um linear model uh module from that
package we import the linear regression
model and then we do um linear
regression.fit
and we pass in our features and our
labels and this is again this is where
that supervised learning is really
coming into play because we're passing
in these labels.
That's really what makes this work,
right? We need those labels to help
guide the model to make those updates.
If you guys remember, the model is
something that looks like this.
So, this was a bunch of different
coefficients
um times the features,
however many we have. Um, and so these
labels are really taking the place of
this and they're helping us um make the
correct updates to these to these
coefficients or sometimes we call them
weights. Um, these B 0, B1, B2. Um, we
find out what the optimal one is to get
the best fit, right? To get the line of
best fit. Um that's what the model
training when we call this ffit ffit
that's really what it's doing in the
background is finding all those
coefficients right to end up with the
line of best fit that has the lowest
amount of error.
Okay so hopefully that makes sense.
That's just a dofit fit um to train our
models. And that's really going to be um
the case for
uh pretty much every single model that
we uh train with scikitlearn. It's
pretty much going to be a fit. We pass
in our training uh features and our
training labels.
Okay, so we had that and this was the
visualization of that where we had our
test points kind of scattered and we see
our line of best fit is the one that
goes through there with that minimal
error. That's that's the whole goal.
Pretty decent predictor.
Okay. And then we talked about
overfitting underfitting. So just to
recap this overfitting is the concept of
our model basically memorizing our
training data. It performs really well
on that training set but it is not able
to generalize outside of that. So it
performs poorly on the test set or data
that it's never seen before. Um and
that's overfitting. So the reason that
it overfits is generally the model is
too complex and it needs to be um it
needs to be simplified a bit. And one of
the things we're going to do today is
see a couple of ways we can alter the
linear regression model um if we are
overfitting to prevent overfitting. Um
so there's going to be ways to handle
this. Um and so we're going to explore
some of those today.
uh underfitting is kind of the reverse
of that. Remember, it's where the model
is not learning enough. So, the
performance is poor even on the training
data. It's not good on the test data
either. Um that is a sign that the model
is probably too simple and maybe we
should use something more complex like
go from a linear regression maybe to use
a polomial regression. Um or maybe use
an entirely different model altogether.
um if we're underfitting, our
performance is poor, it's a good signal
we should try something else. Um
okay,
so we talked about those
and one of the things we also talked
about was evaluations. If you guys
remember, we had different metrics that
we could compute to get a gauge of how
good our model is actually performing.
Um one of those was MSE, which is this
mean squared error function. Um so we
did this example during class last time
on Monday um where we uh were able to
generate the mean squared error. That's
one of our metrics. And we can see what
the mean squared error is on the
training set and see what it is on the
test set by um just passing in our um
training predictions and our training
labels, our test predictions and our
test labels. pass those into this mean
squared error function and it computes
the MSE and that's that's a helpful
function from the scikitlearn metrics
um package um or module I should say and
we'll be using that quite a bit to do
you know evaluation of of especially of
regression right mean squared error is
pretty is probably the most common uh
performance metric we can have and if
you guys remember what it's really doing
is measuring these distances So mean
squared error is kind of like the
average distance away from our our
points to the actual um to the
predictions which the predictions are
all on this line. Um so it's like
measuring on average how how much error
do we have on average right? Um, and the
idea is the closer to zero the better.
Generally means that the distance away
from our prediction to our points is
pretty low. The closer to zero it is.
Um, which is pretty desirable.
So a low MSE is kind of what we're
looking for. Um, closer to zero the
better. And so um if one model has if
one model has um a low lower MSE than
another, it's it's a better performing
model, right? It has less error.
Okay. And then we also looked at the R r
squared or sometimes known as R2 um
score. Um this is another metric that we
could use that measures the the
variability
um of uh the predictions and if our
model is capturing that variability um
well um and so R squar is has a range of
0 to one one is better that means the
model is capturing the the changes in in
the um output it um our predictions
follow along with those same changes um
so they're pretty close um so closer to
one would be a better score. So we have
those kind of metrics. So like on this
data um this would this would show that
this model was underfitting remember
because this
mean this MSE was bad and this MSE was
bad.
Um and what we should think of these in
the units of what our labels are. um
especially if we take the square root of
this the RMSSE that was another metric
we had um the square root of this is
actually in the exact units that we um
have for our labels. So uh in this
example this was the um this was the the
units or the sales versus the TV
products, right? Um and so this would
indicate that on average if we take the
square root of this um
and the square root of this um we have
uh
um we're on average about 11 sales units
off squared. So if we take the square
root of that um it's somewhere around 3
to four um somewhere in between three
and four units off. And this is as well.
Um, and because both of these are still
not close to zero. Um, this would be
under fit. And this shows that as well.
This isn't that close to one. It's
decent, but it's not um not that close
to one. So, we would say and performance
is poor on both training and test sets.
That's the key indicator of
underfitting. It's poor on both.
Yeah. Exactly. High MSE correlates to
underfitting. Yes. Yes. And it what's
key is it's high MSE on both on both the
training and the test sets.
If you have a high MSE on your test set
but a low MSE on your training set,
that's overfitting, right? Where it's
not generalizing from the training set
to the test data that it hasn't seen
before. That's overfitting. So the key
is high MSE on both sets.
All right. So we talked about that. Um
we did polomial regression last time. So
that was um doing
that was uh making a curved graph um by
transforming the features into polomial
features and then doing linear
regression with that. So you guys
remember from Monday we did this where
um we took our features and uh transform
them according to this polomial features
from scikitlearn. So we can go all the
way up to degree whatever degree we
want. So we put in four here, but
there's nothing special about four
really. This is just testing it out. Um
and we generate the the polomial
features and we can fit a linear
regression on those polomial features
and we get a slightly better model,
right? Um it fits the data a little bit
better than just a straight line. this
curved line with the polomial
features um performs a little bit better
and we could see that with the MSE right
we could evaluate the MSE of this um and
it would be lower
it would be lower than the curve line
and that's something we could do um we
would just have to pass in these test
predictions training predictions and
then the the test labels and training
labels and passes into the mean squared
error function and we could compute that
right wouldn't be hard to
All right. And then finally, where we
left off, um, you know, is on our
performance metrics. So, we talked about
mean squared error. That's that average
distance away from the labels to our
predictions. Um, and we take the square
root of that. It's it's basically
measuring the same thing, but it's the
square root of it is um more
interpretable because it's in the same
units as our label.
um mean absolute error is is the average
distance of the absolute value. So it's
not the squared distance formula like a
uklidian distance but it is a absolute
value. So it's a little bit um less
sensitive to outliers. They don't get
magnified as much. Um but it's not
typically used as much as a mean squared
error would be with regression. um we
talked about the last time because um
the distance formula or that distance is
actually what's used to train the model.
So it's a more natural um fit for a
performance metric for it.
All right. And then we had R square. We
just talked about that closer to zero
would be um worse. Closer to one would
be better. That means that the model
explains um all the variability in the
in the predictions. Uh it captures those
predictions um closely to the labels
um very well. So uh one would be better.
Closer to one would be better.
All right. So that's where we left off.
Um we're gonna pick up from there with
cross validation. Um, we've actually
already seen one method of cross
validation. So, we're going to study um
we're going to kind of recap that and
and then um talk about cross validation
in general um and look at some more
sophisticated techniques of it um coming
up next. But before I do that, any
questions about anything we've covered
um to this point in in the recap or
anything from Monday? Any questions on
that?
All right. So let's talk about uh cross
validation. Um now this term cross
validation refers to a technique that
evaluates performance. And what it does
is it divides our data into essentially
um training and test sets which we've
kind of already seen. And then we are
able to train a model on on the training
set, evaluate it on the test set, and
that's where that's where we get the
name cross validation because we're
crossing over our model from one batch
of data used to train it over to another
set of data used to validate those
predictions. Um, and there's actually
different ways to do cross validation.
So cross validation is a bit of an
umbrella term for multiple ways to do
that. We've already seen one way of
doing that um which I'm going to scroll
down to is um known as a hold out cross
validation. So that's um what we've been
doing so far. So this is just um
generating a train and a test set
train um split.
Um that's the that's what's known as the
hold out cross validation method. Um and
and this is exactly what we've been
doing so far, which is you split your
data into some type of split, usually
7030,
um of a train and test
and then you um train your model on this
section of data and then apply it to
this to evaluate performance. Right? So
that's that's what's known as the hold
out method. Um it is uh you know
relatively simple. It's pretty fast to
do. Um, but there are more robust ways
to try to divide up our data a little
bit uh more evenly. Instead of just
having one split, we can actually do
many splits, which is the idea of um the
next kind of cross validation I'll
cover. But hold out method is one that
we've already studied. It's the most
basic type of cross validation you can
have. Um so hold out this is the most
basic
and we we've already been we've already
been uh working with this type. Okay.
So we've we've already seen hold out
method. Let me uh explain to you a more
sophisticated method a little bit more
advanced of a cross validation um which
is known as Kfold cross validation. So
this is um going to be a little bit more
advanced of a technique but this is the
idea of kfold is that you take your data
set
and you split it into k number of what
are called splits or folds. So you take
your data and you let's say it was let's
say k equals 5. So we have five splits
here.
Okay. So let's say k equals 5. we have
five splits. So what we're going to do
is we're going to we're going to train
our model on K minus one of those folds.
So if K was five, we had five splits.
We're going to take our model and train
it on four out of five of those uh
splits. So let's say it's these four.
We train it on these four.
Okay. And then what we do is the one
split that's left over we will we will
test our model against that split. So
we'll test here.
Okay. Now, this sounds very similar to
the hold out method where we're doing a
train test split, but it's a little bit
this kful cross validation is a little
bit more sophisticated because we repeat
this process that I just mentioned over
and over for all combinations of the
splits. So then what we'll do, this is
just one trial that we'll do it again,
but this time we will pick um four
different splits.
So, this time we might pick,
let me do blue. This time we might pick
this one, this one,
um,
this one,
and this one.
And then those four we will train our
data on. And then we will test against
this one. Okay? And we'll do we'll
repeat this
repeat for all combos of the folds.
Okay. So we'll repeat that. So
essentially what we're doing is rotating
through. Every time we rotate through
one of the folds is going to be left out
as a test set. Now this is a little bit
more robust than just a train test
split, right? because we are exposing
our model to more of the data in in
doing this, right? Because we're going
to split it evenly into five or 10
splits. Those are pretty common um
number of folds to use. 10 or five. Um
those are the ones I've most commonly
seen. Um but we're going to by rotating
through which folds are being used for
training, which ones being left out. um
we are exposing our our model to more of
the data this way than just doing a
single train test split. Right? So now
what do we do with with the results is
every time we do this we we generate um
an MSE let's say or some type of
performance metric. So let's say we
generate an MSE from this guy,
we generate an MSE from this version and
we generate an MSE for all combos.
each combo we generate MSE and then what
we do is we average
the metrics
or the in this case uh if we use MSE we
would average those together. So every
time we do a fold combination and we
keep four of them for training, one for
test and we rotate through all those
combinations, we are going to generate
an MSE for every combination
then we're just going to average those
MSE's to get a final. So the final MSE
of cross val of this kffold.
So the final metric
is just the average of the uh
performance on all of the fold
combinations. Okay. So our final MSE, we
just average all those MSE's from all of
our combinations.
Okay.
Now, what's the advantage to doing this?
It's way more robust of a estimate of
the of the performance of the model
because we're exposing it to all
basically all of our data, right? We're
getting a sense of how it performs
across all those different folds. Um
rather than just doing a single train
test split, which is a bit it's basic,
it works, but it's a bit basic. Um so
this is more robust estimate of the
performance.
Now, what's the drawback to doing this
is that it's more intensive. So, if you
have a lot of data, this is going to be
pretty expensive to do because you're
going to have to especially you have a
high number of folds, right? You're
going to have to divide your data into k
number of folds and you're going to have
to do this over and over again. Um, and
if it's a large data set, it might take
your model a long time to train. It's
going to be a little bit more uh
computationally intense than if we just
did a train test split.
Okay, we just did a single like 7030
split. We only do that once. We only
train the model once, right? We train it
on the 70, apply it to the 30% test data
and evaluate performance that way. Um,
so we're only really using the model and
training the model once, but in this
kfold, we're going to do it um, you
know, k number of times essentially
or I should say one for every
combination that we have to work through
of of all the folds.
Okay.
All right. Does that make sense? Any any
questions on Kfold cross validation? So
K K K K K K K K K K K K K K K K K K K K
K K K K K K K K K K K K K K K K K K K K
K is an important uh number here. It
it's how many folds, how many splits do
you have? A typical value for K is going
to be somewhere like five or 10.
So 10 folds or five folds. Those are
pretty pretty standard
from what from what I've seen.
But does the does the concept make sense
or is there any questions on it on in
terms of um you're always going to leave
one fold out. You're going to split it
up into K number of folds. Always leave
one out. Train on the rest of it.
Evaluate on that one that gets left out
and then rotate those through. And
you're going to do that for every
combination and average all those
metrics.
And by the way, there's going to be an
easy function in scikitlearn that will
do this for us. So managing all these
combinations will be really easy. It's
actually just built into scikitlearn. So
we don't have to um we don't have to do
this all by hand. Okay, this will be in
scikitlearn. It'll handle doing all
these combinations of folds for us and
computing the average metric will be
really easy. So um
we don't have to worry about that. We're
going to see an example of this coming
up shortly.
All right, of kfold cross validation.
But this is a this is a really widely
used technique. And again like the
purpose you may be wondering like what's
the purpose ultimately of doing this?
It's to get a sense of if our model is
going to perform well on new data.
That's really what we want to know. like
is the model going to perform well when
I start to use it on new data that it's
never seen before and this kffold is a
decent indicator of that because we are
varying which data it sees across many
different folds right so it's a it's
kind of a good um proxy to exposing it
to different kinds of data each time and
seeing how it performs
right all right because we're working
our way through each one of the folds
there's always going to be one fold left
out. We're going to change which fold
gets left out each time. And um that's
sort of mimicking the idea of we're
going to apply our model to new data and
see how it performs. And it's it's new
data every fold.
um how we know which model is best suits
for which scenario because we have Yeah,
that's a good question. Um, so my we're
going to learn this as we go along
because we haven't covered all the
models yet, but generally the best
advice I can give on that is
you you generally want to start as
simple as you can get and then if it's
not performing well then work your way
up to something more complex.
So we are going to have models that are
simpler. We're going to have models that
are more complex. The rule of thumb is
to start with the most simple model that
works.
So you're usually going to have the same
ones that you're going to try in the
beginning. And linear regression is a
very simple model. It's usually the
first one you want to try for regression
because it's the simplest.
Um, and for classification, we're going
to have a similar like logistic
regression is the simplest kind of
classification model we could have. So
usually want to start with that and then
if it underfits like if we see it's
producing a lot of error then we work
our way up to a more sophisticated
model.
So um that's the way we that's the way
it should usually go is simple to
complex it based on their performance.
So we evaluate it and then we can repeat
the process. If it's not performing well
we can try something different that's
more complex if it's underfitting.
Uh this is a good question. Does a model
reset after training each k minus one
fold? Um yeah, it's essentially like a
blank model every time uh every fold. So
um we imagine like you have a brand you
have a fresh model every um k minus one
combination. Yes.
And the reason the reason it has to be
that way is because you don't want the
other folds influencing the model that
like on on the next combination. You
don't want the previous combination to
influence the results on the next one,
right? Um you want it to be a fresh
evaluation on every combination of
folds.
Okay.
All right. So, let me describe to you a
variation on what we just um talked
about with the K-fold. So, there's
another cross validation known as
stratified K-fold. And um this is the
same exact procedure as kfold except
that when we this is used for
classification.
Um so when we do classification
uh we want to make sure that the
different categories are going to be um
split amongst those folds in a
proportional way. So we don't what we
don't want to happen is um when we split
apart the data. So, let's say we have
let's say we're predicting um spam not
spam. What we don't want to have happen
when we do our splits is we don't want
to have all of the spams end up in one
fold and then every other fold has no
spam, no spam, no spam, no spam, right?
That's not very good. Um because if we
if we train against all these guys, we
have no shot at predicting spam when
they've never seen spam before. So
stratify kayfold is is used in
classification
and it's to um it's to make our splits
ensure that they have basically a
balanced number of categories for each
split. Um so that we don't end up with
certain splits with way more spams than
not spams. Um so we we do what's called
stratifying where we make sure the
proportions are balanced across each uh
split. So this is only really useful in
classification, not really necessary in
regression because we're predicting a
value. But if we were predicting a
category,
like in classification like fraud, not
fraud, we don't want to do the split and
have every single fraud example um by
bad luck in our shuffling and split end
up in one split and every other um every
other split has no examples of fraud.
Right? So we want to stratify this to
spread out those um frauds against all
the other splits. Um so uh again um
scikitlearn will take care of that for
you. Um but if you're doing
classification and you have an
imbalanced data set um you you really
want to make sure you stratify k-fold.
um imbalanced meaning that you have a a
um different number. Like if you're
doing fraud, not fraud, you have way
more not frauds than frauds. Um where
where that category is imbalanced,
you want to make sure it's balanced
across all your splits.
Um so this is this is useful in
classification only, not really
regression, which is what we're talking
about right now. Um but it's just a
variation on this that ensures when we
do those folds um the data is
distributed evenly amongst those folds
as much as we can. The labels are I
should say.
Okay. So that's stratified kfold. It's
the same same procedure once we have our
splits. It's the same where we do k
minus one of them. We train test on that
last fold um and then rotate through all
the folds and and average all the
metrics. the same exact procedure. It's
just the splitting itself um is going to
be balanced in a stratified kfold.
Okay. So, hold out we've already talked
about um is just doing a single train
test split. We've talked about that. One
more variation that is a bit of an
extreme version of K-fold. So it's
actually the same process as Kfold, but
it's an extreme version is if you set K
equal to the number of data points. So
you basically are um this is a really
really extreme kfold where you um
basically are training on all the data.
Um so you're training on all the data
except one point and then you test
against that one point. Um now why would
you ever do this? Um it's mainly so for
this reason here. It's to um maximize
the amount of training data that your
model gets exposed to because instead of
just doing instead of just doing five
splits
um which would be like
you know these four folds are going to
be used and then we um test against one
fold. um we're essentially going to use
99% of the data, right? One point is
going to be left out. 99% of the data
gets used to train. Um and then we're
always going to leave out one point. And
and the issue is we're actually going to
do that over and over and over again and
rotate that one point to cover the whole
data set. So, we're going to train on
99%, leave one that one point out,
and then rotate through every
combination of points until we've left
out every single point, and then average
all those together. Um, so this is a
this is an extreme kfold. Again, the
number of folds is actually equal to the
number of data points in this case. So,
we have every point is its own fold and
we train on everything but one. Test on
that one. This gets you the maximum size
of your training data because you're
basically gonna have every point but one
used in the training.
This gets you the maximum size. However,
it gets you the maximum uh expense
especially for large data sets. This is
going to be usually you're not going to
use this um especially for large data
sets because it's just too extreme. It's
going to take you a really long time to
work through every single point being
left out. um it's just going to take a
while to do.
So, for that reason, the leave one out
um that that's why it's called leave one
out because it's you're leaving one out
every single time. Um is rarely used. I
I have don't really see it used that
often, but it is an extreme version of
kful cross validation.
Okay. But rarely ever actually used. I
think the the ones that get used the
most are definitely the hold out method
with just a regular train test split. Um
and then uh the other one that gets used
quite a bit is is kfold
or stratified kfold if you're if you're
doing classification,
but certainly kfold in the in a
regression case.
Okay.
All right. Um we're going to do an
example with these guys. So we'll do
that next.
um with with the different cross
validation techniques. Um but any
questions on what they are doing
conceptually before we actually do the
code example.
Okay.
Very good.
All right. So, let's see some examples.
Um, let's go into our code and build a
model and do the different cross
validation techniques on it. Um, you're
going to see it's actually going to be
really easy to do and we it sounds
complex like doing the kfold and leaving
one out and testing. It sounds kind of
complex, but I promise you scikitlearn
makes it really easy to do. Um,
and so, uh, we won't need to do too much
besides just use the right, uh, tools
from scikitlearn. Uh, so we're going to
we're going to see that. Um, so here we
have some imports. The, um, primary, uh,
thing that's a little bit new for us is
going to be these, um, different kinds
of cross validation techniques. So we
have our kfold, we have our stratified
kfold, leave one out. Um, which are
those different cross validation
techniques. Um, these are going to be
used in combination with this cross val
score which is going to keep track of
the different um metrics and then
average them
uh while we do one of these um cross
validation techniques. So this guy gets
used in combination with one of these to
um as as we're going to see in the code
uh to average those metrics um doing the
different folds, right? Perform doing
performance against the different folds.
Okay. And then of course we need a model
using a linear regression. That's that's
the one we've studied so far. Um and
then we have just a regular metrics. If
we want to compute those um using maybe
just hold out, right? And hold out um
which which is just a regular train test
split um we could use these guys to
evaluate performance.
But in a more sophisticated kfold style
of cross validation, we're going to use
this to evaluate the the performance.
Okay, let's see.
So, we're going to be working with this
housing with ocean proximity data. Um,
you guys should have this one. Uh,
so you guys should have this one. So, if
you want to follow along and run it
yourself, um, you can load that one in.
Um, I want to make sure that I have it.
Let me pull that one in. So, it should
be this guy.
I'm going to load that in so I can make
sure I run it with you guys.
Um,
let me run this.
Do you guys have that data?
the housing with ocean proximity.
It's another it's another housing data
set. Um
but it it's a little bit different than
the ones we've seen before. It has a a
special feature for how close it is to
the ocean at different locations.
So it looks kind of like this. If we
load it in and do our head, which is
usually what we do, right? We can see um
we can see that it's got these features.
So it's got uh uh bedrooms, total rooms,
um it's got uh median age. Now this is
this is looks a little strange for total
rooms and um uh bedrooms and population
etc. But it's um
it's it's got those uh it's got those
because it's representing an entire
neighborhood. So it's an entire
neighborhood. And we're looking at this
um this is actually going to be our
label is this median house value for the
entire neighborhood. So what's that
median value uh in the neighborhood? And
this is the total number of bedrooms,
total number of rooms, um population,
households. So, how many houses are
there? Um, median income. And of course,
these are scaled. So, these are um
likely times, you know, uh thousands. Um
but um that's our data. We could
describe it.
So we can see the average age, average
median age. Um which sounds a little um
weird, but that's it's because again
this is the median of data within a
neighborhood. Um so the average of those
is about 28 or 29. Um we have
u
total bedrooms. The we can look at the
men. There's some data that only has
one. So, it's likely only one house in
there. Um, which is what this
represents. There's only one house. So,
there there is some neighborhood that
only has one house. Um, and we see the
median um we see the minimum uh median
house values there. And then the maximum
down here um is a pretty big number.
6,000 households is the largest that we
have in any any one of these
neighborhoods.
Okay. So, just a little bit of
description of the data.
Okay. So, then we can run.info. So, this
is um let me ask you guys, were you able
to load this? Were you able to run this?
If you're following along, were you able
to
load it and take a look at
Okay, great. Great.
Okay, so we're able to load that and
then look at head. Perfect. Um
Okay.
Um and then we run describe which gives
us that uh usual kind of statistical
description. Uh so we can see some
interesting stats about those.
What do you guys notice about the info?
Anything interesting that we see from
there?
Is there any missing data
any features that have missing data? Can
we see
object? Yeah, object type usually is
string. If it's an object type, that
usually means string. Python when we
read it into pandas it usually is just a
string.
So that that makes sense like we have
mostly numerical features but then we
have a this ocean proximity which is a
string.
Yeah. Total bedrooms has nles. That's
right. Because you can see here this
does not equal the number of uh rows
that we have. So this is the number of
rows which about 20,000 rows. That's a
good size data set, right? 20,000 rows.
That's decent. Um we're definitely
missing some data here for sure. Um we
could count how much we're missing
exactly by running this is NATO sum. Um
and so we see that total bedrooms is
missing about 200 uh 200 rows are
missing total bedroom uh value.
Okay. And then one thing I wanted to
look at is yes, this is a string. So
what remember what we can do with those?
That's a categorical.
So ocean proximity
is a categorical
string
feature.
So we can take a look at its value
counts, which is usually a good idea to
take a look and see what possible values
that feature could be. So if we look at
our
um what are we calling this? Housing
data.
housing data
ocean
proximity
value counts.
So, here's the different types that that
one can be. So, there's some
neighborhoods that are less than 1 hour
from the ocean. There's some that are
inland. There's some that are near the
ocean. There's some that are near a bay.
There's even five of them that are on an
island. So, these are the different
values of the ocean proximity. So,
remember, you can always do that. If you
see a string feature, you can always
take a look at what its um categories
are. And it looks like most things are
less than 1 hour from the ocean, but
it's kind of evenly distributed here. Um
otherwise
very few islands.
But as you can imagine like this feature
is probably going to be important for
determining um what the value is, right?
Probably going to be important.
Okay. So, um, we need to deal with these
NLES. If we're going to build a model,
right? So, um, this is all of our
typical data prep. If we want to build a
model, we're going to have to deal with
these NLES. What do you guys think we
should do with the NLES? What would you
what do you think for total bedrooms?
What do you think is a good strategy to
do? Keep in mind, we have 20,000 points,
20,000 rows I should say, and about 200
of them are null.
Right. So about 200 are null. Um so what
do you what do you guys think would be
like a good strategy to deal with those
NLES in that case?
average. We can't ignore it because we
can't ignore that column.
We can't ignore the whole column. So,
something needs to go there.
Probably don't want to make it zero.
I think average is a decent average is a
decent idea. Probably don't want to make
it zero because um that would indicate
that there's no bedrooms and yet we
still have a bunch of total rooms. So it
probably doesn't make sense to do zero.
Average, I think average could be a
decent one.
Now in this example, what we're actually
going to do is we're
rows.
We're actually going to drop the rows al
together. Now, why are we doing that?
It's because we have so much data and
only 200 of them are null.
Okay, only 200 of them are null. So,
we're actually just going to drop the
rows. Now, that's a choice.
Um, that's a choice, right? Is that we
could fill in with the average like you
guys are suggesting. What we're actually
going to do is just drop the rows. It it
makes up less. It makes up about 1% of
the whole data. So it's not that much of
it is missing. We can drop those rows.
So that's actually what we're going to
do here is we remove all the roles with
the NLES by doing drop NA. So this just
drops them. So those rows are cut out.
Um, it's arguable that we could replace
it's arguable that we could just replace
it with something and I think you guys
have good thoughts which is the average
a default
um assume total bedrooms. We could we
could try that. Yeah.
Assign a value based on comparable home
value. Yes, you could do that too.
That's a good strategy is to look at the
other rows that are similar to it and
fill in a value. That's absolutely fair.
Um, in this example, we're actually just
going to drop those rows,
but I think that's totally um totally
valid.
This is a choice.
We could fill NA with different values
such as the average
total bedrooms
um derive a value etc. So we could
derive something which I think Brent you
have a good suggestion that's a good
suggestion. Um we could derive something
like that uh and fill in the blank and
that's I think that's totally valid. Um,
we could take the average of the um
bedrooms. Uh, I meant total rooms here.
Sorry, total rooms. Um, we could fill in
we could fill it in with the total rooms
for that category um or for that row.
Um, many options. In this case, we're
actually just going to drop those rows
because they make up such a small
percentage relative to the 20,000 rows
that we have. It's about 1%. Right? 200
rows is about 1% of 20,000.
So, we're just going to drop them. But
that's a choice. We don't have to drop
them. We could fill in with something.
Um, and if we did that, we would use
fill NA rather than drop NA, right?
Uh after dropping the rows, how many? So
it's just so after we drop the rows, um
after we drop the rows, it's just going
to be we still have all our other rows
are intact, right? So if we look at this
now,
we now have um slightly uh slightly less
entries.
So now we have this this many um rather
than rather than this many,
right? We dropped those 200
But they're all filled in. Yeah, they're
So all the other columns are still
filled in. We're just we're we're
cutting out the whole row. So if you
think about our data set, um we have all
these rows and all these columns. What
we're doing is like if there's a null
here, we're just we're just getting rid
of that whole row, right? And so we
still have all the other rows intact.
Uh, we can drop them because we have a
good sample size. Yes,
that's exactly right, Ronald. Yep, we
can drop them because we have we have
20,000 rows and only 200 are missing
values. So, that's totally fine.
Uh, drop a removes all rows that has any
null. Yes, that's true. It it will go
ahead and just drop any row where
there's any null, no matter what column
it's in. Yes,
index. Yeah, the index is not getting
reset. Um, that's true. So, um, what we
what you can always do is you can reset
the index. So, um, if you want to, it's
optional. We we're not really going to
use the index for anything that
important, right? But what we could do
is, uh, reset index.
Uh,
we could do that, right? Which will
reset it.
So now now it gets reset.
But um let me actually I don't I don't
really want to do that. I'm going to
reset this.
Um,
yeah, we could do that.
Okay.
So now importantly there should be uh no
missing data of this of this new one
where we've dropped NAS. Right. So now
this is good. If you now the reason we
had to do this is because if we try to
build a linear regression and we have
NLES in there. Um the the issue is like
how do you build a model where you have
something like this
and these are null? Like what do how do
you multiply a number by a null?
Um we can't really do that, right?
we can't really do that. So, um,
so therefore, uh, we need to get rid of
NLES like the the null is not really
going to work in there. So, uh, we need
to get rid of them for linear regression
to to really have a chance to work,
right? To train it and be able to use
it.
You got to get rid of those nles.
All right,
any questions so far? So, we haven't
done any modeling yet. We're doing some
We're doing some data preparation before
we get to the modeling. And we haven't
done any cross validation yet. We
haven't set that up. We're just doing
our data preparation before we get to
the modeling. Right? So, we've dropped
some NAS. We've checked it. Um, we're
going to do one more prep step, which is
to um change that ocean proximity
feature into something numerical because
again, how do you build a model where
you're inserting a string into those
like beta 1, beta 2, beta 3 times of
features? You can't really do that when
it's a string. Um, so what we're going
to do, I'm going to get rid of this
because I don't think we really need
that. um is we are going to uh run this
get dummies function which is our um our
get dummies function is our usual one to
uh our git dummies one is our usual one
to um
uh get our one hot encoding.
So this is our uh one hot encoding here.
We now are going to have data that's
like this, right? So we have ocean. So
So by the way, this prefix
um this prefix is OP, which which is
short for ocean proximity, right? So we
have ocean proximity uh less than 1 hour
from the ocean, ocean proximity inland,
ocean proximity island, near bay, near
ocean. So these first five rows are near
the bay. Um so they have a one there and
a zero in the other spots. So this is
good. This one hot encodes that feature
into these numerical uh values,
right?
Were you guys able to run that one? they
get dummies.
So the reason that Yeah, that's a great
question. How did it go ocean proximity?
It's because um that is the only uh
string feature we have. That's the only
one we have. So it it's going to look
for any non-numericals and one hot
encode those however many however many
there are. So whatever objects we have
which are strings, it's going to
automatically oneh hot encode those.
Yeah, we could have Right. We could have
went here and did Right. We could have
done ocean
proximity,
but we only have one of those features.
So it's just going to do that to the
whole data frame
uh on that one feature. So what we're
going to do is um go ahead and split it
into an x and a y um which the x is
always what includes our features. The y
is what we are trying to predict which
is the label. Now, um, in order to
separate those out, what we're going to
do is assign X to be the variable that
is, um, our data frame minus this median
house value column. So what this is
doing is um uh it's not permanently
dropping because we're not uh dropping
it in place but it is returning us a
copy of the data frame with the median
house value column left out right it's
dropped. So this is this is uh something
we want to do because that will the rest
of it will contain our features right.
So, um this will temporarily or I should
say return a copy of the DF with um
median house value
dropped,
right? Median house value dropped. Um so
we go ahead and drop that one. Uh now
remember it's not permanent. It's just
giving us uh the remainder of it which
is this housing data. dropping this and
it's assigning that to X and then we're
taking the actual median house value
column from the original data and
assigning that to Y. So this is going to
be our labels,
right? So this is what we are trying to
predict.
Okay, so that is our Y and that's always
how it is. X is our features, Y is our
labels. Um hopefully that makes sense.
What this is doing is this is going to
get rid of that label column and
everything else will be our features and
then this will get rid of this will just
assign the label column to Y.
All right. And then what we can do is
pass X and Y into our train test split
function and this will generate the hold
out set. So if we want to do the hold
out cross validation this is how we
would do it is we would split the data
into X train X test Y train Y test um
using train test split. So this is what
we did last time. This would be this
would be for hold out cross validation
right where we are uh uh just have that
one one set for testing one set for uh
one set for training one test one set
for testing I should say right so this
is pretty standard train test split um
we pass in that x we pass in the y we
use a 30% test size which pretty
standard
and random state so that we get the
consistent shuffling if we were to run
this multiple times. Um we we get that
uh consistent randomization.
Okay,
so we have that and so now our X train
is a percentage um of the data frame of
the 20,000 uh rows and the X test is uh
30% of that. So it's only about 6,000
rows, which is what um the shape of that
is.
Yeah. X. So X is our features. So we're
we're putting all of our data in that is
our features into X. And so the the um
most efficient way of doing that is um
the most efficient way of doing that is
to
uh just take our data and drop the
median house value column because that's
our label column. So we just remove
that. The rest of the data is our
features. So that's what that's what
this X is, right? It's all of our
feature data. All of our columns that is
not the label column essentially is what
that's doing. And then Y is our label
column from our original data,
right? Y is our label column. And so
this this will um contain all of our
labels which is the median house value.
X X contains every column but the one
we're going to so we we ultimately
decide that but X contains um X is
everything that is not our dependent
variable which is what we're predicting.
So we're removing what we are trying to
predict from X. X should be everything
else. That's always how it's going to
be. X is X is always going to be all of
those independent variables that we're
using to predict the median house value.
So we are going to predict the median
house value. We need to remove it from
X.
So we're we're taking everything but
that column.
So it's the whole data frame. It's the
whole data frame minus this one column
with just the dependent variable. Right.
Exactly right. Removing the dependent
variable and keeping all the
independence. That's exactly right.
Exactly right. So think about it in
terms of the model. Let's go back to the
features. Right. Think about it in terms
of the model. We are trying to predict
this this value. We're building a model
to try to predict this. So we are going
to make sure x is everything but this
right. So this is actually just y.
That's our label. That's our dependent
variable. Right? That's y. Everything
else is belongs to x. Everything else
belongs to x including all of these.
Right? We choose this one to be y
because we're building a model to
predict that. That's our label.
All right. So, we have our we use X and
Y to do our train test split. So, we
have our our training features and our
test features and then our training
label and test labels here. Um, pretty
standard there.
Um, okay. So, this is what's new is if
we want to do k-fold uh validation, what
we're going to do is create a kfold
object. So, we have this kfold from
scikitlearn that we already imported. we
are going to create a kfold um where we
are going to specify how many folds we
want. So that is the in uh inslits
parameter as this says um this is going
to be uh uh in this case we're going to
do 10 folds. That's pretty standard. So
I think the typical number of folds that
I've seen and I've worked with in my in
my career is usually five or 10.
Five or 10 folds is the standard.
Okay. So, we're doing 10 folds in this
case and we're setting a random state
because we're going to do shuffling. So,
in order to produce those folds, we're
going to shuffle the data first and then
split it into five folds, right? So,
this this kffold object is going to
manage creating these splits for us,
right? These even splits. I know I I
didn't draw it even, but um it's going
to manage these five folds for us and
it's going to shuffle the data and
assign them to these different folds and
we're and then what we're going to do is
use those to do our training.
We're going to execute the cross
validation using this kfold object.
Okay, so we create the kfold
um we initialize our model as well. So,
of course, in order to train something
uh in the K-folds, we're going to need a
model. In this case, we're using linear
regression, right? Which is which is the
model we've been studying so far. So,
you have a linear regression. Um now,
look how easy it's going to be in order
to execute cross validation. All we need
to do is um all we need to do is create
a cross file score function
um or I should say use the cross file
score function from scikitlearn. So we
use that with the model we want to
train. So our model goes first. So
that's the linear regression object.
Then our data. So our extra our features
and our label for our training.
And then um let me skip over this for a
second. I'll explain what this is in a
second. Um but then we are using uh the
cross validation technique is our
K-fold. So this is where our K-fold
object goes in the CV parameter which is
cross validation. So what cross
validation strategy are you using? We're
using Kfold and the K-fold we're using
is this one we defined up here KF. So
we're putting that right here for this.
And then um in jobs um allows us to
parallelize this. So if we set it to
negative one that's the that that's the
default um it will do it will actually
train across the different combinations
in parallel um which speeds it up. So
you want to you want to keep this to
negative one if you can. So um now let
me describe the scoring. So what this
means is we put in our metric here. Um
and so you can put mean absolute error,
you can put in mean squared error. Um
those are the two that we can use. And
um the reason we it has a negative in
front of it is because we want to find
the one that has the lowest score.
That's going to be our best model is the
one that has the lowest score. So, we
take the absolute value.
I'm sorry. We take the abs the the the
metric and we take the negative of it.
Um because the highest scoring one is
going to be the closest to zero. Um so
it's just a we use the we use the
negative of the of the metric. Um
because on the number line like the the
highest um scoring one should be the
least um or I should say the maximum
negative that we can get. That's going
to be closest to zero. So if here's
zero, this will be like -1 is better
than -10. Right? So something that
scores um the maximum negative uh
absolute error would be closest to zero.
And something that has more is going to
be on this side.
So this is only the reason we need this
is only just to keep track of the scores
of each individual um fold. Okay.
So the one so the reason we can do that
is at the end we can kind of see which
which combination performed the best. um
it's going to be the one that has the
highest uh highest value of the negative
which is closest to zero.
That's just a convention.
Yeah, it's just because um it's because
the cross validation is looking to
maximize the metric. So whatever has the
best score
um whatever has the best score is
considered the best uh performance. Um
but we are using uh something where
lower is better. So we we take the
negative and like the the highest
negative would be closest to zero,
right? The highest negative is going to
be closest to zero.
So that so it's it's just because like
we want the lower score to be the best.
The lowest score should be the best.
So we take the negative of it. Um and so
something that is more negative is going
to be worse. Yeah, that's the reason.
So something that's down this way is
going to be worse.
Okay. So it runs this
and what you can see is if we actually
print this out, if we print out our
k-fold scores, what we should get is 10
different scores.
And you can see um we have 10 different
uh scores here, which are all negative
because we're taking the negative of the
absolute of the mean absolute error. Um
so what we would be looking for here is
um we want to take the average of these
scores but take the absolute value of
them to get the best performance. So
this is capturing like this is the score
on the first fold combination. This is
the score on the second fold
combination. This is the score on the
third fold combination and on and on and
on. And these are the absolute errors.
Okay, these are the absolute errors. Um,
so if we take a look at computing the uh
average, which by the way, we don't need
this import because we're using the
numpy average. So that's fine. Um, we
can take the absolute value of those um
and take a look at the average MSE
or sorry MAE. Now I want you to think
about this this uh average performance.
So this is our performance right here on
the cross validation.
This is our average
M AE across all of our fold
combinations. So that's a that's an
indicator of our performance, right? Um
for the cross validation.
Now what are the units of our original
uh the original median value? They're
already in the thousands, right? So if
we go to that feature, they're already
in these hundreds of thousands. So this
is not a very good error. It's it's kind
of high, right? Because it's in this is
49,000.
Um that's that's how far away we are in
absolute value on average is 49,000 um
dollars on the median value. That's not
very good. So this score
this score is
um not very good. So this model is not
performing that well and we can see that
by comparing this error to our actual uh
data. So this is right around 50,000
and our median uh house values are in
the hundreds of thousands. So on average
we're 50,000 off when we make a
prediction. That's a significant amount,
right? That's a significant amount on
average um when our when our data is in
about the hundreds of thousands here.
So we are um we have a significant
amount of error 50,000 relative to the h
to our units that our our data is in.
Right? Um so this score is not very
good. Um
and so we see that from the cross
validation. So look how easy the cross
valid is. Again we just do cross file
score. We put in our model. We put in
our data. We put in our cross validation
uh strategy here which is kfold. And we
can generate these metrics across all
the fold combinations. So it's this
function is taking care of rotating
those and doing every combo with just
the 10 different combinations here of
the of the folds.
10 different instances where you have
you know 10 different folds are the ones
that are left out for evaluation.
Um so it's managing that for us using
this data right using this training data
here. Um and we uh we generate these um
generate these scores.
Okay. So that's kf fold. It's not hard
to do. All you have to do is um just use
a cross file score. And we could change
this to mean squared error. That's you
know we could do that too. That'd be
pretty easy. Um, so that'd be no issue.
We just happen to be using the absolute
error here. Of course, we could use
squared error.
Were you guys able to get this to run?
K-fold scores.
It produces an array of 10 10 different
scores, which should make sense because
those are these are the um we're
splitting our data into 10 different
folds,
right?
10 different folds. than leaving one out
to do our evaluation on. So the one that
gets left out every time is what's
producing these scores. So it's 10
different ones get left out when we
rotate through all the combinations.
And so we average these scores
and we get this amount. We get about
50,000 in error on average.
Um, what do you think would be what do
you think would be acceptable? So, if
our if we're predicting the price, like
if we're a real estate agent and we're
predicting these prices and they
typically are
Yeah, close to zero would be great.
That'd be fantastic. Closer to zero
would be better. The average is um
206,000.
So 50,000 is a decent percentage of
that. Um so you know you can compute it
as a percentage right. So 50,000 is a
decent percentage of that. Um probably
you want this to be less than 20,000
would be about 10% error. 20,000
right? So maybe like 30,000 somewhere in
there.
Yeah. 10% would be 5% error. 10,000
would be 5% error. That's true. That's
true. So that would be that would be
much better. So being closer to zero,
like the smaller the better, of course.
Of course. Um but yeah, I would say an
acceptable percentage of error is
probably 20%.
Probably 20%, which would be um like
40,000 or less would probably be
acceptable.
Usually when we usually when you build
models um 80% accuracy is usually uh
considered decent.
Usually considered decent
80%. So I'd say 40,000 or less would be
kind of ideal.
Does that make sense
to answer the question?
That's a good question. What value is
acceptable? I think probably less than
40,000 would be ideal. That's right
around 20% error.
All right, so that's K-fold. Um let's do
just a regular hold out now. So this is
just using our training and test data.
Um doing model.fit and calculating an
MSE on the test data. So this is this is
just the um hold out strategy here where
we just have um this is less robust but
it's a lot quicker to do and easier to
set up. Right? So um this is using the
hold out strategy. So just a regular
um train test split.
Are we going to rebuild the model? No,
not necessarily. There's some things we
could do most likely. And like one thing
we did not do was scale our features.
Remember I said that's a pretty
important thing to do is to scale our
features. We did not do that. So that
would be an enhancement to this that
we're going to So I I actually do think
we'll do that later. Yes. So I think we
will actually do that now that I'm
thinking about it. Yes. One of the
things we can do is scale these features
using like a minmax scaler or a standard
scaler. that's actually going to help us
um that's going to help us do better
predictions.
So that that's one thing we could do. Um
but yeah, we will we'll try to see if we
can get better.
It should help it. Yeah, usually you
want to scale you want to scale the
data. That's something we didn't do in
our preparation step. We did a lot of
the things we should do. We removed nles
and we did one hot encoding to the
proximity feature like this one. Um
those are good to do but we didn't scale
any of these other we didn't scale any
of the features right we didn't scale
any of them. Um it you it will have an
effect. It usually when we scale it
it'll be a better model.
It'll it'll learn a little bit better if
we can scale the data. Um so that way
like these
um like ages aren't you know drastically
different than like in scale than total
bedrooms or income
uh those kind of things. So we usually
want these to be in a similar scale
range.
So we'll we will I think we'll scale
them coming up in a bit and it should
help the model.
We've talked about that before, right?
Scaling usually is a good idea to do
when you're prepping your data for
modeling.
No, you want to you want to scale your
test data as well. You're going to do
both. You're going to scale your
training data. You're going to scale it.
So that's actually a good point you
bring up is any transformations you do
on your training to build your model,
you should also do on your test set so
you get an applesto apples comparison.
You should always do the same
transformations.
Yes. Would scaling data impact K? Yeah,
it could. It could make it better. It
could uh Yeah, it should impact it. We
should get a better model. So, when we
do the different folds, we'll get
different we'll get better scores. Yeah,
it it will impact
uh yeah, if they're so that's a good
point. If they're going to use our
model, then yes, they have to scale the
data as well. If they're going to if we
build the model on the assumption that
the input is scaled, then yes, they have
to also scale their data when they're
using it with our model. That's true.
I mean, not really. I'll show you why.
There's something that's actually going
to make it easier um that that will
automate doing the scaling for them. So,
they don't they don't have to do the
scaling manually. it'll just it'll
happen automatically when they use the
model. I'm going to show you something
that's going to automate that which is
going to be called a pipeline.
So that part will be automated and they
won't have to do that. So it won't be
heavy on the user. No, in theory it is,
but
has a really helpful tool to make it
easy to do that. So I'm going to I'm
going to show us that um later on in the
notebook.
No, the data data is not for a single
house. It's for like a neighborhood. So
there's a certain number of households
in the neighborhood. And this is the
we're predicting the median house value
of that neighborhood.
Yeah. So there's a there's certain
number of households. There's there's
like an a median income, a population,
certain number of people that live
there. Um proximity generally of where
that location is. It also has a latitude
and longitude.
So,
and a median age in that neighborhood.
So, yeah, it's not just a single house.
Okay, let's go back to this was the hold
out strategy. So, this is a lot simpler.
This is just model.fit, right? This is
just model.fit on the training uh data.
And then we um can predict on the test
features and generate test predictions.
And then we can compute our error on
those um we can compute our error
amongst the test predictions and our
test uh label. So that's our useful mean
squared error function, right? To to
compute the MSE. Um let's see what the
MSE is. So MSE is right here.
Um now what we could do is we can take
the MSE
and we can take the square root of it.
So let's actually do that. Let's um do
MP. Square root of the
um test
MSE
and we get um 67 we get 67,000.
So that's pretty high on this. So when
we just now look at the difference of
that, right? When we just do a train
test split,
um
when we just do a train test split, we
get a worse score because it's not as
it's not as robust, right? We're not
showing that to many of the other uh
folds. So, we get a lot more error this
way on the test data.
So, this is um actually worse
performance just doing the train test
split.
This is a really higher.
Yeah, we can. We can. I'm going to I'm
going to show us how to how the scaling
will be done automatically. Yes, we can.
Um there's there's a really easy tool to
do that will scale it automatically.
It's going to be later in this notebook.
I'll show us it.
All right. So, just to recap this, this
is fitting the model.
This is fitting the model. This is
making the predictions, right?
Model.predict.
So, this is making the predictions. And
then this is calculating the error, the
mean squared error, which is looking at
our test labels versus our test
predictions, right? And this is
computing the distance, the average
distance away from these values to these
values,
right?
And then we can also compute the R squar
R R squar and we see that it's not a
very good R squar 65 uh is not a very
great model
um because it closer to one would be
better. So this is still this is not
very good.
We know that we knew that from the cross
file score but this is just doing um
this is just doing a hold out uh where
we do a train and test split. Right? So
it's a little bit simpler but it's not
quite as robust. Um,
it's not quite as robust as the cross
valve, but it works. Um, it's, you know,
we can do hold out. Um,
we can do hold out, uh, to to quickly
evaluate a model and see if we need to
make any adjustments.
It's a little bit quicker to run.
Okay. And any questions on it? Does it
make sense what we're doing here?
Model.fit fit to train it predict to get
our predictions. Um this is pretty
standard, right? To train is the
model.fit and then to use the model to
predict we predict on the test features.
Um so this is passing on on all of our
features into this model to generate
predictions for every row. That's
something I also want to point out that
may be a little bit confusing is this is
a data frame. So we're passing in a
bunch of rows of features with columns,
right? So um we're passing in a bunch of
data that looks like this. And what
we're doing is essentially making a
prediction for every row. So this will
generate a prediction. This row will
generate a prediction. This row will
generate a prediction and on and on and
on. So this this predict will predict
for every row. And so we end up with
this collection of predictions here for
each row. and we're comparing those to
the labels that we have for those rows
from our from our supervised learning,
right? From our data set. So that's
truly supervised learning, right? We
have the examples and we're comparing
those to what our model is predicting to
to get our performance.
All right.
So let's uh let's try the other just so
you can see it. The leave one out. Now
the leave one out cross validation is
going to actually work the same way
where we put in the leave one out um
strategy inside of the cross file score.
Now here we don't need to specify how
many folds there are because we know how
many they're going to be. It's going to
be the number of data points, right? So
which is actually going to be quite
large because there's 20,000 rows. So
this is going to be extremely
uh extremely um intensive because we are
doing um you know 20,000 examples and
leaving one example out to be our
validation and then um doing that across
every 20,000 uh examples.
So we could do it though just to see how
it works. Um we have this again leave
one out. We generate our crossfile score
from our model our data and then same
scoring that we had before and but this
time we change our cross file to be
instead of our kfold object we have our
leave one out object which is this
um and then we could run this. We can
compute our average uh across the all
the folds. Now this is going to be a lot
bigger of an array. It's going to be a
20,000 size array and we're going to
compute the average across it.
So, let's do that. It's going to take a
moment because there's lots. So, if you
notice it when you run, it's going to
take a little bit of time to run because
it's running across all 20,000 examples
and leaving one out. So, you have 20,000
and then one left out to uh test
against. So, it's quite intensive. You
can see it's taking a lot more time.
It's still running. It's taking a while.
Okay, just let that run. Still running.
So, if you guys try running this, it's
going to take a little bit of time.
Hopefully, that makes sense why it's
taking so long, right? It's because it's
instead of doing 10 folds, it's it's
putting every data point but one is the
training set and then iterating through
all 20,000 points.
This takes a while to do.
Let's see what our
RAM our memory is a little increased.
Okay,
still running. That's okay. I'll let it
run.
Come back when it's finished.
Yeah, exactly. This is a this is for
this is giving us a performance
evaluation. This is like the average
error across all of our uh different
folds. Um now this is the extreme case
where we have the number of folds equals
the number of points.
Right? So it's an extreme case but yes
it's just like kfold. It's giving us
that performance estimate.
Okay. It's about the same. Right. This
is still around 50,000.
Not much difference, right? Still right
around there. But look how much longer
it took. That took 2 minutes to run. The
other one was pretty instant, right? So
this this took about 2 minutes to run.
So um definitely uh
yeah, definitely don't want to run this
uh too often. I think that it's
generally preferred to do k-fold. If
you're going to do cross validation,
generally want to do k-fold or just the
regular hold out train test split. Uh
generally better than doing leave one
out. It's just going to take too long
and um it results in about the same kind
of score as the kfold.
Okay,
any questions about um the cross
validation that we just did.
Okay,
good. And as it says here that the
stratified kfold is usually used for
classification. Again, we're not doing
classification yet. That's in going to
be in lesson four. So, we don't need to
worry too much about that. Just for
regression, um regular k-fold is
preferred, right? Because we don't need
to um worry about distributing
categories amongst our folds uh in any
regression problems.
And as we see the error is kind of high.
Um there's going to be some things we
can do to improve that which will be uh
later on we'll learn about some more
advanced models. This signals that the
performance is bad. We probably need a
more complex model. Um one thing we
could try before we try a complex model
is to do scaling. We will try to do
scaling. I'm going to show us how we can
do that coming up um in a in a nice
streamlined fashion. Um, but uh outside
of that, if we still had bad
performance, we would likely need to use
a more advanced model. And we'll learn
about more advanced models uh in the
next lesson. And what's great is some of
those advanced models can actually be
used for regression. So they have
variations that can be used for both
classification and regression, which is
pretty cool. So I'll point those out
when we get to them. Um, okay.
So what I want to talk about now is a
way we can combat overfitting. So if we
have overfitting which remember that is
the case where the uh the we see good
performance on the training data but
then um it doesn't generalize over to
the test data. We get poor performance
on the test data. Um there's there's a
drop off there. Um that would signal
overfitting.
overfitting
and one way of um combating overfitting
is to do something called regularization
which we're going to talk about next. So
the key idea in regularization
is to
change our uh the change the way we
train. Essentially, what we're going to
do is modify our training
uh error function or sometimes called
the objective function or loss function.
We're going to change that to add a
penalty to penalize excessive complex
complexity. Essentially the the way that
we're going to penalize is by making
sure the size of the coefficients
doesn't grow too much which should
mitigate overfitting because remember in
linear regression what we are learning
are the coefficients right we're
learning the beta 0 the beta 1 the beta
2 and on and on however many betas there
are beta n we're learning all of those
guys um through the regression error
function we're trying to minimize that
error function. That's how it trains. We
talked about that on Monday.
Um so what we're going to do is um
basically penalize the these guys
growing too big and making sure we kind
of keep them small so that no one
coefficient has a dominant uh effect on
the model. And this should help with
overfitting and complexity. It should
make the model simpler because all the
coefficients are going to be encouraged
to be smaller. They're not going to grow
too big. Um, and this this has the
effect of making the model so basically
make the model simpler.
Make the model simpler is what these
regularization techniques are
essentially trying to achieve is is
remove complexity, make them a little
bit simpler, make these coefficients
smaller so that you can generalize a bit
better and and prevent overfitting. So
we want to prevent
uh overfitting,
right, is what we want to do. Um so
there's going to be a penalty and I'll
show you where that penalty gets added
and kind of what it looks like.
Um but uh to control the level of that
penalty we are actually going to
introduce another parameter to our model
um called alpha.
Alpha is going to scale the penalty. So
if alpha is really high that imposes a
stronger penalty on the coefficients um
which will make the model a lot simpler.
So the higher the alpha the simpler the
model we will get and we the the risk
with that is we actually underfit. So if
alpha is too big we may underfit the
training data
um a bit too much because it will make
the model way too simple. Um and again
I'll show you what this means
mathematically in a moment. Um but on
the other hand if we have a lower alpha
this will have a lower penalty. it's a
weaker penalty term and that'll lead to
a model that is um a bit more complex.
Um which could um risk some level of
overfitting. Um so there's so there's
still the risk of overfitting if you
have a low alpha. And of course if alpha
goes all the way to zero there's no
penalty at all. So you're back to your
original linear regression um which
could risk a lot of overfitting.
Right? So you you generally want to pick
an alpha um effectively and actually
we're going to see h what's the best way
to pick alpha. Um we're actually going
to learn how to do that. I'm going to
show us how doing some tuning techniques
to pick what alpha should be. Um but um
a a pretty industry standard alpha that
most people default to is alpha equals
to one. So just just one which signals
that there should be some penalty. we
just have alpha equal to one is a
standard penalty. We don't want it to be
too high. We don't want it to be too
low. Like we don't want it to be a
fraction. Um but a penalty of one is
usually uh good enough.
Okay, I'm going to show you where that
comes into play in a moment.
Um but the whole purpose of doing this
is to mitigate overfitting, right? Um
that's what and and doing this penalty
is is called regularization. So adding
so going beyond just regular linear
regression adding this extra penalty to
to the training process um to penalize
large weights large coefficients
um is known as regularization.
Okay. Um and there's two common
penalties that are added. Um so there's
actually two different variations on the
penalty. Um we're going to study both of
them and um they're they're known as
lasso. So if you take linear regression
and add a particular type of penalty,
it's known as lasso. If you add another
type of penalty, it's known as ridge
regression. We're going to study both of
those and what their differences are.
But these are the primary two
uh regularization tech uh models that
are used um to take a regular both of
these take regular linear regression and
just modify the training process a
little bit in different ways. Two
different ways. um using that alpha
um to penalize the terms in slightly
different mathematical ways. So we're
going to learn about these two guys.
Lasso regression there. Both of these
are just offshoots of linear regression.
So underlying model is still linear
regression. It just adds different types
of penalties to the training process.
So both of these are still in the family
of linear regression. In fact, in um in
scikitlearn, they both come from they
both are still from the linear model
family in inside of the linear model
module, which is where linear regression
comes from. So there's still linear
regression. They just have different
styles of penalties added to them. Um
which we're going to see.
Okay, so just to recap that
regularization is the process of adding
a penalty to the training to discourage
complexity. In this case, we're going to
discourage large coefficients.
And um this should help prevent
overfitting.
And so uh these are going to lead us to
two different offshoots of linear
regression that have two different
penalties.
lasso and ridge regression, which we're
going to uh study next,
but they they function the same way as
linear regression. They will just have
different penalty terms added onto their
training process um to discourage
uh discourage um again those large
weights.
Okay, any questions about regularization
before we first look at our we're going
to look at our first uh variation on on
our first regularization technique which
is going to be called lasso regression.
Okay, let's look at lasso regression. So
what is lasso regression? It's actually
lasso is short for least absolute
shrinkage and selection operator
regression. Um and this will function by
adding a particular penalty to the
linear regression model. So again, it's
based on linear regression. That's the
underlying model. It's just that during
the training process, we are going to um
add a penalty which has the effect of
shrinkage of the weights. That's why
it's called shrinkage. It encourages
smaller weights through that penalty.
And it also will shrink some of them so
much that they'll become zero. And so it
has has an effect of kind of selection
which means that some of them get wiped
out to zero.
And this means that whatever is left
over is kind of what's selected as our
features because the other ones will
have zero weight applied to them. So
this penalty will really favor small
weights um and penalize really large
weights. In fact, it will favor small
weight so much that some of them will
actually um be shrunk to zero um during
the training process. And the ones that
are left over are the ones that um are
the ones that are what we call selected
because they are the ones that remain in
in the training um after the other ones
get uh coefficients of zero. Um now when
you make some of the coefficient zero
you are inherently making the model
simpler right there's less features
involved in the prediction that or less
features that have an effect on the
prediction. So this definitely makes the
model simpler. This lasso this shrinkage
and selection uh process makes makes the
model simpler for sure. Um
and this is supposed to reduce
overfitting. Right? If you make the
model simpler, it's not as complex. It
has less of a chance of memorizing
training data and not generalizing over
to test data. So our whole goal with uh
regularization is to make our model
better at generalization, right? Over to
test data from the original training
data.
Um so how does this happen? We have to
go back to the
uh training process. If you guys
remember, I I wrote out this equation a
little bit earlier, which is the
distance. This is the sum of squared
distance between our labels and our
prediction.
This is basically the mean squared error
uh calculation that we're trying to
reduce when we build our model using the
training data. Um so this is just in
standard linear regression. This is the
um uh sum of squares uh distance, right?
So this is this is what the model is
trying to minimize when it learns these
coefficients.
So when it learns these coefficients,
it's trying to minimize this guy
minimize. It's trying to find the betas
that minimize this quantity
mathematically. That's what it's doing.
Um and there's there's a algorithm that
will discover what the best betas are
that actually minimize uses that gives
us a line of best fit, right? That's
what we've been talking about for
regression.
Now, in regularization,
here's, by the way, here is that same
thing, but we've just inserted our model
for the predictions. This is our model.
Just a fancy way of writing down our
model, right? It's the beta 0 plus all
of these betas. So, beta 1 x1 plus beta
2 x2
plus on and on and on, right? That's
that's what this uh means. If you're
unfamiliar with the sigma notation, it
just means sum. So it's the sum of all
these guys or this term. Um, so this is
this here is just a regular linear
regression
uh training regular linear regression
training. So we the training process
solves for these parameters, right? It
solves for these weights. We discover
what those are by minimizing this
quantity. That's the whole training
process. Um, but when we do lasso,
we add a penalty which is this.
Here is our penalty.
So basically um we take our linear
regression training which is this and we
add on a penalty which is this. And you
can see exactly what this penalty when
when you minimize this penalty. It's
when these weights are small. So this
encourages
So minimizing this quantity encourages
small weights
encourages small betas
beta I
right you or in this case beta j sorry
this encourages small beta js uh because
we want this thing to be minimized
minimized
so Um, what's going to make this minimal
is of course the line of best fit and
small weights, right? Are going to make
are going to bring this error down the
most.
So, um, and here's our alpha, right?
Here's our alpha. So, you can encourage
a higher penalty with a larger alpha or
a lower penalty. If alpha equals zero,
what happens to that term? It just goes
away. So if alpha equals zero, there's
no penalty and we're back to uh we're
back to regular
linear regression.
We just have regular linear regression
because we have no penalty at that point
when alpha equals zero. So the smaller
alpha is, the less penalty we're
enforcing and in the regularization.
Okay.
Now what happens is in reality when you
train with lasso. So this is lasso is
this particular penalty. This is called
the lasso penalty
or sometimes um people call this the L1
penalty.
Um L1 just comes from the fact that this
is the first power or absolute value. Um
so it's not a squared penalty, it's a
single uh single power penalty
um there. But when you add this lasso
penalty, what can happen is it it does
because the because you're minimizing
this, it does encourage some of these
weights to become zero.
So some if you're really trying to get
the lowest quantity of this,
the lower the better.
What makes this thing lower is of course
if some of these go away if some of
these go to zero then that of course
will lower this as much as we as much as
possible right so what happens during
the training is some of these
coefficients actually they're encouraged
to be small because of this penalty but
some of them will actually become will
actually become zero um in order to get
the best model the best fit some of
these will actually get so small that
they'll basically become zero
And that means that that that feature
basically has no effect anymore. It's
it's been the model has been simplified,
right? That feature no longer really has
an effect.
So just to call out the alpha again, um
if alpha zero some code, uh basically
you have your linear regression, you're
back to linear regression because alpha
0 is just wiping this out and you're
back to linear regression.
um if alpha is infinity. Now if alpha is
infinity that's an extreme. So if alpha
is infinity the only way to make this
minimize is if all your coefficients are
zero. If every beta is zero then this
will lower the the error as as much as
possible. So you basically have no
model. So if all coefficients are zero
you have no model and that's useless. So
you don't want your penalty you don't
want your alpha to be huge is what this
is saying. You also don't want your
alpha to be small. you're basically back
to linear regression. So you want
something in between. Um and the typical
typical value is alpha equals 1.
Typical is alpha equals 1
to have some level of penalty there. So
just a regular kind of regular penalty
term.
But we are actually going to have a way
to test and evaluate which alphas are
the best.
Um,
basically you can yeah you can have a
you can have a penalty that's close to
zero. You can get rid of this if just a
regular linear regression performs
pretty well. You can basically have no
penalty in that case.
Yeah. So nearer zero or like it could be
that adding a little bit of penalty
actually helps the overfitting and it
could be really small. One thing that
we're basically going to do is have a
strategy to try out different alphas.
try different alphas
and evaluate performance
and then we can decide which so that's
what we're going to do is have a
strategy to just plug in different
alphas generate the like train the model
and then see what its performance is and
see if those alphas are good what what
which alpha is the best we can evaluate
that
because we can train the model and see
what it performance is
right.
Yeah. Yeah. So, we'll do that. We'll
practice that.
Okay. Great. Any other questions about
this lasso regression? So, remember this
is linear regression here. This is the
this is how you're training to find the
betas in linear regression. So this is
just linear regression uh um training
function there.
We're adding a penalty which is this is
the lasso penalty
lasso penalty there right we're adding
that this is known as regularization
and the goal of regularization is to
prevent overfitting. So you add a
penalty here this makes the model
simpler which prevents overfitting.
helps you generalize better when it's
simpler.
Any questions conceptually on this?
We're going to do a code example with it
coming up, but any questions on this?
Uh yeah, you you so that's the thing,
Ronald, is you may be willing to
sacrifice some accuracy in order to
generalize to unseen data because
remember that's what we're really trying
to get after is we may be willing to
sacrifice some accuracy on this training
data in order to have it perform better
on the test data, right? we may be
willing to do that. That's a willing
that's an okay sacrifice
as long like if if it generalizes
better. That's what we want. That's what
we're trying to do here is add a
penalty, make the model simpler, and
help it generalize better to new and
unseen data. Right?
That's that picture I've been using with
the with the um train and test split.
Where is the square?
So in the model there's no square. So
remember the model is the model is this
um equation uh that has no squares in
it, right? It's beta 0 plus beta 1 x1
plus beta 2 x2 plus beta n xn.
That's the that's the linear regression
model. This is the now this this is the
model but this is the equation that
helps us train and find the betas. This
is how this is what we find the betas
with. So we'll continue. Um we were
talking about the lasso regression which
uh adds it takes linear regression right
which is this optimization and adds in a
penalty um scaled by the alpha. Um, and
what that does in order to minimize this
whole thing, it encourages these to be
small uh as possible. Um, which makes
the model simpler, right? The weights
don't get overly big and complex. Um,
they they tend to stay small. In fact,
some of them can even go all the way to
zero. Um, which makes the model even
more simpler,
right? Um, so let's practice uh using it
in code. It's actually really easy to
use. It's going to be essentially the
same uh style and and code as linear
regression except we are um just going
to have to uh put in our alpha parameter
um when we use the lasso. So here we are
um from the linear model family right
which makes sense. It's a linear
regression offshoot that has this
penalty in it during the training. um we
are grabbing our lasso regression. Um it
also has a version of the lasso that
we're going to take a look at that is
used for cross validation which is
really um convenient as well. So it has
a cross validation lasso which is a
really convenient um combination of
basically cross val score and lasso um
all in one. So it actually is really
nice to use that way. Um so we'll take a
look at that example. Um, but we are
importing it. The main thing is going to
be the lasso model here. Um, we're going
to be using a different data set for
this one. So, not the ocean uh data, but
this hitters data, which is a baseball
data set. Um, so it has 322 rows um with
20 different columns and it looks like
this. So, you want to download that one.
Um, hopefully you guys have access to
that one.
Um,
so I will upload it into
this.
So give me a moment.
There's that. And then we can run this.
Okay. So we are displaying the data and
so it has um the the hitters names and
then it has a bunch of different
statistics. These are all baseball
statistics.
Um, if you're unfamiliar with with them,
that's okay. It's not a big deal. Um,
but just different baseball stats here.
Okay. Were you guys able to load that?
Um, if you're following along, were you
able to load that? You should have
access to this data. The hitters CSV.
This is the one we're going to use for
the lasso model
to build a lasso model.
Yeah.
Okay. Able to load that one. Perfect.
Okay. So, able to load that one. Um, and
we take a look at the the head. Um, so
we're actually going to uh drop this
unnamed column because we don't care
about their name. it's actually just the
batter's name which is not going to be
useful in modeling. Um so and remember
that's generally true like an ID, a user
ID, like a customer ID, a name, that's
usually not going to be useful in any
kind of modeling. So we're actually just
going to drop that uh column and we're
going to do it in place.
And access equals 1 means we're dropping
that column. Um, so we're going to drop
that and we should no longer have that
column and we have all of these guys
now. So you want to run that. This will
drop that. Um, this will drop drops the
column in place.
Um, and now we can see we have uh all we
have this data where um we have this
data where it's now removed. So, this
that column is now gone and now we have
these guys. Um, do you notice anything
about this
from the info?
Looks like we have a couple categorical
features, a few of them, league and
division
and new league. What do you notice about
this
nullles? Yep. So, there's definitely
some missing data there um that we're
going to have to deal with.
So, it looks like there are uh there are
59.
Um there are 59. Now we could we the
alternative to doing that is we could uh
we could just use our usual code where
we do dfis
uh isnull.
Um and then we do uh dot sum to total
those up across our different columns.
And we can see that uh we have 59 of
those in the salary column. That's this
is the standard way of doing that,
right?
standard way of doing that. And we have
so we have 59 of those.
59 of those. So we have to deal with it.
Any ideas on how to deal with it?
Any ideas on how to deal with it? This
is now this is 59 out of 300.
So,
what do you guys think about that? It's
a little bit different than 200 out of
20,000. A little bit different. We have
We have about 60 out of 300.
There's a decent amount.
Any ideas on how to handle this one?
Replace. Yep, we should replace. What do
you think we should replace with?
It's a float. It's a floating point uh
value.
By the way, something unique about this
that's a little different than usual,
too, is that the uh this is actually the
column we're going to use as our label.
So, we're actually going to predict the
salary based on the uh based on the um
rest of the features. So, we definitely
need to fill in these nles, right?
Because they're actually going to be the
labels.
We're missing some labels uh in our
data.
We definitely need to fill them in.
Yeah. So, we're going to replace them.
All right. So, we'll we will replace
them down below. That's going to be
coming up. Uh we'll come back and
replace them. um before we replace them,
we're actually going to get our uh one
hot encodings for those three different
um features we have. Um so we do uh get
dummies with this. Now um of course we
don't need to do this if we just so this
code we don't need to do if we just pass
in the dype here
um which is uh then we don't need to do
this. So we can comment this out.
Um so now what I want you guys to notice
is this is the alternative to what we
did before where we are purposely just
doing these columns not the whole data
frame but just doing these columns and
then we can um concatenate those these
one hot encodings. We're going to
concatenate back to the data frame.
Right? So if we do our dummies and then
do dummies.info info. Um, we can see
that we end up with six new columns. And
in fact, we can do dummies.head
and take a look at what those are.
Right? So, these are league A, league
uh, N, division E, W, division W, new
league A, new league N.
Okay.
So, um these are uh these are our one
hot encodings for these three different
features which are strings, right? So,
those those features were strings. If
you go back up, those were our only
string features we had. So, we've one
hot encoded those so we can use them in
our model. What we need to do is just
concatenate this back to our data frame.
Right? So, we just need to concatenate
it back into our data.
Okay. So, what we're going to do then is
we're going to grab um we're going to
grab Y as our salary. And of course,
we're going to fill nles on that Y
coming up shortly. But we're going to
grab Y as our salary and X new. Now
before building a full X, we're going to
take a look at X numerical as our data
frame minus these columns. The reason
we're doing minus those is because we
are going to concatenate our dummy
variables back into this that are going
to replace these guys. So we're going to
replace these anyways with our one hot
encodings. We don't want the strings. So
we're going to get rid of those. And
we're also going to get rid of the
salary because that's going to be part
of our that's just a label. So we don't
want that in the X, the eventual X.
Are you guys able to run this one?
Hope I'm not going too fast. You guys
able to run this? And does it make
sense? What we're doing is we're putting
our labels in Y, which is what we
usually do. So, we're going to predict
the salary
and we're getting ready to build the X.
But before we first want to get rid of
those one hot the strings. This is
getting rid of the strings
and this is getting rid of the label.
And that's going to be part of our
features. What we need to do is build
our final X by concatenating our dummies
with this. Do you guys see that? We're
going to concatenate our dummies with
this to build our final X.
But but prior to doing that, we need to
get rid of these string columns here. So
we're dropping those
dropping those from the uh data frame uh
and getting a numerical uh x numerical
here.
You can see the columns of that are just
these guys here. So the the results we
need to concatenate our we need to
concatenate this guy um into this and
then that'll be our full x all of our
features.
Okay. So you can see x is going to be
pd.con
of this with our dummies.
This with our dummies. And um
uh instead of doing this, I'm actually
going to do the full dummies. We don't
need to
um pick just a few columns. We're
actually going to do our full dummies
here and um do x equals 1. Now, the
reason that's the case is because um
this will get rid of one column per
feature and basically assume that if you
have a if you have a zero, the other one
should be a one. If you have a one, the
other one should be a zero. Um so it
basically makes that assumption because
we only have two of them. Um so whenever
there's a one, the other should be zero.
Um, so you can get away with just having
these three, but um I think it makes
more sense to just have to have the full
dummies,
but by process of elimination, you can
get away with just using two of them
because anytime you have a zero, the
other one should be the other feature
would have been would have been a one,
right? And vice versa, when there's a
one, the other feature would have been a
zero.
So we do that one.
And you can see all of our uh all of our
one hot encoding features end up back in
there.
So this is the code that I want you guys
to run. I think it makes more sense. It
follows along what we've been doing.
um which will concatenate our dummies
back to our features here to build out
our full X. So now X is all of our
features. Um remember X
X contains all of our features
now.
So X contains all of our features and so
we have all of this now.
Okay. Were you guys able to run this
one?
Damn. We have y, we have x. We still
need to deal with the nles in y. So that
something we still need to deal with.
But hopefully you have this. Now
all these are numerical.
So that should be good with the model.
That's one thing about X is you should
you our X should have all numerical
features, right? Because it's going to
go into a model to to learn those betas.
So it needs to have all numerical
features,
right? These are going to be all
numerical, which makes sense. We change
we did one hot encoding to change all
those guys to numerical.
Sorry, I'm scrolling down.
Okay, we do fill in the nil later. Okay.
Okay.
Any questions so far? So, we're just
getting our data ready. We haven't
applied the lasso yet, but we're just
doing some prep. Now, hopefully you guys
recognize th these are some standard
steps that we're taking when we do our
modeling. We have to do these data prep
steps. They're necessary. And so, if it
seems like it's a lot of work, that's
because it is. It is work that you do to
prepare your data to get ready for
modeling. You have to do that. Okay.
So, we're doing that here. Um, now we're
going to do our train test split because
we're just going to do uh we're going to
do hold out here. So, we're doing a
train test split with about with a test
size of about 0.25. So, again, anywhere
between 0.2 to.3 would be okay.
Um, so uh
it's our choice. We could do 0 2. We
could do 3. We could do anywhere in
between there. We're doing 0.25. That's
fine. Um, that's okay. So, we we build
our train test split right there. Um, so
pretty pretty simple and we've seen that
a bunch of times with our X and our Y
data frames. There we have our train
test split.
Okay.
Um, now what we're going to do is do our
our scaling. So, we're we didn't do this
last time, but we're going to do this
now as uh because we should get in the
habit of doing that. Um is um we're
going to um go ahead and scale our
features and we're going to use the
standard scaler here uh to do that
scaling. Okay. Now, we could use minmax
scaler. That's fine, too. We're just
going to use the standard scaler here.
Um and remember we are going to uh um
use the standard scaler from sklearn and
we're going to transform our features uh
uh according to our um according to our
training data. So we have our
pre-processing standard scaler here. So
we import that guy and then we um build
our standard scaler and fit it on the
training data only on the numerical
features. Um so that's which is going to
be uh all of these guys. So we're doing
the scaling on all of these guys. Now
something to note is that we are not
scaling all of these one hot encodings
mainly because it doesn't make sense to
scale those really. They're zero or one.
They don't need to be scaled, right?
They're already zero and one. So they're
they don't need to be even if we were
doing minmax scaling, it's going to put
them between zero and one. It wouldn't
affect it really, right? So these one
hot encoding features, we're not going
to scale because they're they're always
going to be zero or one.
There's no need to scale them really.
Um, but we're going to scale all the
other features here that are floats.
So that's these guys here. These
numerical features we're going to scale.
Okay.
Don't need to we don't really need to
scale the one hot encoding. Uh, it's
pretty much already scaled.
Oh, you should change that. Um, go back
and rerun go back and rerun this, but
make sure you have your data type as int
here.
Make sure you add that in there to
change that over to integer and rerun
that and then rerun the rerun the
concatenation.
So, make sure you run this
and then u make sure you rerun this and
rerun the concatenation part which is uh
this
Okay. So, we go ahead and fit the um
scaler to this data and then we're going
to transform our training features,
those numerical features. um we're and
then we're going to uh transform these
features uh uh the test features in the
same way. So we're going to perform the
same transformation from the scaler on
the test data. So that's something
really important I want to note here is
that we always scale both the training
and test data. We always scale both. Of
course, we're going to train the model
on the training data. Um, but we are
going to also test it on the testing
data and it also needs to be scaled
because our model that we build is going
to assume scaled features. The
coefficients that it learns are going to
be assuming scaled features.
So, we need to also scale our test data
in the same way. So, we're doing that as
well.
So, we scale that and now we have our uh
training and testing features have been
scaled.
No, we haven't replaced. We're going to
do that. We have not yet. We're going to
do that coming up in a minute. Yeah, we
haven't done that. Um, it is it is the
label. We definitely need to replace
NLES. We just haven't done it yet
because it's not in the features and
we're doing all of our uh uh
pre-processing to our pre-processing to
our features.
Yeah. So, we're definitely we need to
we're going to in a minute.
Okay. So, if you look at the data now,
it's all been scaled. So, these are all
um zcores. These are all on a much
better scale now. Um, and these are we
still have our one hot encoding features
which are zero or one. So this scaling
should lead to a better model than if we
didn't scale. So scaling is really
important. We can see that here.
Okay.
Now, um, let me ask you guys, were you
able to run the scaling? Are you caught
up to here? If you're following along,
were you able to run the scaling?
Okay, great. Great.
Awesome.
Okay. So, uh what we're going to do now
is replace NLES in the uh replace NLES
by calculating the median of the data.
So, what I want you to notice is that we
are taking the NLES now this is um this
is on purpose is we are purposely taking
the NLES um out of the median
calculation. So we're skipping the NLES
when we compute the median because we
don't want those NLES to affect the
median calculation.
Um so we compute a median salary here
and then we fill our NLES with the
median salary um from the training data.
So this is our choice. This is a choice
um to use the median and it's also a
choice to use the training set median
for both train and test. What we could
have done this is an alternative that we
could have done is use the entire column
and then um use the median of all of the
data to replace. That's really up to us.
Um this is one way of doing it. We could
have done before we did the split. We
could have um filled in with the median
earlier. We chose to do it here mainly
because it doesn't affect the features.
So we could have done this earlier and
did it before we did the split and
filled the NAS. Um really doesn't it's
doesn't matter that much which way we do
it. Um but we do need to fill in NLES.
We cannot have those be null when we
when we put it into our model. So some
way we need to fill in nulls. Um and so
in this strategy we're filling in our y
train um with the median salary from our
training data. And same with this we're
filling in with the median salary of the
training data as well. But that's a
choice. We could fill in with the mean
with the average. Um we could fill in
with the we could do it with all the
data together before we split it. we
could have filled in with all of the the
median across the whole data set. Um
either one works. You can do it either
way, but we we did it um later here to
show that it doesn't really affect the
features. So, we can choose when we do
it, right? It doesn't affect the
features at all. So, we can do all of
our pre-processing on the features and
then do our label uh filling in NLES um
if if we have them.
uh x numerical. Um make sure you're
running uh this
uh x numerical was defined here
when we split it apart um from
uh when we dropped these columns here.
So make sure you're running this. This
is x numerical.
It's defined there.
So go back up this uh this cell
where we split apart the y and we and we
have the x here x numerical.
Make sure you run this.
Make sure you run this. And then you can
run these. Then you run this to build x.
All right.
Are we up to here with this filling in
the labels?
Uh because then we can build our model
once we're up to here. We've scaled
everything. We filled in our NLES.
We've gotten one hot encoding.
Yeah, it is. That's why you know that's
why we spend a lot of uh time on model
on data preparation with pandas, right?
That's why we did all that pandas work
for sure. Yes, there is a lot of work
before we can build a model.
Yes, the mo do you guys notice that like
the modeling is relatively easy. It's
just a fit and predict. The modeling is
actually really easy. It's all the other
work that's that's more involved, right?
more code.
The modeling itself is really easy.
It's just it's just one line of like
ffit.
Yeah, pretty easy to do.
And then you do evaluation which is a
couple lines.
Yep. There's these are all the these are
the common steps. All these steps we're
doing are very very prototypical in
model building is you let's just go back
through this to see what we did, right?
We imported our data. Um we analy we
dropped this name column because it's
not useful to us. So we dropped that. Um
we filled in the NLES eventually. Um but
you know if there were any nulls in our
features we would have to deal with
those as well by replacing them or
dropping the rows like we did earlier.
Um
and then we do one hot encoding because
of course we can't have any string
columns going into our models. We got a
one hot encode.
Um we uh then build our X and Y by
concatenating the one hot encoded back
to the numerical features.
Then we train test split. Right? That's
pretty common. Or we could do cross
validation either way. Um a kfold cross
validation. Then we scale. So we didn't
do this last time, but this is something
we should get in the habit of is scaling
um our features. So we do that. And now
we're ready to model. So now we're ready
to model. Um so that's this part.
Okay. So let's do the model. Um the
model is actually uh pretty easy to do.
So we're going to use a lasso. So we
have a lasso model here. Notice what
we're setting our alpha to. So the big
parameter we really need, ignore this
iterations. We actually don't really
need the we don't really need that
parameter. Um so just ignore it for the
moment. But the big one that we're
setting here is the alpha. So when we
did linear regression, we didn't need
any parameters to go inside the linear
regression object. We didn't need any
parameters, right? Because there are
really no parameters of it. But for
lasso, the important one is the alpha.
And so we need to know what to set alpha
to. Um let's start with alpha equals 1.
That's a good starting place. So a
typical um starting point
for alpha
um
is uh is one. So that's a typical
starting point. And so we can set alpha
equals to one. This max iterations is
the the parameter that governs the
training process because it is
iterative. So if for some reason we we
can't converge to the right betas and
we've run it for 10,000 steps once we
pass 10,000 steps, uh it will stop and
just give us the betas at that point.
But it will likely never hit this
number. It'll converge before then. So
um we don't really need to um specify
it. So, I'm actually just going to get
rid of it. Um, it's not really a big
deal. It should converge before then.
Um, but if if we want to set like a
maximum step size in the optimization,
we definitely could there. Uh, but not
concerned about that too much. But
here's our lasso. And then we're just
going to do a fit on our data. So, look
how easy that is. Just like a linear
regression, lasso.fit,
right? So, we do fit. Um,
oh, I didn't run this. I'm sorry. I got
to run this. Okay. Actually, that's a
good example of what happens when you
don't when you have nles, right? So, it
says our our null contains nan. That's
because I didn't run this. But now that
should be filled in. Now, we should be
able to run this. Okay, perfect. So it
runs.
Okay. So you can see what the intercept
is. Um this is one of our coefficients,
right? The intercept is 457. And look
now what's really interesting about the
coefficients is look at what some of the
coefficients are. Some of them are
actually zero, which is really So some
of them ended up being at zero, which is
very very interesting. that means that
those features get cancelceled out and
they're basically not part of the model
which is really interesting. Um, so we
have all these coefficients and some of
them are zero.
Yeah, negative0 is just because of the
convergence like they started out
negative and worked their way up to
zero. it. Negative Z really just means
zero, but they just were coming from
they were like small negatives and ended
up at zero
during the training process. They were
negative at one point and it ended up
zero. Um
so yeah, negative 0 just obviously means
zero. Um it's still still zero there.
So what's interesting is some of these
features ended up uh being zero which
you don't usually see in a linear
regression. So if we were to train this
using a linear regression we typically
wouldn't see that but some of these
turned out to be zero because again
we're encouraging those betas to be
small. we're encouraging them to be uh
small and so um you know what happens is
some of them can be shrunk all the way
down to zero meaning those features
don't contribute that's a really simple
model at that point right so we've taken
something complex that includes all of
these features and actually reduced it
into something simple that only includes
these features
right
so that's what it does um now we need to
evaluate this to see how good of a model
it is. But that's what this is saying
here in this text is that um a positive
uh coefficient indicates that as the
independent variable increases the
dependent variable also increases.
Negative coefficient means as the
independent variable increases dependent
decreases because it's reducing the
value. Um and lasso is known for feature
selection by shrinking some of them to
zero effectively removing those
variables from the model from the
equation right
um
so that's what happens
some of them end up being zero
were you guys able to run this this
lasso uh fit which is the training of
the lasso Control.
No, it doesn't ensure there's no
overfit, but it helps with overfitting.
It's supposed to help by making the
model simpler. And this is definitely a
simpler model because it's removing some
of the features from the model
essentially, right? Because some of the
features aren't going to contribute.
It's a simpler model.
It doesn't it doesn't mean there's not
going to be any overfitting, but it
helps prevent it. That's what it's
designed to do to help prevent it.
Yeah. So higher coefficient. Yes. The
higher coefficient means it's a more
important feature towards the
prediction. Yes. That's what it means
for sure. The higher the magnitude, the
more of a contributor towards that
prediction. Uh it is. Yes.
And it's not just it's it could be
higher positive or negative there. Like
a higher negative is also a pretty big
factor,
right? So So you want to think about it
in terms of absolute value.
does not guarantee but helps. Yes, it
doesn't guarantee it but it's designed
to help overfitting, help prevent it.
Yes, absolutely.
Okay.
So let's do some evaluation. Um so let's
do in this case we are going to do our
predict
Oh, yeah. I'm not sure why that's the
case.
Interesting.
We could try increasing the um max
iterations.
Okay, that's why. Yeah. So then you get
that result with the with the higher max
iterations.
It doesn't get cut off there.
I think that's why you probably left
this in there,
which is fine. You get about the same
numbers.
Yeah.
All right. Let's evaluate this. So,
we're going to to to do evaluation. I
want you guys to see again. We should
get in the habit of doing evaluation,
which is taking our model and predicting
on the training and predicting on the
test sets, right? So we predict on the
train set and calculate our MSE
and we um calculate our R2 score um or R
squar score I should say. Uh but again
the MSE is the one we're really going to
use mostly. Um but we calculate so we do
our predictions and then we compare that
into our mean squared error with our
labels
and we uh go ahead and do the same thing
with the test. Right? So we do uh
lasso.predict
on our test features and we go ahead and
compare that with the test labels. And
so what we're doing there is generating
our MSE.
So, we we take a look at our MSE and we
get uh 84,000
MSE. Um, and so, of course, we could
take the um what we could do with that
is take a look at the um MSE on the uh
we could do um MP. Square root
and do the square root of the MSE test.
and we get um 340. So this would be in
the units of our label. So, we go back
and look at our label um for some of
those um
so uh we are in 300s and our data is
like right around the 500. So, of
course, if we describe this um we could
see what the statistics are of it. So,
we could do df.describe describe and
generate that. But that doesn't look
like a very good error, right? If these
are in the 400s, um that's that's not a
very good error.
So again, it's not a very great model.
But one thing I want you to see is that
it's it's not overfitting.
Um if anything, it's actually
underfitting, which is what this kind of
um MSE suggests, right? because our
error here is 84 uh excuse me 84,000.
Um
our our area here is 84,000
excuse me and on the test set it's
116,000.
Um so these two errors are both bad. So
it's not overfitting. This is actually
underfitting. So it's not overfitting,
it's actually underfitting. Um, and so
that's the risk with something like
lasso is that it's making the model a
bit too simple and we actually risk
underfitting, which is what happens. We
have too much error across both the
training and the test set. Overfitting
is when we do we have really good
performance on the training set, but bad
performance on the test set. We're not
overfitting.
um we are uh underfitting because our
performance is not good either way. Even
this R squar is pretty low. It's not
even at 50%.
Okay, so that's so we we do the
evaluation and again the evaluation just
comes down to making predictions and
computing our error amongst those
predictions to our labels. That's always
what the uh evaluation is going to be
for MSE.
What's the ideal MSE? What do you think
it should be? What is So, think about it
like this. The MSE represents the
average distance between our predictions
and the labels.
So, if we're getting it right all the
time, what's that distance going to be
if we're always right? What's our
distance from what's our distance from
our predictions to our labels going to
be if we're always getting it right?
Zero. Yeah, there's not going to be any
distance. It's going to be right. It's
going to be perfectly aligned, right?
There's going to be no distance there.
So, yeah, an ideal MSE is zero. That's
an ideal MSE.
So, anything close to like the smaller
the better for MSE. The smaller the
better. Um, for this R squared, uh, it's
it's a scale between 0 to one where one
is the best. So, one would be perfectly
aligned predictions. Um, so, and again,
this this is we actually multiply by 100
to get uh because it's it's a number
between 0 and one. So we get about 47%
which is not good.
Okay.
All right. Any questions on this
evaluation?
All right. I want to show you something
which is
Yeah, this that's true. The scale of it
matters on the data because we should be
you should always interpret your MSE in
the scale of
um your your labels because your labels
like in this case our labels um you know
we could take uh for example we could
easily let's actually do that let's take
the average
let's take the average of our labels on
the training data
and and we could see what those are. Um,
so the average is 500,
right? The average is 500. And look at
what our uh square root of our MSE is,
which is in the same units as our
original. Um, so we have uh quite a bit
of error. 340 when our units are right
around 500.
So that's quite a bit of error.
Yeah, MSE of zero means our our uh our
predictions are nearly identical to the
test labels. Yes, that's what MSE of
zero means. There's zero distance.
So closer to zero, the better.
But we talked about it as you you really
so the rule of thumb should be what is
your RMSSE as a percentage of your
typical value. So your typical value is
in the 500s. Our our RMSSE is 340.
That's just really high. That's over
like 60% of that value.
So that's just a lot. That's too much
error. What we would love this RMSSE to
be is under 20% of the typical value. So
that means on average we are 20% or less
off in our prediction. That would be
good. That would be pretty good. That
means we're like 80% accurate,
right? That'd be pretty ideal. So you
got to think about it in terms of this
RMSSE, which is in the same units as
your labels.
This is the
RMSSE
which is in the same units as the
labels.
So and then to interpret this we have
340
is compared to
typical
um salary unit of 500
right so this is uh quite a bit when the
typical value is 500 and we are off on
average by 340 units
that's so much relative to the typical
value
that's just too. That's a lot of error.
That's not a very good model, right?
It's underfitting. It's definitely
underfitting.
Yeah. So, that's a great question. What
should we do from here? So, um because
we're underfitting
um we should use a more complex model.
So uh we're going to learn about those
in lesson four, but we should use
something different. This linear
regression is still too basic. Even with
lasso, it's still too basic.
Yeah, we're underfitting because we But
it could also be we're underfitting with
a regular linear regression. We should
test that out. Um, and maybe it would be
an exercise for you guys um to test that
out yourself. It shouldn't be hard to
do. Um, you already have all the data
scaled. You So, do you see how you would
do that? You would just come in here and
build a linear regression rather than a
lasso and dofit and then you would
evaluate it the same way with a
dotpredict. It's really easy to do that.
And then we can compare that um to to
this. It shouldn't be that hard to do
that, right?
And something you guys could do for
sure. Um,
is build the linear regression and
actually compare it and see what kind of
difference it makes. I mean, we honestly
we could do it ourselves. We could do it
right now. Maybe it's worth trying that.
So, let's build a linear regression
for comparison.
So we have our linear regression
uh is linear regression and then we do
ffit linear regression.fit fit
right so so this will train it um and
then we can evaluate it so lin MSE is
mean squared error
and then we can do our um let's do our
training let's do the training and then
um let's predict
actually let me do that here
uh y prediction
train
linear
equals um linear regression.predict
and then we're going to predict on our
training features.
Okay, do you guys see what I'm doing?
I'm building a linear regression for
comparison.
I'm doing dofit here to train it and
then I'm making some predictions on the
training set and we're going to evaluate
those. I'm going to replace that here
with y prred
uh train
linear. So these predictions
Okay. So, if you guys want this code, I
can paste it in.
So, let's see what the RMSSE for just a
linear model is.
It's a little bit better. It's better
for sure.
So 289 is better than this 340. It's
better. It's getting closer to zero.
It's still underfitting though,
right? And that's just on the training
set. Let's look at the Let's do the same
thing, but on
Let's change this. Let's swap this out
for um test.
And then let's do test.
And then let's do test
test.
and then
test.
Okay, so this is producing test
predictions on the test set.
We are generating an MSE test
and then we're doing MSE test
which is using the test labels and our
test predictions
and then we take the square root of that
for RMSSE and then we're going to
generate that. So it's still under fit.
I mean this is still high. This is still
high um on the test set and versus on
the training set. So it's still pretty
high. Um, even the basic linear
regression is under is still
underfitting. Still underfitting, right?
Even without the lasso, which is lasso
is supposed to help with overfitting.
It's definitely not overfitting. Um,
it's definitely underfitting,
but this is a signal that it's kind of
overfitting because this is performing
better on the training data and then it
gets worse on the test data.
Definitely gets worse, right?
Did you guys follow?
I'm just running this above I'm running
this above this. It doesn't matter where
you put it. We could uh we could move it
down.
We could move it down to I just ran I
just picked a new cell right here.
and ran it. But we could move it.
Actually, let's do that. Let's move it
down
to
after the lasso evaluation.
Okay. So, I just moved it there.
And then let's move
this down.
So, I just put it here after the um
after this. So this is the um this is
basically the objective function right
of the training process. So during the
algorithm that runs when we call ffit in
scikitlearn it's going to find these
betas right it's actually going to learn
what these best betas are for our model.
Um this is our model here right it's the
combination of betas times our features
um plus an intercept beta. Uh so that's
our model but um we penalize those large
uh weights in absolute value by um
adding a penalty term like this um where
alpha is some level of penalty that we
want to provide. Usually alpha equals 1
is okay. But um actually what we're
going to learn uh to finish out this
section is there's going to be a
systematic way we can test out different
alphas um that represent the level of
penalty we want to uh apply to lasso or
even ridge
uh regression. So that was the lasso and
um if you guys remember using it was
super easy. Uh we worked through this
problem with this um baseball data um
and we had uh
let's see scrolling down we um split out
our numerical data and we did uh we one
hot encoded our our categorical data
combined it back together. Hopefully
that um rings a bell there. Um and we
actually scaled our data which is pretty
standard to do is we do some type of
scaling to our features especially our
numerical features right want to scale
those in some way whether it's minmax
scale or standard scaler um want to do
that and so we did that for this example
and then we um ran the lasso regression
which is pretty easy to use. You just
use the lasso object and you pick an
alpha here. Um, again, we are going to
have a way to test out different alphas
that could be candidates and we can see
which one's the best. Um, so I'm going
to show us that today coming up shortly.
But that was that was the lasso. If you
guys remember, we did that. Um, this it
we compared that to a basic linear
regression which is just this pretty
straightforward just a fit and then
predict and then we can generate mean
squared error. Um, still not a very good
mean squared error on this data, it's
still fairly large. Um, so it's still
not, no matter which model we use, it's
still not very good, but at least we can
practice doing that comparison. That's
what we did last time. We did this on
Wednesday.
Um
and then
we saw that the effect of different
alphas we had a lasso um
we had a lasso uh cross validation
example here. So beyond just using a
regular lasso model that um scikitlearn
has a lasso cv which allows you to try
out different alphas uh with cross
validation and um figure out what the
best alpha is. Um, now we're actually
going to have a different strategy
that'll instead of just picking random
ones, we can actually um supply multiple
parameters that we may want to test um
as many as the models may support. And
in some more complex models will have
more than one parameter like lasso only
has the alpha. Um, technically it also
has its max iterations, but really the
only one that matters is this alpha.
Other models have many more
hyperparameters that we can um uh change
and so we want a way to systematically
test out those different combinations
and to see which one leads to the best
uh version of that model. Let's say the
best results. So um we're going to
explore that coming up. So we had lasso.
Um now this is where we ended last time.
We had ridge regression. If you guys
remember, this one is just a slightly
different penalty. Um,
it takes the it I drew it out for us. It
takes the same penalty we had before.
So, it has that um residual sum of
squares error, which is the main one we
used for linear regression, but it has a
penalty with an alpha and then it has
the sum of the beta squares
beta squares. So it penalizes it has a
penalty but it penalizes slightly
differently where it uses the square not
the absolute value. That's the ridge
regression. And this has the similar
effect of you don't in order to minimize
this right because our goal in training
a model was to minimize this thing
minimize this um quantity and find the
best betas that minimize this. Um so
generally yes you want to encourage
lower values but the um once you get
values that are a fraction if you square
them they actually get smaller. Um so uh
it's it's not um it's not necessary to
shrink them all the way to zero. They
will get smaller as soon as they're kind
of below one. Um so they don't encourage
it to completely go away uh like the
absolute value does. It's just slightly
different minimization. Um so what we
see with the ridge is we don't see the
features kind of get wiped out
completely like we do with a lasso. In
lasso they get encouraged to be um to
become zero because that's kind of the
only way to minimize an absolute value.
But with squares they can keep getting
smaller and smaller and smaller um
fractions and they don't have to become
zero. It's not as harsh of a of a
penalty.
Um,
so, uh, the ridge was easy to use as
well. Um, and it also has an alpha that
we can set. So, it's literally the same
exact code, just a different model.
There's slightly different penalty and
it results in different coefficients.
You notice that none of them are exactly
zero. Like with the lasso, you can get
ones that are exactly zero. We don't see
that with the ridge. You remember that?
Um so we we s pointed out that last
time. Notice the coefficients aren't
zero. Um and then we can evaluate it. So
we did our MSE calculation which is a
pretty standard thing where we use our
model to predict on a training set,
predict on a test set, evaluate those um
by computing the metric like the mean
squed error and we can see if we're
overfitting, underfitting. This is
definitely the same kind of story we've
seen with all these models is
underfitting because the error is so big
across both sets
across training and tests. So it's it's
definitely underfitting.
Um
and same thing as lasso, it has a cross
validation uh variation on it that
allows you to try out different alphas
and um do different folds. So 10fold,
fivefolds, whatever and compute the um
try to find the best alpha that way.
Okay.
All right.
Any questions on this so far from last
time from reviewing that a little bit?
Hopefully that uh hopefully that is
jogging your memory a little bit on
ridge and lasso. Um, you know, where
we're going to pick it up today is to
finish out this lesson with one more
model,
which is going to be a combination of
ridge and lasso. So, you can actually
combine them together
um in a linear fashion, those penalties.
So, you can actually have both
penalties, the absolute value and the
square. And when you have both penalties
um that's a special model called the
elastic net uh regression or elastic net
model. Um so this is a combination of
lasso and ridge together. So you have
lasso, you have ridge and then you have
elastic net which combines both of those
penalties. Um let me show you the
equation.
So here is the uh so here is the the
model. This is the same that we've
always had. This is our usual u model
fitting for linear. This is a basic
linear regression um loss function or
objective function that we're trying to
minimize to find the betas. Notice how
we have both of our penalties though
this time. So instead of just having one
of the penalties, we actually have both.
So we have the lasso penalty
and then we have the ridge penalty here.
So we actually use both of them and um
try to find a balance of minimizing
those two uh those two penalties.
Okay. And notice how they instead of
just a single alpha, we kind of have a
balance on both of them.
So we can actually weight the lasso one
more. We can weight the ridge one more.
We can weight them the same. Uh we can
um change that around as much as we
want. So they have two different weights
there um that they could be.
Um now what happens in reality is uh
we're going to see this in the model is
that um usually what happens is these
get combined into a fraction. So there's
usually a ratio of lambda 1 to lambda 2
and this is known as the um this is
sometimes known as the L1 ratio
and this is a this is a a parameter
inside the model that we'll be able to
set um along with alpha. So we'll be
able to set an alpha and then this
ratio. Um the idea is is that um the
ratio will uh allow us to control which
one is more dominant. So if this number
is bigger the um this lasso penalty will
will be weighted more. If this ratio is
smaller if it's less than one for
example that means that the um ridge
regression is more uh dominant. Um but
the so we'll have this we'll have really
this and this at our disposal and alpha
is um
alpha is kind of like a a you can think
of it as a scale that is um so lambda 1
kind of like lambda 1 plus lambda 2 um
combined to equal alpha.
So it's like our total level of penalty
um our total level of penalty and we can
set that equal to one. We can set it
equal to whatever we want. Um and so
these will be in this ratio and there'll
be a total level of penalty that we can
apply. So the model will actually use
these two parameters when we when we do
it. But that's how they're that's how
they're all related.
Okay. So ridge uses both penalties.
That's the only difference between lasso
or sorry elastic net uses both
penalties. Um so one thing I want you to
notice is that uh if we um if we want we
could set this L1 ratio all the way to
zero.
Um which uh if we do that um the only
way this L1 ratio could be zero would be
if lambda 1 is zero. So it would just
revert back to ridge regression. So it
complet if if this is zero this will
wipe out this term and we'll be back to
ridge if the L1 ratio is zero.
Okay.
All right. So we have a elastic net
model. Um now it's used the exact same
way as we did the other models in the
code. So we have elastic net um uh from
the scikitlearn linear model family just
exactly where we had linear regression
lasso ridge all of those came from this
linear model um elastic net also comes
from there and then the cross validation
version also comes from there um
so let's see so when we build our model
it's going to be um very very simple
easy stuff because it's the same code
that we always have um we just use the
elastic net. We set an alpha alpha
equals 1 is pretty standard um just like
it is in in the last one ridge that's
industry standard is one and then an L
L1 ratio of.5
that's pretty standard as well. What the
L1 ratio of.5 is is kind of a um
uh kind of a that means that the lambda
1 to lambda 2 ratio is 1/2. Um, so
that's that's a pretty standard uh ratio
as well. But again, we could set this
equal to one and they'd be kind of
equally weighted. Um, L1 ratio of a half
means that the uh ridge regard the the
ridge penalty is a little bit more
weighted uh in that in that situation.
Okay.
So uh once we have this model um we can
do ffit and we can run that on our
training data and we can um get we can
figure out what our parameters are like
our coefficients and our intercepts. Our
model will have that but more
importantly we can use our model to
predict right so we can predict on the
test set. Um let me go back and load our
data and actually run this.
So, we're going to be using the same
data that we did for uh lasso,
which is the I'm scrolling back up so I
can load it. It's the baseball data
here.
Um,
just run it from there.
It's this hitters.csv. So, hopefully you
have that one.
Let me load this.
Okay, so we loaded that and then that
should load.
Drop that unnamed column.
We will get our dummies
and then concatenate those split
scale. I'm just rerunning things. I'm
rerunning things so we can see our model
one more time.
So rerun that. Take a look at that. That
looks good. and then
fill in the nles on the on those.
Okay. So, we should be able to run our
uh elastic net now.
Okay. So, let's import that and then
let's build our model. So, there we go.
We build our model and the intercept is
that. Now, of course, we can look at our
coefficients. Let's look at that.
Look at our coefficients. So remember
the coefficients are the uh betas. These
are our betas that are in our model. Um
so we can take a look at those. Now um
they're it's somewhere in between. It's
not a full lasso where we're going to
see some of these be zero. It's not a
full ridge. Um so the coefficients we
get are different. They're somewhere in
between there those two models that
we've already built. So not quite the
same um somewhere in between there.
Um and then we can use our model to make
predictions and and compute the MSE
uh or the RMSSE I should say as well. So
we can take the mean squared error, pass
that into the square root and comput the
RMSSE. So still pretty bad. Um this is
right around that 300 range of what
we've gotten for our other RMSSE. So,
it's not like elastic net is any better
than those other like linear or lasso or
ridge. And that's not surprising because
it's just adding those extra penalties.
We don't expect it to magically get
better. It's actually a more complex
um when we add when we add those in,
we're actually reducing it and making it
simpler. And we need something more
complex, I should say. So, we're making
it simpler um by by making penalizing
our weights a little bit more. And so
it's still not a good fit. That's not
really surprising, right? It's still not
really a great fit.
And we can we can even double check
that. We know our RMSSE is pretty bad.
Um but we can double check it with this
R2 score. And it's, you know, still not
good. Remember, a one would be really
good. Um that'd be like a perfect linear
model. This is um still pretty bad.
Okay, so as we said, the alpha controls
the overall strength. Um, so the higher
the alpha, the more overall penalty
we're supplying, which makes the model
simpler. Um,
uh, but the L1 um ratio determines the
mix or that ratio of the lambdas, the
lasso to the ridge. Um, if you have it
be um exactly zero, you you revert all
the way back to um if you if you put it
at zero, you revert all the way back to
ridge one would be all the way to pure
lassos. Somewhere in between like one
half is is good.
Okay,
so this is another example of trying out
different values of alpha in the CV to
see which one works. Now again, I'm
going to show us in a minute a
systematic way to do this, but this is
just trying out um different alphas that
we set up in this uh in this um
uh range. So we have different uh values
between minus2 and two um
logarithmically.
Um so these are uh logarithm values that
are between this between minus2 and two
and we choose 100 different alphas and
then we choose 100 different um L1
ratios between 0.01 and one and we run
that we run this um cross validation
with 10 folds. So this is quite a bit.
So we're doing 10 folds and we're trying
out 100 different um options. Uh every
time we do an option we're trying out 10
folds to evaluate it. So, it's going to
take a minute to run.
It's still running here. But again, what
this is doing is trying out different
alphas and it's it's going to do a cross
validation. And you guys remember the
t-fold cross validation is where we take
our data and we divide it into 10 folds
and then we um train on nine of those
and then test on the remaining fold and
then we rotate all the folds 10 times.
and that we average those mean squared
error metrics together um against those
10 different uh fold options to generate
a basically like an average performance
for that value of alpha. And we're doing
that 100 times for all these different
100 alphas that there are and 100
different L1 ratios that we're trying
with them.
So that's quite a bit of processing but
uh it did finish.
So we can see what our best alpha is and
our best one ratio. So we get the best
alpha is this best one ratio is this. Um
and therefore we can uh build a model
with those with just these two guys as
the alpha and the L1 and um see how that
performs.
We build that model and then we predict
on the test set and we generate the
RMSSE. It's just a little bit better.
It's still not It's just a little bit
better, but it's still not good, right?
It's still 338. It is just way too big.
Remember, this is RMSSE, so it's in the
units of our uh target variable. So,
it's in the units of, if we go back to
our data, actually, I can just print it
out here.
um this RMSSE.
If I just do this, we could take a look
at um DF
or I could look at Y test
and you can see some of these values.
These are these salary values in the
hundreds, right? Some of them are in the
thousands. Um but an error of like 338
is just too big. That's a really big
error. That means we would be off by an
average of 300 when our our values if we
just do the mean
um
is only 550 as on average is 550 but we
have this amount of error on average um
so that's just a way too big of a
proportion of error right it's not a
very good model and again we can verify
that by looking at this R2 for.
So if we go down here,
still not very good.
Here's some of our coefficients. So
remember, you can always take your
coefficients and line them up to your
your data columns. Uh so that you can
get a sense of what coefficient belongs
with what feature. So that's all we're
doing here is just creating a series
where those coefficients instead of just
printing out the coefficients, we're
actually lining them up to the columns.
So this tells us um remember the larger
it is the more influence it kind of has
on the on the final result. Um either
way, so like this has a big negative
influence. Um, this has a large positive
influence.
Okay, let me pause there. Any questions
about the
elastic net model?
This is a really this model is a really
good one to use when you are building a
linear regression and it's performing
well, but it's overfitting. This is a
really good one to use because you can
balance
lasso and ridge you can get the best of
both worlds. So the the main strategy is
if you are using a linear regression and
you see overfitting
um meaning that it's performing decently
so on the training set
it's performing okay but then on the
test set like you know it's it's not
underfitting. it's performing pretty
well on the training set, but then on
the test set it's um performance is much
worse. That's overfitting. If you're
overfitting, then this is a great model
to use because we can try basically by
by rotating through different alphas and
different L1 ratios, we can try out
different strengths of penalty and
different variations on lasso and ridge
together. This is a really good model to
to use for those overfitting cases where
linear regression is doing decently
um but it's overfitting.
Right? So far we haven't ran that case
because so far no matter what model
we've used it's always underfit. So
anytime we have those underfitting cases
it signals that we should likely just
use a more complex model and we haven't
learned about those yet.
um we will coming up in lesson four, but
um that's for this data. That's ultim
ultimately what we'd want to do is
probably use a more advanced model
because it's underfitting um just using
a linear regression and and then using
the the overfitting variations of linear
regression like lasso ridge and elastic
net.
Okay.
Any questions on this on elastic then
the TV? Yeah, we Yeah, I think I have
it. I can share it with you.
I said that and now I can't find it. I
thought I had it.
I don't have it. I thought I had it, but
I don't.
If anyone does have that one.
Yeah, I'll look one more time. I thought
I had that one.
Um,
yeah, it's not in there. I had it. Let
me see.
Yeah, I don't have it either. I thought
I had it in here.
Yeah, I don't have that one.
I don't have that one. and I'll have to
find it. Uh I have this marketing data.
I don't think this is the same one.
I have this marketing data. I don't
think that's the right one, but you can
take a look at it.
No, we're using So, for this example,
we're using the same hitters data set
that we used earlier for lasso.
No, that's an earlier one.
That's from the uh very beginning of the
notebook. So that's the that's from this
one.
Oh, this Oh, this is where it is. Sorry.
This is where it is. You can find it
here.
That's right. It was from a URL.
It was used in the very beginning of the
notebook.
And we did we did this.
Okay.
That's right. That's why I didn't have
it downloaded.
Okay.
All right. Any other questions on the
elastic net before I move I'm going to
move on to uh finding those a systematic
way to find the best hyperparameters.
Um, I'm going to show you a couple
strategies to doing that. Um, so far
we've just ran CV with some random
choices. Um, I'm going to show you a
better, more systematic approach. That's
kind of the industry standard for doing
tuning. Um, so I'm going to I'm going to
show you that next. But any questions on
the elastic net?
Okay. And again like you know
scikitlearn makes it really easy for you
guys because it just behaves the same
way as any other model. You use the
object and then you do ffit and predict
right? So the ffit is going to train it
um and the predict is going to allow you
to use that model to predict. It's it's
super easy that way. Every scikitlearn
model is like that dofitit and predict.
So it provides a really simple way to
use basically every model.
Okay,
let's talk about let's finish up this
lesson with a couple things. Um, one of
those things is going to be
hyperparameter tuning. So what is this?
The hyperparameter tuning is a
systematic way to find the best
parameters in a machine learning model.
So a lot of machine learning models have
what are called hyperparameters.
These are not the betas that we learn
during the training that's learned from
the data. These are settings that we set
ahead of time like the alpha. That's a
perfect example like alpha L1 ratio in
in the elastic net. We set those up
ahead of time and depending on what we
pick for those we get different
performance, right? And so what we
really need is a systematic way to find
the best settings for those
hyperparameters as we are training our
models. Um the the the main like idea
behind this process though is going to
be to systematically try out different
combinations as many as we want to try.
And so we're we're basically going to
have a strategy for tuning that is going
to exhaust all the combinations of those
hyperparameters that we want to try
until we find the one that performs the
best. Um and and that strategy is known
as grid search. Um and essentially what
it does is it sets up a grid um where
which is basically like a matrix to say
okay which parameters do you want to
try? I want to try um alpha and I want
to try L1 ratio
um L1 ratio like let's say I want to try
these two. So we set these up in a grid
where we say, "Okay, I want to try this
value. I want to try this value. I want
to try this value. This one, this one,
this one, and on and as many as we want
to try." So we could set up set those up
systematically like a linear um a
linearly spaced like I want to try every
alpha between between 0 and 10 spaced by
one. Um whatever, you know, we can set
up different ranges of those, but that's
going to be in this grid. And then the
L1 ratio, same thing. We can try out
different values of these that we want
to try. Let's say there's many of those.
Um maybe every um tenth between 0 to one
I want to try out. Um so you set up your
parameters and you can set up as many as
you want in the grid. And then
essentially what you're going to do to
do grid search is you're going to work
your way through every combination of
those. You're going to try out this
combo. You're going to try out this
combo. You're going to try out this
combo.
this combo. So the first value of alpha
with every possible L1 ratio, then go to
the next, try out the next value of
alpha with every L1 ratio, and on and on
and on. So we're going to try
all combos
in the grid.
We're going to try all combos and we're
going to find the lowest MSE
combination. find lowest
MSE
combo.
So whatever leads to the best model is
going to be the um parameters that are
that are deemed to be the best. And the
idea is once we have found those we know
that we can use we can go ahead and
train a model with those best alpha and
len ratio and on and on and on.
Yeah, when you get an So this goes back
to the error. Remember that for a
regression,
the error is this measurement of how far
off we are, right? So if we have a bunch
of points and we draw we fit a line
through there, the the MSE is measuring
this distance, right? So what do you
think is a good distance? Like if our
model is perfect,
what's the best distance from our
predictions to the actual points? Zero.
Yes. So the lower the better. The lower
the better. Um so for an R RMSSE, the
lower the closer to zero the better.
However, the RMSSE can be it's its units
are interpreted in the units of our
target.
So what is deemed to be good is relative
to our target. Like let's say our target
is in the thousands. Like it averages in
the thousands. If we produce an MSE of
50 or sorry an RMSSE of 50, that's
pretty good, right? Because our units
are in the thousands
and we're only on average we are off by
50 units,
right? Our distance away is about 50
units. That's pretty good. So the RMSSE
is relative to your target variable.
Does that make sense? Yeah. It depends
on the target. It depends on what you're
trying to predict.
So that's why we got RMSSE that were in
the 300s for those hitters, but the
average was the average of the target
was in the 500s. So that's a really bad
proportion of error relative to the
average target value, right? If our
RMSSE was 300,
but the target was sitting in the 500s,
that's just too much error. Way too much
error, right? That's just too big of a
value. Um, our predictions are just off
way too much in terms of that distance.
So, this would be this was a bad model.
It was underfit.
We know that from the the RMSSE. So
yeah, the RMSSE closer to zero, no
matter what is good,
zero is being perfect. Um, but it to
know what's good, you need to know what
your target is on average and then think
of this as kind of a ratio to that
average target. I think that's the best
way to think about it.
Okay, so going back to this grid idea is
so the grid is just basically laying out
all possible parameter combinations and
trying them all out by fitting and
predicting until and generating an a
metric like an MSE
until we find the one with the lowest
MSE. So find the lowest MSE combination
and that will be the best
that will be the best combo. And then if
we once we know that best combo we can
use that we can use that alpha we can
use that L1 ratio and use that model
going forward. We can we can use those
parameters in our model. So this
strategy it has a name. It's known as
grid search. So it is a hyperparameter
tuning process that tries out all
combinations.
So what's the what's the uh benefit to
this is that we get to test out a lot of
different combo combos of those
parameters like the alpha and L1. So we
can be confident what the best model is,
right? So we can pick the alpha and L1
perfectly because we're trying out a
bunch of different combinations on the
data to see which one's the best. What's
the downside?
It's expensive, right? It's an
exhaustive search. So, if you have many
different parameters and you're trying
out many different combinations, it can
get exponentially
expensive
to perform this search. Okay, so grid
search is great except for the fact that
it can be expensive if you have many
parameters with with very wide ranges
that you're searching over because that
that's a lot of combinations you have to
test, right? And especially if you have
a lot of data, that's going to be
expensive
um to do.
So uh we're going to practice doing grid
search, but that is that's the pro and
con. The pro is that we get to try out
all these combinations and see which
one's the best. The downside is it can
be expensive to do that if you have a
lot of parameters um that you want to
tune for your model um and you have very
uh many different choices that you're
trying to evaluate for those and it just
creates a really big um collection of
combinations that you have to try out,
right? Um that's the only downside to
grid search.
Now on the opposite end of the spectrum
of that is a randomized search or random
search and this will basically just um
do a sampling of those parameters from
um kind of fixed uh specified
distribution. So essentially what you do
is similarly you define your range. So
you say I want to look at alphas um
between zero or sorry between let's say
yeah 0 to 10. I want to look at a bunch
of different alphas um and I want to
look at a bunch of different L1 ratios
that are between 0 to 1
0 to one and um what we do is we say
okay I'm going to restrict only testing
20 30 40 times. I'm not going to do all
possible combinations. I'm just going to
randomly sample something in this range
and randomly sample something in this
range. And so, and I'm going to perform
that experiment a fixed number of times.
So, let's say I set the uh sampling
where I'm only going to do um 20
evaluations.
And so, 20 times we're going to pick a
combo randomly. So, I'm going to pick an
alpha and I'm going to pick an L1 ratio.
L1 ratio.
And um we are we are just going to uh
sample those randomly from this range.
Um and we're going to use those and test
those out. And then it's but otherwise
it's the same as grid search. Whatever
is the lowest MSE. Um, so whatever is
the lowest MSE is the best.
So we evaluate those, we sample, we
train the model, evaluate it. Whatever
is the lowest MSE
is the best is the best combo. Now
what's the benefit to this is it's a
much more controlled experiment in the
sense that we um aren't going to iterate
through every possible combination in
the grid. where we basically set up a
fixed number of times we're going to try
out stuff.
The risk to doing this is that you're
not
you're not exploring all combinations,
right? Because you're randomly sampling,
you may get unlucky and you may not
stumble into the best. You you can make
um samples and figure out what's the
best amongst your samples, but you may
not be covering all the combinations.
Does that make sense? The grid search is
going to try every combo. The random
search is going to randomly sample those
combos.
So, it's not going to try every single
one. It's going to try a limited number,
however many you set up. Now, if you set
that number really, really, really high,
now you're starting to approach a grid
search because now you're sampling so
many of those combos that you basically
are trying them all at that point,
right?
Um
so so that's the way the random search.
So by the way both of these use cross
validation in the sense that when you
evaluate accommodation you're actually
doing it with cross validation. So when
you do an evaluation, you're going to do
probably 10 or five folds where you
split your data and then you test it on
the rest of the folds and evaluate or
train it on the rest of the folds,
evaluate it on one of them and generate
an average MSE to get your evaluation.
So every evaluation is using cross
validation. That's why that's and
hopefully you can see why this would be
so expensive for a really big grid,
right? because you're trying out many
different combinations
and every combination is going to do a
cross validation procedure. So, it's
going to train 10 times and test against
10 different folds and average those
together. It's going to be a pretty
expensive operation
for a really big grid, right, of of
parameters.
Um, but these are the two kind of
systematic approaches we have at trying
out different hyperparameters. Remember
those those things are called
hyperparameters. These are those choices
that we have before we train our model.
Um those choices we have that affect the
performance of the model like the
alphas, the L1 ratios, those kind of
things. Um we have control over what
they're going to be. This is a
systematic approach to find out what the
best
uh value of those parameters is going to
be on our data,
right?
Okay. So, before we practice this, we're
going to practice a grid search first.
Um,
any questions?
Uh, I don't know if it has a built-in
That's a good question. by time limit. I
don't know if it has a built-in way of
doing it, but you could certainly set up
like a a a loop um to like to wrap
around, you know what I mean? Like you
could set up a loop where you check the
time. If it's if if the time elapsed as
you're doing the search, if the time
elapsed is greater than the the time
limit, then you can kind of break early.
Um so it's not hard to implement that,
but I don't know if it has that built
in. I don't think it does
because I don't think it really cares
how long every evaluation takes. It's
just going to exhaust all those
especially in a grid search.
But um yeah, I there's probably a way to
manually kind of set up a time time
loop.
So hyperparameters are um settings that
we have on the model itself and a really
good example of this is like the alpha
and L1 ratio in the in the elastic net.
So they're not things that we um learn
from the data directly like the betas in
the model like those get trained
directly by doing the um lease squares
process right um by doing that gradient
descent and all that optimization.
Um so these are not learned from that.
They're actually set ahead of time. And
so what we're saying is
the best way to understand the effects
of those is to try out different
combinations of those until we land on
the best one. Right? So hyperparameters
are those options we have in the model
like the alpha like the alpha and l1
ratio in the uh elastic net. Many models
have hyperparameters. Um we're actually
going to see that in in future models
that we study. they have options that
you can set that affect their
performance.
And so this this is just a strategy to
evaluate those different options to see
which one's the best.
Yeah. So again, hyperparameters, those
are settings on the model itself um that
affect the performance of it.
And basically we have the two two
strategies here. We can set up an
exhaustive grid and search through all
of those until we find the lowest MSE uh
option or we can randomly sample
potential options, try them out and see
which one's the lowest as well. And do
that a fixed number of times. Um
sort of like a fixed number of trials
almost. um which has a risk of not
trying out every option but but
hopefully you try out enough that you've
explored the space a bit and you get
some quality choices there but no
guarantees right no guarantees you try
everything which is what a grid search
will do it will try everything
okay now luckily per usual scikitlearn
has something to manage this process for
us in terms of grid search. Um so in
that way we will not need to manage this
process ourselves. We can just rely on
scikitlearn. And so if you're doing
hyperparameter tuning um this is going
to come from the model selection module
inside of sklearn. So we're going to
import from from skarn the model
selection module. We have our grid
search cross validation.
Okay. That's what the CV stands for.
grid search cross validation. So, this
is going to do that grid search
strategy. Um, we're going to set it up
with our dictionary essentially of
choices. So, we're going to say, hey,
here's the alphas I want to try. Here's
the L1 ratios I want to try. Um, and
here's my other settings like uh how
many folds I want to use, what my random
state is for the shuffling. So, we'll
set all that up. Um,
and then we'll just run the grid search.
And then what should come out of that is
the best options for our parameters from
the grid. And then we can use those
going forward in the we can build a
model with those best options, right? So
we're really doing some evaluation here
of what is going to be those best
alphas, those best1 ratios on our data
set, right? And the only way to really
know that is to evaluate them because
they're not things that are learned
during the training. Hopefully that
makes sense, right? They're not things
that we learn directly from training.
They're things that we have to set and
then kind of evaluate and see how they
affect things.
Okay. So, we have grid search CV. That's
going to be our primary um tool to do
the evaluations of the different
hyperparameter options.
Grid search CV. Um we're going to set up
our cross validation uh object here. Now
I want you to pay attention to this is
that um it's a slightly different
version than the kfold we had earlier.
So we've used kfold before with a
certain number of folds. This would be
10 folds and we can set a random state
for the shuffling um that happens in the
folds.
But this is actually a slight different
variation on it where it is a repeated
kfold where we do three repeated trials.
Now why would we do that? It's to be
extra extra extra careful with the
shuffling.
So this what this means is we do three
different shuffles. So we do k-fold, we
actually repeat it three times with
three different shufflings. That's all
that means. So the repeated k-fold is
actually a bit beyond the just basic
kfold. What basic kfold will do will
we'll shuffle and then do our splits
into 10 splits and then train on nine of
those. test on the other one and rotate
through all the splits.
We're actually going to do that process
three different times with three
different shuffles. So this and we're
going to average 30 results instead of
just 10. So repeated kfold is just going
ab above and beyond to do extra to
repeat the kfold three different times.
In this case only three. We could do
more.
But um now is that necessary to do? You
could argue not necessarily. Um but it
just provides extra robustness
uh beyond just our single shuffle and
then split and then rotation of those
folds. Right? We're doing it actually
three different shuffles. Um so we're
repeating our kfold three times uh for
every now is the thing is we're doing
that for every evaluation. So it is
going to be more expensive than just a
basic kfold.
So we have three different K-fold trials
that we're doing essentially.
Okay, hopefully that makes sense. This
is the repeated K-fold. We haven't
really seen that before. We've only
worked with the Kfold, which would get
rid of this repeats option and only have
uh 10 splits in a random state for the
for the single shuffle that we do. So we
can um recreate that same shuffle every
time. Um but now we're actually going to
do three random shuffles, uh three
different trials. So one shuffle creates
the and then create the 10 splits,
evaluate, then go back and do another
shuffle, another new 10 splits. So, one
thing that should be um clear is that we
get different splits every time because
we're going to shuffle once, right?
We're going to shuffle once and generate
our splits
and then we're going to shuffle again,
generate these splits which are going to
be different and then shuffle one more
time for for three different times,
right? And then get get these splits and
then we're going to get 10 metrics here,
10 metrics here, 10 metrics here, and
then average all of those together.
So, it's a bit more just going up extra
above and beyond for a kfold. Okay.
All right. So, here comes the fun of
when you do grid search. Now, the grid
is actually just a dictionary. It's a
Python dictionary where you declare what
your parameters are going to be inside
the dictionary and you set up a range of
values that you're go or a list. It can
be a list. It can be a range
but some declaration of what you are
going to test and evaluate inside of
your grid search. So the grid is
initialized as an empty dictionary.
And then what we do is we say okay in my
grid I want to check different alphas.
So we're going to add a collection of
alphas in here that we're going to test.
So let me make a comment there. We add a
add a range of alphas to test. And this
range is a this is just like the Python
range. Um
this is just like a Python range um uh
operator here where this is going to be
uh every so it's going to be um every
uh value between
zero and one um uh steps with a step
size
of 0.1. So it's going to try a bunch of
different alphas um between uh zero and
0.1
sorry 0 and one stepping by 0.1. So it's
going to try zero.1
2.3 point 4.5 6 right all the way up to
one.
So that's what this will do and it's a
numpy range. So it's just all those
decimals between 0 to one.
It you can use either that's valid.
Yeah, you can do you can do that to
create a dictionary or you can use the
keyword um dict. You can use either one.
Either one works.
Whatever whatever you want to use.
They're the same.
Yeah. The the reason people prefer
dictionary is because um sets are
created with the same braces.
So it it makes it clear what you're
creating as a dictionary. If you use if
you use this, that's the only advantage
is it's just plainly obvious what you're
making. Uh because technically you can
make a set with the curly braces as
well.
Yeah,
no worries. Um okay, so we have our
alphas here. So what I want you to
notice is that we are going to try out
different alphas and we are that's the
only parameter we are going to try in
our ridge regression. So we're going to
we're going to try ridge but just try
different alphas in the in this range um
in our grid search. So the grid search
CV takes in a model. It takes in our
grid dictionary which is really
critical. We need that dictionary to
declare what we're going to try.
um we need a scoring to say to find the
best. Now remember it uses the negative
to find the lowest which is going to be
the the least negative option.
Um otherwise it wouldn't um just based
on the optimization it would look for
the highest value. Um so the highest
would be closest to zero in this
situation. Um so we use negative and
again we could use squared error. It's
using absolute. We could use um squared
uh either either one works.
Um more typical would probably be
squared error, but um absolute is fine.
Here's where we have our repeated kfold.
So we pass in our um how we're doing CV.
That can be it can be a kfold object. It
can actually just be an integer, which
is say I just want to do 10 splits or
five splits um to to do every
evaluation. But these are the bare
minimum that you need just really the
model and the grid and your CV. Um what
metric you're using to evaluate what's
going to be the best. And then this end
jobs is to parallelize. If you have it
set to minus one, it's going to it's
going to try out all the grid options in
parallel. Um which is nice. It's going
to help speed up the overall search.
Okay. So let me mark that down as n
jobs equals minus one.
tries out the combos in parallel.
So in this situation, we actually don't
have more than one parameter. We only
have the alpha. So we're really just
going to be systematically working our
way through every alpha and evaluating
which one's the best right with this.
And notice that in order to use this
grid search, all we have to do is call
search.fit. So it works kind of like
every other model does, right? It's the
grid search.fit.
And we pass in our data.
And we um once we're once this prints
out the results, you get a results
object
um which has a best score and then a
dictionary with your best parameters.
So, whatever your best grid member was
or grid members, um it prints that out
and you can So, for from that, we can um
grab our best alpha, which it which
let's confirm what that ends up being.
Oops. We need to import repeated kfold.
So we'll import that.
Oh, I didn't. Let's do from
sklearn.linear
model import ridge.
Okay.
Okay. So, it completed the search and
what we found is this is the best score
is 238 for the mean absolute error and
the best alpha that we got was 0.9. So,
the best alpha that worked here, the one
that gave us the best score was actually
0.9 as the alpha. So what it did is it
tried out everything between this range
and 0.9 was the best. So it did cross
validation tried out every single combo
in our grid.
So if we want we could actually print
out
print our grid so we can see
what our combinations were.
So, it tried out all of these guys and
the best one that we had was 0.9.
Okay, so pretty cool how that works. And
if we had other parameters, like if we
were doing a elastic net, we could add
those into our dictionary and it would
do all combinations of those. So if we
did um so for instance to add to our
grid we could do grid
um L1 ratio
this would be for like an elastic net
right now the ridge regression by itself
doesn't have an L1 ratio parameter but
just as an example um we could try out
different ranges
um similar range different one um maybe
an exact list whatever we want to do. So
this is going to try out different ones
between 0 to one as well. And so it's
going to try out every combination of
these from this grid.
Okay, if we did that. But again, this
the ridge regression doesn't have an L1
ratio. The elastic net does. So that the
ridge regression only has an alpha to to
as a hyperparameter. So we're only
testing out that one.
Okay, so that's grid search CV.
Pretty useful. This is pretty useful in
doing parameter tuning again when you
want to try out ranges of different
values and you can evaluate those to see
which one is your best and then we can
use that best going forward. So we can
for instance this is what this code does
below it is it fetches the best. Um you
can do it this way or you can do it um
the alternative is to do results.b best
params
and then you can just grab it like this
alpha.
Either way you can do get or like this
um and it this is just a dictionary,
right? And you can grab your alpha. So
that's the 0.9 um and we can pass that
alpha into the ridge regression and go
back and refit it to our data um and
then use that model going forward. So
the grid search really just evaluates
those different options, tells you
what's the best according to this score,
right?
And you should, by the way, you should
interpret this score in the positive
sense. It's only negative because we're
purposely making it negative to find out
what the lowest option is, right?
Because the lower is the better. So we
we purposely make it negative to make it
whatever is the least negative is the
winner. Um more negative is worse.
So it's the really positive version of
it is the is the true result for the
error. Um and they are a tool from
scikitlearn to put together your model
with your pre-processing steps. So they
kind of get automated together. Um and
they combine everything into kind of a
streamline process. You're going to see
what that looks like, but it's a really
nice um feature of scikitlearn. Um why
would we care about pipelines? They help
organize our code um so that we ensure
that we basically always run the
pre-processing steps before we train and
use a model to with the predictions. Um,
so it bundles those steps together,
minimizes the risk of forgetting a step
because one of the things that can
happen is when you do pre-processing, if
you're doing it on the training set, you
have to do it on new test data as well
when you put it through your model
because your model is training against
that pre-processed data.
So in order to make sure you never
forget that, you can bundle it all
together in a pipeline which is going to
make things really really easy to use
and and make sure that those steps
happen in a repeatable way. Um and it
makes things easier to uh deploy that
model as well because everything is
together in one pipeline. So in the in
the industry, I've seen this a lot. Um
you know, people will do their initial
exploration steps and initial model
building. They may not use pipelines
right away, but as they found their
model, um they'll generally move it into
a pipeline and all their steps into a
pipeline so that it's uh easier to work
with um when you're when you're
deploying it and actually using it uh in
in the real world. Um, so here's what a
pipeline generally looks like. It's from
scikitlearn. It's this pipeline object.
Um, and the pipeline is made up of steps
that we're going to see that that are
various um uh basically um kinds of
pre-processing we've seen before like a
scaler or um filling in missing values.
Those kind of things we can put here in
the steps which is basically a list. um
steps is just going to be a list of
scikitlearn functions that we can apply
to data. One of those being a model. Um
and then whenever we use the pipeline,
it's basically um you know it's going to
be something like pipeline.fit
or pipeline.predict.
So the pipeline kind of behaves like a
model. It's just going to contain many
more steps than that like the
pre-processing steps we've worked with
before. Um, and it also has some
capabilities for caching. So you can
like uh cache some of the data in
memory. Um, so that if you're reusing
the predictions, it kind of goes faster.
Um, so there's some options for that
too. I'm not too concerned about that at
this stage, but the main thing is going
to be filling out our steps and then
using the pipeline.
Okay.
Um, so some important bits of
information about the pipeline is that
it is going to be a sequence of data
transformations that will have at the
very end of the pipeline the model
because of course we're going to do
transformations and then train a model
or predict with a model. So every
Oh, can you guys hear me? Okay,
not able to hear me. Thanks for letting
me know. Can you guys were you able to
hear me so far?
Okay. Make sure. Yeah, it might be on
your internet or your your uh Yeah, it
seems like seems like it's good. So,
no, you can't hear me. Check your
volume. Check your headphones if you're
wearing headphones. Oh, no issues. Okay,
perfect.
Okay. Yeah, local internet issue. Yeah.
Okay.
Always let me know. always let me know
because it could be the case that it is
me. So, um always always make sure to
let me know. Um but sounds like yeah,
you may want to check on that. Um
so, okay, what I was saying is every
pipeline is going to have a uh a
sequence of steps that go first and then
the model at the end. Um so, the order
really matters. Um
uh so the order matters in the sense
that we want our transformations to go
first. Things like scaling, things like
filling in missing values, we want those
to be first and then we want our uh
model to be last because we want those
transformations to happen prior to
training or prior to prediction. So
usually what you'll see in these
pipelines is a model at the end, right?
a model that's going to be at the end of
the pipeline because we want basically
our processing steps then our training
or processing steps then our
predictions. Um so everything in the
pipeline though is going to be from
scikitlearn. Uh that's how it gets
automated in the sense that all of those
things are going to have fit and
transform functions built into them so
the pipeline can use them. Uh, and then
the last step is going to be a model
that has a fit and a predict. So it's
pretty standard that the last part of
the pipeline is just going to be a
model. Um,
uh, so we can um, as we do more
modeling, we're going to play around
with the pipelines quite a bit and see
how we can change up some of the
parameters. like if we want to change a
model's parameter um we can actually
adjust it to do things like uh grid
search or cross validation. So um we're
going to see some examples of some
pipelines but for right now mostly what
we're going to see is how to build one
and then how to use one. And then as we
get into lesson four, we'll get some
more practice with pipelines because
we're going to start using them quite a
bit uh to build our models rather than
do manual steps uh all the manual
pre-processing
um and then kind of building a model
from there. We'll just include all of it
together in a pipeline.
Okay, so the example we're going to do
is with this housing with ocean
proximity. So we've actually looked at
this data set before. Um so we have uh
this ocean proximity data set that has
the feature of like how close it is to
the ocean like the bay or the less than
1 hour. Remember we had that and it had
the median house value for different
neighborhoods. Um so we're going to work
with that one again. Let me make sure I
have that one uploaded.
You guys should have this one. It should
be in your uh data sets.
Um, I'll I can upload it here in case
you don't have it though.
Does this use multi-threading? I think
it does. Yeah, I think in order to do it
can do uh um I think it can do
processing in parallel for some of the
pipeline steps. Um, now does it use that
all the time? Not necessarily because
some of it is sequential in nature where
you have to do one step and then you do
the next step and then you do the next
step. So it's not like you can do them
in parallel.
Um in terms of the like you need to know
the output of one step to compute the
the output of the next step. Um so it
can but it it doesn't always lend itself
well. The thing that will use
multi-threading is is like the training
process can be parallelized
like the fit um can be and for some
models it can be parallelized not every
model
it. So long answer is or the short
answer is that it depends
depends on what kind of transforms
you're doing and what kind of model
you're using. if you can really take
advantage of that.
Okay, so we load our data here and take
a look at that. Um, do you guys have
this data set? Are you able to load it
in? If you're following along, are you
able to load it?
Okay.
And and again, we've worked with this
data before, so hopefully it's somewhat
familiar. Remember, every row represents
a neighborhood, and it has a we're going
to end up trying to predict this median
house value as our target um variable,
our dependent variable. Um and we're
going to use the rest of these features.
Remember that um this feature is in
particular going to need to be one hot
encoded,
right? It's going to be one hot encoded
because it is currently a string. and we
need to turn that into a numerical
feature which is the one hot encoded
feature. So we're gonna have to do that
but we're going to do that as part of
our pipeline.
Okay. So we'll be able to include that
in our pipeline steps uh to to do one
hot encoding which is nice.
All right. So we're going to split apart
our data um as we normally do. So we're
going to uh create our feature uh data
frame which is everything but this
median house value. So we go ahead and
drop that column and then our target is
the median house value. So it is just
that column here. Pretty standard. Um
and then we're going to train test split
and um split it into 30%
uh test data. And again random state you
can choose whatever you want to be. that
just affects the shuffling. Um, so
whatever doesn't really matter what it
is. It's just so that when you rerun
this, you get the same result in the in
the shuffle.
Okay, so we have our train and our test.
So you want to make sure you run those.
All right, so what we're going to do is
take a look at our data
and see if we have any null values. Um,
if you guys remember this data actually
did have null values. You can see it
here in this this guy and exactly how
many there are is from this the sum. So
we have um 162 nles in in this data. Uh
and this is just a training data. So of
course you know the test data could have
that in there as well. Um so that's
something we're going to want to make
sure we fill in the blanks on any data
set we use whether we're using the
training or test set. Um, like if we're
doing training, we want to make sure
that gets filled in. If we're doing
predictions with the test set, want to
make sure that gets filled in. Um, so we
we should be doing that. Um, now
what we're going to do is use this data
to help uh train our pipeline or or use
with our pipeline. We need to construct
our pipeline. So far, we've just split
apart our data. We haven't done anything
with our processing steps in our model
yet. Um so roughly
it this should be the flow of our
pipeline. What should happen is we
should be doing some type of feature
scaling
um some type of uh feature um
manipulation. So that could be
engineering, that could be um that could
be uh doing the one hot encoding. Um so
extracting new features like one hot
encoding,
one hot encoding. Um we are going to be
doing that and and by the way this is
split up into this is when we use our
pipeline for training.
Um it's going to look like this where we
do our scaling, we do one hot encoding.
Um we have our model here. Um so that
could be a linear regression, that could
be a lasso, that could be a ridge, it
could be elastic net. Whatever model we
end up using is going to be last in the
pipeline. And we're going to run this
pipeline. Ultimately, we're going to run
pipeline.fit,
right? We're going to run a fit
function. and we get a fitted model as
the result of this pipeline.
Then when we use it when we use our
model for prediction,
we use our model for prediction in this
lower part, it's the same pipeline, same
exact pipeline, but it's this model has
now been trained.
So we now have a trained model here. So
the great thing about the pipeline is
it's the same this is the same pipeline
that we're using here. So it's just
going to it's going to repeat those same
transformations. It's going to do our
scaling. It's going to do our one hot
encoding. It's going to use our model
and it's going to generate predictions
and generate uh we can we can do
predictions. We can do evaluation like
an across validation. Um we can use it
however we want to use it. Uh but notice
that the pipeline makes it consistent
between training and test. We're using
the exact same transformations
and the model is last. It's it's either
being trained or it's being used for
prediction but it's last. Our
transformations are up front which are
things like our scaling, things like our
one hot encoding, right? Those happen
first. No matter what data we put
through there, we put our training data
through there, we put our test data
through there, they're going to go
through the same steps,
right?
So that's that's the design of the
pipeline. That's what it's supposed to
do. So our job is to create those steps.
So we need to create those relevant
steps and then put them together into
this pipeline. Okay, so that's going to
be the code we're going to see coming up
is we're going to build out these steps
and then put them together into the
pipeline.
Um, any questions on this diagram? Does
it make sense what we're trying to do
with this pipeline? We want to have
repeatable steps during the training,
during a prediction process.
Okay.
All right.
All right. So, um, a couple of things
we're going to need is, uh, to first of
all, let's jot down what steps we're
actually going to do. We're going to
need to deal with missing values. So,
we're going to fill in we're going to
need a pre-process pre-processing step
that fills in any nles. We always need
that, right? So, if there's nles, we're
going to fill them in somehow.
We're going to define how we do that in
our in our step. Um, and we also need to
one hot encode. And we need to scale,
right? Those are pretty standard steps
that we've dealt with whenever we're
building these models, right? So, pretty
standard things. fill in nles one hot
encode any categorical data whatever
however much we have and then go ahead
and um standardize which is the scaling.
So this this just is the same word for
scaling our numeric features. So we're
going to we're going to define those. Um
so that's why we're going to go ahead
and import from pre-processing. We're
going to import our scaler. Um again we
could use minmax scaler here. We're
going to use standard scaler. Um but we
could use minmax. Um we have our oneh
hot encoder here. Now usually when we do
oneh hot encoding we use pd.get dummies.
This does the same thing as that. But
because we're going to be building a
pipeline we actually want the
scikitlearn version of git dummies. So
this is the scikitlearn version of git
dummies here. And it and we have to use
that version in the pipeline because
everything in the pipeline needs to be
an sklearn object. It needs to be an
sklearn tool or object.
So um instead of using pandas get
dummies, we're using one hot encoder
which is does the same thing. Okay. In
fact, it just this basically just uses
pd.get dummies um under the hood.
Okay. So, it just uses that. Uh,
anyways, it's just code that builds on
builds on that.
Now, what's really nice here is we're
also going to use from sklearn.impute,
we're going to use a simple imper. Now,
what this is is an automated way to fill
in missing values. So, this is a fancy
way of basically doing the the fill na
on a data frame. So, simple imputer um
we are going to basically fill in the
blanks. What we're going to do when we
create this object is give it a strategy
of how to fill in blanks. Should you use
the average? Should you use the median?
Should you use the max? Should use the
men? Should you use a default value?
We're going to tell it what to do in
this object.
Okay. So, we're going to we're so we're
going to use this as our automated tool
for filling in missing values. So,
that's really nice. it has. So this is
going to be a critical part of our
pipeline an imputer that's going to fill
in missing values.
So we have that
yeah coding to reduce coding. Exactly.
Uh we have our pipeline now. So we have
our pipeline. So our pipeline is going
to hold everything. So we need the
pipeline object um to hold everything
and that comes from sklearn.pipeline.
Um, so everything's going to actually go
into a pipeline object. We're going to
see how that looks. Um, and finally,
we're going to from skarn.compose, we're
going to use a column transformer. The
reason we're going to do this is because
we are going to specify for some columns
like the numerical features, we should
be scaling.
For some columns like the categorical
features, we should be one hot encoding.
So the column transformer will allow us
to map different transformations to
different sections of columns which is
really useful. So this is actually going
to be a critical part of our pipeline to
apply to make sure we only apply this to
numerical features and only apply this
to categorical features. Right? So this
column transformer will help us um to to
apply pre-processing to particular
columns. Um like that ocean proximity is
the only one that really needs this but
every other column is going to need this
all the numerical features.
So we're going to use this column
transformer and again we're going to see
how this looks but just trying to give
you an idea of why we're importing all
these things.
Okay, so let's import those.
Uh, this mentions about the column
transformer. We just talked about it. It
allows us to have a particular column or
group of columns get the right
transformation. So again, uh, looking
ahead to our pipeline, the numerical
features are the ones that are going to
need scaling, but the categorical
features are the ones that are going to
need one hot encoding. However many
categoricals there are, in this case,
there's really only one, which is that
ocean proximity. to go back to our data.
Um, you can even see that in the info,
there's just that one. Um, and we see
that here, right? Just this one string
column that should be one hot encoded.
All these other guys should be scaled,
right? They should all be uh uh standard
scaled.
So, this will allow us to specify those
distinctions.
All right. So, let's get started
building our pipeline. So, this is going
to be really cool. We're going to build
out the pipeline. Um, let's extract our
numerical data and our categorical data.
Now, this is a really neat way of doing
that that I'm not sure we've seen
before. Um, so what this does is we'll
take our data frame, particularly our
training data frame, and select our
data.
That's what this select dtypes does is
select data from it. Um, which includes
only the object type columns. So only
the object types. Now what's that? The
object type is the string, right? So
this should select only this column
because it's in the include.
We go here include only object types in
the result. And so this should only have
our one categorical column which is the
ocean proximity. So housing cat is going
to have a reference to our uh it's going
to be a list that has a a basically just
our ocean proximity feature because this
select dtypes will make sure we only
pick object types and um
grab those columns. So this is a way to
neatly grab um our categorical features
here by including the object types. Now
on the flip side we can exclude object
types and get everything else. So this
is going to be all other columns which
is excluding the object. So this is
excluding this meaning we should get all
of our numerical features that way. So
this will be all of our numericals
by excluding the object type and this
will be our housing num which is short
for numerical. So this excludes
the uh object type
meaning all numerical
features
right all numerical features there.
Okay.
So, if we were to uh let's double check
this. Let's sanity check this. If we
were to print out the housing
cat, um this should be just the ocean
proximity feature, which it is. So, just
that one. If we were to print out the
housing num, this should be all the
numerical features, which are all these
guys. So it's just a reference to those
columns so that we can uh use those
later when we're mapping uh this
transform needs to go to this column
like the one hot encoding needs to go to
this column and the scaling needs to go
to these columns right so we have those
uh names of those columns already at our
disposal. So, we're just doing that.
And this is just a
simple check.
Uh, are you guys able to run this?
If you're following along, let me pause
there. Make sure I'm not going too fast.
Uh it so the the issue with a specific
data type like that is none of these are
ants. They're actually all floats. So we
did float. I think that should work. But
yes, that's the idea.
Great. I'm glad to hear that right there
with me. Great. Glad to hear that.
Okay. So, we have our columns picked out
here, which we're going to use later.
Okay.
All right. So, let's go ahead and build
out our steps for each of these types.
So, um for our numerical features, let's
build out our pipeline steps. So what
we're going to do is build out a
numerical pipeline. And it's going to be
a pipeline with a list
of tupils. And the reason these are
tupils is because every tupil has a
name. So here this is a name that we can
it can be whatever we want it to be. So
we're calling it imputer. We could call
it anything we want. We could call it
fill in the blanks. We could call it
null filling. Call it whatever you want.
We're calling it imputer because it's
that's a pretty um easy name for it. An
accurate name to what it's doing. Um but
the important thing is after the name
you give it, you put in the scikitlearn
object that you are going to use to
operate on your data. So in this case,
we're using a simple impery
of median. Now that's a choice. We could
use a strategy of mean, max.
Um, we could provide it a constant
default value. But what this means is we
are going to fill any blanks we find in
those columns with the median value of
that column. That's the strategy for the
imper. So that's pretty cool. This is
kind of an automated way to fill in the
blanks using for any column using its
median,
right? And so we could change that. We
could put mean here or max or min or
whatever. Um
but we are filling in the blank on any
column with its median. And the reason
this works is because we are going to
apply this pipeline only to these
numerical features. So that is fine.
We're we're not going to apply it to the
categorical features. We're going to
apply it to only those numerical. So it
should have a median value, right? So
that that's totally fine. So we're going
to now look at how we're constructing
the steps. We have a list of tupils.
Here's one tupil
which is the imper with a simple imper
of strategy median. And then we can have
as many tupils as we want which
represent processing steps. So every let
me write that down. Every tupil
represents
a pre-processing
step on our data.
Okay, so we have an imputer step named
imputer and the reason it has a name is
just so you can reference it in the
pipeline if you need to. So you so it
has like a a reference name um that you
give it. Um but this is the more
important part is the actual scikitlearn
object that's doing the processing. So
in this case a simple computer but
notice that we have a secondary step
which is our scaling. Now this makes
sense. This is something we should be
doing to our features is we should be
scaling them. So here we we say okay
let's fill in any blanks first.
By the way order
matters.
So, and what I mean by that is the
simple imper
is before the scaler. Now, that's
important because what that means is we
should be filling in any blanks before
we attempt scaling.
So, that order actually matters. We're
going to fill in blanks first in this
list. That's first. We're going to fill
in blanks. Then we are going to scale
right then we scale which makes sense
right so we we fill in blanks first then
we apply the scaler to scale our
features so those are our two steps
so so pretty simple um we are building
out our two steps now this is just one
piece of the puzzle we are going to put
this pipeline together with our one hot
encoding that's going to be coming up
next and build out our final pipeline.
But this is um a a pipeline that has two
steps that will actually be used with a
larger pipeline coming up where we we do
one hot encoding to our categoricals and
then we put a model in there at the end
to train and and use for prediction. So
um pipelines can actually be composed is
is uh something to realize there is that
we can have a pipeline that contains a
few steps. We can have another pipeline
over here that contains a few steps and
we can actually um kind of put them
together into a final pipeline that has
both pipelines uh kind of merged
together. Okay. So we're going to see
that coming up when we construct our
final one. Our final one, as you can
imagine, needs to handle this mapping of
basically saying, let's do one hot
encoding to these guys and then do this
pipeline here to these numerical
features. That's what our final pipeline
needs to handle, and it will. We're
going to build that out.
But let me pause here. Um, were you guys
able to run this? Are you with me on
this this pipeline here?
Does that make sense? Those two steps
one is filling in blanks with a median
whatever column. So where so this is
this is what's so amazing about this is
this is going to automatically search
for nles and if you come across a column
with a null, it's going to use the
median of that column
to fill in the blank, right? To fill in
those nles.
Okay,
great. Glad to hear. Glad to hear.
Okay.
All right. So we are going to now um put
this together with a column transformer
to basically say what steps are going to
be mapped to what columns.
Um so now you can see what we're doing
here is using the column transformer
which is going to be a list of tupils
again. So this is another um list of
tupils.
But the important thing is um
each tupil
has a name
followed by so it has a name uh which
again is is generic. You can say
whatever you want it to be. So here
we're kind of shortening this to
numerical. This is short for
categorical. But the important thing is
it's followed by a pipeline
slashstep
followed by a pipeline slashstep
um followed by a uh followed by a list
of columns that it applies to. So you
can see that pattern here. What we're
saying is we're going to apply that
numerical pipeline we just defined. So
this is saved in a numerical pipeline
object here. We're going to apply that
to those numerical features. So this is
that list
of numerical features here. So that's
how we do the mapping. We have a tupil
here that says okay apply these steps to
these columns.
Those go together in that tupole, right?
Apply these steps to this uh these
columns. And then apply this step. Now
what is the step? This is a one hot
encoder
which is going to uh uh encode um those
features and it's going to uh ignore um
basically nles for now. That's a choice
but it's going to ignore um uh basically
ignore nles and and uh skip over them
for now. We now we know there's no NLES
because we already did an is NA from
before and we know there's not any NLES
in that ocean proximity. So this isn't
going to be an issue. But that's what
that would do.
But we have a one hot encoder here which
we're going to apply to our categorical
features. Now of course that's just the
ocean proximity feature but that but
again you see the pattern in the tupil
is apply this transform which is a one
hot encoding to this column apply these
numerical transforms which is a whole
pipeline. So it's two steps in a
pipeline of um
uh an imputer and a scaler are going to
be applied to this
really nice. So those are going to be
all together in this column transformer.
And that is our way to signal that for
these numerical features, use these
steps. For our categorical features, use
this step. And and you know, if we had
more than one step, we were applying to
categorical. We could build a pipeline
for the categorical and it would and do
the same thing. We have more than one
step here. And so it's good practice
when you have more than one step to just
put that in a pipeline because we have
more than one step. We'll just put that
in this list inside of the pipeline and
we can map that pipeline to those
features. Here we only have one step. So
it's okay to just put that there um and
apply that to the categorical features.
But if we had more than one step um it
would be good practice to put that in a
pipeline
which is what we do here. Right? This
pipeline is being mapped to these
features. This step is being applied to
this feature.
Okay,
how about that? Are you guys able to run
that one? Does that make sense what we
have set up so far? So, we're almost
there. We almost have our final
pipeline. We have our pre-processing
basically done to say our numerical
features should be processed with that
other pipeline and our categorical
features should be one hot encoded.
We're getting close. The only thing
we're really missing here is a model.
The only thing we're really missing is
to have our final model training
pipeline is to actually include a model
which should come at the end.
Right? So it should we should be doing
these steps first
then doing modeling which we know right
we we've done that uh many times. We've
done our pre-processing and then we do
our modeling.
Any questions on that?
Okay.
Fantastic.
All right.
So, if we wanted to uh see if we wanted
to test this so far, um we could. So, we
could run the pre-processing and
actually run a fit transform on our data
and this will um basically apply that
pipeline to the data. Now, this would be
a sanity check. This is a good This is a
good kind of um This is a good sanity
check that our pre-processing
works. So, it's doing what we expected
to do. It's not our final pipeline
because we don't have our model in there
yet, but this is just to ensure that all
of the features are kind of behaving as
we expect. So, we can uh we can do that
and we can take a look at the um
results. This looks pretty good. this
all of our numerical features ended up
scaled
which is pretty good and we have one hot
encoded features for that ocean
proximity over here.
Okay, so this looks pretty this looks
reasonable of those steps being applied
to the right columns. But this is a good
kind of sanity check to just run our fit
transform on our data to ensure those
steps are actually happening and they
are. You can see here the result of the
scaling and the uh the one hot encoding.
So that that all looks pretty
reasonable,
right? And uh what we should also do is
make sure there are no nulls in this
which there shouldn't be because we did
the imputer. So we should be doing uh is
na dot
sum
And there is no NLES anymore. So that
looks pretty good, right? Those got
filled in uh by doing our steps. Our
pipeline steps executed really nicely on
our training data. Um and and we were
off and running. And there's nothing
unique about the training data. We could
do this to our test data as well
and verify that those steps are running
and they would, right? There's nothing
really that special about running it on
the training data. Um, it should also
work on the test features as well, and
it does. You can check that for
yourself.
Okay.
All right. So, that's pretty cool. We
can uh verify all that's working.
Any questions on that?
We're almost there with our full
pipeline. This this is this is not the
full pipeline, but this is something
that will run during our full pipeline.
Of course, our features are going to be
transformed according to those steps and
then it will be uh put into our model to
either predict or train with. Um
so let's do that. Let's actually build
out our final uh model here. So it's
actually going to be really easy to do.
All we need to do is um put in our
model. So here we're going to import the
ridge model here. Now we could use any
we could use linear regression, we could
use lasso, we could use elastic net. Um
we're just going to use ridge um uh um
just to test it out. And um we are going
to uh now put in a final pipeline. So
we're going to use our pipeline. And so
we're going to create a new one here.
and map our pre-processing to our
pre-processing that we've already built.
So, this is a column transformer that
already has all of our steps. And then
notice what comes after it is just the
model. Now, that's pretty pretty basic,
but it makes sense that it should come
after that model. Um, and of course,
this is a generic name. We could we can
name it whatever we want to. um model
ridge is pretty reasonable um to because
it is a ridge uh regression but uh of
course we could we could change that.
Okay, so that builds out our uh final um
pipeline. So now we have a pipeline and
what's great about that is this signals
that all of these steps should be
completed prior to doing anything with
this model. So all of those processing
steps are going to run and then we're
going to do ffit orpredict and that so
that's really great. It ensures that all
those steps are running together every
single time we call predict with this
with this model. So we're just going to
use the pipeline in place of the model
to ensure that all of those steps are
running together. And this is our this
is kind of our final pipeline that we
would use uh with like something like
ffit or predict.
So let me make that uh a note of that.
Now we can use this final pipeline just
like a regular model i.e. pipeline.fit
or pipeline.predict.
We could use it in ei in either fashion
uh to to train the pipeline would be
this guy and then use the pipeline to
predict would be this. And what we
should realize is under the hood these
steps are running first and then we
train it or these steps run first then
we use it for prediction.
Okay.
Questions on that? Does that make sense?
On this final pipeline here, it's just
now it it's really cool because we have
a pipeline
made up of a of a pipeline really,
right? A pipeline made up of a pipeline.
But that's scikitlearn allows you to do
that to compose pipelines in this way.
That's that's pretty uh pretty uh normal
there.
Okay,
what I want to show you is we can
actually use this pipeline in a grid
search. So that's pretty amazing. We can
use this pipeline in any way we can use
a mo like a regular model. It's just
that now our pre-processing steps have
kind of been packaged together with our
model to ensure that they always run
anytime we do any processing with this
model. Um so for instance we can do a
grid search just like we did with a
regular with with just a model right
with just this. Um we can do the same
thing with the whole pipeline. Um, so
the only catch is that you want to make
sure in your grid you name things in the
appropriate way inside of your your uh
keys in your dictionary. So uh for
instance um inside of the grid uh we're
going to set up the alpha that would be
used with this ridge regression by
referencing its name. So this is model
ridge is this is the name of the model
inside of the pipeline. So you want to
make sure that goes first.
And then what scikitlearn does is it
recognizes parameters that belong with
this model by using a double underscore.
So the so you have underscore underscore
alpha um here. So the double
uh underscore
signals a parameter
belonging to model ridge. in the in the
pipeline.
Okay, so we have a model ridge is just a
reference to the model in our pipeline.
That's the one we're going to test out
these parameters with. And
underscore_pha is just a way to say this
alpha belongs to this model. Okay, it
belongs so it's going to be used with
that model in our pipeline. Um otherwise
it's going to work exactly the same way.
It's just we need to line up this naming
convention of of scikitlearn.
You just have to reference this to
whatever name you provided here and then
underscore parameter. So L1 ratio alpha
whatever right would go there.
Okay. So there is a range from 0.1 to
two uh step size of 0.1
um and then we do our grid search CV. So
this is exactly the same setup as we had
before. It's just that our model is now
the pipeline. So our pipeline is going
in there. Um we have our grid going in
there. We have our scoring is the same,
you know, negative absolute error. Um
we're using five-fold cross validation
and we're parallelizing that search. Um,
so we're going to search through these
alphas and uh basically fit this to our
um data and find the best um find the
best alpha.
So it's going to try out all those
combinations and try to come up with the
best alpha.
So looks like the best alpha was 0.1 for
the ridge.
Okay, is the best. So then um if we
wanted to we could uh then predict using
the model um which would be doing
something like this. Um and we could
also go back and do something like so we
could
now use um this param. So we could do
model
um equals ridge
and then we could put in our alpha
um alpha is our results our best
parameters and then we get that model
ridge alpha and then we just rebuild our
our pipeline
equals um pipeline and then we uh put in
this new model here. So we could do
this. This would be going back and just
um putting in our best alpha here for
this model and then uh ensuring that's
part of our our pipeline. So we're just
overwriting that pipeline with the best
model there
to get the best model in our pipeline.
Okay,
so that's all this is doing is just
initializing a new um let me actually I
can put this code in here.
This is actually just getting this is
just getting a model with the best alpha
and then reinserting that into our our
uh we're just overwriting our final
pipeline there with the best model that
we have.
So pretty cool that pipeline can be used
basically exactly like a model, right?
It's it's going right here in the grid
search and being used uh entirely like a
basic model. So we do ffit
um and that allows us to use it. We
could dopredict. We could even do
pipeline.predict once we we could go
back and do final pipeline.fit
um with this and then final
pipeline.predict with this and evaluate
Okay,
so pretty cool that pipeline can be used
uh basically exactly like how a model
would be any way we' use a model.fit
model.predict, we can use a pipeline.
So grid search is for instance something
that can use a model in there. Um but
instead of just a model, we're ensuring
we have our pre-processing steps kind of
bundled with that model in this
pipeline.
Any
questions on
uh this example so far?
Were you guys able to run it up to here?
Were you able to run the grid search?
Okay, great.
Okay.
Okay. So, this is this is uh just
showing you what's actually happening
underneath the hood is uh you know,
we're doing some scaling. We're doing
some one hot encoding um
and we're doing some uh we're doing a
model here. And that's all part of our
pipeline. Um, and then we can use the
pipeline however we want. So for
example, I know it's not here, but for
an example, we could use um once we do
once we have this final pipeline um we
can can use the final um
pipeline to predict. So we can do um
predictions
equals final
pipeline.predict
and then we can pass in our test data.
Now what happens on this is once we have
ran our our pipeline.fit we have a
trained pipeline and then when we run
this final pipeline.predict uh this data
is going to be transformed.
It's going to go through those
transformation steps and then we would
apply our model to it at the end uh to
to make those predictions and then we
can evaluate those predictions which is
what we're doing kind of here.
Right?
Okay.
All right. So in conclusion uh we have
gone through a lot of stuff here. Um,
we've gone through regression, we've
done the regularization on regression.
So hopefully we have a good foundation
on regression. Um, what we're going to
do in a little bit is actually do some
additional practice with regression on a
new problem. We're going to do a
capstone problem and do some additional
regression work with that. Um, so we'll
do that next. Um but the other thing we
learned is how to evaluate the
regression using things like mean
squared error, RMSSE which is square
root of that. Um which is which is
really cool. So we have a sense of that
error which is our distance from our
prediction to the actual value. That's
always what these uh that's always what
these things are doing like this, right?
This mean absolute error metric from
scikitlearn is computing the average
distance from these predictions to these
test labels that we have, right? Those
actual values. Um, and that gives us a
sense of on average how far away are our
predictions
um to see how good of a model that we
have, right? And we should be evaluating
that error generally
um against the scale of our targets to
see, you know,
uh how far off we typically are.
Okay. Any questions at all on this
lesson on regression? Uh anything we
covered up to this point? We're going to
do some more practice with the next
we'll do the capstone. So we get so we
just do some more regression problems.
Yeah, it's a that's another bad score.
It's a little bit hard to interpret this
though because it's m ae. Um, so one
thing we could do is is compute mean
squared error and then take the square
root of it to get the RMSSE which is a
much better uh evaluation metric in
terms of our target. Um so we could
actually run that. Uh if we go back here
and um we could generate for instance we
could generate the MSE which is the mean
squared
error
and it's it's the same exact function uh
of using our predictions.
Um
and then we could just print that out.
Mean squared error.
So we have mean squared error and then
what we can do is let's take the um MP.
square root of that.
So that way we can generate the RMSSE.
So yeah that I mean that's pretty bad.
That's uh pretty bad. Uh now let's let's
go back and look at our
uh data though. So let's take a look at
the average for our y. Um remember one
thing we should be doing is taking a
look at um what our uh let's take a look
at y test mean
to get an average value. So the average
value is in the 200,000s. So
this isn't this isn't awful. This is
70,000. It's still a decent amount of
error. It's not as bad as the models we
have before though, right? This is an
average median price of the house is in
the 206,000 range and our error is off
by like 70,000,
right?
So, it's not good. Um, but it's not
hor like as bad as the it's not as
horrible as we've seen so far. Right.
This is a little bit better of a model.
A little bit better. closer to zero
would be better, right? Um but the
smaller the better. But uh remember this
is the um these even the mean absolute
error is is technically in similar units
as the as the uh um
as the target. So 50,000 60,000 here
70,000 it's still a decent amount of
error in terms of 200,000.
Uh so far we only come up with models
and test their accuracy with available
data. We haven't used a model to make
completely new predictions on No, we
haven't done that. Uh except we know how
to do that. Um it would so to make
predictions on new data would be exactly
how we're making them on our available
data because we actually do that all the
time. If we go back down to our model
building,
um it's it looks just like this, right?
where we take so for instance we do
predictions all the time on test data
that was never involved in the training.
So it's it's as if this data mimics new
data that we've never seen before. So if
we had new raw data it would just it
would be the same exact process. the new
now with our pipeline it makes it a
little bit easier because with the
pipeline
um the raw data will go through those
transformations which it should right
the raw data should because if it's
missing data it needs to be filled in if
it has categoricals it needs to be one
hot encoded so that's the purpose of the
pipeline actually is to make sure that
if we're dealing with raw data um those
steps can happen on the data before it
goes into the model. Right?
So we so
that's kind of the purpose of the
pipeline
is to ensure that we run those steps
ahead of using it using a model with it.
But but ultimately that's how it uh any
scikitlearn model is going to be doing
the predict even if it's a pipeline
right it's going to be uh we just go
back down here it's going to be um
predict it's always going to be that on
new data
>> hello everyone in this session we will
cover all the important supervised and
unsupervised learning algorithms that
are widely used with hands-on
demonstrations in Python our instructors
with rich experience in machine learning
will take us through this course but
before Before we begin, make sure to
subscribe to the SimplyLearn channel and
hit the bell icon to never miss an
update.
So, we will start by understanding the
basics of machine learning from a short
animated video followed by the
difference between supervised and
unsupervised learning. We will then jump
into learning the various algorithms
from scratch. So, we will understand
linear regression, logistic regression,
decision tree, and random forest. We'll
then look at support vector machines and
K nearest neighbor algorithm with a
hands-on demonstration in Python.
Finally, we'll get an idea about
unsupervised learning algorithms such as
K means clustering and principal
component analysis. We will conclude
this session with regularization in
machine learning. So let's get started.
>> We know humans learn from their past
experiences and machines follow
instructions given by humans.
But what if humans can train the
machines to learn from their past data
and do what humans can do and much
faster? Well, that's called machine
learning. But it's a lot more than just
learning. It's also about understanding
and reasoning. So today we will learn
about the basics of machine learning. So
that's Paul. He loves listening to new
songs.
He either likes them or dislikes them.
Paul decides this on the basis of the
song's tempo, genre, intensity, and the
gender of voice. For simplicity, let's
just use tempo and intensity for now.
So, here tempo is on the x-axis, ranging
from relaxed to fast, whereas intensity
is on the y-axis, ranging from light to
soaring. We see that Paul likes the song
with fast tempo and soaring intensity
while he dislikes the song with relaxed
tempo and light intensity. So now we
know Paul's choices. Let's say Paul
listens to a new song. Let's name it as
song A. Song A has fast tempo and a
soaring intensity. So it lies somewhere
here. Looking at the data, can you guess
whether Paul will like the song or not?
Correct. So Paul likes this song. By
looking at Paul's past choices, we were
able to classify the unknown song very
easily, right? Let's say now Paul
listens to a new song. Let's label it as
song B. So song B lies somewhere here
with medium tempo and medium intensity.
Neither relaxed nor fast, neither light
nor soaring. Now, can you guess whether
Paul likes it or not? Not able to guess
whether Paul will like it or dislike it.
Are the choices unclear? Correct. We
could easily classify song A. But when
the choice became complicated as in the
case of song B. Yes. And that's where
machine learning comes in. Let's see
how. In the same example for song B, if
we draw a circle around the song B, we
see that there are four votes for like
whereas one vote for dislike. If we go
for the majority votes, we can say that
Paul will definitely like the song.
That's all. This was a basic machine
learning algorithm also. It's called K
nearest neighbors. So this is just a
small example in one of the many machine
learning algorithms quite easy right
believe me it is but what happens when
the choices become complicated as in the
case of song B that's when machine
learning comes in it learns the data
builds the prediction model and when the
new data point comes in it can easily
predict for it more the data better the
model higher will be the accuracy there
are many ways in which the machine
learns it could be either supervised
learning unsupervised learning or
reinforcement learning. Let's first
quickly understand supervised learning.
Suppose your friend gives you 1 million
coins of three different currencies. Say
1 rupee, 1 and 1 dirham. Each coin has
different weights. For example, a coin
of 1 rupee weighs 3 g. 1 euro weighs 7 g
and 1 dirham weighs 4 g. Your model will
predict the currency of the coin. Here
your weight becomes the feature of coins
while currency becomes their label. When
you feed this data to the machine
learning model, it learns which feature
is associated with which label. For
example, it will learn that if a coin is
of 3 g, it will be a 1 rupee coin. Let's
give a new coin to the machine. On the
basis of the weight of the new coin,
your model will predict the currency.
Hence, supervised learning uses labeled
data to train the model. Here, the
machine knew the features of the object
and also the labels associated with
those features. On this note, let's move
to unsupervised learning and see the
difference. Suppose you have cricket
data set of various players with their
respective scores and wickets taken.
When we feed this data set to the
machine, the machine identifies the
pattern of player performance. So, it
plots this data with the respective
wickets on the x-axis while runs on the
y-axis. While looking at the data,
you'll clearly see that there are two
clusters. The one cluster are the
players who scored high runs and took
less wickets while the other cluster is
of the players who scored less runs but
took many wickets. So here we interpret
these two clusters as batsmen and
bowlers. The important point to note
here is that there were no labels of
batsmen and bowlers. Hence the learning
with unlabeled data is unsupervised
learning. So we saw supervised learning
where the data was labeled and the
unsupervised learning where the data was
unlabeled. And then there is
reinforcement learning which is a
reward-based learning or we can say that
it works on the principle of feedback.
Here let's say you provide the system
with an image of a dog and ask it to
identify it. The system identifies it as
a cat. So you give a negative feedback
to the machine saying that it's a dog's
image. The machine will learn from the
feedback and finally if it comes across
any other image of a dog, it'll be able
to classify it correctly. That is
reinforcement learning. To generalize
machine learning model, let's see a
flowchart. Input is given to a machine
learning model which then gives the
output according to the algorithm
applied. If it's right, we take the
output as a final result. Else we
provide feedback to the training model
and ask it to predict until it learns. I
hope you've understood supervised and
unsupervised learning. So let's have a
quick quiz. You have to determine
whether the given scenarios uses
supervised or unsupervised learning.
Simple, right? Scenario one. Facebook
recognizes your friend in a picture from
an album of tagged photographs.
Scenario two, Netflix recommends new
movies based on someone's past movie
choices.
Scenario three, analyzing bank data for
suspicious transactions and flagging the
fraud transactions. Think wisely and
comment below your answers. Moving on,
don't you sometimes wonder how is
machine learning possible in today's
era? Well, that's because today we have
humongous data available. Everybody's
online either making a transaction or
just surfing the internet and that's
generating a huge amount of data every
minute. And that data my friend is the
key to analysis. Also, the memory
handling capabilities of computers have
largely increased which helps them to
process such huge amount of data at hand
without any delay. And yes, computers
now have great computational powers. So
there are a lot of applications of
machine learning out there. To name a
few, machine learning is used in
healthcare where diagnostics are
predicted for doctor's review. The
sentiment analysis that the tech giants
are doing on social media is another
interesting application of machine
learning. Fraud detection in the finance
sector and also to predict customer
churn in the e-commerce sector. While
booking a cab, you must have encountered
search pricing often where it says the
fair of your trip has been updated.
Continue booking. Yes, please. I'm
getting late for office. Well, that's an
interesting machine learning model which
is used by global taxi giant Uber and
others where they have differential
pricing in real time based on demand,
the number of cars available, bad
weather, rush hour, etc. So they use the
search pricing model to ensure that
those who need a cab can get one. Also,
it uses predictive modeling to predict
where the demand will be high with a
goal that drivers can take care of the
demand and search pricing can be
minimized. Great. Hey Siri, can you
remind me to book a cab at 6 p.m. today?
>> Okay, I'll remind you.
>> Thanks.
>> No problem.
>> Comment below some interesting everyday
examples around you where machines are
learning and doing amazing jobs. Hi
guys, this is Acha from SimplyLearn and
we're going to talk about the two types
of machine learning supervised and
unsupervised learning their types and
applications. But before we talk about
them, let's quickly understand what is
machine learning. These days
applications use artificial intelligence
in machine learning to optimize speech
recognition. I usually ask Siri things I
want to know like hey Siri how far is
the nearest subway? So whenever we ask
something to Siri, a powerful speech
recognition kicks off and converts the
audio into its corresponding textual
form which is then sent to the Apple
servers for further processing. Then
neural language processing algorithms
are run to understand the user's intent
and then finally Siri tells you the
answer. Well, this is what machine
learning is all about. making the
machines learn and act like humans by
feeding them with data and information
without being explicitly programmed. As
we saw in the previous example, when the
data comes in, machines immediately
starts analyzing the data and eventually
gets trained on it and learns it. Now
when a new data point comes in, machine
accurately makes prediction and
decisions based on the past data. Now
that you know what is machine learning,
let's talk about supervised and
unsupervised learning. Supervised
learning as the name suggests works
under supervision that is it's a
learning in which machine is trained
with data which is welllabeled and then
predicts with the help of the label data
set. But what is a label data set? Data
for which you already know the target
answer is called a label data. Like I
show you an image and tell you that it's
a dog then it's a label data. While if I
show you an image without telling you
what exactly it is, then it's an
unlabelled data. Now let's say we have
images which are labeled as spoon or
knife. We then feed it to the machine
which analyzes and learns the
association of these images with its
labels based on its features such as
shape, size, sharpness etc. Now when a
new image is fed to the machine without
any label with the help of the past data
the machine is able to predict
accurately and tell that it's a spoon.
Hence in supervised machine learning the
algorithm teaches the model to learn
from the labeled example that we
provide. So supervised learning can be
further divided into classification and
regression. It is a classification
problem when the output variable is
categoric such as red or blue, disease
or no disease, male or female. Whereas
it's a regression problem when the
output variable is a real or continuous
value. For example, salary based on work
experience, weight based on height. So
it creates a predictive model showing
trends in data. So now if I say will I
get a salary raise or not, that's
classification. But if I say how much
salary raise will I get, that is
regression. Now let's understand
classification with the help of an
example. In order to predict whether an
email is a spam or not, first we need to
teach our machine what a spam mail looks
like. This is done based on a lot of
spam filters like firstly reviewing the
content of the email then review the
email header and search if it contains
any falsified information. This is done
based on some keywords like free lottery
prize claim etc. Then general blacklist
filters to stop emails that come from
already blacklisted known spammers and
etc. So all these filters scores the
email which is known as spam score. The
lower the total spam score of the email,
it is more likely that the email will
land in subscribers inboxes. So based on
the content labels and spam score of the
new incoming mail, the algorithm decides
whether it should land in inbox or the
spam folder. Now let's quickly
understand regression. So let's say we
have two variables that is temperature
and humidity where temperature is the
independent variable and humidity is the
dependent variable such that as the
temperature increases humidity decreases
hence they're correlated. When we feed
this data to a regression model it will
understand the relationship between
these two variables and how one variable
depends on the other. After the machine
is trained it can easily predict the
humidity based on the given temperature.
Well, that was about regression. Now,
let's see some real life applications
where supervised learning is used. So,
supervised learning is used in risk
assessment to assess risk in financial
services or an insurance domain to
minimize the risk portfolio of the
companies. Image classification.
Facebook recognizes your friend in a
picture from an album of tact photos.
So, image classification is one of the
key use cases of demonstrating
supervised machine learning algorithms.
Obviously a lot more goes into all these
like convolutional neural networks etc.
Fraud detection whether the transactions
made by the user are authentic or not
and visual recognition the ability of a
machine learning model to identify
objects places people and actions in
images. Now let's quickly talk about
unsupervised learning. In unsupervised
learning, there is no supervision that
is no training will be given to the
machine allowing it to act on the data
which is not labeled. Hence, machine
tries to identify patterns and gives the
response. Let's take a similar example
as before. But this time we do not tell
the machine whether it's a spoon or a
knife. The machine identifies patterns
from the given set and groups them based
on their patterns, similarities, etc.
Again unsupervised learning can be
further grouped into clustering and
association. Clustering is basically
where the machine forms groups based on
the behavior of the data. Secondly,
association. It is a rule-based machine
learning to discover interesting
relation between variables in large data
sets. For example, which customer made
similar product purchases is clustering.
Whereas association is which products
were purchased together. Now let's
understand clustering with the help of
an example. To reduce their churn rate,
a telecom company studies the behavior
of the customers based on average call
duration and internet usage and observes
that while some customers call duration
is quite high, others have heavy
internet usage. The customers are
grouped based on their observed behavior
and a strategy is adopted to minimize
churn rate and maximize profit via
suitable promotions and campaigns. As
you can see in the chart on the right
hand side, customers in group A uses
more data and also have high
qualuration. Group B customers are heavy
internet users while group C customers
have high qualation. So group B will be
given more data benefit plans while
group C will be given cheaper call rates
to buy their loyalty. So this was the
example of clustering. Now let's
understand association with another
example. Let's say customer one goes to
a supermarket and buys these products.
say bread, milk, fruits, wheat. Then
customer two goes and buys bread, milk,
rice and butter. Now when customer three
goes and buys bread, it is highly likely
that he will also buy milk. Hence
relationship is established based on
customer behavior and recommendations
are made. Now let's look at some real
life applications of unsupervised
learning. Market basket analysis is a
machine learning model based on the
algorithm that if you buy a certain
group of items, you are less or more
likely to buy another group of items.
Semantic clustering. Semantically
similar words share similar context.
People post their queries on websites in
their own ways. Semantic clustering
groups all responses in a cluster with
same meaning to ensure that the customer
finds the information they want quickly
and easily. It plays an important role
in information retrieval. Good browsing
experience and comprehension. Delivery
store optimization. Machine learning
models are used to predict the demand
and keep up with the supply also to open
stores where demand is more and
optimizing routes for more efficient
deliveries according to past data and
behavior. We can also use unsupervised
machine learning models to identify
accidentprone areas based on the
intensity of those accidents and the
area in order to introduce safety
measures. By now I hope you've
understood supervised and unsupervised
learning. For a quick recap, let's see a
few differences between the two. The
most fundamental difference is that
supervised learning uses known and
labelled data and unsupervised learning
uses unlabelled data as their input.
Secondly, supervised learning follows a
feedback mechanism while unsupervised
learning does not. Also, the most
commonly used algorithms in supervised
learning are decision tree, logistic
regression, support vector machine, etc.
And in unsupervised learning there K
means clustering, hierarchical
clustering, a priori algorithm and many
more.
>> Welcome to linear regression. My name is
Richard Kersner. I'm with SimplyLearn.
Let's look at an example of a common use
for linear regression, profit estimation
of a company. If I was going to invest
in a company, I would like to know how
much money I could expect to make. So
we'll take a look at a venture
capitalist firm and try to understand
which companies they should invest in.
So we'll take the idea that we need to
decide the companies to invest in. We
need to predict the profit the company
makes and we're going to do it based on
the company's expenses and even just a
specific expense. In this case we have
our company, we have the different
expenses. So we have our R&D which is
your research and development. We have
our marketing. Uh we might have the
location. We might have what kind of
administration it's going through. Based
on all this different information, we
would like to calculate the profit. Now,
in actuality, there's usually about 23
to 27 different markers that they look
at if they're a heavy duty investor.
We're only going to take a look at one
basic one. We're going to come in and
for simplicity, let's consider a single
variable, R&D, and find out which
companies to invest in based on that.
So, we take our R&D and we're plotting
the profit based on the R&D expenditure,
how much money they put into the
research and development. And then we
look at the profit that goes with that.
We can predict a line to estimate the
profit. So we can draw a line right
through the data. And when you look at
that, you can see how much they invest
in the R&D is a good marker as to how
much profit they're going to have. We
can also note that companies spending
more on R&D make good profit. So let's
invest in the ones that spend a higher
rate in their R&D. What's in it for you?
First, we'll have an introduction to
machine learning followed by machine
learning algorithms. These will be
specific to linear regression and where
it fits into the larger model. Then
we'll take a look at applications of
linear regression, understanding linear
regression, and multiple linear
regression. Finally, we'll roll up our
sleeves and do a little programming in
use case profit estimation of companies.
Let's go ahead and jump in. Let's start
with our introduction to machine
learning along with some machine
learning algorithms and where that fits
in with linear regression. Let's look at
another example of machine learning.
Based on the amount of rainfall, how
much would be the crop yield? So we here
we have our crops, we have our rainfall,
and we want to know how much we're going
to get from our crops this year. So
we're going to introduce two variables,
independent and dependent. The
independent variable is a variable whose
value does not change by the effect of
other variables and is used to
manipulate the dependent variable. It is
often denoted as X. In our example,
rainfall is the independent variable.
This is a wonderful example because you
can easily see that we can't control the
rain, but the rain does control the
crop. So we talk about the independent
variable controlling the dependent
variable. Let's define dependent
variable as a variable whose value
change when there is any manipulation in
the values of the independent variables.
It is often denoted as y. And you can
see here our crop yield is dependent
variable and it is dependent on the
amount of rainfall received. Now that
we've taken a look at a real life
example, let's go a little bit into the
theory and some definitions on machine
learning and see how that fits together
with linear regression. numerical and
categorical values. Let's take our data
coming in and this is kind of random
data from any kind of project. We want
to divide it up into numerical and
categorical. So numerical is numbers,
age, salary, height, where categorical
would be a description, the color, a
dog's breed, gender. Categorical is
limited to very specific items where
numerical is a range of information. Now
that you've seen the difference between
numerical and categorical data, let's
take a look at some different machine
learning definitions. When we look at
our different machine learning
algorithms, we can divide them into
three areas. Supervised, unsupervised,
reinforcement. We're only going to look
at supervised today. Unsupervised means
we don't have the answers and we're just
grouping things. Reinforcement is where
we give positive and negative feedback
to our algorithm to program it. and it
doesn't have the information till after
the fact. But today we're just looking
at supervised because that's where
linear regression fits in. In supervised
data, we have our data already there and
our answers for a group. And then we use
that to program our model and come up
with an answer. The two most common uses
for that is through the regression and
classification. Now, we're doing linear
regression. So, we're just going to
focus on the regression side. And in the
regression we have simple linear
regression, we have multiple linear
regression and we have polomial linear
regression. Now on these three simple
linear regression is the examples we've
looked at so far where we have a lot of
data and we draw a straight line through
it. Multiple linear regression means we
have multiple variables. Remember where
we had the rainfall and the crops. We
might add additional variables in there
like how much food do we give our crops?
When do we harvest them? Those would be
additional information add into our
model and that's why it' be multiple
linear regression. And finally we have
polomial linear regression that is
instead of drawing a line we can draw a
curved line through it. Now that you see
where regression model fits into the
machine learning algorithms and we're
specifically looking at linear
regression. Let's go ahead and take a
look at applications for linear
regression. Let's look at a few
applications of linear regression.
Economic growth used to determine the
economic growth of a country or a state
in the coming quarter can also be used
to predict the GDP of a country. Product
price can be used to predict what would
be the price of a product in the future.
We can guess whether it's going to go up
or down or should I buy today. Housing
sales to estimate the number of houses a
builder would sell and what price in the
coming months. Score predictions.
Cricket fever to predict the number of
runs a player would score in the coming
matches based on the previous
performance. I'm sure you can figure out
other applications you could use linear
regression for. So let's jump in and
let's understand linear regression and
dig into the theory. Understanding
linear regression. Linear regression is
the statistical model used to predict
the relationship between independent and
dependent variables by examining two
factors. The first important one is
which variables in particular are
significant predictors of the outcome
variable. And the second one that we
need to look at closely is how
significant is the regression line to
make predictions with the highest
possible accuracy. If it's inaccurate,
we can't use it. So, it's very important
we find out the most accurate line we
can get. Since linear regression is
based on drawing a line through data,
we're going to jump back and take a look
at some uklitian geometry. The simplest
form of a simple linear regression
equation with one dependent and one
independent variable is represented by y
= m * x + c. And if you look at our
model here, we plotted two points on
here. Uh x1 and y1, x2 and y2. y being
the dependent variable, remember that
from before. And x being the independent
variable. So y depends on whatever x is.
M in this case is the slope of the line
where M equals the difference in the Y2
- Y1 and X2 - X1. And finally we have C
which is the coefficient of the line or
where happens to cross the zero axis.
Let's go back and look at an example we
used earlier of linear regression. We're
going to go back to plotting the amount
of crop yield based on the amount of
rainfall. And here we have our rainfall.
Remember, we cannot change rainfall. And
we have our crop yield, which is
dependent on the rainfall. So, we have
our independent and our dependent
variables. We're going to take this and
draw a line through it as best we can
through the middle of the data. And then
we look at that. We put the red point on
the y ais is the amount of crop yield
you can expect for the amount of
rainfall represented by the green dot.
So, if we have an idea what the rainfall
is for this year and what's going on,
then we can guess how good our crops are
going to be. and we've created a nice
line right through the middle to give us
a nice mathematical formula. Let's take
a look and see what the math looks like
behind this. Let's look at the intuition
behind the regression line. Now, before
we dive into the math and the formulas
that go behind this and what's going on
behind the scenes, I want you to note
that when we get into the case study and
we actually apply some Python script
that this math that you're going to see
here is already done automatically for
you. You don't have to have it
memorized. It is, however, good to have
an idea what's going on so if people
reference the different terms, you'll
know what they're talking about. Let's
consider a sample data set with five
rows and find out how to draw the
regression line. We're only going to do
five rows because if we did like the
rainfall with hundreds of points of
data, that would be very hard to see
what's going on with the mathematics.
So, we'll go ahead and create our own
two sets of data. And we have our
independent variable x and our dependent
variable y. And when x was 1, we got y =
2. When x was uh 2, y was 4. And so on
and so on. If we go ahead and plot this
data on a graph, we can see how it forms
a nice line through the middle. You can
see where it's kind of grouped going
upwards to the right. The next thing we
want to know is what the means is of
each of the data coming in, the x and
the y. The means doesn't mean anything
other than the average. So, we add up
all the numbers and divide by the total.
So, 1 + 2 + 3 + 4 + 5 over 5 equals 3.
And the same for y, we get four. If we
go ahead and plot the means on the
graph, we'll see we get 3a 4, which
draws a nice line down the middle, a
good estimate. Here, we're going to dig
deeper into the math behind the
regression line. Now, remember before I
said you don't have to have all these
formulas memorized or fully understand
them, even though we're going to go into
a little more detail of how it works.
And if you're not a math wiz and you
don't know if you've never seen the
sigma character before, which looks a
little bit like an e that's opened up,
that just means summation. That's all
that is. So, when you see the sigma
character, it just means we're adding
everything in that row. And for
computers, this is great because as a
programmer, you can easily iterate
through each of the XY points and create
all the information you need. So in the
top half, you can see where we've broken
that down into pieces. And as it goes
through the first two points, it
computes the squared value of X, the
squared value of Y, and X * Y. And then
it takes all of X and adds them up. All
of Y adds them up. All of X squar adds
them up. And so on and so on. And you
can see we have the sum of equal to 15.
The sum is equal to 20. all the way up
to x * y where the sum equals 66. This
all comes from our formula for
calculating a straight line where y
equals the slope* x plus the coefficient
c. So we go down below and we're going
to compute more like the averages of
these and we're going to explain exactly
what that is in just a minute and where
that information comes from. It's called
the square means error, but we'll go
into that in detail in a few minutes.
All you need to do is look at the
formula and see how we've gone about
computing it line by line instead of
trying to have a huge set of numbers
pushed into it. And down here you'll see
where the slope m equals and then the
top part if you read through the
brackets you have the number of data
points times the sum of x * y which we
computed one line at a time there. And
that's just the 66. and take all that
and you subtract it from the sum of x
times the sum of y and those have both
been computed. So you have 15 * 20. And
on the bottom we have the number of
lines times the sum of x^2 easily
computed as 86 for the sum minus I'll
take all that and subtract the sum of
x^2. And we end up as we come across
with our formula. You can plug in all
those numbers which is very easy to do
on the computer. You don't have to do
the math on a piece of paper or
calculator. And you'll get a slope of 6
and you'll get your C coefficient. If
you continue to follow through that
formula, you'll see it comes out as
equal to 2.2. Continuing deeper into
what's going behind the scenes, let's
find out the predicted values of y for
corresponding values of x using the
linear equation where m=6 and c = 2.2.
We're going to take these values and
we're going to go ahead and plot them.
We're going to predict them. So y =6 * x
= 1 + 2.2 = 2.8 so on and so on. And
here the blue points represent the
actual y values and the brown points
represent the predicted yv values based
on the model we created. The distance
between the actual and predicted values
is known as residuals or errors. The
best fit line should have the least sum
of squares of these errors also known as
equare. If we put these into a nice
chart where you can see X and you can
see Y what the actual values were and
you can see Y predicted you can easily
see where we take Y minus Y predicted
and we get an answer. What is the
difference between those two and if we
square that Y - Y prediction squared we
can then sum those squared values.
That's where we get the 64 plus the.36 +
1 all the way down until we have a
summation equals 2.4. So the sum of
squared errors for this regression line
is 2.4. We check this error for each
line and conclude the best fit line
having the least e value. In a nice
graphical representation, we can see
here where we keep moving this line
through the data points to make sure the
best fit line has the least squared
distance between the data points and the
regression line. Now we only looked at
the most commonly used formula for
minimizing the distance. There are lots
of ways to minimize the distance between
the line and the data points like sum of
squared errors, sum of absolute errors,
root mean square error, etc. What you
want to take away from this is whatever
formula is being used, you can easily
using a computer programming and
iterating through the data calculate the
different parts of it. That way, these
complicated formulas you see with the
different summations and absolute values
are easily computed one piece at a time.
Up until this point, we've only been
looking at two values, X and Y. Well, in
the real world, it's very rare that you
only have two values when you're
figuring out a solution. So, let's move
on to the next topic, multiple linear
regression. Let's take a brief look at
what happens when you have multiple
inputs. So, in multiple linear
regression, we have uh well, we'll start
with the simple linear regression where
we had y = m + x + c and we're trying to
find the value of y. Now with multiple
linear regression we have multiple
variables coming in. So instead of
having just x we have x1 x2 x3 and
instead of having just one slope each
variable has its own slope attached to
it. As you can see here we have m1 m2 m3
and we still just have the single
coefficient. So when you're dealing with
multiple linear regression you basically
take your single linear regression and
you spread it out. So you have y = m1 *
x1 + m2 * x2 so on all the way to m x to
the nth and then you add your
coefficient on there. Implementation of
linear regression. Now we get into my
favorite part. Let's understand how
multiple linear regression works by
implementing it in Python. If you
remember before we were looking at a
company and just based on its R&D trying
to figure out its profit. We're going to
start looking at the expenditure of the
company. We're going to go back to that.
We're going to predict his profit, but
instead of predicting it just on the
R&D, we're going to look at other
factors like administration costs,
marketing costs, and so on. And from
there, we're going to see if we can
figure out what the profit of that
company is going to be. To start our
coding, we're going to begin by
importing some basic libraries. And
we're going to be looking through the
data before we do any kind of linear
regression. We're going to take a look
at the data to see what we're playing
with. Then we'll go ahead and format the
data to the format we need to be able to
run it in the linear regression model.
And then from there we'll go ahead and
solve it and just see how valid our
solution is. So let's start with
importing the basic libraries. Now I'm
going to be doing this in Anaconda
Jupyter notebook, a very popular IDE. I
enjoy it because it's such a visual to
look at and so easy to use. Um just any
ID for Python will work just fine for
this. So break out your favorite Python
IDE. So, here we are in our Jupyter
notebook. Let me go ahead and paste our
first piece of code in there. And let's
walk through what libraries we're
importing. First, we're going to import
numpy as np. And then I want you to skip
one line and look at import pandas as
pd. These are very common tools that you
need with most of your linear
regression. The numpy, which stands for
number python, is usually denoted as np,
and you have to almost have that for
your sklearn toolbox. So, you always
import that right off the beginning.
pandas. Although you don't have to have
it for your sklearn libraries, it does
such a wonderful job of importing data,
setting it up into a data frame so we
can manipulate it rather easily and it
has a lot of tools also in addition to
that. So we usually like to use the
pandas when we can and I'll show you
what that looks like. The other three
lines are for us to get a visual of this
data and take a look at it. So we're
going to import mapplot library.pipplot
as plt and then seabour as sns. Seabor
works with the mattplot library. So you
have to always import mapplot library
and then seabour sits on top of it. And
we'll take a look at what that looks
like. You could use any of your own
plotting libraries you want. There's all
kinds of ways to look at the data. These
are just very common ones. And the
seabor is so easy to use. It just looks
beautiful. It's a nice representation
that you can actually take and show
somebody. And the final line is the
amberigned mattplot library inline. That
is only because I'm doing an inline IDE.
My interface in the Anaconda Jupiter
notebook requires I put that in there or
you're not going to see the graph when
it comes up. Let's go ahead and run
this. It's not going to be that
interesting because we're just setting
up variables. In fact, it's not going to
do anything that we can see, but it is
importing these different libraries and
setup. The next step is load the data
set and extract independent and
dependent variables. Now, here in the
slide, you'll see companies equals PD
read CSV. And it has a long line there
with the file at the end. 10,00
companies.csv. You're going to have to
change this to fit whatever setup you
have. And the file itself, you can
request. Just go down to the commentary
below this video and put a note in there
and SimplyLearn will try to get in
contact with you and supply you with
that file so you can try this coding
yourself. So, we're going to add this
code in here. And we're going to see
that I have companies equals
PD.reader_csv.
And I've changed this path to match my
computer. C/simplylearn/1000
companies.csv. And then below there,
we're going to set the x equals to
companies under the i location. And
because this is companies is a pd data
set, I can use this nice notation that
says take every row, that's what the
colon the first colon is, comma, except
for the last column. That's what the
second part is where we have a colon
minus one and we want the values set
into there. So x is no longer a data set
a pandas data set but we can easily
extract the data from our pandas data
set with this notation and then y we're
going to set equal to the last row. Well
the question is going to be what are we
actually looking at? So let's go ahead
and take a look at that and we're going
to look at the companies.
Which lists the first five rows of data
and I'll open up the file in just a
second so you can see where that's
coming from. But let's look at the data
in here as far as the way the pandas
sees it. When I hit run, you'll see it
breaks it out into a nice setup. This is
what pandas, one of the things pandas is
really good about is it looks just like
an Excel spreadsheet. You have your rows
and remember when we're programming, we
always start with zero. We don't start
with one. So it shows the first five
rows 0 1 2 3 4 and then it shows your
different columns. R&D spend,
administration, marketing spend, state,
profit. It even notes that the top are
column names. It was never told that,
but Pandas is able to recognize a lot of
things that they're not the same as the
data rows. Why don't we go ahead and
open this file up in a CSV so you can
actually see the raw data. So here I've
opened it up as a text editor. And you
can see at the top we have R&D spend,
administration, marketing spend, state,
profit, carriage return. I don't know
about you, but I'd go crazy trying to
read files like this. That's why we use
the pandas. You could also open this up
in an Excel and it would separate it
since it is a comma separated variable
file. But we don't want to look at this
one. We want to look at something we can
read rather easily. So let's flip back
and take a look at that top part, the
first five row. Now, as nice as this
format is where I can see the data, to
me it doesn't mean a whole lot. Maybe
you're an expert in business and
investments and you understand what
$165,34920
compared to the administration cost of
$136,897.80
so on so on helps to create the profit
of $192,26183.
That makes no sense to me whatsoever. No
pun intended. So let's flip back here
and take a look at our next set of code
where we're going to graph it so we can
get a better understanding of our data
and what it mean. So at this point we're
going to use a single line of code to
get a lot of information so we can see
where we're going with this. Let's go
ahead and paste that into our uh
notebook and see what we got going. And
so we have the visualization and again
we're using SNS which is pandas. As you
can see, we imported the mapplot
library.pipplot as plt, which then the
seabor uses and we imported the seabour
as sns. And then that final line of code
helps us show this in our um inline
coding. Without this, it wouldn't
display and you could display it to a
file and other means. And that's the map
plot library in line with the amber sign
at the beginning. So here we come down
to the single line of code. Seabor is
great because it actually recognizes the
panda data frame. So I can just take the
companies.core
for coordinates and I can put that right
into the seaborn. And when we run this,
we get this beautiful plot. And let's
just take a look at what this plot
means. If you look at this plot on mine,
the colors are probably a little bit
more purplish and blue than the original
one. Uh we have the columns and the
rows. We have R and D spending. We have
administration. We have marketing
spending and profit. And if you cross
index any two of these, since we're
interested in profit, if you cross-index
profit with profit, it's going to show
up, if you look at the scale on the
right, way up in the dark. Why? Because
those are the same data. They have an
exact correspondence. So R&D spending is
going to be the same as R&D spending.
And the same thing with administration
costs. But right down the middle, you
get this dark row or dark um diagonal
row that shows that this is the highest
corresponding data. That's exactly the
same. And as it becomes lighter, there's
less connections between the data. So we
can see with profit, obviously profit is
the same as profit. And next, it has a
very high correlation with R&D spending,
which we looked at earlier. And it has a
slightly less connection to marketing
spending and even less to how much money
we put into the administration. So now
that we have a nice look at the data,
let's go ahead and dig in and create
some actual useful linear regression
models so that we can predict values and
have a better profit. Now that we've
taken a look at the visualization of
this data, we're going to move on to the
next step. Instead of just having a
pretty picture, we need to generate some
hard data, some hard values. So let's
see what that looks like. We're going to
set up our linear regression model in
two steps. The first one is we need to
prepare some of our data so it fits
correctly. And let's go ahead and paste
this code into our Jupyter notebook. And
what we're bringing in is we're going to
bring in the sklearn pre-processing
where we're going to import the label
encoder and the one hot encoder. To use
the label encoder, we're going to create
a variable called label encoder and set
it equal to capital L label capital E
encoder. This creates a class that we
can reuse for transferring the labels
back and forth. Now about now you should
ask what labels are we talking about.
Let's go take a look at the data we
processed before and see what I'm
talking about here. If you remember when
we did the companies.head and we printed
the top five rows of data. We have our
columns going across. We have column
zero which is R&D spending, column one
which is administration, column two
which is marketing spending and column
three is state. And you'll see under
state we have New York, California,
Florida. Now to do a linear regression
model, it doesn't know how to process
New York. It knows how to process a
number. So the first thing we're going
to do is we're going to change that New
York, California, and Florida. And we're
going to change those to numbers. That's
what this line of code does here. X
equals and then it has the colon, 3 in
brackets. The first part, the colon,
comma, means that we're going to look at
all the different rows. So we're going
to keep them all together. But the only
row we're going to edit is the third
row. And in there, we're going to take
the label coder and we're going to fit
and transform the x also the third row.
So, we're going to take that third row,
we're going to set it equal to a
transformation. And that transformation
basically tells it that instead of
having a uh New York, it has a zero or a
one or a two. And then finally, we need
to do a one hot encoder, which equals
one hot encoder categorical features
equals three. And then we take the X and
we go ahead and do that equal to one hot
encoder fit transform X to array. This
final transformation preps our data for
us. So it's completely set the way we
need it as just a row of numbers. Even
though it's not in here, let's go ahead
and print X and just take a look what
this data is doing. You'll see I have an
array of arrays and then each array is a
row of numbers. And if I go ahead and
just do row zero, you'll see I have a
nice organized row of numbers that the
computer now understands. We'll go ahead
and take this out there because it
doesn't mean a whole lot to us. It's
just a row of numbers. Next on setting
up our data, we have avoiding dummy
variable trap. This is very important.
Why? Because the computer's
automatically transformed our header
into the setup and it's automatically
transformed all these different
variables. So when we did the encoder,
the encoder created two columns. And
what we need to do is just have the one
because it has both the variable and the
name. That's what this piece of code
does here. Let's go ahead and paste this
in here. And we have x= x colon, one
colon. All this is doing is removing
that one extra column we put in there
when we did our one hot encoder and our
label encoding. Let's go ahead and run
that. And now we get to create our
linear regression model. And let's see
what that looks like here. And we're
going to do that in two steps. The first
step is going to be in splitting the
data. Now, whenever we create a uh
predictive model of data, we always want
to split it up. So, we have a training
set and we have a testing set. That's
very important. Otherwise, we'd be very
unethical without testing it to see how
good our fit is. And then we'll go ahead
and create our multiple linear
regression model and train it and set it
up. Let's go ahead and paste this next
piece of code in here. And I'll go ahead
and shrink it down a size or two so it
all fits on one line. So from the
sklearn module selection, we're going to
import train test split. And you'll see
that we've created four completely
different variables. We have capital X
train capital X test smallercase Y train
smallerase Y test. That is the standard
way that they usually reference these
when we're doing different uh models.
usually see that a capital X and you see
the train and the test and the lowercase
Y. What this is is X is our data going
in. That's our R&D spin, our
administration, our marketing. And then
Y, which we're training, is the answer.
That's the profit because we want to
know the profit of an unknown entity. So
that's what we're going to shoot for in
this tutorial. The next part, train,
test, split. We take X and we take Y.
We've already created those. X has the
columns with the data in it and Y has a
column with profit in it. And then we're
going to set the test size equals 0.2.
That basically means 20%. So 20% of the
rows are going to be tested. We're going
to put them off to the side. So since
we're using a thousand lines of data,
that means that 200 of those lines we're
going to hold off to the side to test
for later. And then the random state
equals zero. We're going to randomize
which ones it picks to hold off to the
side. We'll go ahead and run this. It's
not overly exciting because it's setting
up our variables. But the next step is
the next step we actually create our
linear regression model. Now that we got
to the linear regression model, we get
that next piece of the puzzle. Let's go
ahead and put that code in there and
walk through it. So here we go. We're
going to paste it in there. And let's go
ahead and uh since this is a shorter
line of code, let's zoom up there so we
can get a good look. And we have from
the sklearn.linear_model,
we're going to import linear regression.
Now, I don't know if you recall from
earlier when we were doing all the math.
Let's go ahead and flip back there and
take a look at that. Do you remember
this where we had this long formula on
the bottom and we were doing all this
summization and then we also looked at
setting it up with the different lines
and then we also looked all the way down
to multiple linear regression where
we're adding all those formulas
together. All of that is wrapped up in
this one section. So what's going on
here is I'm going to create a variable
called regressor. And the regressor
equals the linear regression. That's a
linear regression model that has all
that math built in. So we don't have to
have it all memorized or have to compute
it individually. And then we do the
regressor.fit.
In this case, we do xrain and y train
because we're using the training data. X
being the data in and y being profit
what we're looking at. And this does all
that math for us. So within one click
and one line, we've created the whole
linear regression model and we fit the
data to the linear regression model. And
you can see that when I run the
regressor, it gives an output linear
regression. It says copy X equals true,
fit intercept equals true, in jobs equal
1, normalize equals false. It's just
giving you some general information on
what's going on with that regressor
model. Now that we've created our linear
regression model, let's go ahead and use
it. And if you remember, we kept a bunch
of data aside. So, we're going to do a Y
predict variable and we're going to put
in the X test. And let's see what that
looks like. Scroll up a little bit.
Paste that in here. Predicting the test
set results. So, here we have Y predict
equals regressor.predict
X test going in. And this gives us Y
predict. Now, because I'm in Jupiter in
line, I can just put the variable up
there. And when I hit the run button,
it'll print that array out. I could have
just as easily done print y predict. So
if you're in a different IDE that's not
an inline setup like the Jupyter
notebook, you can do it this way. Print
y predict. And you'll see that for the
200 different test variables we kept off
to the side, it's going to produce 200
answers. This is what it says the profit
are for those 200 predictions. But let's
don't stop there. Let's keep going and
take a couple look. We're going to take
just a short detail here and calculating
the coefficients and the intercepts.
This gives us a quick flash at what's
going on behind the line. We're going to
take a short detour here and we're going
to be calculating the coefficient and
intercepts. So you can see what those
look like. What's really nice about our
regressor we created is it already has
the coefficients for us. And we can
simply just print regressor.coefficient
underscore. When I run this, you'll see
our coefficients here. And if we can do
the regressor coefficient, we can also
do the regressor intercept. And let's
run that and take a look at that. This
all came from the multiple regression
model. And we'll flip over so you can
remember where this is going into where
it's coming from. You can see the
formula down here where y = m1 * x1 + m2
* x2 and so on and so on plus c the
coefficient. So these variables fit
right into this formula. Y equ= slope 1
* column 1 variable plus slope 2 *
column 2 variable all the way to the m
into the n and x to the n + c the
coefficient or in this case you have -
8.89 8 9 to the power of two etc etc
times the first column and the second
column and the third column and then our
intercept is the minus one3009
point. Boy, it gets kind of complicated
when you look at it. This is why we
don't do this by hand anymore. This is
why we have the computer to make these
calculations easy to understand and
calculate. Now, I told you that was a
short detour and we're coming towards
the end of our script. As you remember
from the beginning, I said if we're
going to divide this information, we
have to make sure it's a valid model,
that this model works and understand how
good it works. So calculating the R
squar value, that's what we're going to
use to predict how good our prediction
is. And let's take a look at what that
looks like in code. And so we're going
to use this from sklearn.metrics.
We're going to import R2 score. That's
the R squared value. We're looking at
the error. So in the R2 score, we take
our Y test versus our Y predict. Y test
is the actual values we're testing. That
was the one that was given to us. So we
know are true. The Y predict of those
200 values is what we think it was true.
And when we go ahead and run this, we
see we get a 9352.
That's the R2 score. Now, it's not
exactly a straight percentage. So it's
not saying it's 93% correct, but you do
want that in the upper 90s. O and higher
shows that this is a very valid
prediction based on the R2 score. And if
R squar value of N1 or 92 as we got on
our model remember it does have a random
generation involved. This proves the
model is a good model which means
success. Yay. We successfully trained
our model with certain predictors and
estimated the profit of the companies
using linear regression. What is
logistic regression? Let's say we have
to build a predictive model or a machine
learning model to predict whether the
passengers of the Titanic ship have
survived or not the shipwreck. So how do
we do that? So we use logistic
regression to build a model for this.
How do we use logistic regression? So we
have the information about the
passengers, their ID, whether they have
survived or not, their class and name
and so on and so forth. And we use this
information where we already know
whether the person has survived or not.
That is the labeled information and we
help the system to train based on this
information with based on this labeled
data. This is known as labeled data. And
during the process of building the
model, we probably will remove some of
the non-essential parameters or
attributes here. We only take those
attributes which are really required to
make these predictions. And once we
train the model, we run new data through
it whereby the model will predict
whether the passenger has survived or
not. All right. What is logistic
regression? As I mentioned earlier,
logistic regression is an algorithm for
performing binary classification. So
let's take an example and see how this
works. Let's say your car has not been
serviced for quite a few years and now
you want to find out if it is going to
break down in the near future. So this
is like a classification problem. Find
out whether your car will break down or
not. So how are we going to perform this
classification? So here's how it looks.
If we plot the information along the X
and Y axis, X is the number of years
since the last service was performed and
Y is the probability of your car
breaking down. And let's say this
information was this data rather was
collected from several car users. It's
not just your car but several car users.
So that is our labeled data. So the data
has been collected and um for for the
number of years and when the car broke
down and what was the probability and
that has been plotted along x and y
axis. So this provides an idea or from
this graph we can find out whether your
car will break down or not. We'll see
how. So first of all the probability can
go from 0 to one. As you all aware
probability can be between 0 and one.
And as we can imagine it is intuitive as
well. As the number of years are on the
lower side maybe 1 year, 2 years or 3
years till after the service the chances
of your car breaking down are very
limited. Right? So for example, chances
of your car breaking down or the
probability of your car breaking down
within 2 years of your last service are
0.1 probability. Similarly 3 years is
maybe.3 and so on. But as the number of
years increases let's say if it was 6 or
7 years there is almost a certainty that
your car is going to break down. That is
what this graph shows. So this is an
example of a application of the
classification algorithm and we will see
in little details how exactly logistic
regression is applied here. One more
thing needs to be added here is that the
dependent variables outcome is discrete.
So if we are talking about whether the
car is going to break down or not. So
that is a discrete value. The y that we
are talking about the dependent variable
that we are talking about what we are
looking at is whether the car is going
to break down or not yes or no that is
what we are talking about. So here the
outcome is discrete and not a continuous
value. So this is how the logistic
regression curve looks. Let me explain a
little bit what exactly and how exactly
we are going to uh determine the class
the outcome rather. So for a logistic
regression curve a threshold has to be
set saying that because this is a
probability calculation remember this is
a probability calculation and the
probability itself will not be zero or
one but based on the probability we need
to decide what the outcome should be. So
there has to be a threshold like for
example 0.5 can be the threshold let's
say in this case. So any value of the
probability below 0.5 is considered to
be zero and any value above.5 is
considered to be one. So an output of
let's say8
will mean that the car will break down.
So that is considered as an output of 1
and let's say an output of 29 is
considered as zero which means that the
car will not break down. So that's the
way logistic regression works. Now let's
do a quick comparison between logistic
regression and linear regression because
they both have the term regression in
them. So it can cause confusion. So
let's try to remove that confusion. So
what is linear regression? Linear
regression is a process is once again an
algorithm for supervised learning.
However, here you're going to find a
continuous value. You're going to
determine a continuous value. It could
be the price of a real estate property.
It could be your hike, how much hike
you're going to get or it could be a
stock price. These are all continuous
values. These are not discrete compared
to a yes or a no kind of a response that
we are looking for in logistic
regression. So this is one example of a
linear regression. Let's say the HR team
of a company tries to find out what
should be the salary hike of an
employee. So they collect all the
details of their existing employees,
their ratings and their salary hikes,
what has been given and that is the
labeled information that is available
and the system learns from this. It is
trained and it learns from this labeled
information so that when a new employees
information is fed based on the rating
it will determine what should be the
high. So this is a linear regression
problem and a linear regression example.
Now salary is a continuous value. You
can get 5,000, 5,500,
5,600. It is not discrete like a cat or
a dog or an apple or a banana. These are
discrete or a yes or a no. These are
discrete values, right? So this where
you're trying to find continuous values
is where we use linear regression. So
let's say just to extend on this
scenario, we now want to find out
whether this employee is going to get a
promotion or not. So we want to find out
that is a discrete problem, right? A yes
or no kind of a problem. In this case,
we actually cannot use linear regression
even though we may have labeled data. So
this is the label data. So based on the
employee rating these are the ratings
and then some people got the promotion
and this is the ratings for which people
did not get promotion that is a no and
this is the rating for which people got
promotion we just plotted the data about
whether a person has got an employee has
got promotion or not yes no right so
there is nothing in between and what is
the employees rating okay and ratings
can be continuous that is not an issue
but the output is discrete in In this
case whether employee got promotion yes
no okay so if we try to plot that and we
try to find a straight line this is how
it would look and as you can see it
doesn't look very right because looks
like there will be lot of errors this
root mean square error if you remember
for linear regression would be very very
high and also the the values cannot go
beyond zero or beyond one. So the graph
should probably look somewhat like this
clipped at 0 and one. But still the
straight line doesn't look right.
Therefore instead of using a linear
equation we need to come up with
something different and therefore the
logistic regression model looks somewhat
like this. So we calculate the
probability and if we plot that
probability not in the form of a
straight line but we need to use some
other equation. And we will see very
soon what that equation is. Then it is a
gradual process. Right? So you see here
people with some of these ratings are
not getting any promotions and then
slowly uh at certain rating they get
promotion. So that is a gradual process
and uh this is how the math behind
logistic regression looks. So we are
trying to find the odds for a particular
event happening and this is the formula
for finding the odds. So the probability
of an event happening divided by the
probability of the event not happening.
So P if it is the probability of the
event happening probability of the
person getting a promotion and divided
by the probability of the person not
getting a promotion that is 1 minus P.
So this is how you measure the odds. Now
the values of the odds range from 0 to
infinity. So when this probability is
zero then the odds will the value of the
odds is equal to zero and when the
probability becomes 1 then the value of
the odds is 1 by 0 that will be infinity
but the probability itself remains
between 0 and 1. Now this is how an
equation of a straight line looks. So y
is equal to beta 0 plus beta 1x where
beta 0 is the y intercept and beta 1 is
the slope of the line. If we take the
odds equation and take a log of both
sides, then this would look somewhat
like this. And the term logistic is
actually derived from the fact that we
are doing this. We take a log of px by 1
minus px. This is an extension of the
calculation of odds that we have seen,
right? And that is equal to beta 0 plus
beta 1x which is the equation of the
straight line. And now from here if you
want to find out the value of px we will
see we can take the exponential on both
sides and then if we solve that equation
we will get the equation of px like this
px is equal to 1 by 1 + e ^ of minus
beta 0 + beta 1x and recall this is
nothing but the equation of the line
which is equal to y is equal to beta 0 +
beta 1x. So that this is the equation
also known as the sigmoid function and
this is the equation of the logistic
regression alg. All right and if this is
plotted this is how the sigmoid curve is
obtained. So let's compare linear and
logistic regression how they are
different from each other. Let's go
back. So linear regression is solved or
used to solve regression problems and
logistic regression is used to solve
classification problems. So both are
called regression. But linear regression
is used for solving regression problems
where we predict continuous values.
Whereas logistic regression is used for
solving classification problems where we
have had to predict discrete values. The
response variables in case of linear
regression are continuous in nature.
Whereas here they are categorical or
discrete in nature. And the linear
regression helps to estimate the
dependent variable when there is a
change in the independent variable.
Whereas here in case of logistic
regression it helps to calculate the
probability or the possibility of a
particular event happening. And linear
regression as the name suggests is a
straight line. That's why it's called
linear regression. Whereas logistic
regression is a sigmoid function and the
curve is the shape of the curve is S.
It's an S-shaped curve. This is another
example of application of logistic
regression in weather prediction.
Whether it's going to rain or not rain.
Now keep in mind both are used in
weather prediction. If we want to find
the discrete values like whether it's
going to rain or not rain that is a
classification problem. We use logistic
regression. But if we want to determine
what is going to be the temperature
tomorrow, then we use linear regression.
So just keep in mind that in weather
prediction, we actually use both. But
these are some examples of logistic
regression. So we want to find out
whether it's going to be rain or not,
it's going to be sunny or not, whether
it's going to snow or not. These are all
logistic regression examples. A few more
examples. Classification of objects.
This is a again another example of
logistic regression. Now here of course
one distinction is that these are
multiclass classification. So logistic
regression is not used in its original
form but it is used in a slightly
different form. So we say whether it is
a dog or not a dog. I hope you
understand. So instead of saying is it a
dog or a cat or elephant we convert this
into saying so because we need to keep
it to binary classification. So we say
is it a dog or not a dog? Is it a cat or
not a cat? So that's the way logistic
regression can be used for classifying
objects. Otherwise there are other
techniques which can be used for
performing multiclass classification. In
healthcare logistic regression is used
to find the survival rate of a patient.
So they take multiple parameters like
trauma score and age and so on and so
forth and they try to predict the rate
of survival. All right. Now finally
let's take an example and see how we can
apply logistic regression to predict the
number that is shown in the image. So
this is actually a live demo. I will
take you into Jupyter notebook and u
show the code. But before that let me
take you through a couple of slides to
explain what we're trying to do. So
let's say you have an 8x8 image and the
the image has a number 1 2 3 4 and you
need to train your model to predict what
this number is. So how do we do this? So
the first thing is obviously in any
machine learning process you train your
model. So in this case we are using
logistic regression. So and then we
provide a training set to train the
model and then we test how accurate our
model is with the test data which means
that like any machine learning process.
We split our initial data into two parts
training set and test set. With the
training set we train our model and then
with the test set we we test the model
till we get good accuracy and then we
use it for for inference. Right? So that
is typical methodology of uh uh
training, testing and then deploying of
machine learning models. So let's uh
take a look at the code and uh see what
we are doing. So I'll not go line by
line but just take you through some of
the blocks. So first thing we do is
import all the libraries and then we
basically take a look at the images and
see what is the total number of images.
We can display using mattplot lip some
of the images or a sample of these
images and um then we split the data
into training and test as I mentioned
earlier and we can do some exploratory
analysis and uh then we build our model.
We train our model with the training set
and then we test it with our test set
and find out how accurate our model is
using the confusion matrix the heat map
and use heat map for visualizing this
and uh I will show you in the code what
exactly is the confusion matrix and how
it can be used for finding the accuracy
in our example we got we get an accuracy
of about 94 which is pretty good or 94%
which is pretty good all right so what
is the confusion matrix. This is an
example of a confusion matrix and uh
this is used for identifying the
accuracy of a classification model or
like a logistic regression model. So the
most important part in a confusion
matrix is that first of all this as you
can see this is a matrix and the size of
the matrix depends on how many outputs
uh we are expecting right. So the the
most important part here is that the
model will be most accurate when we have
the maximum numbers in its diagonal like
in this case that's why it has almost 93
94% because the diagonals should have
the maximum numbers and the others other
than diagonals the cells other than the
diagonals should have very few numbers.
So here that's what is happening. So
there is a two here. There are there's a
one here. But most of them are along the
diagonal. This what does this mean? This
means that the number that has been fed
is zero and the number that has been
detected is also zero. So the predicted
value and the actual value are the same.
So along the diagonals that is true.
Which means that let's let's take this
diagonal right. If if the maximum number
is here that means that uh like here in
this case it is 34 which means that 34
of the images that have been fed or
rather actually there are two
mclassifications in there. So 36 images
have been fed which have number four and
out of which 34 have been predicted
correctly as number four and one has
been predicted as number eight and
another one has been predicted as number
nine. So these are two mclassifications.
Okay. So that is the meaning of saying
that the maximum number should be in the
diagonal. So if you have all of them so
for an ideal model which has let's say
100% accuracy everything will be only in
the diagonal. There will be no numbers
other than zero in all other cells. So
that is like a 100% accurate model.
Okay. So that's the gist of how to use
this matrix. How to use this uh
confusion matrix. So I know the name uh
is a little funny sounding confusion
matrix but actually it is not very
confusing. It's very straightforward. So
you are just plotting what has been
predicted and what is the labeled
information or what is the actual data
that's also known as the ground truth
sometimes. Okay, these are some fancy
terms that are used. So predicted label
and the actual label that's all it is.
Okay. Yeah. So we are showing a little
bit more information here. So 38 have
been predicted and here you will see
that all of them have been predicted
correctly. There have been 38 zeros and
the predicted value and the actual value
is is exactly the same. Whereas in this
case right it has uh there are I think
37 + 5 yeah 42 have been fed the images
42 images are of digit three and uh the
accuracy is only 37 of them have been
accurately predicted. Three of them have
been predicted as number seven and two
of them have been predicted as number
eight and so on and so forth. Okay. All
right. So with that let's go into
Jupyter notebook and see how the code
looks. So this is the code in in Jupyter
notebook for logistic regression. In
this particular demo, what we are going
to do is train our model to recognize
digits which are the images which have
digits from let's say 0 to 5 or 0 to 9
and um and then we will see how well it
is trained and whether it is able to
predict these numbers correctly or not.
So let's get started. So the first part
is as usual we are importing some
libraries that are required and uh then
the last line in this block is to load
the digits. So let's go ahead and run
this code. Then here we will visualize
the shape of these uh digits. So we can
see here if we take a look this is how
the shape is 1797 by 64. These are like
8 by8 images. So that's that's what is
reflected in this uh shape. Now from
here onwards we are basically once again
importing some of the libraries that are
required like numpy and map plot and we
will take a look at uh some of the
sample images that we have loaded. So th
this one for example creates a figure uh
and then we go ahead and take a few
sample images to see how they look. So
let me run this code and so that it
becomes easy to understand. So these are
about five images sample images that we
are looking at 0 1 2 3 4. So this is how
the images this is how the data is.
Okay. And uh based on this we will
actually train our logistic regression
model and then we will test it and see
how well it is able to recognize. So the
way it works is the pixel information.
So as you can see here this is an 8x 8
pixel kind of a image and uh the each
pixel whether it is activated or not
activated that is the information
available for each pixel. Now based on
the pattern of this activation and
non-activation of the various pixels
this will be identified as a zero for
example right similarly as you can see
so overall each of these numbers
actually has a different pattern of the
pixel activation and that's pretty much
that our model needs to learn for which
number what is the pattern of the
activation of the pixels right so that
is what we are going to train our model.
Okay. So the first thing we need to do
is to split our data into training and
test data set. Right? So whenever we
perform any training, we split the data
into training and test. So that the
training data set is used to train the
system. So we pass this probably
multiple times. Uh and then we test it
with the test data set. And the split is
usually in the form of there and there
are various ways in which you can split
this data. It is up to the individual
preferences. In our case here we are
splitting in the form of 23 and 77. So
when we say test size as 2023
that means 23% of the entire data is
used for testing and the remaining 77%
is used for training. So there is a
readily available function which is uh
called train test split. So we don't
have to write any special code for the
splitting. It will automatically split
the data based on the proportion that we
give here which is test size. So we just
give the test size automatically
training size will be determined and uh
we pass the data that we want to split
and the the results will be stored in x
train and y train for the training data
set. And what is x train? This are these
are the features right which is like the
independent variable and y train is the
label right so in this case what happens
is we have the input value which is or
the features value which is in x train
and since this is a labeled data for
each of them each of the observations we
already have the label information
saying whether this digit is a zero or a
one or a two so that this this is what
will be used for comparison to find out
whether the the system is able to
recognize it correctly or there is an
error for each observation it will
compare with this right so this is the
label so the same way x train y train is
for the training data set x test y test
is for the test data set okay so let me
go ahead and execute this code as well
and then we can go and check quickly
what is the how many entries are there
and in each of this so x train the shape
is 1383x
64 and y train has 1383 because there is
uh nothing like the second part is not
required here and then x test shape we
see is 414 so actually there are 414
observations in test and 1383
observations in train so that's
basically what these four lines of code
are are saying okay then we import the
uh logistic regression
library and uh which is a part of
scikitlearn. So we we don't have to
implement the logistic regression
process itself. We just call these uh
the function and uh let me go ahead and
execute that so that uh we have the
logistic regression library imported.
Now we create an instance of logistic
regression. Right? So logistic regr is a
is an instance of logistic regression
and then we use that for training our
model. So let me first execute this
code. So these two lines. So the first
line basically creates an instance of
logistic regression model and then the
second line is where we are passing our
data the training data set. Right? This
is our the the predictors and uh this is
our target. We are passing this data set
to train our model. All right. So once
we do this in this case the data is not
large but by and large uh the training
is what takes usually a lot of time. So
we spend in machine learning activities
in machine learning projects we spend a
lot of time for the training part of it.
Okay. So here the data set is relatively
small so it was pretty quick. So all
right so now our model has been trained
using the training data set and uh we
want to see how accurate this is. So
what we'll do is we will test it out in
probably faces. So let me first try out
how well this is working for one image.
Okay, I will just try it out with one
image my the first entry in my test data
set and see whether it is uh correctly
predicting or not. So and in order to
test it so for training purpose we use
the fit method. There is a method called
fit which is for training the model and
once the training is done if you want to
test for uh a particular value new input
you use the predict method. Okay. So
let's run the predict method and we pass
this particular image and uh we see that
the shape is or the prediction is four.
So let's try a few more. Let me see for
the next 10 uh seems to be fine. So let
me just go ahead and test the entire
data set. Okay, that's basically what we
will do. So now we want to find out how
accurately this has u performed. So we
use the score method to find what is the
percentages of accuracy and we see here
that it has performed up to 94%
accurate. Okay. So that's uh on this
part. Now what we can also do is we can
um also see this accuracy using what is
known as confusion matrix. So let us go
ahead and uh try that as well. Uh so
that we can also visualize how well uh
this model has uh done. So let me
execute this piece of code which will
basically import some of the libraries
that are required and um we we basically
create a confusion matrix an instance of
confusion matrix by running confusion
matrix and passing these uh values. So
we have so this confusion_matrix
method takes two parameters one is the y
test and the other is uh the prediction.
So what is a y test? These are the
labeled values which we already know for
the test data set and predictions are
what the system has predicted for the
test data set. Okay. So this is known to
us and this is what the system has uh
the model has generated. So we kind of
create the confusion matrix and we will
print it. And uh this is how the
confusion matrix looks. As the name
suggests it is a matrix and um the key
point out here is that the accuracy of
the model is determined by how many
numbers are there in the diagonal. The
more the numbers in the diagonal, the
better the accuracy is. Okay. And first
of all, the total sum of all the numbers
in this whole matrix is equal to the
number of observations in the test data
set. That is the first thing, right? So
if you add up all these numbers, that
will be equal to the number of
observations in the test data set. And
then out of that, the maximum number of
them should be in the diagonal. That
means the accuracy is pretty good. If
the the numbers in the diagonal are less
and in all other places there are a lot
of numbers uh which means the accuracy
is very low. The diagonal indicates a
correct prediction that this means that
the actual value is same as the
predicted value. Here again actual value
is same as the predicted value and so
on. Right? So the moment you see a
number here that means the actual value
is something and the predicted value is
something else. Right? Similarly here
the actual value is something and the
predicted value is something else. So
that is basically how we read the
confusion matrix. Now how do we find the
accuracy? You can actually add up the
total values in the diagonal. So it it's
like 38 + 44 + 43 and so on and divide
that by the total number of test
observations that will give you the
percentage accuracy using a confusion
matrix. Now let us visualize this
confusion matrix in a slightly more
sophisticated way uh using a heat map.
So we will create a heat map with some
we'll add some colors as well. It's uh
it's like a more visually visually more
appealing. So that's the whole idea. So
if we let me run this piece of code and
this is how the heat map looks. Uh and
as you can see here the diagonals again
are all the values are here most of the
values. So which means reasonably this
seems to be reasonably accurate and yeah
basically the accuracy score is 94%.
This is calculated as I mentioned by
adding all these numbers divided by the
total test values or the total number of
observations in test data set. Okay. So
this is the confusion matrix for
logistic regression.
All right. So now that we have seen the
confusion matrix, let's take a quick
sample and see how well uh the system
has classified and we will take a a few
examples of the data. So if we see here
we we picked up randomly a few of them.
So this is uh number four which is the
actual value and also the predicted
value both are four. This is an image of
zero. So the predicted value is also
zero. Actual value is of course zero.
Then this is the image of nine. So this
has also been predicted correctly 9 and
actual value is 9. And this is the image
of one. And again this has been
predicted correctly as like the actual
value. Okay. So this was a quick demo of
logistic regression. How to use logistic
regression to identify images.
>> What is a decision tree? Let's go
through a very simple example before we
dig in deep. Decision tree is a
treeshaped diagram used to determine a
course of action. Each branch of the
tree represents a possible decision or
occurrence or reaction. Let's start with
a simple question. How to identify a
random vegetable from a shopping bag?
So, we have this group of vegetables in
here. And we can start off by asking a
simple question. Is it red? And if it's
not, then it's going to be the purple
fruit to the left, probably an eggplant.
If it's true, it's going to be one of
the red fruits. Is the diameter greater
than two? If false, it's going to be a
what looks to be a red chili. And if
it's true, it's going to be a bell
pepper from the capsicum family. So,
it's a capsicum.
Problems that decision tree can solve.
So, let's look at the two different
categories the decision tree can be used
on. It can be used on the
classification, the true false, yes, no,
and it can be used on regression where
we figure out what the next value is in
a series of numbers or a group of data.
In classification, the classification
tree will determine a set of logical if
then conditions to classify problems.
For example, discriminating between
three types of flowers based on certain
features. In regression, a regression
tree is used when the target variable is
numerical or continuous in nature. We
fit the regression model to the target
variable using each of the independent
variables. Each split is made based on
the sum of squared error. Before we dig
deeper into the mechanics of the
decision tree, let's take a look at the
advantages of using a decision tree and
we'll also take a glimpse at the
disadvantages. The first thing you'll
notice is that it's simple to
understand, interpret, and visualize. It
really shines here because you can see
exactly what's going on in a decision
tree. Little effort is required for data
preparation. So, you don't have to do
special scaling. There's a lot of things
you don't have to worry about when using
a decision tree. It can handle both
numerical and categorical data as we
discovered earlier and nonlinear
parameters don't affect its performance.
So even if the data doesn't fit an easy
curved graph, you can still use it to
create an effective decision or
prediction. If we're going to look at
the advantages of a decision tree, we
also need to understand the
disadvantages of a decision tree. The
first disadvantage is overfitting.
Overfitting occurs when the algorithm
captures noise in the data. That means
you're solving for one specific instance
instead of a general solution for all
the data. High variance. The model can
get unstable due to small variation in
data. Low bias tree. A highly
complicated decision tree tends to have
a low bias which makes it difficult for
the model to work with new data.
Decision tree important terms. Before we
dive in further, we need to look at some
basic terms. We need to have some
definitions to go with our decision tree
in the different parts we're going to be
using. We'll start with entropy. Entropy
is a measure of randomness or
unpredictability in the data set. For
example, we have a group of animals in
this picture. There's four different
kinds of animals. And this data set is
considered to have a high entropy. You
really can't pick out what kind of
animal it is based on looking at just
the four animals as a big clump of of uh
entities. So as we start splitting it
into subgroups, we come up with our
second definition which is information
gain. Information gain it is a measure
of decrease in entropy after the data
set is split. So in this case based on
the color yellow, we've split one group
of animals on one side as true and those
who aren't yellow as false. As we
continue down the yellow side, we split
based on the height. True or false
equals 10. And on the other side, height
is less than 10. True or false? And as
you see as we split it, the entropy
continues to be less and less and less.
And so our information gain is simply
the entropy E1 from the top and how it's
changed to E2 in the bottom. And we'll
look at the uh deeper math, although you
really don't need to know a huge amount
of math when you actually do the
programming in Python because it'll do
it for you. But we'll look on the actual
math of how they compute entropy.
Finally, we want to know the different
parts of our tree and they call the leaf
node. Leaf node carries the
classification or the decision. So it's
the final end at the bottom. The
decision node has two or more branches.
This is where we're breaking the group
up into different parts. And finally,
you have the root node. The topmost
decision node is known as the root node.
How does a decision tree work? Wonder
what kind of animals I'll get in the
jungle today? Maybe you're the hunter
with the gun. Or if you're more into
photography, you're a photographer with
a camera. So let's look at this group of
animals and let's try to classify
different types of animals based on
their features using a decision tree. So
the problem statement is to classify the
different types of animals based on
their features using a decision tree.
The data set is looking quite messy and
the entropy is high in this case. So
let's look at a training set or a
training data set and we're looking at
color. We're looking at height and then
we have our different animals. We have
our elephants, our giraffes, our
monkeys, and our tigers. And they're of
different colors and shapes. Let's see
what that looks like. And how do we
split the data? We have to frame the
conditions that split the data in such a
way that the information gain is the
highest. Note, gain is the measure of
decrease in entropy after splitting. So
the formula for entropy is the sum
that's what this symbol looks like. That
looks like kind of like a uh e funky e
of k where i equals 1 to k. K would
represent the number of animal the
different animals in there where value
or P value of I would be the percentage
of that animal times the log base 2 of
the same the percentage of that animal.
Let's try to calculate the entropy for
the current data set and take a look at
what that looks like. And don't be
afraid of the math. You don't really
have to memorize this math. Just be
aware that it's there and this is what's
going on in the background. And so we
have three giraffes, two tigers, one
monkey, two elephants, a total of eight
animals gathered. And if we plug that
into the formula, we get an entropy that
equals 3 over8. So we have three
giraffes, a total of 8 times the log.
Usually they use base 2 on the log. So
log base 2 of 3 over8 plus in this case,
let's say it's the elephants, 2 over 8.
Two elephants over total of 8 time log
base 2 2 over 8 plus one monkey over
total of 8. log base 2 1 over 8 and plus
2 over 8 of the tigers log base 2 over 8
and if we plug that into our computer or
calculator I obviously can't do logs in
my head we get an entropy equal to.571
the program will actually calculate the
entropy of the data set similarly after
every split to calculate the gain now
we're not going to go through each set
one at a time to see what those numbers
are just want you to be aware that this
is a formula or the mathematics behind
It gain can be calculated by finding the
difference of the subsequent entropy
values after a split. Now we will try to
choose a condition that gives us the
highest gain. We will do that by
splitting the data using each condition
and checking that the gain we get out of
them. The condition that gives us the
highest gain will be used to make the
first split. Can you guess what that
first split will be just by looking at
this image? As a human, it's probably
pretty easy to split it. Let's see if
you're right. If you guessed the color
yellow, you're correct. Let's say the
condition that gives us the maximum gain
is yellow. So we will split the data
based on the color yellow. If it's true,
that group of animals goes to the left.
If it's false, it goes to the right. The
entropy after the splitting has
decreased considerably. However, we
still need some splitting at both the
branches to attain an entropy value
equal to zero. So we decide to split
both the nodes using height as a
condition. Since every branch now
contains single label type, we can say
that entropy in this case has reached
the least value. And here you see we
have the giraffes, the tigers, the
monkey and the elephants all separated
into their own groups. This tree can now
predict all the classes of animals
present in the data set with 100%
accuracy. That was easy. Use case loan
repayment prediction. Let's get into my
favorite part and open up some Python
and see what the programming code and
the scripting looks like. In here, we're
going to want to do a prediction. And we
start with this individual here who's
requesting to find out how good his
customers are going to be, whether
they're going to repay their loan or not
for this bank. And from that, we want to
generate a problem statement to predict
if a customer will repay loan amount or
not. And then we're going to be using
the decision tree algorithm in Python.
Let's see what that looks like. And
let's dive into the code. In our first
few steps of implementation, we're going
to start by importing the necessary
packages that we need from Python. and
we're going to load up our data and take
a look at what the data looks like. So,
the first thing I need is I need
something to edit my Python and run it
in. So, let's flip on over. And here I'm
using the Anaconda Jupiter notebook.
Now, you can use any Python IDE you like
to run it in, but I find the Jupyter
Notebook's really nice for doing things
on the fly. And let's go ahead and just
paste that code in the beginning. And
before we start, let's talk a little bit
about what we're bringing in. And then
we're going to do a couple things in
here. where I have to make a couple
changes as we go through this first part
of the import. The first thing we bring
in is numpy as np. That's very standard
when we're dealing with mathematics,
especially with uh very complicated
machine learning tools. You'll almost
always see the numpy come in for your
num your numbers. It's called number
python. It has your mathematics in
there. In this case, we actually could
take it out, but generally you'll need
it for most of your different things you
work with. And then we're going to use
pandas as pd. That's also a standard.
The pandas is a dataf frame setup and
you can liken this to uh taking your
basic data and storing it in a way that
looks like an Excel spreadsheet. So as
we come back to this when you see np or
pd those are very standard uses you'll
know that that's the pandas and I'll
show you a little bit more when we
explore the data in just a minute. Then
we're going to need to split the data.
So I'm going to bring in our train test
and split and this is coming from the
sklearn package cross validation. In
just a minute, we're going to change
that and we'll go over that, too. And
then there's also the sktree import
decision tree classifier. That's the
actual tool we're using. Remember, I
told you don't be afraid of the
mathematics. It's going to be done for
you. Well, the decision tree classifier
has all that mathematics in there for
you, so you don't have to figure it back
out again. And then we have
sklearn.metrics
for accuracy score. We need to score our
our setup. That's the whole reason we're
splitting it between the training and
testing data. And finally, we still need
the sklearn import tree. And that's just
the basic tree function that's needed
for the decision tree classifier. And
finally, we're going to load our data
down here. And I'm going to run this and
we're going to get two things on here.
One, we're going to get an error. And
two, we're going to get a warning. Let's
see what that looks like. So the first
thing we had is we have an error. Why is
this error here? Well, it's looking at
this. It says I need to read a file. And
when this was written, the person who
wrote it, this is their path where they
stored the file. So let's go ahead and
fix that.
And I'm going to put in here my file
path. I'm just going to call it full
file name. And you'll see it's on my C
drive. And there's this very lengthy
setup on here where I stored the data
2.csv file.
Don't worry too much about the full path
because on your computer it'll be
different. The data.2 CSV file was
generated by SimplyLearn. If you want a
copy of that, you can comment down below
and request it here in the YouTube.
And then if I'm going to give it a name,
full file name, I'm going to go ahead
and change it here to full
file name. So let's go ahead and run it
now and see what happens.
And we get a warning
when you're coding. Understanding these
different warnings and these different
errors that come up is probably the
hardest lesson to learn. So let's just
go ahead and take a look at this and use
this as a uh opportunity to understand
what's going on here. If you read the
warning, it says the cross validation is
depreciated. So it's a warning on it's
being removed and it's going to be moved
in favor of the model selection. So if
we go up here, we have
sklearn.crossvalidation.
And if you research this and go to
sklearn site, you'll find out that you
can actually just swap it right in there
with model selection.
And so when I come in here and I run it
again, that removes a warning. What
they've done is they've had two
different developers develop it in two
different branches and then they decided
to keep one of those and eventually get
rid of the other one. That's all that is
and very easy and quick to fix.
Before we go any further, I went ahead
and opened up the data from this file.
Remember the the data file we just
loaded on here, the data_2.c
CSV. Let's talk a little bit more about
that and see what that looks like both
as a text file because it's a
commaepparated variable file and in a
spreadsheet. This is what it looks like
as a basic text file. You can see at the
top they've created a header and it's
got 1 2 3 4 five columns and each column
has data in it. And let me flip this
over cuz we're also going to look at
this uh in an actual spreadsheet so you
can see what that looks like. And here
I've opened it up in the open office
calc, which is pretty much the same as
um Excel and zoomed in. And you can see
we've got our columns and our rows of
data. A little easier to read in here.
We have a result, yes, yes, no. We have
initial payment, last payment, credit
score, house number. If we scroll way
down,
we'll see that this occupies a 101 lines
of code or lines of data with uh the
first one being a column and then 1,000
lines of data.
Now, as a programmer,
if you're looking at a small amount of
data, I usually start by pulling it up
in different sources so I can see what
I'm working with.
But in larger data, you won't have that
option. it would just be um too too
large. So you need to either bring in a
small amount that you can look at it
like we're doing right now or we can
start looking at it through the Python
code. So let's go ahead and move on and
take the next couple steps to explore
the data using Python. Let's go ahead
and see what it looks like in Python to
print the length and the shape of the
data. So let's start by printing the
length of the database. We can use a
simple lin function from Python. And
when I run this, you'll see that it's a
thousand long. And that's what we
expected. There's a thousand lines of
data in there. If you subtract the
column head, and this is one of the nice
things when we did the uh balance data
from the panda read CSV, you'll see that
the header is row zero. So, it
automatically removes a row and then
shows the data separate. It does a good
job sorting that data out for us. And
then we can use a different function.
And let's take a look at that. And
again, we're going to utilize the tools
in Panda.
And since the balance data was loaded as
a Panda data frame,
we can do a shape on it. And let's go
ahead and run the shape and see what
that looks like.
What's nice about the shape is not only
does it give me the length of the data,
we have a th00and lines, it also tells
me there's five columns. So when we were
looking at the data, we had five columns
of data. And then let's take one more
step to explore the data using Python.
And now that we've taken a look at the
length and the shape, let's go ahead and
use the uh pandas module for head.
Another beautiful thing in the data set
that we can utilize. So let's put that
on our sheet here. And we have print
data set and balance data.head.
And this is a pandas print statement of
its own. So it has its own print feature
in there. And then we went ahead and
gave a label for our print job here of
data set. Just a simple print statement.
And we run that. And let's just take a
closer look at that. Let me zoom in
here.
There we go.
Pandas does such a wonderful job of
making this a very clean readable data
set. So you can look at the data, you
can look at the column headers, you can
have it uh when you put it as the head,
it prints the first five lines of the
data. And we always start with zero. So
we have five lines. We have 0 1 2 3 4
instead of 1 2 3 4 5. That's a standard
scripting and programming set is you
want to start with the zero position.
And that is what the data head does. It
pulls the first five rows of data. Puts
it in a nice format that you can look at
and view. Very powerful tool to view the
data. So instead of having to flip and
open up an Excel spreadsheet or open
Office Cal or trying to look at a word
doc where it's all scrunched together
and hard to read, you can now get a nice
open view of what you're working with.
We're working with a shape of a thousand
long, five wide. So we have five columns
and we do the full data head. You can
actually see what this data looks like.
The initial payment, last payment,
credit scores, house number. So let's
take this now that we've explored the
data and let's start digging into the
decision tree. So in our next step,
we're going to train and build our data
tree. And to do that, we need to first
separate the data out. We're going to
separate into two groups so that we have
something to actually train the data
with. And then we have some data on the
side to test it to see how good our
model is. Remember with any of the
machine learning, you always want to
have some kind of test set to to weigh
it against so you know how good your
model is when you distribute it. Let's
go ahead and break this code down and
look at it in pieces. So first we have
our X and Y.
Where do X and Y come from? Well, X is
going to be our data and Y is going to
be the answer or the target. You can
look at it source and target. In this
case, we're using X and Y to denote the
data in and the data that we're actually
trying to guess what the answer is going
to be. And so to separate it, we can
simply put in X equals the balance of
the data values. The first brackets
means that we're going to select all the
lines in the database. So, it's all the
data. And the second one says we're only
going to look at columns 1 through five.
Remember, always start with zero. Zero
is a yes or no. And that's whether the
loan went default or not. So, we want to
start with one. If we go back up here,
that's the initial payment and it goes
all the way through the house number.
Well, if we want to look at uh 1 through
five, we can do the same thing for y,
which is the answers. And we're going to
set that just equal to the zero row. So,
it's just the zero row and then it's all
rows going in there. So, now we've
divided this into two different data
sets. One of them with the
data going in and one with the answers.
Next, we need to split the data.
And here you'll see that we have it
split into four different parts. The
first one is your X training, your X
test, your Y train, your Y test.
Simply put, we have X going in where
we're going to train it and we have to
know the answer to train it with. And
then we have X test where we're going to
test that data and we have to know in
the end what the Y was supposed to be.
And that's where this train test split
comes in that we loaded earlier in the
modules. This does it all for us. And
you can see they set the test size equal
to.3. So that's roughly 30% will be used
in the test. And then we use a random
state. So it's completely random which
rows it takes out of there. And then
finally we get to actually build our
decision tree. And they've called it
here CLF entropy. That's the actual
decision tree or decision tree
classifier. And in here, they've added a
couple variables which we'll explore in
just a minute. And then finally, we need
to fit the data to that. So, we take our
CLF entropy that we created and we fit
the X train. And since we know the
answers for X-ray or the Y train, we go
ahead and put those in. And let's go
ahead and run this. And what most of
these sklearn modules do is when you set
up the variable, in this case, when we
set the CLF entropy equal decision tree
classifier, it automatically prints out
what's in that decision tree. There's a
lot of variables you can play with in
here. And it's quite beyond the scope of
this tutorial to go through all of these
and how they work. But we're working on
entropy. That's one of the options.
We've added that it's completely a
random state of 100, so 100%. And we
have a max depth of three. Now, the max
depth, if you remember above when we
were doing the different graphs of
animals, means it's only going to go
down three layers before it stops. And
then we have minimal samples of leaves
is five. So, it's going to have at least
five leaves at the end. So, I'll have at
least three splits or have no more than
three layers and at least five end
leaves with the final result at the
bottom. Now that we've created our
decision tree classifier, not only
created it, but trained it, let's go
ahead and apply it and see what that
looks like. So, let's go ahead and make
a prediction and see what that looks
like. We're going to paste our predict
code in here. And before we run it,
let's just take a quick look at what's
doing here. We have a variable y predict
that we're going to do. And we're going
to use our variable CLF entropy that we
created.
And then you'll see predict. And it's
very common in the sklearn modules that
their different tools have the predict
when you're actually running a
prediction. In this case, we're going to
put our X test data in here. Now, if you
delivered this for use, an actual
commercial use, and distributed it, this
would be the new loans you're putting in
here to guess whether the person's going
to be uh pay them back or not. In this
case though, we need to test out the
data and just see how good our sample
is, how good of our tree does at
predicting the loan payments. And
finally, since Anaconda Jupyter notebook
is works as a command line for Python,
we can simply put the y predict en to
print it. I could just as easily have
put the print
and put brackets around y predict en to
print it out. We'll go ahead and do
that. It doesn't matter which way you do
it. And you'll see right here that it
runs a prediction. This is roughly 300
in here. Remember, it's 30% of a
thousand. So, you should have about 300
answers in here. And this tells you
which each one of those lines of ourh
test went in there. And this is what our
y predict came out. So, let's move on to
the next step where we're going to take
this data and try to figure out just how
good a model we have. So, here we go.
Since sklearn does all the heavy lifting
for you and all the math, we have a
simple line of code to let us know what
the accuracy is. And let's go ahead and
go through that and see what that means
and what that looks like. Let's go ahead
and paste this in. And let me zoom in a
little bit. There we go.
So you have a nice full picture. And
we'll see here. We're just going to do a
print accuracy is.
And then we do the accuracy score. And
this was something we imported um
earlier. If you remember at the very
beginning, let me just scroll up there
real quick so you can see where that's
coming from. That's coming from here
down here from sklearn.metrics metrics
import accuracy score. And you could
probably run a script, make your own
script to do this very easily. How
accurate is it? How many out of 300 do
we get right? And so we put in our y
test. That's the one we ran the predict
on. And then we put in our y predict en
that's the answers we got. And we're
just going to multiply that by 100
because this is just going to give us an
answer as a decimal and we want to see
it as a percentage. And let's run that
and see what it looks like. And if you
see here, we got an accuracy of
93.666667.
So when we look at the number of loans
and we look at how good our model fit,
we can tell people it has about a 93.6
fitting to it. So just a quick recap on
that. We now have accuracy set up on
here. And so we have created a model
that uses the decision tree algorithm to
predict whether a customer will repay
the loan or not. The accuracy of the
model is about 94.6%.
The bank can now use this model to
decide whether it should approve the
loan request from a particular customer
or not. And so this information is
really powerful. We may not be able to
as individuals understand all these
numbers because they have thousands of
numbers that come in, but you can see
that this is a smart decision for the
bank to use a tool like this to help
them to predict how good their uh
profit's going to be off of the loan
balances and how many are going to
default or not. We're going to be
looking at random forest, one of the
many powerful tools in the machine
learning library. Before we dive into
the topic, let's start by looking at a
few of the uses for random forest.
Currently today, it's used in remote
sensing. Uh for example, they're used in
the ETM devices. If you're a space buff,
that's the enhanced thermatic mapper
they use on satellites which see uh far
outside the human spectrum for looking
at land masses. and they acquire images
of the earth's surface. The accuracy is
higher and training time is less than
many other machine learning tools out
there. Also, object detection,
multiclass object detection is done
using random forest algorithms. A good
example is a traffic where you're trying
to sort out the different cars, buses,
and things. And it provides a better
detection in complicated environments.
They're very complicated up there. And
then we have uh another example connect.
And let's take a little closer look at
connect. Connect. They use a random
forest as part of the game console and
what it does is it tracks a body
movements and it recreates it in the
game and let's see what that looks like.
Uh we have a user who performs a step.
In this case it looks like Elvis Presley
going there that is then recorded so
that connect registers the movement and
then it marks the user based on
accuracy. And it looks like we have uh
Prince going on this one from Elvis
Presley to Prince. It's great. Uh so it
marks user base on the accuracy. If we
look at that a little closer, we have a
training set to identify body parts.
Where are the hands? Where are the feet?
Uh what's going on with the body? That
then goes into a random forest
classifier that learns from it. Once
we've trained the classifier, it then
identifies the body parts while the
person's dancing. It's able to represent
that in a computer format. And then
based on that, it scores the game and
how accurate you are as being Elvis
Presley or Prince in your dancing. So
why random forest? It's always important
to understand why we use this tool over
the other ones. What are the benefits
here? And so with the random forest, the
first one is there's no overfitting. If
you use of multiple trees, reduce the
risk of overfitting. Training time is
less. Overfitting means that we have fit
the data so close to what we have as our
sample that we pick up on all the weird
parts and instead of predicting the
overall data, you're predicting the
weird stuff which you don't want. High
accuracy runs efficiently on large
database. For large data, it produces
highly accurate predictions. In today's
world of uh big data, this is really
important. And this is probably where it
really shines. This is where Y random
forest really comes in. It estimates
missing data. Data in today's world is
very messy. So when you have a random
forest, it can maintain the accuracy
when a large proportion of the data is
missing. What that means is if you have
data that comes in from uh five or six
different areas and maybe they took one
set of statistics in one area and they
took a slightly different set of
statistics in the other. So they have
some of the sh same shared data, but one
is missing like the uh number of
children in the house if you're doing
something over demographics. and the
other one is missing the size of the
house. It will look at both of those
separately and build two different trees
and then it can do a very good job of
guessing which one fits better even
though it's missing that data. Let us
dig deep into the theory of exactly how
it works. And let's look at what is
random forest. Random forest or random
decision forest is a method that
operates by constructing multiple
decision trees. The decision of the
majority of the trees is chosen by the
random forest as the final decision. And
let's uh we have some nice graphics
here. We have a decision tree and they
actually use a real tree to denote the
decision tree which I love. And given a
random some kind of picture of a fruit.
This decision tree decides that the
output is it's an apple. And we have a
decision tree too where we have that
picture of the fruit goes in and this
one decides that it's a lemon. And the
decision three tree gets another image
and it decides it's an apple. And then
this all comes together in what they
call the random forest. And this random
forest then looks at it and says, "Okay,
I got two votes for apple, one vote for
lemon. The majority is apples. So the
final decision is apples." To understand
how the random forest works, we first
need to dig a little deeper and take a
look at the random forest and the actual
decision tree and how it builds that
decision tree. In looking closer at how
the individual decision trees work,
we'll go ahead and continue to use the
fruit example since we're talking about
trees and forests. A decision tree is a
treerehaped diagram used to determine a
course of action. Each branch of the
tree represents a possible decision,
occurrence, or reaction. So in here we
have a bowl of fruit and if you look at
that it looks like um they switch from
lemons to oranges. So we have oranges,
cherries, and apples. And the first
decision of the decision tree might be
is a diameter greater than or equal to
three. And if it says false, it knows
that they're cherries because everything
else is bigger than that. So all the
cherries fall into that decision. So we
have all that data we're training. We
can look at that. We know that that's
what's going to come up. Is the color
orange? Well, goes, hm, orange or red?
Well, if it's true, then it comes out as
the orange. And if it's false, that
leaves apples. So in this example, it
sorts out the fruit in the bowl or the
images of the fruit. A decision tree.
These are very important terms to know
because these are very central to
understanding the decision tree and when
working with them. The first is entropy.
Everything on the decision tree and how
it makes a decision is based on entropy.
Entropy is a measure of randomness or
unpredictability in the data set. uh
then they also have information gain,
the leaf node, the decision node and the
root node. We'll cover these other four
terms as we go down the tree, but let's
start with entropy. So starting with
entropy, we have here a high amount of
randomness. What that means is that
whatever is coming out of this decision,
if it was going to guess based on this
data, it wouldn't be able to tell you
whether it's a lemon or an apple. it
would just say it's a fruit. Uh so the
first thing we want to do is we want to
split this apart and we take the initial
data set. We're going to set create a
data set one and a data set two. We just
split it in two. And if you look at
these new data sets after splitting
them, the entropy of each of those sets
is much less. So for the first one,
whatever comes in there, it's going to
sort that data and it's going to say,
okay, if this data goes this direction,
it's probably an apple. And if it goes
into the other direction, it's probably
a lemon. So that brings us up to
information gain. It is the measure of
decrease in the entropy after the data
set is split. What that means in here is
that we've gone from one set which has a
very high entropy to two lower sets of
entropy and we've added in the values of
E1 for the first one and E2 for the
second two which are much lower. And so
that information gain is increased
greatly in this example. And so you can
find that the information grain simply
equals uh decision E1 minus E2. As we're
going down our list of uh definitions,
we'll look at the leaf node. And the
leaf node carries the classification or
the decision. So we look down here to
the leaf node. We finally get to our set
one or our set two. When it comes down
there and it says, "Okay, this object's
gone into set one." If it's gone into
set one, it's going to be split by some
means and we'll either end up with
apples on the leaf node or a lemon on
the leaf node. And on the right, it'll
either be an apple or lemons. Those leaf
nodes are those final decisions or
classifications. Uh that's the
definition of leaf node in here. If
we're going to have a final leaf where
we make the decision, we should have a
name for the nodes above it. And they
call those decision nodes. A decision
node. decision node has two or more
branches and you can see here where we
have the uh five apples and one lemon
and in the other case the five lemons
and one apple. They have to make a
choice of which tree it goes down based
on some kind of measurement or
information given to the tree. And that
brings us to our last definition. The
root node, the topmost decision node is
known as the root node. And this is
where you have all of your data and you
have your first decision. it has to make
or the first split in information. So
far, we've looked at a very general
image um with the fruit being split.
Let's look and see exactly what that
means to split the data and how do we
make those decisions on there. Uh let's
go in there and find out how does a
decision tree work. So let's try to
understand this and let's use a simple
example and we'll stay with the fruit.
We have a bowl of fruit and so let's
create a problem statement and the
problem is we want to classify the
different types of fruits in the bowl
based on different features. The data
set in the bowl is looking quite messy
and the entropy is high in this case. So
if this bowl was our decision maker, it
wouldn't know what choice to make. It
has so many choices. Which one do you
pick? Apple, grapes, or lemons. And so
we look in here. We're going to start
with a d a training set. So this is our
data that we're training our data with
and we have a number of options here. We
have the color and under the color we
have red yellow purple uh we have a
diameter uh 331 331 and we have a label
apple lemon grapes apple lemon grapes
and how do we split the data? We have to
frame the conditions to split the data
in such a way that the information gain
is the highest. It's very key to note
that we're looking for the best gain. We
don't want to just start sorting out the
smallest piece in there. We want to
split it the biggest way we can. And so
we measure this decrease in entropy.
That's what they call it, entropy.
There's our entropy after splitting. And
now we'll try to choose a condition that
gives us the highest gain. We will do
that by splitting the data using each
condition and checking the gain that we
get out of them. The conditions that
give us the highest gain will be used to
make the first split. So let's take a
look at these different conditions. We
have color, we have diameter, and if we
look underneath that, we have a couple
different values. is we have diameter
equals 3, color equals yellow, red,
diameter equals 1. And when we look at
that, you'll see over here we have 1 2 3
4 threes. That's a pretty hardy
selection. So let's say the condition
gives us a maximum gain of three. So we
have the most pieces fall into that
range. So our first split from our
decision node is we split the data based
on the diameter. Is it greater than or
equal to three? If it's not, that's
false. It goes into the grape bowl. And
if it's true, it goes into a bowl fold
of lemon and apples. The entropy after
splitting has decreased considerably. So
now we can make two decisions. If you
look at they're very uh much less chaos
going on there. This node has already
attain an entropy value of zero. As you
can see, there's only one kind of label
left for this branch. So no further
splitting is required for this node.
However, this node on the right is still
requires a split to decrease the entropy
further. So, we split the right node
further based on color. If you look at
this, if I split it on color, that
pretty much cuts it right down the
middle. And it's the only thing we have
left in our choices of color and
diameter, too. And if the color is
yellow, it's going to go to the right
bowl. And if it's false, it's going to
go to the left bowl. So, the entropy in
this case is now zero. So, now we have
three bowls with zero entropy. There's
only one type of data in each one of
those bowls. So, we can predict a lemon
with 100% accuracy. And we can predict
the apple also with 100% accuracy along
with our grapes up there. So, we've
looked at kind of a basic tree in our
forest. But what we really want to know
is how does a random forest work as a
whole. So to begin our um random forest
classifier, let's say we already have
built three trees. And we're going to
start with the first tree that looks
like this. Just like we did in the
example, this tree looks at the
diameter. If it's greater than or equal
to three, it's true. Otherwise, it's
false. So one side goes to the smaller
diameter, one side goes to larger
diameter. And if the color is orange,
it's going to go to the right. True.
We're using oranges now instead of
lemons. And if it's red, it's going to
go to the left. False. We build a second
tree very similar, but it's split
differently. Instead of the first one
being split by a diameter, uh this one
when they created it, if you look at
that first bowl, it has a lot of red
objects. So it says, is the color red?
Because that's going to bring our
entropy down the fastest. And so, of
course, if it's true, it goes to the
left. If it's false, it goes to the
right. And then it looks at the shape,
false or true, and so on and so on. And
tree three is the diameter equal to one.
And it came up with this because there's
a lot of cherries in this bowl. So that
would be the biggest split on there is
is the diameter equal to one. That's
going to drop the entropy the quickest.
And as you can see, it splits it into
true. If it goes false, and they've
added another category, does it grow in
the summer? And if it's false, it goes
off to the left. If it's true, it goes
off to the right. Let's go ahead and
bring these three trees so you can see
them all in one image. So this would be
three completely different trees
categorizing a fruit. And let's take a
fruit. Now let's try this. And this
fruit, if you look at it, we've
blackened it out. You can't see the
color on it. So it's missing data.
Remember one of the things we talked
about earlier is that a random forest
works really good if you're missing
data, if you're missing pieces. So this
fruit has an image, but maybe it's a
person had a black and white camera when
they took the picture. And we're going
to take a look at this. And it's going
to have um they put the color in there,
so ignore the color down there. But the
diameter equals three. We find out it
grows in the summer equals yes. And the
shape is a circle. And if you go to the
right, you can look at what one of the
decision trees did. This is the third
one. Is the diameter greater than equal
to three? Is a color orange? Well, it
doesn't really know on this one, but it
if you look at the value, it' say true,
and it go to the right. Tree two
classifies it as cherries. Is a color
equal red? Is the shape a circle? True.
It is a circle. So, this would look at
it and say, "Oh, that's a cherry." And
then we go to the other classifier and
it says, "Is the diameter equal one?"
Well, that's false. Does it grow in the
summer? True. So, it goes down and looks
at as oranges. So, how does this random
forest work? The first one says it's an
orange. The second one said it was a
cherry. And the third one says, hm, it's
an orange. And you can guess that if you
have two oranges and one says it's a
cherry, uh, when you add that all
together, the majority of the vote says
orange. So, the answer is it's
classified as an orange, even though we
didn't know the color and we're missing
data on it. I don't know about you, but
I'm getting tired of fruit. So, let's
switch. And I did promise you we'd start
looking at a case example and get into
some Python coding. Today, we're going
to use the case the iris flower
analysis.
This is the exciting part as we roll up
our sleeves and actually look at some
Python coding. Before we start the
Python coding, we need to go ahead and
create a problem statement. Wonder what
species of iris do these flowers belong
to? Let's try to predict the species of
the flowers using machine learning in
Python. Let's see how it can be done. So
here we begin to go ahead and implement
our Python code. And you'll find that
the first half of our implementation is
all about organizing and exploring the
data coming in. Let's go ahead and take
this first step, which is loading the
different modules into Python. And let's
go ahead and put that in our favorite
editor, whatever your favorite editor
is. In this case, I'm going to be using
the Anaconda Jupiter Notebook, which is
one of my favorites. Certainly, there's
Notepad++ and Eclipse and dozens of
others, or just even using the Python
terminal window. any of those will work
just fine to go ahead and explore this
Python coding. So, here we go. Let's go
ahead and flip over to our Jupyter
notebook. And I've already opened up a
new page for Python 3 code. And I'm just
going to paste this right in there. And
let's take a look and see what we're
bringing into our Python. The first
thing we're going to do is from the
sklearn.data sets import load iris. Now,
this isn't the actual data. So this is
just the module that allows us to bring
in the data, the load iris. And the iris
is so popular. It's been around since
1936 when Ronald Fiser published a paper
on it. And they're measuring the
different parts of the flower. And based
on those measurements, predicting what
kind of flower it is. And then if we're
going to do a random forest classifier,
we need to go ahead and import a random
forest classifier from the sklearn
module. So sklearn.semble
import random forest classifier. And
then we want to bring in two more
modules. Um, and these are probably the
most commonly used modules in Python and
data science with any of the um, other
modules that we bring in. And one is
going to be pandas. We're going to
import pandas as pd. PD is the common
term used for pandas. And pandas is
basically creates a data format for us
where when you create a pandas data
frame, it looks like an Excel
spreadsheet. And you'll see that in a
minute when we start digging deeper into
the code. Panda is just wonderful
because it plays nice with all the other
modules in there. And then we have
Numpy, which is our numbers Python. And
the numbers Python allows us to do
different mathematical sets on here.
We'll see right off the bat, we're going
to take our NP and we're going to go
ahead and seed the randomness with it
with zero. So NP.random seed is seeding
that as zero. This code doesn't actually
show anything. We're going to go ahead
and run it because I need to make sure I
have all those loaded. And then let's
take a look at the next module on here.
The next six slides, including this one,
are all about exploring the data.
Remember, I told you half of this is
about looking at the data and getting it
all set. So, let's go ahead and take
this code right here, the script, and
let's get that over into our Jupyter
notebook. And here we go. We've gone
ahead and uh run the imports. Now I'm
going to paste the code down here
and let's take a look and see what's
going on. The first thing we're doing is
we're actually loading the iris data.
And if you remember up here, we loaded
the module that tells it how to get the
iris data. Now we're actually assigning
that data to the variable iris. And then
we're going to go ahead and use the df
to define dataf frame. And that's going
to equal pd. And if you remember that's
pandas as pd. So that's our pandas and
panda dataf frame. And then we're
looking at iris data and columns equals
iris feature names. And we're going to
do the DF head. And let's run this so
you can understand what's going on here.
The first thing you want to notice is
that our DF has created uh what looks
like an Excel spreadsheet. And in this
Excel spreadsheet, we have set the
columns. So up on the top, you can see
the four different columns. And then we
have the data iris.data down below. It's
a little confusing without knowing where
this data is coming from. So let's look
at the bigger picture and I'm going to
go print. I'm just going to change this
for a moment and we're going to print
all of Iris and see what that looks
like. So when I print all of Iris I get
this long list of information. And you
can scroll through here and see all the
different titles on there. What's
important to notice is that first off
there's a brackets at the beginning. So
this is a Python dictionary
and in a Python dictionary you'll have a
key or a label and this label pulls up
whatever information comes after it. So
feature names which we actually used
over here under columns is equal to an
array of sele length sele width pedal
length pedal width. These are the
different names they have for the four
different columns. And if you scroll
down far enough you'll also see data
down here. Oh goodness, it came up right
towards the top. And uh data is equal to
the different data we're looking at.
Now, there's a lot of other things in
here like target. We're going to be
pulling that up in a minute. And there's
also the names uh the target names which
is further down. And we'll show you that
also in a minute. Let's go ahead and set
that back to the head. And this is one
of the neat features of pandas and panda
dataf frames is when you do df.ad or the
panda dataf frame. head. It'll print the
first five lines of the data set in
there along with the headers if you have
them. In this case, we have the column
headers set to iris features. And in
here, you'll see that we have 0 1 2 3 4.
In Python, most arrays always start at
zero. So, when you look at the first
five, it's going to be 0 1 2 3 4, not 1
2 3 4 5. So, now we've got our iris data
imported into a data frame. Let's take a
look at the next piece of code in here.
And so in this section here of the code,
we're going to take a look at the
target. And let's go ahead and get this
into our notebook, this piece of code,
so we can discuss it a little bit more
in detail. So here we are in our Jupyter
notebook. I'm going to put the code in
here. And before I run it, I want to
look at a couple things going on. So we
have uh DF species. And this is
interesting because right here you'll
see where I have DF species in brackets
which is uh the key code for creating
another column. And here we have
iris.target.
Now these are both in the pandas setup
on here. So in pandas we can do either
one. I could have just as easily done
iris and then in brackets target
depending on what I'm working on. Both
are um acceptable. Let's go ahead and
run this code and see how this changes.
And what we've done is we've added the
target from the iris data set as another
column on the end.
Now what species is this is what we're
trying to predict. So we have our data
which tells us the answer for all these
different pieces. And then we've added a
column with the answer. So that way when
we do our final setup, we'll have the
ability to program our our neural
network to look for these this different
data and know what a satossa is or a
veraricolor which we'll see in just a
minute or virginica. Those are the three
that are in there. And now we're going
to add one more column. I know we're
organizing all this data over and over
again. It's kind of fun. There's a lot
of ways to organize it. What's nice
about putting everything onto one data
frame is I can then do a print out and
it shows me exactly what I'm looking at.
And I'll show you where you where that's
different where you can alter that and
do it slightly differently. But let's go
ahead and put this into our script up to
now. And here we go. We're going to put
that down here and we're going to run
that. And let's talk a little bit about
what we're doing. Now we're exploring
data. And one of the challenges is
knowing how good your model is. Did your
model work? And to do this, we need to
split the data. And we split it into two
different parts. They usually call it
the training and the testing. And so in
here, we're going to go ahead and put
that in our database so you can see it
clearly. And we've set it df. And
remember, you can put brackets. This is
creating another column. Is train. So
we're going to use part of it for
training. And this equals np. Remember
that stands for numpy.random.uniform.
So we're generating a random number
between zero and one. And we're going to
do it for each of the rows. That's where
the length df comes from. So each row
gets a generated number. And if it's
less than 75, it's true. And if it's
greater than 75, it's false. This means
we're going to take 75% of the data
roughly because there's a randomness
involved. And we're going to use that to
train it. And then the other 25% we're
going to hold off to the side and use
that to test it later on. So let's flip
back on over and see what the next step
is. So now that we've labeled our
database for which is training and which
is testing, let's go ahead and sort that
into two different variables, train and
test. And let's take this code and let's
bring it into our project. And here we
go. Let's paste it on down here. And
before I run this, let's just take a
quick look at what's going on here. is
we have up above we created remember
there's our def head which prints the
first five rows and we've added a column
is train at the end and so we're going
to take that we're going to create two
variables we're going to create two new
data frames one's called train one's
called test 75% in train 25% in test and
then to sort that out we're going to do
that by doing df our main original data
frame with the iris data in it and if df
F is train equals true, that's going to
go in the train. And if DF is train
equals false, it goes in the test. And
so when I run this, we're going to print
out the number in each one. Let's see
what that looks like. And you'll see
that it puts 118 in the training module
and it puts 32 in the testing module,
which lets us know that there was 150
lines of data in here. So if you went
and looked at the original data, you
could see that there's 150 lines and
that's roughly 75% in one and 25% for us
to test our model on afterward. So let's
jump back to our code and see where this
goes. In the next two steps, we want to
do one more thing with our data, and
that's make it readable to humans. Um, I
don't know about you, but I hate looking
at zeros and ones. So, let's start with
the features and let's go ahead and take
those and make those readable to humans
and let's put that in our code.
Let's see. Here we go. Paste it in. And
you'll see here we've done a couple very
basic things. We know that the columns
in our data frame, again, this is a
panda thing, the DF columns, and we know
the first four of them, 01, 2, 3, that'd
be the first four are going to be the
features or the titles of those columns.
And so when I run this, you'll see down
here that it creates an index, sea
length, sea width, pedal length, and
pedal width. And this should be familiar
because if you look up here, here's our
column titles going across. And here's
the first four. One thing I want you to
notice here is that when you're in a
command line, whether it's Jupyter
notebook or you're running command line
in the uh terminal window, if you just
put the name of it, it'll print it out.
This is the same as doing print
features.
And the shortand is you just put
features in here. If you're actually
writing a code and saving the script and
running it by remote, you really need to
put the print in there. But for this,
when I run it, you'll see it gives me
the same thing.
But for this, we want to go ahead and
we'll just leave it as features because
it doesn't really matter. And this is
one of the fun thing about Jupyter
Notebooks is I'm just building the code
as we go. And then we need to go ahead
and create the labels for the other
part. So, let's take a look and see what
that for. Our final step in prepping our
data before we actually start running
the training and the testing is we're
going to go ahead and convert the
species on here into something the
computer understands. So, let's put this
code into our script and see where that
takes us.
All right, here we go. We've set y equal
to pd.factorize
train species of zero. So, let's break
this down just a little bit. We have our
pandas right here. PD factoriize. What
is factorized doing? I'm going to come
back to that in just a second. Let's
look at what train species is and why
we're looking at the group zero on
there. And let's go up here. And here is
our species.
Remember this on that? We created this
whole column here for species. And then
it has satossa, satossa, satossa,
satossa. And if you scroll down enough,
you'd also see virginica and
veraricolor.
We need to convert that into something
the computer understands. Zeros and
ones. So the train species of zero
because this is in the format of a of an
array of arrays. So you have to have the
zero on the end. And then species is
just that column. Factoriize goes in
there and looks at the fact that there's
only three of them. So when I run this,
you'll see that Y generates an array
that's equal to, in this case, it's the
training set, and it's zeros, ones, and
twos representing the three different
kinds of flowers we have. So now we have
something the computer understands, and
we have a nice table that we can read
and understand. And now finally we get
to actually start doing the predicting.
So here we go. Uh we have two lines of
code. Oh my goodness, that was a lot of
work to get to two lines of code. But
there is a lot in these two lines of
code. So let's take a look and see
what's going on here and put this into
our full script that we're running. And
let's paste this in here. And let's take
a look and see what this is. We have
we're creating a variable CLF. And we're
going to set this equal to the random
forest classifier. And we're passing two
variables in here. And there's a lot of
variables you can play with. As far as
these two are concerned, they're very
standard. In jobs, all that does is to
prioritize it. Not something to really
worry about. Usually when you're doing
this on your own computer, you do end
jobs equals 2. If you're working in a
larger or big data and you need to
prioritize it differently, this is what
that number does is it changes your
priorities and how it's going to run
across the system and things like that.
And then the random state is just how it
starts. Zero is fine for here.
But uh let's go ahead and run this.
We also have clf.fit train features, y.
And before we run it, let's talk about
this a little bit more. CLF.fit.
So, we're fitting, we're training it. We
are actually creating our random forest
classifier right here. This is the code
that does everything. And we're going to
take our training set. Remember, we kept
our test off to the side. And we're
going to take our training set with the
features. And then we're going to go
ahead and put that in. And here's our
target, the Y. So, the Y is 0, 1, and
two that we just created. And the
features is the actual data going in
that we put into the training set. And
let's go ahead and run that.
And this is kind of an interesting thing
because it printed out the random force
classifier
and everything around it. And so when
you're running this in your terminal
window or in a script like this, this
automatically treats this like just like
when we were up here and I typed in y
and it printed out y instead of print y.
This does the same thing. It treats this
as a variable and prints it out. But if
you were actually running your code,
that wouldn't be the case. And what is
printed out is it shows us all the
different variables we can change. And
if we go down here, you can actually see
in jobs equals 2. You can see the random
state equals zero. Those are the two
that we sent in there. You would really
have to dig deep to find out all these
different meanings of all these
different settings on here. Some of them
are self-explanatory if you kind of
think about it a little bit. Like max
features is auto. So all the features
that we're putting in there, it's just
going to automatically take all four of
them. Whatever we send it, it'll take.
Some of them might have so many features
because you're processing words. There
might be like 1.4 million features in
there because you're doing legal
documents and that's how many different
words are in there. At that point, you
probably want to limit the maximum
features that you're going to process.
And leaf nodes, that's the end nodes.
Remember, we had the fruit and we're
talking about the leaf nodes. Like I
said, there's a lot in this. We're
looking at a lot of stuff here. So you
might have uh in this case there's
probably only think three leaf nodes,
maybe four. You might have thousands of
leaf nodes at which point you do need to
put a cap on that and say, "Okay, you
can only go so far and then we're going
to use all of our resources on
processing this." And that really is
what most of these are about is limiting
the process and making sure we don't uh
overwhelm a system. And there's some
other settings in here. Again, we're not
going to go over all of them. Warm start
equals false. or start as if you're
programming it one piece at a time
externally since we're not we're not
going to have like we're not going to
continually to train this particular
learning tree and again like I said
there's a lot of things in here that
you'll want to look up more detail from
the sklearn and if you're digging in
deep and running a major project on here
for today though all we need to do is
fit or train our features and our target
Y. So now we have our training model.
What's next? If we're going to create a
model,
we now need to test it. Remember, we set
aside the test features, test group, 25%
of the data. So let's go ahead and take
this code and let's put it into our uh
script and see what that looks like.
Okay, here we go. And we're going to run
this.
And it's going to come out with a bunch
of zeros, ones, and twos, which
represents the three type of flowers,
the satossa, the virginica, and the
versa color. And what we're putting into
our predict is the test features. And I
always kind of like to know what it is I
am looking at. So, real quick, we're
going to do test
features. And remember, features is an
array
of sele
width, pedal length, pedal width. So
when we put it in this way, it actually
loads all these different columns that
we loaded into features. So if we did
just features, let me just do features
in here so you can see what features
looks like. This is just playing with
the with Panda's data frames. You'll see
that it's an index. So when you put an
index in like this
into test features into test, it then
takes those columns and creates a Panda
data frames from those columns. And in
this case, we're going to go ahead and
put those into our predict. So, we're
going to put each one of these lines of
data, the 5.0, 3.4, 1.5, point2, and
we're going to put those in, and we're
going to predict what our new um forest
classifier is going to come up with. And
this is what it predicts. It predicts uh
0000121122.
and and uh again this is the flower type
satossa vica and versa color. So now
that we've taken our test features let's
explore that. Let's see exactly what
that data means to us. So the first
thing we can do with our predicts is we
can actually generate a different
prediction model. When I say different,
we're going to view it differently. It's
not that the data itself is different.
So let's take this next piece of code
and put it into our script.
So we're pasting it in here and you'll
see that we're doing uh predict and
we've added underscore proba for
probability. So there's our clff.predict
probability. So we're we're running it
just like we ran it up here, but this
time with this we're going to get a
slightly different result and we're only
going to look at the first 10. So you'll
see down here instead of looking at all
of them uh which was uh what 27 you'll
see right down here that this generates
a much larger field on the probability
and let's take a look and see what that
looks like and what that means. So when
we do the predict underscore probaba for
probability it generates three numbers.
So we had three leaf nodes at the end
and if you remember from all the theory
we did this is the predictors. The first
one is predicting a one for satossa. It
predicts a zero for virginica. And it
predicts a zero for versol. And so on
and so on and so on. And let's um you
know what? I'm going to change this just
a little bit. Let's look at 10
to 20 just because we can.
And we start to get in a little
different of data. And you'll see right
down here it gets to this one. This line
right here. And this line has zero 0.5
0.5.
And so if we're going to vote and we
have two equal votes, it's going to go
with the first one. So it says uh
Satossa gets zero votes, virginica
gets.5 votes, VersaColor gets.5 votes,
but let's just go with the virginica
since these two are equal and so on and
so on down the list. You can see how
they vary on here. So now we've looked
at both how to do a basic predict of the
features and we've looked at the predict
probability. Let's see what's next on
here. So now we want to go ahead and
start mapping names for the plants. We
want to attach names so that it makes a
little more sense for us. And that's
what we're going to do in these next two
steps. We're going to start by setting
up our predictions and mapping them to
the name. So let's see what that looks
like. And let's go ahead and paste that
code in here and run it. And this goes
along with the next piece of code. So
we'll skip through this quickly and then
come back to it a little bit. So, here's
iris.target
names.
And uh if you remember correctly, this
was the the names that we've been
talking about this whole time, the
Satossa, Vica, VersaColor. And then
we're going to go ahead and do the
prediction again. We've run it. We could
have just set a variable equal to this
instead of rerunning it each time, but
we're going ahead and run it again.
CLF.predict test features. Remember that
returns the zeros, the ones, and the
twos. And then we're going to set that
equal to predictions. So this time we're
actually putting it in a variable. And
when I run this,
it distributes and it comes out as an
array. And the array is satossa,
satossa, satossa, satossa, satossa.
We're only looking at the first five. We
could actually do let's do the first 25
just so we can see a little bit more on
there. And you'll see that it starts
mapping it to all the different flower
types, the versa color and the virginica
in there. And let's see how this goes
with the next one. So, let's take a look
at the top part of our species in here.
And we'll take this code and put it in
our script.
And let's put that down here and paste
it. There we go. And we'll go ahead and
run it. And let's talk about both these
sections of code here and how they go
together. The first one is our
predictions. And I went ahead and did uh
predictions through 25. Let's just do
five.
And so we have stosis, satossis, stosis,
satossis. That's what we're predicting
from our test model. And then we come
down here and we look at test species.
And remember, I could have just done
test.species.head.
And you'll see it says Satossa, Satossa,
Satossa, Satossa. And they match. So the
first one is what our forest is doing
and the second one is what the actual
data is. Now is we need to combine these
so that we can understand what that
means. We need to know how good our
forest is, how good it is at predicting
the features. So that's where we come up
to the next step, which is lots of fun.
We're going to use a single line of code
to combine our predictions and our
actuals so we have a nice chart to look
at. And let's go ahead and put that in
our script in our Jupyter notebook here.
Let's see. Let's go ahead and paste that
in. And then I'm going to because I'm on
the Jupyter notebook, I can do a control
minus so we can see the whole line
there.
There we go. resize it and let's take a
look and see what's going on here. We're
going to create in pandas. Remember PD
stands for pandas and we're doing a
cross tab. This function takes two sets
of data and creates a chart out of them.
So when I run it, you'll get a nice
chart down here. And we have the
predicted species.
So across the top you'll see the satossa
versus color virginica and the actual
species satossa versus color virginica.
And so the way to read this chart and
let's go ahead and take a look on how to
read this chart here. When you read this
chart, you have satossa where they meet,
you have versolar where they meet, and
you have virginica where they meet. And
they're meeting where the actual and the
predicted agree. So this is the number
of accurate predictions. So in this
case, it equals 30. If you add 13 + 5 +
12, you get 30. And then we notice here
where it says virginica, but it was
supposed to be versol. This is
inaccurate. So now we have two two
inaccurate predictions and 30 accurate
predictions. So we'll say that the model
accuracy is 93. That's just 30 divided
by 32. And if we multiply it by 100, we
can say that it is 93% accurate. So we
have a 93% accuracy with our model. I
did want to add one more quick thing in
here on our scripting before we wrap it
up. So let's flip back on over to my
script. in here. We're going to take
this uh line of code from up above. I
don't know if you remember it, but
predicts equals the iris.target_names.
So, we're going to map it to the names
and we're going to run the prediction.
And we read it on test features. But,
you know, we're not just testing it. We
want to actually deploy it. So, at this
point, I would go ahead and change this.
And this is an array of arrays. This is
really important when you're running
these to know that. So, you need the
double brackets. And I could actually
create data. Maybe let's let's just do
two flowers. So maybe I'm processing
more data coming in. And we'll put two
flowers in here. And then uh I actually
want to see what the answer is. So let's
go ahead and type in PRS and print that
out. And when I run this, you'll see
that I've now predicted two flowers that
maybe I measured in my front yard as
VersaColor and VersaColor.
Not surprising since I put the same data
in for each one. This would be the
actual uh end product going out to be
used on data that you don't know the
answer for.
So that's going to conclude our
scripting part of this. Introducing
naive base classifier. Have you ever
wondered how your mail provider
implements spam filtering or how online
news channels perform news text
classification or how companies perform
sentimental analysis of their audience
on social media? All of this and more is
done through a machine learning
algorithm called naive bay classifier.
Welcome to Naive Bay tutorial. My name
is Richard Kersner. I'm with the
SimplyLearn team. That's
www.simplearn.com.
Get certified get ahead. What's in it
for you? We'll start with what is naive
bays? A basic overview of how it works.
We'll get into naive bays and machine
learning where it fits in with our other
machine learning tools. Why do we need
naive bays and understanding naive bays
classifier a much more in-depth of how
the math works in the background?
Finally, we'll get into the advantages
of the naive bay classifier in the
machine learning setup. And then we'll
roll up our sleeves and do my favorite
part. We'll actually do some Python
coding and do some text classification
using the naive bays. What is naive
bays? Let's start with a basic
introduction to the bay theorem named
after Thomas Bae from the 1700s who
first coined this in the western
literature. Naive bay classifier works
on the principle of conditional
probability as given by the bay theorem.
Before we move ahead, let us go through
some of the simple concepts in the
probability that we will be using. Let
us consider the following example of
tossing two coins. Here we have two
quarters and if we look at all the
different possibilities of what they can
come up as, we get that they could come
up as head heads. come up as head, tail,
tail, head and tell tail. When doing the
math on probability, we usually denote
probability as a P, a capital P. So the
probability of getting two heads equals
1/4. You can see in our data set, we
have two heads and this occurs once out
of the four possibilities. And then the
probability of at least one tail occurs
three/arters of the time. You'll see on
three of the coin tosses, we have tails
in them. And out of four, that's
three/4s. And then the probability of
the second coin being a head given the
first coin is tail is 1/2. And the
probability of getting two heads given
the first coin is a head is 1/2. We'll
demonstrate that in just a minute and
show you how that math works. Now when
we're doing it with two coins, it's easy
to see. But when you have something more
complex, you can see where these pro
these formulas really come in and work.
So the base theorem gives us the
conditional probability of an event A
given another event B has occurred. In
this case, the first coin toss will be B
and the second coin toss A. This could
be confusing because we've actually
reversed the order of them and go from B
to A instead of A to B. You'll see this
a lot when you work in probabilities.
The reason is we're looking for event A,
we want to know what that is. So, we're
going to label that A since that's our
focus. And then given another event B
has occurred. In the Baze theorem, as
you can see on the left, the probability
of A occurring given B has occurred
equals the probability of B occurring
given A has occurred times the
probability of A over the probability of
B. This simple formula can be moved
around just like any algebra formula.
And we could do the probability of A
after given B times probability of B
equals the probability of B given A
times probability of A. You can easily
move that around and multiply it and
divide it out. Let us apply B theorem to
our example. Here we have our two
quarters and we'll notice that the first
two probabilities of getting two heads
and at least one tail we compute
directly off the data. So you can easily
see that we have one example hh out of
four 1/4 and we have three with tails in
them giving us three quarters or 3/4
75%. The second condition the second uh
set three and four we're going to
explore a little bit more in detail.
Now, we stick to a simple example with
two coins because you can easily
understand the math. The probability of
throwing a tail doesn't matter what
comes before it. And the same with the
heads. So, it's still going to be 50% or
1/2. But when that come when that
probability gets more complicated, let's
say you have a d6 dice or some other
instance, then this formula really comes
in handy. But let's stick to the simple
example for now. In this sample space,
let A be the event that the second coin
is head and b be the event that the
first coin is tails. Again, we reversed
it because we want to know what the
second event's going to be. So, we're
going to be focusing on A. And we write
that out as the probability of A given
B. And we know this from our formula
that that equals the probability of B
given A times the probability of A over
the probability of B. And when we plug
that in, we plug in the probability of
the first coin being tails given the
second coin is heads and the probability
of the second coin being heads given the
first coin being over the probability of
the first coin being tails. When we plug
that data in and we have the probability
of the first coin being tails given the
second coin is heads times the
probability of the second coin being
heads over the probability of the first
coin being tails. You can see it's a
simple formula to calculate. We have 1/2
* 1/2 over 1/2 or 1/2 =.5 or 1/4. So the
B theorem basically calculates the
conditional probability of the
occurrence of an event based on prior
knowledge of conditions that might be
related to the event. We will explore
this in detail when we take up an
example of online shopping further in
this tutorial. Understanding naive bays
and machine learning. Like with any of
our other machine learning tools, it's
important to understand where the naive
bays fits in the hierarchy. So under the
machine learning, we have supervised
learning and there is other things like
unsupervised learning. There's also
reward system. This falls under the
supervised learning. And then under the
supervised learning, there's
classification. There's also regression.
But we're going to be in the
classification side. And then under
classification is your naive bays. Let's
go ahead and glance into where is naive
bays used. Let's look at some of the use
scenarios for it. As a classifier, we
use it in face recognition. Is this
Cindy or is it not Cindy or whoever? Or
it might be used to identify parts of
the face that they then feed into
another part of the face recognition
program. This is the eye. This is the
nose. This is the mouth. Weather
prediction. Is it going to be rainy or
sunny? Medical recognition. News
prediction. It's also used in medical
diagnosis. We might diagnose somebody as
either as high risk or not as high risk
for cancer or heart disease or other
ailments. And news classification you
look at the Google news and it says well
is this political or is this world news
or a lot of that's all done with the
naive bays. Understanding naive bay
classifier. Now we already went through
a basic understanding with the coins and
the two heads and two tails and head
tail tail heads etc. We're going to do
just a quick review on that and remind
you that the naive bay classifier is
based on the bay theorem which gives a
conditional probability of event A given
event B. And that's where the
probability of A given B equals the
probability of B given A times
probability of A over probability of B.
Remember this is an algebraic function
so we can move these different entities
around. We could multiply by the
probability of B. So it goes to the left
hand side and then we could divide by
the probability of A given B and just as
easily come up with a new formula for
the probability of B. To me staring at
these algebraic functions kind of gives
me a slight headache. It's a lot better
to see if we can actually understand how
this data fits together in a table. And
let's go ahead and start applying it to
some actual data so you can see what
that looks like. So, we're going to
start with the shopping demo problem
statement. And remember, we're going to
solve this first in a table form so you
can see what the math looks like. And
then we're going to solve it in Python.
And in here, we want to predict whether
the person will purchase a product. Are
they going to buy or don't buy? Very
important. If you're running a business,
you want to know how to maximize your
profits or at least maximize the
purchase of the people coming into your
store. And we're going to look at a
specific combination of different
variables. In this case, we're going to
look at the day, the discount, and the
free delivery. And you can see here
under the day we want to know whether
it's uh on the weekday, you know,
somebody's working, they come in after
work or maybe they don't work. Weekend,
you can see the bright colors coming
down there celebrating not being in work
or holiday. And did we offer a discount
that day? Yes or no. Did we offer free
delivery that day? Yes or no. And from
this, we want to know whether the
person's going to buy based on these
traits so we can maximize them and find
out the best system for getting somebody
to come in and purchase our goods and
products from our store. Now, having a
nice visual is great, but we do need to
dig into the data. So, let's go ahead
and take a look at the data set. We have
a small sample data set of 30 rows.
We're showing you the first 15 of those
rows for this demo. Now, the actual data
file you can request. Just type in below
under the comments on the YouTube video
and we'll send you some more information
and send you that file. As you can see
here, the file is very simple columns
and rows. We have the day, the discount,
the free delivery, and did the person
purchase or not. And then we have under
the day whether it was a weekday, a
holiday, was it the weekend? This is a
pretty simple set of data. And long
before computers, people used to look at
this data and calculate this all by
hand. So let's go ahead and walk through
this and see what that looks like when
we put that into tables. Also note in
today's world, we're not usually looking
at three different variables and 30
rows. Nowadays, because we're able to
collect data so much, we're usually
looking at 27, 30 variables across
hundreds of rows. The first thing we
want to do is we're going to take this
data and uh based on the data set
containing our three inputs day,
discount, and free delivery, we're going
to go ahead and populate that to
frequency tables for each attribute. So,
we want to know if they had a discount,
how many people buy and did not buy. Uh
did they have a discount? Yes or no. Do
we have a free delivery? Yes or no. On
those days, how many people made a
purchase and how many people didn't? And
the same with the three days of the
week. Was it a weekday, a weekend, a
holiday? And did they buy? Yes or no? As
we dig in deeper to this table for our
bay theorem, let the event buy be a. Now
remember when we looked at the coins, I
said we really want to know what the
outcome is. Did the person buy or not?
And that's usually event A is what
you're looking for. And the independent
variables, discount, free delivery, and
day be B. So we'll call that probability
of B. Now let us calculate the
likelihood table for one of the
variables. Let's start with day, which
includes weekday, weekend, and holiday.
And let us start by summing all of our
rows. So, we have the uh weekday row,
and out of the weekdays, there's 9 plus
2, so there's 11 weekdays. There's eight
weekend days and 11 holidays. Wow,
that's a lot of holidays. And then we
want to sum up the total number of days.
So, we're looking at a total of 30 days.
Let's start pulling some information
from our chart and see where that takes
us. And when we fill in the chart on the
right, you can see that nine out of 24
purchases are made on the weekday, 7 out
of 24 purchases on the weekend, and
eight out of 24 purchases on a holiday.
And out of all the people who come in,
24 out of 30 purchase. You can also see
how many people do not purchase. On the
weekday, it's two out of six didn't
purchase and so on and so on. We can
also look at the totals and you'll see
on the right, we put together some of
the formulas. The probability of making
a purchase on the weekend comes out 11
out of 30. So out of the 30 people who
came into the store throughout the
weekend, weekday and holiday, 11 of
those purchases were made on the
weekday. And then you can also see the
probability of them not making a
purchase. And this is done for doesn't
matter which day of the week. So we call
that probability of no buy would be 6
over 30 or 0.2. So there's a 20% chance
that they're not going to make a
purchase no matter what day of the week
it is. And finally, we look at the
probability of B if A. In this case,
we're going to look at the probability
of the weekday and not buying. Two of
the no buys were done out of the weekend
out of the six people who did not make
purchases. So when we look at that,
probability of the week day without a
purchase is going to be.33 or 33%. Let's
take a look at this at different
probabilities. And uh based on this
likelihood table, let's go ahead and
calculate conditional probabilities as
below. The first three we just did. The
probability of making a purchase on the
weekday is 11 out of 30 or roughly 36 or
37%
367. The probability of not making a
purchase at all doesn't matter what day
of the week is roughly.2 or 20%. And the
probability of a weekday no purchase is
roughly two out of six. So two out of
six of our no purchases were made on the
weekday. And then finally we take our P
of A. If you looked we've kept the
symbols up there. So we got P of
probability of B, probability of A,
probability of B if A. We should
remember that the probability of A if B
is equal to the first one times the
probability of no per buys over the
probability of the weekday. So we could
calculate it both off the uh table we
created. We can also calculate this by
the formula and we get the.367
which equals or.33
*2 over.367 which equals.179
or roughly uh 17 to 18%. And that'd be
the probability of no purchase done on
the weekday. And this is important
because we can look at this and say as
the probability of buying on the weekday
is more than the probability of not
buying on the weekday, we can conclude
that customers will most likely buy the
product on a weekday. Now, we've kept
our chart simple and we're only looking
at one aspect. So, you should be able to
look at the table and come up with the
same information or the same conclusion.
That should be kind of intuitive at this
point. Next, we can take the same setup.
We have the frequency tables of all
three independent variables. Now we can
construct the likelihood tables for all
three of the variables we're working
with. We can take our day like we did
before. We have weekday, weekend, and
holiday. And we filled in this table.
And then we can come in and also do that
for the discount. Yes or no. Did they
buy? Yes or no. And we fill in that full
table. So now we have our probabilities
for a discount and whether the discount
leads to a purchase or not. And the
probability for free delivery. Does that
lead to a purchase or not? And this is
where it starts getting really exciting.
Let us use these three likelihood tables
to calculate whether a customer will
purchase a product on a specific
combination of day, discount, and free
delivery or not purchase. Here, let us
take a combination of these factors. Day
equals holiday, discount equals yes,
free delivery equals yes. Let's dig
deeper into the math and actually see
what this looks like. And we're going to
start with looking for the probability
of them not purchasing on the following
combinations of days. We are actually
looking for the probability of A equal
no buy. No purchase. And our probability
of B we're going to set equal to is it a
holiday? Did they get a discount? Yes.
And was it a free delivery? Yes. Before
we go further, let's look at the
original equation. the probability of a
if b equals the probability of b given
the condition a and the probability
times probability of a over the
probability of b occurring. Now this is
basic algebra so we can multiply this
information together. So when you see
the probability of a given b in this
case the condition is b c and d or the
three different variables we're looking
at. And when you see the probability of
B, that would be the conditions. We're
actually going to multiply those three
separate conditions out. Probability of
you'll see that in just a second in the
formula times the full probability of A
over the full probability of B. So here
we are back to this and we're going to
have let A equal no purchase. And we're
looking for the probability of B on the
condition A where A sets for three
different things. Remember that equals
the probability of A given the condition
B. And in this case, we just multiply
those three different variables
together. So we have the probability of
the discount times the probability of
free delivery times the probability is
the day equal a holiday. Those are our
three variables of the probability of A
if B. And then that is going to be
multiplied by the probability of them
not making a purchase. And then we want
to divide that by the total
probabilities and they're multiplied
together. So we have the probability of
a discount, the probability of a free
delivery, and the probability of it
being on a holiday. When we plug those
numbers in, we see that one out of six
were no purchase on a discounted day,
two out of six were a no purchase on a
free delivery day, and three out of six
were a no purchase on a holiday. Those
are our three probabilities of A of B
multiplied out. And then that has to be
multiplied by the probability of a no
purchase. And remember the prob
probability of a noby is across all the
data. So that's where we get the 6 out
of 30. We divide that out by the
probability of each category over the
total number. So we get the 20 out of 30
had a discount, 23 out of 30 had a yes
for free delivery, and 11 out of 30 were
on a holiday. We plug all those numbers
in, we get.178.
So in our probability math, we have
a.178
if it's a no-by for a holiday, a
discount, and a free delivery. Let's
turn that around and see what that looks
like if we have a purchase. I promise
this is the last page of math before we
dig into the Python script. So here
we're calculating the probability of the
purchase using the same math we did to
find out if they didn't buy. Now we want
to know if they did buy. And again,
we're going to go by the day equals a
holiday, discount equals yes, free
delivery equals yes, and let a equal
buy. Now, right about now, you might be
asking, why are we doing both
calculations? Why why would we want to
know the no buys and buys for the same
data going in? Well, we're going to show
you that in just a moment, but we have
to have both of those pieces of
information so that we can figure it out
as a percentage as opposed to a
probability equation. And we'll get to
that normalization here in just a
moment. Let's go ahead and walk through
this calculation. And as you can see
here, the probability of A on the
condition of B, B being all three
categories, did we have a discount with
a purchase, did we have a free delivery
with a purchase, and did we is a day
equal to holiday. And when we plug this
all into that formula and multiply it
all out, we get our probability of a
discount, probability of a free
delivery, probability of the day being a
holiday times the overall probability of
it being a purchase divided by again
multiplying the three variables out. The
full probability of there being a
discount, the full probability of being
a free delivery, and the full
probability of there being a day equal
holiday. And that's where we get this 19
over 24 * 21 over 24 * 8 over 24 * the p
of a 24 over 30 divided by the
probability of the discount the free
delivery times the day or 20 over 30 23
over 30 * 11 over 30 and that gives us
our 986.
So what are we going to do with these
two pieces of data we just generated?
Well, let's go ahead and go over them.
We have a probability of purchase
equals.986.
We have a probability of no purchase
equals.178.
So finally we have a conditional
probabilities of purchase on this day.
Let us take that we're going to
normalize it and we're going to take
these probabilities and turn them into
percentages. This is simply done by
taking the sum of probabilities which
equals 98686 plus.178
and that equals the 1.164.
If we divide each probability by the
sum, we get the percentage. And so the
likelihood of a purchase is 84.71%.
And the likelihood of no purchase is
15.29%
given these three different variables.
So it's if it's on a holiday, if it's a
with a discount and has free delivery,
then there's an 84.71%
chance that the customer is going to
come in and make a purchase. Hooray,
they purchased our stuff. We're making
money. If you were owning a shop, that's
like is the bottom line is you want to
make some money so you can keep your
shop open and have a living. Now, I
promised you that we were going to be
finishing up the math here with a few
pages. So, we're going to move on and
we're going to do two steps. The first
step is I want you to understand why you
want to why you want to use the naive
bays. What are the advantages of naive
bays? And then once we understand those
advantages, we just look at that
briefly. Then we're going to dive in and
do some Python coding. Advantages of
naive bay classifier. So let's take a
look at the six advantages of the naive
bay classifier. And we're going to walk
around this lovely wheel. Looks like an
origami folded paper. The first one is
very simple and easy to implement.
Certainly you could walk through the
tables and do this by hand. You got to
be a little careful because the
notations can get confusing. You have
all these different probabilities and I
certainly mess those up as I put them
on, you know, is it on the top or the
bottom? We got to really pay close
attention to that. When you put it into
Python, it's really nice because you
don't have to worry about any of that.
You let the Python handle that, the
Python module. But understanding it, you
can put it on a table and you can easily
see how it works. And it's a simple
algebraic function. It needs less
training data. So if you have smaller
amounts of data, this is great powerful
tool for that. Handles both continuous
and discrete data. It's highly scalable
with number of predictors and data
points. So, as you can see, you can just
keep multiplying different probabilities
in there and you can cover not just
three different variables or sets. You
can now expand this to even more
categories. Number five, it's fast. It
can be used in real time predictions.
This is so important. This is why it's
used in a lot of our predictions on
online shopping carts, uh, referrals,
spam filters, is because there's no time
delay as it has to go through and figure
out a neural network or one of the other
mini setups where you're doing
classification. And certainly there's a
lot of other tools out there in the
machine learning that can handle these,
but most of them are not as fast as the
naive bays. And then finally, it's not
sensitive to irrelevant features. So it
picks up on your different
probabilities. And if you're short on
data on one probability, you can kind of
it automatically adjusts for that. Those
formulas are very automatic. And so you
can still get a very solid
predictability even if you're missing
data or you have overlapping data for
two completely different areas. We see
that a lot in doing census and studying
of people and habits where they might
have one study that covers one aspect
and another one that overlaps and
because the two overlap they can then
predict the unknowns for the group that
they haven't done the second study on or
vice versa. So it's very powerful in
that it is not sensitive to the
irrelevant features and in fact you can
use it to help predict features that
aren't even in there. So now we're down
to my favorite part. We're going to roll
up our sleeves and do some actual
programming. We're going to do the use
case text classification. Now, I would
challenge you to go back and send us a
note on the notes below underneath the
video and request the data for the
shopping cart. So, you can plug that
into Python code and do that on your own
time. So, you can walk through it since
we walk through all the information on
it. But, we're going to do a Python code
doing text classification. Very popular
for doing the naive bays. So, we're
going to use our new tool to perform a
text classification of news headlines
and classify news into different topics
for a news website. As you can see here,
we have a nice image of the Google News
and then related on the right subgroups.
I'm not sure where they actually pulled
the actual data we're going to use from.
It's one of the standard sets, but
certainly this can be used on any of our
news headlines in classification. So,
let's see how it can be done using the
naive base classifier. Now, we're at my
favorite part. We're actually going to
write some Python script, roll up our
sleeves, and we're going to start by
doing our imports. These are very basic
imports, including our news group. And
we'll take a quick glance at the target
names. Then we're going to go ahead and
start training our data set and putting
it together. We'll put together a nice
graph because it's always good to have a
graph to show what's going on. And once
we've trained it and we've shown you a
graph of what's going on, then we're
going to explore how to use it and see
what that looks like. Now I'm going to
open up my favorite editor or inline
editor for Python. You don't have to use
this. You can use whatever your editor
that you like, whatever uh interface IDE
you want. This just happens to be the
Anaconda Jupiter notebook. And I'm going
to paste that first piece of code in
here so we can walk through it. Let's
make it a little bigger on the screen so
you have a nice view of what's going on.
Uh and we're using Python 3, in this
case 3.5. So this would work in any of
your 3X if you have it set up correctly.
should also work in a lot of the 2x. You
just have to make sure all of the the
versions of the modules match your
Python version. And in here, you'll
notice the first line is your percentage
mattplot library in line. Now, three of
these lines of code are all about
plotting the graph. This one lets the
notebook know and is inline setup that
we want the graphs to show up on this
page. Without it, in a notebook like
this, which is an explorer interface, it
won't show up. Now, a lot of IDEs don't
require that. A lot of them, like on if
I'm working on one of my other setups,
it just has a popup and the graph pops
up on there. So, you have a that setup
also. But for this, we want the mattplot
library in line. And then we're going to
import numpy as np. That's number
python, which has a lot of different
formulas in it that we use for both of
our sklearn module. And we also use it
for any of the upper math functions in
python. And it's very common to see that
as NP numpy as NP. The next two lines
are all about our graphing. Remember I
said three of these were about graphing.
Well, we need our mattplot
library.pipplot
as plt. And you'll see that plt is a
very common setup as is the sns and just
like the np. And we're going to import
seabor as sns and we're going to do the
sns set. Now seabor sits on top of
pipplot and it just makes a really nice
heat map. It's really good for heat
maps. And if you're not familiar with
heat maps, that just means we give it a
color scale. The term comes from the
brighter red it is, the hotter it is in
some form of data. And you can set it to
whatever you want. And we'll see that
later on. So those you'll see that those
three lines of code here are just
importing the graph function so we can
graph it. And as a data scientist, you
always want to graph your data and have
some kind of visual. It's really hard
just to shove numbers in front of people
and they look at it and it doesn't mean
anything. And then from the sklearn data
sets, we're going to import the fetch 20
news groups. Very common one for
analyzing tokenizing words and setting
them up and exploring how the words work
and how do you categorize different
things when you're dealing with
documents. And then we set our data
equal to fetch 20 news groups. So our
data variable will have the data in it.
And we're going to go ahead and just
print the target names. data.target
names. And let's see what that looks
like. And you'll see here we have alt
atheism comp graphics composs
windows.mmiscellaneous
and it goes all the way down to talk
politics.mmiscellaneous talk
religion.mmiscellaneous. These are the
categories they've already assigned to
this news group and it's called fetch 20
because you'll see there's I believe
there's 20 different topics in here or
20 different categories as we scroll
down. Now, we've gone through the 20
different categories and we're going to
go ahead and start defining all the
categories and set up our data. So,
we're actually getting here going to go
ahead and get it get the data all set up
and take a look at our data. And let's
move this over to our Jupyter notebook.
And let's see what this code does.
First, we're going to set our
categories. Now, if you noticed up here,
I could have just as easily set this
equal to data.target_names target names
because it's the same thing, but we want
to kind of spell it out for you so you
can see the different categories. It
kind of makes it more visual so you can
see what your data is looking like in
the background. Once we've created the
categories, we're going to open up a
train set. So this training set of data
is going to go into fetch 20 news groups
and it's a subset in there called train
and categories equals categories. So
we're pulling out those categories that
match. And then if you have a train set,
you should also have the testing set. We
have test equals fetch 20 news group
subset equals test and categories equals
categories. Let's go down one size so it
all fits on my screen. There we go. And
just so we can really see what's going
on, let's see what happens when we print
out one part of that data. So it creates
train and under train, it creates train
data. And we're just going to look at
data piece number five. And let's go
ahead and run that and see what that
looks like. And you can see when I print
train.data data number five under train.
It prints out one of the articles. This
is article number five. You can go
through and read it on there. And we can
also go in here and change this to test,
which should look identical because it's
splitting the date up into different
groups. Train and test. And we'll see
test number five is a a different
article, but it's another article in
here. And maybe you're curious and you
want to see just how many articles are
in here. We could do length of train.
data. And if we run that, you'll see
that the training data has 11,314
articles. So, we're not going to go
through all those articles. That's a lot
of articles, but um we can look at one
of them just so you can see what kind of
information is coming out of it and what
we're looking at. And we'll just look at
number five for today. And here we have
it. Rewarding the Second Amendment IDs,
VTT, line 58, lines 58 in article, uh
etc. And you can scroll all the way down
and see all the different parts to
there. Now, we've looked at it and
that's pretty complicated when you look
at one of these articles to try to
figure out how do you weight this. If
you look down here, we have different
words and maybe the word from. Well,
from is probably in all the articles.
So, it's not going to have a lot of
meaning as far as trying to figure out
whether this article fits one of the
categories or not. So, trying to figure
out which category it fits in based on
these words is where the challenge comes
in. Now that we've viewed our data,
we're going to dive in and do the actual
predictions. This is the actual naive
bays. And we're going to throw another
model at you or another module at you
here in just a second. We can't go into
too much detail, but it deals
specifically working with words and text
and what they call tokenizing those
words. So, let's take this code and
let's uh skip on over to our Jupyter
notebook and walk through it. And here
we are in our Jupyter notebook. Let's
paste that in there. And I can run this
code right off the bat. It's not
actually going to display anything yet,
but it has a lot going on in here. So
the top we had the print module from the
earlier one. I didn't know why that was
in there. So we're going to start by
importing our necessary packages. And
from the sklearn features
extraction.ext,
we're going to import TF IDF vectorzer.
I told you we're going to throw a module
at you. We can't go too much into the
math behind this or how it works. You
can look it up. The notation for the
math is usually TF.idf.
And that's just a way of weighing the
words. and it weighs the words based on
how many times are used in a document,
how many times or how many documents
they're used in. And it's a well-used
formula. It's been around for a while.
It's a little confusing to put this in
here. Uh, but let's let them know that
it just goes in there and weights the
different words in the document for us.
That way, we don't have to wait. And if
you put a weight on it, if you remember,
I was talking about that up here
earlier. If these are all emails, they
probably all have the word from in them.
From probably has a very low weight. It
has very little value in telling you
what this document's about. Same with
words like in an article in articles in
cost of un maybe cost might or where
words like criminal weapons destruction
these might have a heavier weight
because they describe a little bit more
what the article is doing. Well, how do
you figure out all those weights in the
different articles? That's what this
module does. That's what the TF
vectorizer is going to do for us. And
then we're going to import our
sklearn.na naive bays and that's our
multinnomial NB multinnomial naive bay
pretty easy to understand that where
that comes from and then finally we have
the skyarn pipeline import make pipeline
now the make pipeline is just a cool
piece of code because we're going to
take the information we get from the TF
vectorizer and we're going to pump that
into the multinnomial NB. So, a pipeline
is just a way of organizing how things
flow. It's used commonly. You probably
already guessed what it is. If you've
done any businesses, they talk about the
sales pipeline. If you're on a work crew
or project manager, you have your
pipeline of information that's going
through or your projects and what has to
be done in what order. That's all this
pipeline is. We're going to take the
TFID vectorzer and then we're going to
push that into the multinnomial inb. Now
we've designated that as the variable
model. We have our pipeline model and
we're going to take that model and this
is just so elegant. This is done in just
a couple lines of code. model.fit and
we're going to fit the data. And first
the train data and then the train
target. Now the train data has the
different articles in it. You can see
the one we were just looking at and the
train.target target is what category
they already categorized that that
particular article as. And what's
happening here is the train data is
going into the TF ID vectorizer. So when
you have one of these articles, it goes
in there, it weights all the words in
there. So there's thousands of words
with different weights on them. I
remember once running a model on this
and I literally had 2.4 million tokens
go into this. So when you're dealing
like large document bases, you can have
a huge number of different words. It
then takes those words, gives them a
weight, and then based on that weight,
based on the words and the weights, and
then puts that into the multinnomial NB.
And once we go into our naive bay, we
want to put the train target in there.
So the train data that's been mapped to
the TFID vectorzer is now going through
the multinnomial NB. And then we're
telling it, well, these are the answers.
These are the answers to the different
documents. So this document that has all
these words with these different weights
from the first part is going to be
whatever category it comes out of. Maybe
it's the um talk show or the article on
religion miscellaneous. Once we fit that
model, we can then take labels and we're
going to set that equal to
model.predict. Most of the sklearn use
the term.predict to let us know that
we've now trained the model and now we
want to get some answers. And we're
going to put our test data in there
because our test data is the stuff we
held off to the side. We didn't train it
on there and we don't know what's going
to come up out of it and we just want to
find out how good our labels are. Do
they match what they should be? Now,
I've already run this through. There's
no actual output to it to show. This is
just setting it all up. This is just
training our model, creating the labels
so we can see how good it is, and then
we move on to the next step to find out
what happened. To do this, we're going
to go ahead and create a confusion
matrix and a heat map. So, the confusion
matrix, which is confusing just by its
very name, is basically going to ask how
confused is our answer. Did it get it
correct or did it miss some things in
there or have some missed labels? And
then we're going to put that on a heat
map so we have some nice colors to look
at to see how that plots out. Let's go
ahead and take this code and see how
that uh take a walk through it and see
what that looks like. So, back to our
Jupyter notebook. I'm going to put the
code in there and let's go ahead and run
that code. Take it just a moment. And
remember, we had the inline. That way,
my graph shows up on the inline here.
And let's walk through the code and then
we'll look at this and see what that
means. So, make it a little bit bigger.
There we go. No reason not to use the
whole screen. Too big. So, we have here
from sklearn metrics import confusion
matrix. And that's just going to
generate a set of data that says I the
prediction was such the actual truth was
either agreed with it or was something
different. And it's going to add up
those numbers so we can take a look and
just see how well it worked. And we're
going to set a variable Matt equal to
confusion matrix. We have our test
target, our test data that was not part
of the training. Very important in data
science, we always keep our test data
separate. Otherwise, it's not a valid
model if we can't properly test it with
new data. And this is the labels we
created from that test data. These are
the ones that we predict it's going to
be. So, we go in and we create our SN
heat map. The SNS is our seaborn which
sits on top of the piplot. So, we create
a SNS.heet map. We take our confusion
matrix and it's going to be uh matt.t.
And then we have other variables that go
into the SNS heat map. We're not going
to go into detail what all the variables
mean. The annotation equals true. That's
what tells it to put the numbers here.
So you have the 166, the one, the 00001.
Format D and C bar equals false have to
do with the uh format. If you take those
out, you'll see that some things
disappear. And then the X tick labels
and the Y tick labels. Those are our
target names. And you can see right
here, that's the alt atheism comp
graphics composs windows.mmiscellaneous.
And then finally we have our plt.xl
label. Remember the SNS or the seabor
sits on top of our mattplot library our
plt. And so we want to just tell it x
label equals a true is is true. The
labels are true. And then the y label is
prediction label. So when we say a true,
this is what it actually is. And the
prediction is what we predicted. And
let's look at this graph because that's
probably a little confusing the way I
rattled through it. And what I'm going
to do is I'm going to go ahead and flip
back to the slides because they have a
black background they put in there that
helps it shine a little bit better so
you can see the graph a little bit
easier. So in reading this graph, what
we want to look at is how the color
scheme has come out. And you'll see a
line right down the middle diagonally
from upper left to bottom right. What
that is is if you look at the labels, we
have our predicted label on the left and
our true label on the right. Those are
the numbers where the prediction and the
true come together. And this is what we
want to see is we want to see those lit
up. That's what that heat map does. As
you can see that it did a good job of
finding those data. And you'll notice
that there's a couple of red spots on
there where it missed. You know, it it's
a little confused when we talk about
talk religion miscellaneous versus talk
politics miscellaneous, social religion
Christian versus alt atheism. It
mislabeled some of those. And those are
very similar topics. so you could
understand why it might mislabel them.
But overall, it did a pretty good job.
If we're going to create these models,
we want to go ahead and be able to use
them. So, let's see what that looks
like. To do this, let's go ahead and
create a definition, a function to run.
And we're going to call this function.
Let me just expand that just a notch
here. There we go. I like mine in big
letters. Predict category. So, we want
to predict the category. We're going to
send it as a string. And then we're
sending it train equals train. We have
our training model. And then we had our
pipeline model equals model. This way we
don't have to resend these variables
each time. The definition knows that
because I said train equals train and I
put the equal for model. And then we're
going to set the prediction equal to the
model.predict s. So it's going to send
whatever string we send to it. It's
going to push that string through the
pipeline, the model pipeline. It's going
to go through and uh tokenize it and put
it through the TF IDF, convert that into
numbers and weights for all the
different documents and words. And then
it'll put that through our naive bay.
And from it, we'll go ahead and get our
prediction. We're going to predict what
value it is. And so we're going to
return train.target names predict of
zero. And remember that the train.target
names, that's just categories. I could
have just as easily put uh categories in
there.predict of zero. So we're taking
the prediction which is a number and
we're converting it to an actual
category. We're converting it from um I
don't know what the actual numbers are.
Let's say zero equals alt atheism. So
we're going to convert that zero to the
word or uh one maybe it equals comp
graphics. So we're going to convert
number one into comp graphics. That's
all that is. And then we got to go ahead
and and then we need to go ahead and run
this. So I load that up. And then once I
run that, we can start doing some
predictions. Let me go ahead and type in
predict category. And let's just do
predict category, Jesus Christ. And it
comes back and says it's social,
religion, Christian. That's pretty good.
Now note, I didn't put print on this.
One of the nice things about the Jupiter
notebook editor and a lot of inline
editors is if you just put the name of
the variable out, it's returning the
variable train.target_ames, target
names. It'll automatically print that
for you. In your own IDE, you might have
to put in print. Let's see where else we
can take this. And maybe you're a space
science buff. So, how about sending load
to international
space station.
And if we run that, we get science
space. Or maybe you're a uh automobile
buff. And let's do um Oh, they were
gonna tell me Audi is better than BMW,
but I'm going to do BMW is better than
an Audi. So maybe our car buff. And we
run that. And you'll see it says
recreational. I'm assuming that's what
RECC stands for. Autos. So I did a
pretty good job labeling that one. How
about uh if we have something like a
caption running through there, President
of India. And if we run that, it comes
up and says talk politics miscellaneous.
So when we take our definition or our
function and we run all these things
through, kudos, we made it. We were able
to correctly classify texts into
different groups based on which category
they belong to using the naive base
classifier. Now we did throw in the
pipeline, the TF IDF vectorzer, we threw
in the graphs. Those are all things that
you don't necessarily have to know to
understand the naive base setup or
classifier, but they're important to
know. One of the main uses for the naive
bays is with the TF IDF tokenizer
vectorzer where it tokenizes a word and
has labels and we use the pipeline
because you need to push all that data
through and it makes it really easy and
fast. You don't have to know those to
understand naive bays but they certainly
help for understanding the industry and
data science. And we can see our
categorizer, our naive base classifier.
We were able to predict the category
religion, space, motorcycles, autos,
politics, and properly classify all
these different things we pushed into
our prediction and our trained model.
Before we dive into the SVM, let's take
a look at applications of the support
vector machine, at least some general
ones that are commonly used with it.
face detection, text and hypertext
categorization, classification of
images, and bioinformatics.
These are only but a few of those that
are used with this SVM. As we go through
this lesson, see if you can figure out
what other ones you could apply it to,
and also what you would want to use some
other tools for. So, in this example,
last week, my son and I visited a fruit
shop. Dad, is that an apple or a
strawberry? So, the question comes up,
what fruit did I just pick up from the
fruit stand? After a couple of seconds,
you can figure out that it was a
strawberry. So, let's take this model a
step further and let's uh why not build
a model which can predict an unknown
data. And in this, we're going to be
looking at some sweet strawberries or
crispy apples. We wanted to be able to
label those two and decide what the
fruit is. And we do that by having data
already put in. So, we already have a
bunch of strawberries. We know our
strawberries and they're already labeled
as such. We already have a bunch of
apples. We know our apples and are
labeled as such. Then once we train our
model, that model then can be given the
new data and the new data is this image.
In this case, you can see a question
mark on it and it comes through and goes
it's a strawberry. In this case, we're
using the support vector machine model.
SVM is a supervised learning method that
looks at data and sorts it into one of
two categories. And in this case, we're
sorting the strawberry into the
strawberry side. At this point, you
should be asking the question, how does
the prediction work? Before we dig into
an example with numbers, let's apply
this to our fruit scenario. We have our
support vector machine. We've taken it
and we've taken labeled sample of data,
strawberries and apples, and we draw on
a line down the middle between the two
groups. This split now allows us to take
new data, in this case an apple and a
strawberry, and place them in the
appropriate group based on which side of
the line they fall in. And that way we
can predict the unknown. As colorful and
tasty as the fruit example is, let's
take a look at another example with some
numbers involved. And we can take a
closer look at how the math works. In
this example, we're going to be
classifying men and women. And we're
going to start with a set of people with
a different height and a different
weight. And to make this work, we'll
have to have a sample data set of female
where we have their height and weight
174, 65, 174, 88, and so on. And we'll
need a sample data set of the male. They
have a height 179, 90, 180 to 80 and so
on. Let's go ahead and put this on a
graph so we have a nice visual. So you
can see here we have two groups based on
the height versus the weight. And on the
left side we're going to have the women,
on the right side we're going to have
the men. Now if we're going to create a
classifier, let's add a new data point
and figure out if it's male or female.
So before we can do that, we need to
split our data first. We can split our
data by choosing any of these lines. In
this case, we draw in two lines through
the data in the middle that separates
the men from the women. But to predict
the gender of a new data point, we
should split the data in the best
possible way. And we say the best
possible way because this line has a
maximum space that separates the two
classes. Here you can see there's a
clear split between the two different
classes. And in this one, there's not so
much a clear split. This doesn't have
the maximum space that separates the
two. That is why this line best splits
the data. We don't want to just do this
by eyeballing it. And before we go
further, we need to add some technical
terms to this. We can also say that the
distance between the points in the line
should be as far as possible. In
technical terms, we can say the distance
between the support vector and the hyper
plane should be as far as possible. And
this is where the support vectors are
the extreme points in the data set. And
if you look at this data set, they have
circled two points which seem to be
right on the outskirts of the women and
one on the outskirts of the men. And
hyper plane has a maximum distance to
the support vectors of any class. Now
you'll see the line down the middle and
we call this the hyper plane because
when you're dealing with multiple
dimensions, it's really not just a line
but a plane of intersections. And you
can see here where the support vectors
have been drawn in dashed lines. The
math behind this is very simple. We take
D+ the shortest distance to the closest
positive point which would be on the
men's side and D minus is the shortest
distance to the closest negative point
which is on the women's side. The sum of
D plus and D minus is called the
distance margin or the distance between
the two support vectors that are shown
in the dashed lines. And then by finding
the largest distance margin, we can get
the optimal hyper plane. Once we've
created an optimal hyper plane, we can
easily see which side the new data fits
in. And based on the hyper plane, we can
say the new data point belongs to the
male gender. Hopefully that's clear how
that works on a visual level. As a data
scientist, you should also be asking
what happens if the hyper plane is not
optimal. If we select a hyper plane
having low margin, then there is a high
chance of mclassification. This
particular SVM model, the one we
discussed so far, is also called
referred to as the LSVM.
So far so clear, but a question should
be coming up. We have our sample data
set. But instead of looking like this,
what if it looked like this where we
have two sets of data, but one of them
occurs in the middle of another set. You
can see here where we have the blue and
the yellow and then blue again on the
other side of our data line. In this
data set, we can't use a hyper plane. So
when you see data like this, it's
necessary to move away from a 1D view of
the data to a two-dimensional view of
the data. And for the transformation, we
use what's called a kernel function. The
kernel function will take the 1D input
and transfer it to a two-dimensional
output. As you can see in this picture
here, the 1D when transferred to a
two-dimensional makes it very easy to
draw a line between the two data sets.
What if we make it even more
complicated? How do we perform an SVM
for this type of data set? Here you can
see we have a two-dimensional data set
where the data is in the middle
surrounded by the green data on the
outside. In this case, we're going to
segregate the two classes. We have our
sample data set and if you draw a line
through, it's obviously not an optimal
hyper plane in there. So to do that, we
need to transfer the 2D to a 3D array.
And when you translate it into a
three-dimensional array using the
kernel, you can see where you can place
a hyper plane right through it and
easily split the data. Before we start
looking at a programming example and
dive into the script, let's look at the
advantage of the support vector machine.
We'll start with highdimensional input
space or sometimes referred to as the
curse of dimensionality. We looked at
earlier one dimension, two dimension,
three dimension. When you get to a
thousand dimensions, a lot of problems
start occurring with most algorithms
that have to be adjusted for. The SVM
automatically does that in
highdimensional space. One of the
highdimensional space, one
highdimensional space that we work on is
sparse document vectors. This is where
we tokenize the words in documents so we
can run our machine learning algorithms
over them. I've seen ones get as high as
2.4 million different tokens. That's a
lot of vectors to look at. And finally,
we have regularization parameter. The
realization parameter or lambda is a
parameter that helps figure out whether
we're going to have a bias or
overfitting of the data. Whether it's
going to be overfitted to very specific
instance or it's going to be biased to a
high or low value. With the SVM, it
naturally avoids the overfitting and
bias problems that we see in many other
algorithms. These three advantages of
the support vector machine make it a
very powerful tool to add to your
repertoire of machine learning tools.
Now, we did promise you a use case
study. We're actually going to dive in
to some Python programming. And so we're
going to go into a problem statement and
start off with the zoo. So in the zoo
example, we have um family members going
to the zoo and we have the young child
going, "Dad, is that a group of
crocodiles or alligators?" Well, that's
hard to differentiate. And zoos are a
great place to start looking at science
and understanding how things work,
especially as a young child. And so we
can see the parents sitting here
thinking, well, what is the difference
between a crocodile and an alligator?
Well, one, crocodiles are larger in
size. Alligators are smaller in size.
Snout width. The crocodiles have a
narrow snout and alligators have a wider
snout. And of course, in the modern day
and age, the father's sitting here is
thinking, "How can I turn this into a
lesson for my son?" And he goes, "Let a
support vector machine segregate the two
groups." I don't know if my dad ever
told me that, but that would be funny.
Now, in this example, we're not going to
use actual measurements and data. We're
just using that for imagery. And that's
very common in a lot of machine learning
algorithms and setting them up. But
let's roll up our sleeves and we'll talk
about that more in just a moment as we
break into our Python script. So here we
arrive in our actual coding and I'm
going to move this into a Python editor
in just a moment. But let's talk a
little bit about what we're going to
cover. First, we're going to cover in
the code the setup, how to actually
create our SVM. And you're going to find
that there's only two lines of code that
actually create it. And the rest of it
is done so quick and fast that it's all
here in the first page. and we'll show
you what that looks like as far as our
data because we're going to create some
data. I talked about creating data just
a minute ago. And so we'll get into the
creating data here and you'll see this
nice correction of our two blobs and
we'll go through that in just a second.
And then the second part is we're going
to take this and we're going to bump it
up a notch. We're going to show you what
it looks like behind the scenes. But
let's start with actually creating our
setup. I like to use the Anaconda
Jupyter notebook because it's very easy
to use, but you can use any of your
favorite Python editors or setups and go
in there. But let's go ahead and switch
over there and see what that looks like.
So here we are in the Anaconda Python
notebook or Anaconda Jupyter notebook
with Python. We're using Python 3. I
believe this is 3.5, but it should be
work in any of your 3x versions. And uh
you'd have to look at the sklearn and
make sure if you're using a 2x version,
an earlier version. Let's go and put our
code in there. And one of the things I
like about the Jupyter notebook is I can
go up to view and I'm going to go ahead
and toggle the line numbers on to make
it a little bit easier to talk about.
And we can even increase the size
because this is edited in in this case
I'm using Google Chrome explorer and
that's how it opens up for the editor.
Although anyone any like I said any
editor will work. Now the first step is
going to be our imports and we're going
to import four different parts. The
first two I want you to look at are line
one and line two are numpy as np and
mapplot library.pipplot
as plt. Now these are very standardized
imports when you're doing work. The
first one is the numbers python. We need
that because part of the platform we're
using uses that for the numpy array. And
I'll talk about that in a minute so you
can understand why we want to use a
numpy array versus a standard python
array. And normally it's pretty standard
setup to use NP for numpy. The map plot
library is how we're going to view our
data. So this has uh you do need the NP
for the sklearn module, but the map plot
library is purely for our use for
visualization. And so you really don't
need that for the SVM, but we're going
to put it there so you have a nice
visual aid and we can show you what it
looks like. That's really important at
the end when you finish everything so
you have a nice display for everybody to
look at. And then finally, we're going
to I'm going to jump one ahead to line
number four. That's the sklearn.datas
sets.samples generator import make
blobs. And I told you that we were going
to make up data. And this is a tool
that's in the sklearn to make up data. I
personally don't want to go to the zoo,
get in trouble for jumping over the
fence, and probably get eaten by the
crocodiles or alligators as I work on
measuring their snouts and width and
length. Instead, we're just going to
make up some data. And that's what that
make blobs is. It's a wonderful tool. If
you're ready to test your your uh setup
and you're not sure about what data
you're going to put in there, you can
create this blob and it makes it really
easy to use. And finally, we have our
actual SVM, the sklearn import SVM on
line three. So that covers all our
imports. We're going to create, remember
I used the make blobs to create data.
And we're going to create a capital X
and a lowercase Y equals make blobs in
samples equals 40. So we're going to
make 40 lines of data. It's going to
have two centers with a random state
equals 20. So each each each group's
going to have 20 different pieces of
data in it. And the way that looks is
that we'll have under X um an XY plane.
So I have two numbers under X and Y will
be 01. That's the two different centers.
So we have yes or no in this case
alligator crocodile. That's what that
represents. And then I told you that the
actual sklearn or the SVM is in two
lines of code. And we see it right here
with CLF equals SVM. SVC kernel equals
linear. And I set C equal to one.
Although in this example, since we are
not uh regularizing the data because we
want it to be very clear and easy to
see, I went ahead. You can set it to a
th00and a lot of times when you're not
doing that. But for this thing linear,
because it's a very simple linear
example, we only have the two dimensions
and it'll be a nice linear hyper plane.
It'll be a nice linear line instead of a
full plane. So we're not dealing with a
huge amount of data. And then all we
have to do is do clff.fit
x, y. And that's it. CLF has been
created. And then we're going to go
ahead and display it. And I'm going to
talk about this display here in just a
second. But let me go ahead and run this
code. And this is what we've done is
we've created two blobs. You'll see the
blue on the side and then kind of an
orang-ish uh on the other side. That's
our two sets of data. They represent one
represents crocodiles and one represents
alligators. And then we have our
measurements. In this case, we have like
the width and length of the snout. And I
did say I was going to come up here and
talk just a little bit about our plot.
And you'll see plt. That's what we
imported. We're going to do a scatter
plot. That means we're just putting dots
on there. And then look at this
notation. I have the capital X and then
in brackets I have a colon, 0ero. That's
from numpy. If you did that in a regular
array, you'll get an error in a Python
array. You have to have that in a numpy
array. It turns out that our make blobs
returns a numpy array. And this notation
is great because what it means is the
first part is the colon means we're
going to do all the rows. That's all the
data in our blob we created under
capital X. And then the second part has
a comma 0ero. We're only going to take
the first value. And then if you notice,
we do the same thing, but we're going to
take the second value. Remember, we
always start with zero and then one. So
we have column zero and column one. And
you can look at this as our XY plots.
The first one is the xplot and the
second one is the y plot. So the first
one is on the bottom 0 2 4 6 8 and 10.
And then the second one x of the one is
the 4 5 6 7 8 9 10 going up the left
hand side. S= 30 is just the size of the
dots. We can see them instead of real
tiny dots. And then cmap equals
plt.cm.paired.
And you'll also see the c equals y.
That's the color. We're using two colors
01. And that's why we get the nice blue
and the two different colors for the
alligator and the crocodile. Now you can
see here that we did this the actual fit
was done in two lines of code. A lot of
times there'll be a third line where we
regularize the data. We set it between
like minus one and one and we reshape
it. But for this it's not necessary and
it's also kind of nice because you can
actually see what's going on. And then
if we wanted to we wanted to actually
run a prediction. Let's take a look and
see what that looks like. And to predict
some new data and we'll show this again
as we get towards the end of digging in
deep. You can simply assign your new
data. In this case I am giving it a uh
width and length 34 and a width and
length 56. And note that I put the data
as a set of brackets and then I have the
brackets inside. And the reason I do
that is because when we're looking at
data it's designed to process a large
amount of data coming in. We don't want
to just process one line at a time. And
so in this case, I'm processing two
lines. And then I'm just going to print
and you'll see clf.predict new data. So
the CLF and the predict part is going to
give us an answer. And let's see what
that looks like. And you'll see 01. So
predicted the first one, the 34 is going
to be on the one side and the 56 is
going to be on the other side. So one
came out as a alligator and one came out
as a crocodile. Now that's pretty short
explanation for the setup, but really we
want to dug in and see what it's going
on behind the scenes. and let's see what
that looks like. So, the next step is to
dig in deep and find out what's going on
behind the scenes and also put that in a
nice pretty graph. We're going to spend
more work on this than we did actually
generating the original model. And
you'll see here that we go through a few
steps and I'm I'll move this over to our
editor in just a second. We come in, we
create our original data. It's exactly
identical to the first part and I'll
explain why we redid that and show you
how not to redo that. And then we're
going to go in there and add in those
lines. We're going to see what those
lines look like and how to set those up.
And finally, we're going to plot all
that on here and show it. And you'll get
a nice graph with the what we saw
earlier when we were going through the
theory behind this where it shows the
support vectors and the hyper plane. And
those are done where you can see the
support vectors as the dash lines and
the solid line which is the hyper plane.
Let's get that into our Jupyter
notebook. Before I scroll down to a new
line, I want you to notice line 13. It
has plot show. And we're going to talk
about that here in just a second. But
let's scroll down to a new line down
here. And I'm going to paste that code
in. And you'll see that the plot show
has moved down below. Let's scroll up a
little bit. And if you look at the top
here of our new section, 1 2 3 and four
is the same code we had before. And
let's go back up here and take a look at
that. We're going to fit the values on
our SVM. And then we're going to plot
scatter it. And then we're going to do a
plot show. So you should be asking why
are we redoing the same code. Well, when
you do the plot show, that blanks out
what's in the plot. So once I've done
this plot show, I have to reload that
data. Now, we could do this simply by
removing it up here, rerunning it, and
then coming down here, and then we
wouldn't have to rerun these first four
lines of code. Now, in this, it doesn't
matter too much. And you'll see the plot
show is down here and then removed right
there on line five. I'll go ahead and
just delete that out of there because we
don't want to blank out our screen. We
want to move on to the next setup. So,
we can go ahead and just skip the first
four lines because we did that before.
And let's take a look at the ax=
plt.gca.
Now, right now, we're actually spending
a lot of time just graphing. That's all
we're doing here. Okay. So, this is how
we display a nice graph with our results
and our data. AX is very standard not
used variable when you're talking about
PLT and it's just setting it to that
axis the last axis in the PLT. It can
get very confusing if you're working
with many different layers of data on
the same graph and this makes it very
easy to reference the ax. So this
reference is looking at the PLT that we
created and we already mapped out our
two blobs on. And then we want to know
the limits. So we want to know how big
the graph is. And we can find out the x
limit and the y limit simply with the
get x limit and get yimit commands which
is part of our metplot library. And then
we're going to create a grid. And you'll
see down here we have we've set the
variable xx equal to npines space ximit
0 ximit 1a 30. And we've done the same
thing for the yspace. And then we're
going to go in here and we create a mesh
grid. And this is a numpy command. So
we're back to our numbers python. Let's
go through what these numpy commands
mean with the line space in the mesh
grid. We've taken xx small xx= np line
space. And we have our x limit zero and
our x limit one and we're going to
create 30 points on it. And we're going
to do the same thing for the y axis. Now
this has nothing to do with our
evaluation. It's uh all we're doing is
we're creating a grid of data. And so
we're creating a set of points between
zero and the x limit. We're creating 30
points. And the same thing with the y.
And then the mesh grid loops those all
together. So it forms a nice grid. So if
we were going to do this say between the
limit 0 and 10 and do 10 points, we
would have a 0 0 1 1 0 1 02 03 04 to 10
and so on. You can just imagine a point
at each corner one of those boxes. And
the mesh grid combines them all. So we
take the y and the xx we created and
creates the full grid. And we've set
that grid into the y coordinates and the
xx coordinates. Now remember, when we're
working with Numbi in Python, we like to
separate those. We like to have instead
of it being x comma 1, you know, x comma
y and then x2 comma y2 and in the next
set of data, it would be a column of x's
and a column of y's. And that's what we
have here is we have a column of y's. We
put it as a capital y y and a column of
x's, capital xx with all those different
points being listed. And finally, we get
down to the numpy vstack. Just as we
created those in the mesh grid, we're
now going to put them all into one
array, XY array. Now that we've created
the stack of data points, we're going to
do something interesting here. We're
going to create a value Z. And the Z
equals the CLF. That's our uh that's our
support vector machine we created and
we've already trained. And we have a
decision function. And we're going to
put the XY in there. So here we have all
this data. We're going to put that XY in
there, that data, and we're going to
reshape it. And you'll see that we have
the xx.shape in here. This literally
takes the xx, resets it up, connected to
the y, and the zvalue lets us know
whether it is the left hand side. It's
going to generate three different
values. The zvalue does, and it'll tell
us whether that data is a support vector
to the left, the hyper plane in the
middle, or the support vector to the
right. So it generates three different
values for each of those points. And
those points have been reshaped so
they're right on a line on those three
different lines. So we've set all of our
data up. We've labeled it to three
different areas and we've reshaped it.
And we've just taken 30 points in each
direction. If you do the math, you have
30 * 30. So that's 900 points of data.
And we separated it between the three
lines and reshaped it to fit those three
lines. We can then go back to our map
plot library where we've created the AX
and we're going to create a contour. And
you'll see here where we have contour,
capital XX, capital Y, Y. These have
been reshaped to fit those lines. Z is
the labels. So now we have the three
different points with the labels in
there. And we can set the colors equals
K. And I told you we had three different
labels, but we have uh three levels of
data. The alpha is just makes it kind of
see-through. So it's only uh 0.5 of the
value in there. So when we graph it, the
data will show up from behind it,
wherever the lines go. And finally, the
line styles. This is where we set the
two support vectors to be dash dash
lines and then a single one is just a
straight line. That's what all that
setup does. And then finally, we take
our ax.scatter. We're going to go ahead
and plot the support vectors, but we've
programmed it in there so that they look
nice like the dash dash line and the
dash line on that grid. And you can see
here when we do the CLFS support
vectors, we are looking at column zero
and column one. And then again we have
the S equals 100. So we're going to make
them larger. And the line width equals
1, face colors equals none. Let's take a
look and see what that looks like when
we show it. And you can see when we get
down to our end result, it creates a
really nice graph. We have our two
support vectors and dash lines. And they
have the near data. So you can see those
two points or in this case the four
points where those lines nicely cleave
the data. And then you have your hyper
plane down the middle which is as far
from the two different points as
possible creating the maximum distance.
So you can see that we have our nice
output for the size of the body and the
width of the snout and we've easily
separated the two groups of crocodile
and alligator. Congratulations. You've
done it. We've made it. Of course, these
are pretend data for our crocodiles and
alligators. But this hands-on example
will help you to encounter any support
vector machine projects in the future.
And you can see how easy they are to set
up and look at in depth. We're going to
cover the K nearest neighbors a lot
referred to as KNN. And KNN is really a
fundamental place to start in the
machine learning. It's a basis of a lot
of other things and just the logic
behind it is easy to understand and
incorporated in other forms of machine
learning. So today, what's in it for
you? Why do we need KNN? What is KN&N?
How do we choose the factor K? When do
we use KNN? How does KN&N algorithm
work? And then we'll dive in to my
favorite part, the use case. Predict
whether a person will have diabetes or
not. That is a very common and popular
used data set as far as testing out
models and learning how to use the
different models in machine learning. By
now, we all know machine learning models
make predictions by learning from the
past data available. So we have our
input values. Our machine learning model
builds on those inputs of what we
already know and then we use that to
create a predicted output. Is that a
dog? Little kid looking over there and
watching the black cat cross their path.
No, dear. You can differentiate between
a cat and a dog based on their
characteristics.
Cats. Cats have sharp claws, uses to
climb, smaller length of ears, meows and
purr. Doesn't love to play around. dogs.
They have dull claws, bigger length of
ears, barks, loves to run around. You
usually don't see a cat running around
people, although I do have a cat that
does that where dogs do. And we can look
at these. We can say uh we can evaluate
the sharpness of the claws. How sharp
are their claws? And we can evaluate the
length of the ears. And we can usually
sort out cats from dogs based on even
those two characteristics. Now, tell me
if it is a cat or a dog. Not question.
Usually little kids know cats and dogs
by now. unless you live a place where
there's not many cats or dogs. So, if we
look at the sharpness of the claws, the
length of the ears, and we can see that
the cat has smaller ears and sharper
claws than the other animals. Its
features are more like cats. It must be
a cat. Sharp claws, length of ears, and
it goes in the cat group. Because KN&N
is based on feature similarity, we can
do classification using KN&N classifier.
So, we have our input value, the picture
of the black cat. It goes into our
trained model and it predicts that this
is a cat coming out. So what is knn?
What is the kn&n algorithm? K nearest
neighbors is what that stands for. Is
one of the simplest supervised machine
learning algorithms mostly used for
classification. So we want to know is
this a dog or it's not a dog? Is it a
cat or not a cat? It classifies a data
point based on how its neighbors are
classified. KN&N stores all available
cases and classifies new cases based on
a similarity measure. And here we've
gone from cats and dogs right into wine.
Another favorite of mine. KN&N stores
all available cases and classifies new
cases based on a similarity measure. And
here you see we have a measurement of
sulfur dioxide versus the chloride level
and then the different wines they've
tested and where they fall on that graph
based on how much sulfur dioxide and how
much chloride. K and K&N is a perimeter
that refers to the number of nearest
neighbors to include in the majority of
the voting process. And so if we add a
new glass of wine there, red or white,
we want to know what the neighbors are.
In this case, we're going to put K
equals 5. We'll talk about K in just a
minute. A data point is classified by
the majority of votes from its five
nearest neighbors. Here, the unknown
point would be classified as red since
four out of five neighbors are red. So,
how do we choose K? How do we know K
equals 5? I mean that's was the value we
put in there. I said we're going to talk
about it. How do we choose the factor K?
KN&N algorithm is based on feature
similarity. Choosing the right value of
K is a process called parameter tuning
and is important for better accuracy. So
at K equals 3, we can classify we have a
question mark in the middle as either a
as a square or not. Is it a square or is
it in this case a triangle? And so if we
set K equals to three, we're going to
look at the three nearest neighbors.
We're going to say this is a square. And
if we put k equals a 7, we classify as a
triangle depending on what the other
data is around it. And you can see as
the k changes depending on where that
point is, that drastically changes your
answer. And uh we jump here. We go, how
do we choose the factor of k? You'll
find this in all machine learning.
Choosing these factors, that's the face
you get. It's like, oh my gosh, did I
choose the right K? Did I set it right
my values in whatever machine learning
tool you're looking at? so that you
don't have a huge bias in one direction
or the other. And in terms of KNN, the
number of K, if you choose it too low,
the bias is based on it's just too
noisy. It's it's right next to a couple
things and it's going to pick those
things and you might get a skewed
answer. And if your K is too big, then
it's going to take forever to process.
So you're going to run into processing
issues and resource issues. So what we
do the most common use and there's other
options for choosing k is to use the
square root of n. So n is a total number
of values you have you take the square
root of it. In most cases you also if
it's an even number so if you're using
uh like in this case squares and
triangles if it's even you want to make
your k value odd. That helps it select
better. So in other words you're not
going to have a balance between two
different factors that are equal. So
usually take the square root of n and if
it's even you add one to it or subtract
one from it and that's where you get the
k value from that is the most common use
and it's pretty solid. It works very
well. When do we use kn? We can use kn
when data is labeled. So you need a
label on it. We know we have a group of
pictures with dogs cats cats. Data is
noisefree. And so you can see here when
we have a class and we have like
underweight 140 23 Hello kitty normal
that's pretty confusing. We have a a
high variety of data coming in. So it's
very noisy and that would cause an
issue. Data set is small. So we're
usually working with smaller data sets
where you might get into gig of data if
it's really clean. It doesn't have a lot
of noise because KN&N is a lazy learner.
I.e. it doesn't learn a discriminative
function from the training set. So it's
very lazy. So if you have very
complicated data and you have a large
amount of it, you're not going to use
the kn. But it's really great to get a
place to start. Even with large data,
you can sort out a small sample and get
an idea of what that looks like using
the KN&N and also just using for smaller
data sets. KN&N works really good. How
does the KN&N algorithm work? Consider a
data set having two variables, height in
centimeters and weight in kilograms. And
each point is classified as normal or
underweight. So we can see right here we
have two variables, you know, true
false. They're either normal or they're
not. They're underweight. On the basis
of the given data, we have to classify
the below set as normal or underweight
using KN&N. So if we have new data
coming in that says 57 kg and 177 cm, is
that going to be normal or underweight?
To find the nearest neighbors, we'll
calculate the ukitian distance.
According to the uklitian distance
formula, the distance between two points
in the plane with the coordinates xy and
ab is given by distance d equals the
square root of x - a^2 + y - b^2. And
you can remember that from the two edges
of a triangle. We're computing the third
edge since we know the x side and the y
side. Let's calculate it to understand
clearly. So we have our unknown point
and we placed it there in red. And we
have our other points where the data is
scattered around. The distance d1 is the
square of 170 minus 167^ squar + 57
- 51^ 2ar which is about 6.7 and
distance 2 is about 13 and distance 3 is
about 13.4. Similarly, we will calculate
the ukitian distance of unknown data
point from all the points in the data
set. And because we're dealing with
small amount of data, that's not that
hard to do and it's actually pretty
quick for a computer and it's not a
really complicated math. You can just
see how close is the data based on the
uklidian distance. Hence, we have
calculated the uklidian distance of
unknown data point from all the points
as shown where x1 and y1 equal 57 and
170 whose class we have to classify. So
now we're looking at that. We're saying
well here's the ukitian distance. Who's
going to be their closest neighbors? Now
let's calculate the nearest neighbor at
k equals 3. And we can see the three
closest neighbors puts them at normal.
And that's pretty self-evident when you
look at this graph. It's pretty easy to
say okay what you know we're just voting
normal normal normal. Three votes for
normal. This is going to be a normal
weight. So majority of neighbors are
pointing towards normal. Hence as per
KN&N algorithm the class of 571 170
should be normal. So a recap of KN&N
positive integer K is specified along
with a new sample. We select the K
entries in our database which are
closest to the new sample. We find the
most common classification of these
entries. This is the classification we
give to the new sample. So, as you can
see, it's pretty straightforward. We're
just looking for the closest things that
match what we got. So, let's take a look
and see what that looks like in a use
case in Python. So, let's dive into the
predict diabetes use case. So, use case,
predict diabetes. The objective, predict
whether a person will be diagnosed with
diabetes or not. We have a data set of
768 people who were or were not
diagnosed with diabetes. And let's go
ahead and open that file and just take a
look at that data. And this is in a
simple spreadsheet format. The data
itself is commaepparated. Very common
set of data. And it's also a very common
way to get the data. And you can see
here we have columns A through I. That's
what 1 2 3 4 5 6 7 8. um eight columns
with a particular attribute and then the
ninth column which is the outcome is
whether they have diabetes. As a data
scientist, the first thing you should be
looking at is insulin. Well, you know,
if someone has insulin, they have
diabetes because that's why they're
taking it. And that could cause issue in
some of the machine learning packages,
but for very basic setup, this works
fine for doing the KNN. And the next
thing you notice is it it didn't take
very much to open it up. Um I can scroll
down to the bottom of the data. There's
768.
It's pretty much a small data set. You
know, at 769, I can easily fit this into
my RAM on my computer. I can look at it.
I can manipulate it. And it's not going
to really tax just a regular desktop
computer. You don't even need an
enterprise version to run a lot of this.
So, let's start with importing all the
tools we need. And before that, of
course, we need to discuss what IDE I'm
using. Certainly, you can use any uh
particular editor for Python, but I like
to use for doing uh very basic visual
stuff. the Anaconda, which is great for
doing demos with the Jupyter Notebook.
And just a quick view of the Anaconda
Navigator, which is the new release out
there, which is really nice. You can see
under home, I can choose my application.
We're going to be using Python 3.6. I
have a couple different uh versions on
this particular machine. If I go under
environments, I can create a unique
environment for each one, which is nice.
And there's even a little button there
where I can install different packages.
So, if I click on that button and open
the terminal, I can then use a simple
pip install to install different
packages I'm working with. Let's go
ahead and go back under home and we're
going to launch our notebook. And I've
already, you know, kind of like uh the
old cooking shows, I've already prepared
a lot of my stuff. So, we don't have to
wait for it to launch because it takes a
few minutes for it to open up a browser
window. In this case, I'm going to it's
going to open up Chrome because that's
my default that I use. And since the
script is pre-done, you'll see I have a
number of windows open up at the top,
the one we're working in. And uh since
we're working on the KN&N predict
whether a person will have diabetes or
not. Let's go and put that title in
there. And I'm also going to go up here
and click on cell. Actually, we want to
go ahead and first insert a cell below.
And then I'm going to go back up to the
top cell. And I'm going to change the
cell type to markdown. That means this
is not going to run as Python. It's a
markdown language. So if I run this
first one, it comes up in nice big
letters, which is kind of nice. Remind
us what we're working on. And by now you
should be familiar with doing all of our
imports. We're going to import the
pandas as pd import numpy is np. Pandas
is the uh pandas data frame and numpy is
a number array. Very powerful tools to
use in here. So we have our imports. So
we've brought in our pandas or numpy our
two general python tools. And then you
can see over here we have our train test
split. By now you should be familiar
with splitting the data. We want to
split part of it for training our thing
and then training our particular model
and then we want to go ahead and test
the remaining data to see how good it
is. Pre-processing a standard scaler
pre-processor so we don't have a bias of
really large numbers. Remember in the
data we had like number of pregnancies
isn't going to get very large where the
amount of insulin they take and get up
to 256. So 256 versus six that will skew
results. So we want to go ahead and
change that so they're all uniform
between minus1 and one. And then the
actual tool. This is the K neighbors
classifier we're going to use. And
finally, the last three are three tools
to test. All about testing our model.
How good is it? We just put down test on
there. And we have our confusion matrix,
our F1 score, and our accuracy. So we
have our two general Python modules
we're importing. And then we have our
six modules specific from the sklearn
setup. And then we do need to go ahead
and run this. So these are actually
imported. There we go. And then move on
to the next step. And so in this set,
we're going to go ahead and load the
database. We're going to use pandas.
Remember pandas is pd. And we'll take a
look at the data in Python. We looked at
it in a simple spreadsheet, but usually
I like to also pull it up so that we can
see what we're doing. So here's our data
set equals PD read CSV. That's a pandas
command. And the diabetes folder I just
put in the same folder where my IPython
script is. If you put in a different
folder, you'd need the full length on
there. We can also do a quick length of
uh the data set. That is a simple Python
command. Leen for length. We might even
let's go ahead and print that. We'll go
print. And if you do it on its own line,
length data set in the Jupyter notebook,
it'll automatically print it. But when
you're in most of your different setups,
you want to do the print in front of
there. And then we want to take a look
at the actual data set. And since we're
in pandas, we can simply do data set
head. And again, let's go ahead and add
the print in there. If you put a bunch
of these in a row, you know that data
set one head, data set two head, it only
prints out the last one. So, I usually
always like to keep the print statement
in there. But because most projects only
use one data frame, Panda's data frame,
doing it this way doesn't really matter.
The other way works just fine. And you
can see when we hit the run button, we
have the 768 lines, which we knew, and
we have our pregnancies. It's
automatically given a label on the left.
Remember the head only shows the first
five lines. So we have zero through
four. And just a quick look at the data.
You can see it matches what we looked at
before. We have pregnancy, glucose,
blood pressure all the way to age. And
then the outcome on the end. And we're
going to do a couple things in this next
step. We're going to create a list of
columns where we can't have zero.
There's no such thing as zero skin
thickness or zero blood pressure, zero
glucose. Uh any of those, you'd be dead.
So, not a really good factor if they
don't if they have a zero in there
because they didn't have the data. And
we'll take a look at that because we're
going to start replacing that
information with a couple of different
things. And let's see what that looks
like. So, first we create a nice list.
As you can see, we have the values
talked about glucose, blood pressure,
skin thickness. Uh, and this is a nice
way when you're working with columns is
to list the columns you need to do some
kind of transformation on. Uh, very
common thing to do. And then for this
particular setup, we certainly could use
the there's some Panda tools that will
do a lot of this where we can replace
the NA, but we're going to go ahead and
do it as a data set column equals data
set column.replace. This is this is
still pandas. You can do a direct.
There's also one that that you look for
your nan. A lot of different options in
here. But the nan numpan is what that
stands for is non doesn't exist. So the
first thing we're doing here is we're
replacing the zero with a numpy none.
There's no data there. That's what that
says. That's what this is saying right
here. So put the zero in and we're going
to replace zeros with no data. So if
it's a zero, that means the person's
well hopefully not dead. Hopefully they
just didn't get the data. The next thing
we want to do is we're going to create
the mean which is the in integer from
the data set from the column mean where
we skip NAS. We can do that. That is a
pandas command there, the skip na. So
we're going to figure out the mean of
that data set. And then we're going to
take that data set column and we're
going to replace all the npnan
with the means. Why did we do that? And
we could have actually just uh taken
this step and gone right down here and
just replace zero and skip anything
where except you could actually there's
a way to skip zeros and then just
replace all the zeros. But in this case,
we want to go ahead and do it this way.
So you could see that we're switching
this to a non-existent value. Then we're
going to create the mean. Well, this is
the average person. So if we don't know
what it is, if they did not get the data
and the data is missing, one of the
tricks is you replace it with the
average. What is the most common data
for that? This way you can still use the
rest of those values to do your
computation and it kind of just brings
that particular value or those missing
values out of the equation. Let's go
ahead and take this and we'll go ahead
and run it. Doesn't actually do
anything. So we're still preparing our
data. If you want to see what that looks
like, we don't have anything in the
first few lines, so it's not going to
show up. But we certainly could look at
a row. Let's do that. Let's go into our
data set. Let's print a data set. And
let's pick in this case, let's just do
glucose. And if I run this, this is
going to print all the different glucose
levels going down. And we thankfully
don't see anything in here that looks
like missing data, at least on the ones
it shows. You can see it skipped a bunch
in the middle because that's what it
does. If you have too many lines in
Jupyter notebook, it'll skip a few and
and go on to the next in a data set. Let
me go and remove this. And we'll just
zero out that. And of course, before we
do any processing, before proceeding any
further, we need to split the data set
into our train and testing data. That
way, we have something to train it with
and something to test it on. And you're
going to notice we did a little
something here with the uh pandas
database code. There we go. My drawing
tool. We've added in this right here off
the data set. And what this says is that
the first one in pandas, this is from
the PD pandas. It's going to say within
the data set, we want to look at the eye
location and it is all rows. That's what
that says. So we're going to keep all
the rows, but we're only looking at
zero, column 0 to 8. Remember column 9.
Here it is right up here. We printed it
in here is outcome. Well, that's not
part of the training data. That's part
of the answer. Yeah, it's column 9, but
it's listed as eight. Number eight. So 0
to eight is nine columns. So uh eight is
the value. And when you see it in here,
zero, this is actually 0 to 7. It
doesn't include the last one. And then
we go down here to Y, which is our
answer. And we want just the last one,
just column 8. And you can do it this
way with this particular notation. And
then if you remember, we imported the
train test split that's part of the
sklearn right there. And we simply put
in our X and our Y. We're going to do
random state equals zero. You don't have
to necessarily seed it. That's a seed
number. I think the default is one when
you seated it. I'd have to look that up.
And then the test size. Test size is
0.2. That simply means we're going to
take 20% of the data and put it aside so
that we can test it later. That's all
that is. And again, we're going to run
it. Not very exciting. So far, we
haven't had any print out other than to
look at the data. But that is a lot of
this is prepping this data. Once you
prep it, the actual lines of code are
quick and easy. And we're almost there.
But the actual writing of our KN&N, we
need to go ahead and do a scale the
data. If you remember correctly, we're
fitting the data in a standard scaler,
which means instead of the data being
from, you know, five to 303 in one
column and the next column is 1 to six,
we're going to set that all so that all
the data is between minus1 and one.
That's what that standard scaler does.
Keeps it standardized. And we only want
to fit the scaler with the training set,
but we want to make sure the testing set
is the X test going in is also
transformed. So it's processing it the
same. So here we go with our standard
scaler. We're going to call it sc__x for
the scaler. And we're going to import
the standard scaler into this variable.
And then our xrain equals sc_x.fit
transform. So we're creating the scaler
on the x-ra variable. And then our x
test, we're also going to transform it.
So we've trained and transformed the
x-ra. And then the x test isn't part of
that training. It isn't part of that of
training the transformer. it just gets
transformed. That's all it does. And
again, we're going to go and run this.
And if you look at this, we've now gone
through these steps, all three of them.
We've taken care of replacing our zeros
for key columns that shouldn't be zero,
and we've replaced that with the means
of those columns. That way, that they
fit right in with our data models. We've
come down here, and we split the data.
So, now we have our test data and our
training data. And then we've taken and
we've scaled the data. So all of our
data going in. No, no, we don't tra we
don't train the Y part, the Y train and
Y test that never has to be trained.
It's only the data going in. That's what
we want to train in there. Then define
the model using K neighbors classifier
and fit the train data in the model. So
we do all that data prep. And you can
see down here we're only going to have a
couple lines of code where we're
actually building our model and training
it. That's one of the cool things about
Python and how far we've come. It's such
an exciting time to be in machine
learning because there's so many
automated tools. Let's see. Before we do
this, let's do a quick length of and
let's do y. We want let's just do length
of y. And we get 768. And if we import
math, we do math dot square root. Let's
do y train. There we go. It's actually
supposed to be x train. Before we do
this, let's go ahead and do import math
and do math square root length of y
test. And when I run that, we get
12.409.
I want to see show you where this number
comes from. We're about to use 12 is an
even number. So if you know if you're
ever voting on things, remember the
neighbors all vote. Don't want to have
an even number of neighbors voting. So
we want to do something odd. And let's
just take one away. We'll make it 11.
Let me delete this out of here. That's
one of the reasons I love Jupyter
Notebook because you can flip around and
do all kinds of things on the fly. So,
we'll go ahead and put in our
classifier. We're creating our
classifier now and it's going to be the
K neighbors classifier. In neighbors
equal 11. Remember, we did 12 - 1 for
11. So, we have an odd number of
neighbors. P= 2 because we're looking
for is it are they diabetic or not? And
we're using the ukitian metric. There
are other means of measuring the
distance. You could do like square
square means value. There's all kinds of
measure this, but the uklidian is the
most common one and it works quite well.
It's important to evaluate the model.
Let's use the confusion matrix to do
that. And we're going to use the
confusion matrix. Wonderful tool. And
then we'll jump into the F1 score. And
finally, accuracy score, which is
probably the most commonly used quoted
number when you go into a meeting or
something like that. So, let's go ahead
and paste that in there. And we'll set
the CM equal to confusion matrix. Y
test, Y predict. So those are the two
values we're going to put in there. And
let me go ahead and run that and print
it out. And the way you interpret this
is you have the Y predicted, which would
be your title up here. You can do uh
let's just do Predicted
across the top and actual going down.
Actual. It's always hard to to write in
here. Actual. That means that this
column here down the middle, that's the
important column. And it means that our
prediction said 94 and prediction in the
actual agreed on 94 and 32. This number
here, the 13 and the 15, those are what
was wrong. So you could have like three
different if you're looking at this
across three different variables instead
of just two. You'd end up with a third
row down here in the column going down
the middle. So in the first case, we
have the the and I believe the zero is a
94 people who don't have diabetes. The
prediction said that 13 of those people
did have diabetes and were at high risk.
And the 32 that had diabetes had
correct, but our prediction said another
15 out of that 15, it classified as
incorrect. So you can see where that
classification comes in and how that
works on the confusion matrix. Then
we're going to go ahead and print the F1
score. Let me just run that. And you see
we get a 69 in our F1 score. The F1
takes into account both sides of the
balance of false positives where if we
go ahead and just do the accuracy
account and that's what most people
think of is it looks at just how many we
got right out of how many we got wrong.
So a lot of people when you're a data
scientist and you're talking to other
data scientists they're going to ask you
what the F1 score the Fore is. If you're
talking to the general public or the uh
decision makers in the business, they're
going to ask what the accuracy is. And
the accuracy is always better than the
F1 score. But the F1 score is more
telling. It lets us know that there's
more false positives than we would like
on here. But 82% not too bad for a quick
flash look at people's different
statistics and running an sklearn and
running the KNN, the K nearest neighbor
on it. So we have created a model using
KN&N which can predict whether a person
will have diabetes or not or at the very
least whether they should go get a
checkup and have their glucose checked
regularly or not. The print accuracy
score we got the 0818 was pretty close
to what we got and we can pretty much
round that off and just say we have an
accuracy of 80%. Tells us it is a pretty
fair fit in the model.
>> So what is game means clustering? C
means clustering is an unsupervised
learning algorithm. In this case, you
don't have labeled data unlike in
supervised learning. So you have a set
of data and you want to group them and
as the name suggests, you want to put
them into clusters which means objects
that are similar in nature, similar in
characteristics need to be put together.
So that's what K means clustering is all
about. The term K is basically is a
number. So we need to tell the system
how many clusters we need to perform. So
if K is equal to two, there will be two
clusters. If K is equal to three, three
clusters and so on and so forth. That's
what the K stands for. And of course
there is a way of finding out what is
the best or optimum value of K for a
given data. We will look at that. So
that is K means clustering. So let's
take an example. C means clustering is
used in many many scenarios but let's
take an example of cricket the game of
cricket let's say you received data of a
lot of players from maybe all over the
country or all over the world and this
data has information about the runs
scored by the people or by the player
and the wickets taken by the player and
based on this information we need to
cluster this data into two clusters
batsmen and bowlers. So this is an
interesting example. Let's see how we
can perform this. So we have the data
which consists of primarily two
characteristics which is the runs and
the wickets. So the bowlers basically
take wickets and the batsmen score runs.
There will be of course a few bowlers
who can score some runs and similarly
there will be some batsmen who will who
would have taken a few wickets. But with
this information, we want to cluster
this players into batsmen and bowlers.
So how does this work? Let's say this is
how the data is. So there are
information there is information on the
y-axis about the run scored and on the
x-axis about the wickets taken by the
players. So if we do a quick plot, this
is how it would look. And um when we do
the clustering, we need to have the
clusters like shown in the third diagram
out here. We need to have a cluster
which consists of people who have scored
high runs which is basically the
batsmen. And then we need a cluster with
people who have taken a lot of wickets
which is typically the bowlers. There
may be a certain amount of overlap but
we will not talk about it right now. So
with K means clustering we will have
here that means K is equal to two and we
will have two clusters which is batsmen
and bowlers. So how does this work? The
way it works is the first step in K
means clustering is the allocation of
two centroidids randomly. So two points
are assigned as so-called centrids. So
in this case we want two clusters which
means K is equal to two. So two points
have been randomly assigned as centrids.
Keep in mind these points can be
anywhere. There are random points. They
are not initially they are not really
the centroidids. Centr means it's a
central point of a given data set. But
in this case when it starts off it's not
really the centroid. Okay. So these
points though in our presentation here
we have shown them one point closer to
these data points and another closer to
these data points. They can be assigned
randomly anywhere. Okay. So that's the
first step. The next step is to
determine the distance of each of the
data points from each of the randomly
assigned centrids. So for example we
take this point and find the distance
from this centr and the distance from
this cent. This point is taken and the
distance is found from this centroid and
this c and so on and so forth. So for
every point the distance is measured
from both the centroids and then
whichever distance is less that point is
assigned to that centroid. So for
example in this case visually it is very
obvious that all these data points are
assigned to this centroid and all these
data points are assigned to this
centroid and that's what is represented
here in blue color and in this yellow
color. The next step is to actually
determine the central point or the
actual centrid for these two clusters.
So we have this one initial cluster,
this one initial cluster. But as you can
see these points are not really the
centroid. Centroid means it should be
the central position of this data set.
Central position of this data set. So
that is what needs to be determined as
the next step. So the central point of
the actual centrid is determined and the
original randomly allocated centr is
repositioned to the actual centroid of
this new clusters and this process is
actually repeated. Now what might happen
is some of these points may get
reallocated. In our example that is not
happening probably but it may so happen
that the distance is found between each
of these data points once again with
these centroidids. And if there is if it
is required some points may be
reallocated. We will see that in a later
example but for now we will keep it
simple. So this process is continued
till the centrid repositioning stops and
that is our final cluster. So this is
our so after iteration we come to this
position this situation where the
centroid doesn't need any more
repositioning and that means our
algorithm has converged convergence has
occurred and we have the cluster two
clusters we have the clusters with a
centroid. So this process is repeated.
The process of calculating the distance
and repositioning the centrid is
repeated till the repositioning stops
which means that the algorithm has
converged and we have the final cluster
with the data points and the
centroidids. So this is what you're
going to learn from this session. We
will talk about the types of clustering.
What is K means clustering? application
of K means clustering. C means
clustering is done using distance
measure. So we will talk about the
common distance measures and then we
will talk about how K means clustering
works and go into the details of K means
clustering algorithm and then we will
end with a demo and a use case for K
means clustering. So let's begin. First
of all, what are the types of
clustering? There are primarily two
categories of clustering. hierarchical
clustering and then partitional
clustering and each of these categories
are further subdivided into elomerative
and divisive clustering and K means and
fuzzy C means clustering. Let's take a
quick look at what each of these types
of clustering are. In hierarchical
clustering, the clusters have a treelike
structure and hierarchical clustering is
further divided into elomerative and
divisive. Elomemerative clustering is a
bottomup approach. We begin with each
element as a separate cluster and merge
them into successively larger clusters.
So for example, we have A B CDE E F. We
start by combining B and C form one
cluster. D and E form one more. Then we
combine D, E and F one more bigger
cluster and then add BC to that and then
finally A to it. Compared to that
divisive clustering or divisive
clustering is a top- down approach. We
begin with the whole set and proceed to
divide it into successively smaller
clusters. So we have ABCDE E F. We first
take that as a single cluster and then
break it down into A B C D E and F. Then
we have partitional clustering split
into two subtypes. K means clustering
and fuzzy C means. In K means clustering
the objects are divided into the number
of clusters mentioned by the number K.
That's where the K comes from. So if we
say K is equal to two, the objects are
divided into two clusters C1 and C2. And
the way it is done is the features or
characteristics are compared and all
objects having similar characteristics
are clubed together. So that's how K
means clustering is done. We will see it
in more detail as we move forward. And
fuzzy C means is very similar to K means
in the sense that it clubs objects that
have similar characteristics together.
But while in K means clustering two
objects cannot belong to or any object a
single object cannot belong to two
different clusters in C means objects
can belong to more than one cluster. So
that is the primary difference between K
means and fuzzy C means. So what are
some of the applications of K means
clustering? C means clustering is used
in a variety of examples or variety of
business cases in real life starting
from academic performance, diagnostic
systems, search engines and wireless
sensor networks and many more. So let us
take a little deeper look at each of
these examples. Academic performance. So
based on the scores of the students,
students are categorized into A, B, C
and so on. Clustering forms a backbone
of search engines. When a search is
performed, the search results need to be
grouped together. The search engines
very often use clustering to do this.
And similarly, in case of wireless
sensor networks, the clustering
algorithm plays the role of finding the
cluster heads which collects all the
data in its respective cluster. So
clustering especially K means clustering
uses distance measure. So let's take a
look at what is distance measure. So
while these are the different types of
clustering in this video we will focus
on K means clustering. So distance
measure tells how similar some objects
are. So the similarity is measured using
what is known as distance measure and
what are the various types of distance
measures. There is ukidian distance.
There is Manhattan distance. Then we
have squared ukitian distance measure
and cosine distance measure. These are
some of the distance measures supported
by k means clustering. Let's take a look
at each of these. What is ukidian
distance measure? This is nothing but
the distance between two points. So we
have learned in high school how to find
the distance between two points. This is
a little sophisticated formula for that.
But we know a simpler one is square
roo of y2 - y1 square + x2 - x1
square. So this is an extension of that
formula. So that is the ukidian distance
between two points. What is the squared
ukidian distance measure? It's nothing
but the square of the ukidian distance
as the name suggests. So instead of
taking the square root, we leave the
square as it is. And then we have
Manhattan distance measure. In case of
Manhattan distance, it is the sum of the
distances across the x-axis and the
y-axis. And note that we are taking the
absolute value so that the negative
values don't come into play. So that is
the Manhattan distance measure. Then we
have cosine distance measure. In this
case, we take the angle between the two
vectors formed by joining the points
from the origin. So that is the cosine
distance measure. Okay. So that was a
quick overview about the various
distance measures that are supported by
K means. Now let's go and check how
exactly K means clustering works. Okay.
So this is how K means clustering works.
This is like a flowchart of the whole
process. There is a starting point and
then we specify the number of clusters
that we want. Now there are couple of
ways of doing this. We can do by trial
and error. So we specify a certain
number maybe k is equal to 3 or four or
five to start with and then as we
progress we keep changing until we get
the best clusters or there is a
technique called elbow technique whereby
we can determine the value of k. What
should be the best value of k? How many
clusters should be formed? So once we
have the value of K we specify that and
then the system will assign that many
centrids. So it picks randomly that to
start with randomly that many points
that are considered to be the centrids
of these clusters and then it measures
the distance of each of the data points
from these centroidids and assigns those
points to the corresponding centr from
which the distance is minimum. So each
data point will be assigned to the
centrid which is closest to it and
thereby we have k number of initial
clusters. However this is not the final
clusters. The next step it does is for
the new groups for the clusters that
have been formed it calculates the mean
position thereby calculates the new
centroid position. the position of the
centrid moves compared to the randomly
allocated one. So it's an iterative
process. Once again the distance of each
point is measured from this new centroid
point and if required the data points
are reallocated to the new centroidids
and the mean position or the new centrid
is calculated once again. If the centrid
moves then the iteration continues which
means the convergence has not happened.
The clustering has not converged. So as
long as there is a movement of the
centrid this iteration keeps happening.
But once the centrid stops moving which
means that the cluster has converged or
the clustering process has converged
that will be the end result. So now we
have the final position of the centroid
and the data points are allocated
accordingly to the closest centrid. I
know it's a little difficult to
understand from this simple flowchart.
So let's do a little bit of
visualization and see if we can explain
it better. Let's take an example. If we
have a data set for a grocery shop. So
let's say we have a data set for a
grocery shop and now we want to find out
how many clusters this has to be spread
across. So how do we find the optimum
number of clusters? There is a technique
called the elbow method. So when these
clusters are formed, there is a
parameter called within sum of squares.
And the lower this value is, the better
the cluster is. That means all these
points are very close to each other. So
we use this within sum of squares as a
measure to find the optimum number of
clusters that can be formed for a given
data set. So we create clusters or we
let the system create clusters of a
variety of numbers maybe of 10 10
clusters and for each value of K the
within SS is measured and the value of K
which has the least amount of within SS
or WSS that is taken as the optimum
value of K. So this is the diagrammatic
representation. So we have on the y-axis
the within sum of squares or wss and on
the x-axis we have the number of
clusters. So as you can imagine if you
have k is equal to one which means all
the data points are in a single cluster
the within ss value will be very high
because they are probably scattered all
over. The moment you split it into two
there will be a drastic fall in the
within ss value and that's what is
represented here. But then as the value
of K increases the decrease the rate of
decrease will not be so high. It will
continue to decrease but probably the
rate of decrease will not be high. So
that gives us an idea. So from here we
get an idea for example the optimum
value of K should be either two or three
or at the most four but beyond that
increasing the number of clusters is not
dramatically changing the value in WSS
because that pretty much gets
stabilized. Okay. Now that we have got
the value of K and let's assume that
these are our delivery points. The next
step is basically to assign two centrids
randomly. So let's say C1 and C2 are the
centrids assigned randomly. Now the
distance of each location from the
centrid is measured and each point is
assigned to the centrid which is closest
to it. So for example these points are
very obvious that these are closest to
C1 whereas this point is far away from
C2. So these points will be assigned
which are close to C1 will be assigned
to C1 and these points or locations
which are close to C2 will be assigned
to C2. And then so this is the how the
initial grouping is done. This is part
of C1 and this is part of C2. Then the
next step is to calculate the actual
centrid of this data because remember C1
and C2 are not the centrids. They've
been randomly assigned points and only
thing that has been done was the data
points which are closest to them have
been assigned to them. But now in this
step the actual centroid will be
calculated which may be for each of
these data sets somewhere in the middle.
So that's like the main point that will
be calculated and the centr will
actually be positioned or repositioned
there. Same with C2. So the new centroid
for this group is C2. this new position
and C1 is in this new position. Once
again, the distance of each of the data
points is calculated from these
centroids. Now remember, it's not
necessary that the distance still
remains the or each of these data points
still remain in the same group. By
recalculating the distance, it may be
possible that some points get
reallocated like so. You see this? So
this point earlier was closer to C2
because C2 was here. But after
recalculating repositioning it is
observed that this is closer to C1 than
C2. So this is the new grouping. So some
points will be reassigned. And again the
centrid will be calculated and if the
centroid doesn't change so that is a
repetative process, iterative process.
And if the centroid doesn't change once
the centroid stops changing that means
the algorithm has converged and this is
our final cluster with this as the
centroid C1 and C2 as the centroids
these data points as a part of each
cluster. So I hope this helps in
understanding the whole process
iterative process of K means clustering.
So let's take a look at the K means
clustering algorithm. Let's say we have
x1, x2, x3, n number of points as our
inputs and we want to split this into k
clusters or we want to create k
clusters. So the first step is to
randomly pick k points and call them
centroidids. They are not real centrids
because centr is supposed to be a center
point but they are just called centrids.
And we calculate the distance of each
and every input point from each of the
centroidids. So the distance of X1 from
C1 from C2 C3 each of the distances we
calculate and then find out which
distance is the lowest and assign X1 to
that particular random centroid. Repeat
that process for X2. calculate its
distance from each of the centroid C1,
C2, C3 up to CK and find which is the
lowest distance and assign X2 to that
particular centroid. Same with X3 and so
on. So that is the first round of
assignment that is done. Now we have K
groups because there are we have
assigned the value of K. So there are K
centroids and uh so there are K groups.
All these inputs have been split into K
groups. However, remember we picked the
centrids randomly. So they are not real
centrids. So now what we have to do, we
have to calculate the actual centroids
for each of these groups which is like
the mean position which means that the
position of the randomly selected
centrids will now change and they will
be the main positions of these newly
formed K groups. And once that is done,
we once again repeat this process of
calculating the distance. Right? So this
is what we are doing as a part of step
four. We repeat step two and three. So
we again calculate the distance of X1
from the centroid C1, C2, C3 and then
see which is the lowest value and assign
X1 to that. Calculate the distance of X2
from C1, C2, C3 or whatever up to CK and
find whichever is the lowest distance
and assign X2 to that centroid and so
on. In this process there may be some
reassignment. X1 was probably assigned
to cluster C2 and after doing this
calculation maybe now X1 is assigned to
C1. So that kind of reallocation may
happen. So we repeat the steps two and
three till the position of the centrids
don't change or stop changing and that's
when we have convergence. So let's take
a detailed look at at each of these
steps. So we randomly pick K cluster
centers. We call them centroidids
because they are not initially they are
not really the centrids. So we let us
name them C1 C2 up to CK. And then step
two, we assign each data point to the
closest center. So what we do, we
calculate the distance of each X value
from each C value. So the distance
between X1 C1 distance between X1 C2 X1
C3 and then we find which is the lowest
value. Right? That's the minimum value
we find and assign X1 to that particular
centroid. Then we go next to x2. Find
the distance of x2 from c1, x2 from c2,
x2 from c3 and so on up to ck. And then
assign it to the point or to the
centroid which has the lowest value and
so on. So that is step number two. In
step number three, we now find the
actual centr for each group. So what has
happened as a part of step number two?
We now have all the points, all the data
points grouped into K groups because we
we wanted to create K clusters, right?
So we have K groups. Each one may be
having a certain number of input values.
They need not be equally distributed. By
the way, based on the distance, we will
have K groups. But remember the initial
values of the C1 C2 were not really the
centrids of these groups, right? we
assign them randomly. So now in step
three, we actually calculate the centr
of each group which means the original
point which we thought was the centrid
will shift to the new position which is
the actual centrid for each of these
groups. Okay? And we again calculate the
distance. So we go back to step two
which is what we calculate again the
distance of each of these points from
the newly positioned centroidids and if
required we reassign these points to the
new centroidids. So as I said earlier
there may be a reallocation. So we now
have a new set or a new group. We still
have K groups but the number of items
and the actual assignment may be
different from what was in step two
here. Okay, so that might change. Then
we perform step three once again to find
the new centroid of this new group. So
we have again a new set of clusters, new
centroidids and new assignments. We
repeat this step two again. Once again
we find and then it is possible that
after iterating through three or four or
five times the centrid will stop moving
in the sense that when you calculate the
new value of the centrid that will be
same as the original value or there will
be very marginal change. So that is when
we say convergence has occurred and that
is our final cluster. That's the
formation of the final cluster. All
right. So let's see a couple of demos of
uh K means clustering. We will actually
see some live demos in uh Python
notebook using Python notebook. But
before that let's find out what's the
problem that we are trying to solve. The
problem statement is let's say Walmart
wants to open a chain of stores across
the state of Florida and uh it wants to
find the optimal store locations. Now
the issue here is if they open too many
stores close to each other obviously the
they will not make profit but if they if
the stores are too far apart then they
will not have enough sales. So how do
they optimize this? Now for an
organization like Walmart which is an
e-commerce giant they already have the
addresses of their customers in their
database. So they can actually use this
information or this data and use K means
clustering to find the optimal location.
Now before we go into the Python
notebook and show you the live code, I
wanted to take you through very quickly
a summary of the code in the slides and
then we will go into the Python
notebook. So in this block we are
basically importing all the required
libraries like numpy, mattplot lib and
so on and we are loading the data that
is available in the form of let's say
the addresses for simplicity sake we
will just take them as some data points.
Then the next thing we do is quickly do
a scatter plot to see how they are
related to each other with respect to
each other. So in the scatter plot we
see that there are a few distinct groups
already being formed. So you can
actually get an idea about how the
cluster would look and how many clusters
what is the optimal number of clusters
and then starts the actual K means
clustering process. So we will assign
each of these points to the centrids and
then check whether they are the optimal
distance which is the shortest distance
and assign each of the points data
points to the centroidids and then go
through this iterative process till the
whole process converges and finally we
get an output like this. So we have four
distinct clusters and um which is we can
say that this is how the population is
probably distributed across Florida
state and uh these centroidids are like
the location where the store should be
the optimum location where the store
should be. So that's the way we
determine the best locations for the
store and that's how we can help Walmart
find the best locations for their stores
in Florida. So now let's take this into
Python notebook. Let's see how this
looks when we are learning running the
code live. All right. So this is the
code for K means clustering in Jupyter
notebook. We have a few examples here
which we will demonstrate how K means
clustering is used and even there is a
small implementation of K means
clustering as well. Okay. So let's get
started. Okay. So this block is
basically importing the various
libraries that are required like
mattplot lib and numpy and so on and so
forth which would be used as a part of
the code. Then we are going and creating
blobs which are similar to clusters. Now
this is a very neat feature which is
available in scikitlearn. Make blobs is
a nice feature which creates clusters of
data sets. So that's a wonderful
functionality that is readily available
for us to create some test data kind of
thing. Okay. So that's exactly what we
are doing here. We are using make blobs
and we can specify how many clusters we
want. So centers we are mentioning here.
So it will go ahead and so we just
mentioned four. So it will go ahead and
create some test data for us. And this
is how it looks. As you can see visually
also we can figure out that there are
four distinct classes or clusters in
this data set. And that is what make
blobs actually provides. Now from here
onwards we will basically run the
standard K means functionality that is
readily available. So we really don't
have to implement K means itself. The C
means functionality or the the function
is readily available. You just need to
feed the data and we'll create the
clusters. So this is the code for that.
We import k means and then we create an
instance of k means and we specify the
value of k. This n_clusters is the value
of k. Remember K means in K means K is
basically the number of clusters that
you want to create and it is a integer
value. So this is where we are
specifying that. So we have K is equal
to four and so that instance is created.
We take that instance and as with any
other machine learning functionality fit
is what we use the function or the
method rather fit is what we use to
train the model. Here there is no real
training uh kind of thing but that's the
call. Okay. So we are calling fit and
what we are doing here we are just
passing the data. So x has these values
the data that has been created right. So
that is what we are passing here and uh
this will go ahead and create the
clusters and uh then we are using
after doing uh fit we run the predict
which basically assigns for each of
these observations which cluster it
belongs to. All right. So it will name
the clusters. Maybe this is cluster one.
This is two, three and so on. Or will
actually start from zero, cluster 0, 1,
2 and 3 maybe. And then for each of the
observations it will assign based on
which cluster it belongs to it will
assign a value. So that is stored in y_k
means when we call predict that is what
it does. And we can take a quick look at
these uh y_k means or the cluster
numbers that have been assigned for each
observation. So this is the cluster
number assigned for observation one.
Maybe this is for observation two,
observation three and so on. So we have
how many about I think 300 samples
right? So all the 300 samples there are
300 values here. Each of them the
cluster number is given and the cluster
number goes from 0 to three. So there
are four clusters. So the numbers go
from 0 1 2 3. So that's what is seen
here. Okay. Now, so this was a quick
example of generating some dummy data
and then clustering that. Okay. And this
can be applied if you have proper data.
You can just load it up into X for
example here and then run the K. So this
is the central part of the K means
clustering program example. So you
basically create an instance and you
mention how many clusters you want by
specifying this parameter n_clusters and
that is also the value of k and then
pass the data to get the values. Now the
next section of this code is the
implementation of a k means. Now this is
kind of a a rough implementation of the
k means algorithm. So we will just walk
you through I will walk you through the
code uh at each step what it is doing
and then we will see a couple of more
examples of how K means clustering can
be used in maybe some real life examples
real life use cases. All right. So in
this case here what we're doing is
basically implementing K means
clustering and there is a function or a
library calculates for a given two pairs
of points it will calculate the the
distance between them and see which one
is the closest and so on. So this is
like this is pretty much like what K
means does right. So it calculates the
distance of each point or each data set
from predefined centroid and then based
on whichever is the lowest this
particular data point is assigned to
that centroid. So that is basically
available as a standard function and we
will be using that here. So as explained
in the slides the first step that is
done in case of C means clustering is to
randomly assign some centrides. So as a
first step we randomly allocate a couple
of centrids which we call here we're
calling as centers
and then we put this in a loop and we
take it through an iterative process.
For each of the data points, we first
find out using this function pair-wise
distance argument. For each of the
points, we find out which one which
center or which uh randomly selected
centrid is the closest and accordingly
we assign that data or the data point to
that particular centrid or cluster. And
once that is done for all the data
points, we calculate the new centr by
finding out the mean position with the
the center position. Right? So we
calculate the new centroid and then we
check if the new centroid is the
coordinates or the position is the same
as the previous centroid. The positions
we will compare and if it is the same
that means the process has converged. So
remember we do this process till the
centroidids or the centrid doesn't move
anymore right so the centroid gets
relocated each time this reallocation is
done so the moment it doesn't change
anymore the position of the cent doesn't
change anymore we know that convergence
has occurred so till then so you see
here this is like an infinite loop while
true is an infinite loop it only breaks
when the centers are the same the new
center and old center positions are the
name and once that is uh done uh we
return the centers and the labels. Now
of course as explained this is not a
very sophisticated and advanced
implementation very basic implementation
because one of the flaws in this is that
sometimes what happens is the centroid
the position will keep moving but in the
change will be very minor. So in that
case also that is actually convergence
right. So for example the change is 0.1
we can consider that as convergence
otherwise what will happen is this will
either take forever or it will be never
ending. So that's a small flaw here. So
that is something additional checks may
have to be added here. But again as
mentioned this is not the most
sophisticated uh implementation. This is
like a kind of a rough implementation of
the k means clustering. Okay. So if we
execute this code this is what we get as
the output. So this is the definition of
this particular function and then we
call that find_clusters and we pass our
data x and the number of clusters which
is four and if we run that and plot it
this is the output that we get. So this
is of course each cluster is represented
by a different color. So we have a
cluster in green color, yellow color and
so on and so forth. And these big points
here these are the centroidids is the
final position of the centroidids. And
as you can see visually also this
appears like a kind of a center of all
these points here. Right? Similarly this
is like the center of all these points
here and so on. So this is the example
or this is an example of a
implementation of K means clustering and
uh next we will move on to see a couple
of examples of how K means clustering is
used in maybe some real life scenarios
or use cases. In the next example or
demo, we are going to see how we can use
K means clustering to perform color
compression. We will take a couple of
images. So there will be two examples
and uh we will try to use C means
clustering to compress the colors. This
is a common situation in image
processing when you have an image with
millions of uh colors but then you
cannot render it on some devices which
may not have enough memory. Uh so that
is the scenario where where something
like this can be used. So before again
we go into the Python notebook let's
take a look at quickly the the code. As
usual we import the libraries and then
we import the image and uh then we will
flatten it. So the reshaping is
basically we have the image information
is stored in the form of pixels and uh
if the image is like for example 427x
640 and it has three colors. So that's
the overall dimension of the of the
initial image. we just reshape it and um
then feed this to our algorithm and this
will then create clusters of only 16
clusters. So this this colors there are
millions of colors and now we need to
bring it down to 16 colors. So we use k
is equal to 16 and u this is how when we
visualize this is how it looks. There
are these are all about 16 million
possible colors. The input color space
has 16 million possible colors and we
just sub compress it to 16 colors. So
this is how it would look when we
compress it to 16 colors. And this is
how the original image looks. And after
compression to 16 colors, this is how
the new image looks. As you can see,
there is not a lot of information that
has been lost. though the image quality
is definitely reduced a little bit. So
this is an example which we are going to
now see in Python notebook. Let's go
into the Python and once again as always
we will import some libraries and load
this image called flower.jpg.
Okay. So let we'll load that and this is
how it looks. This is the original image
which has I think 16 million colors and
uh this is the shape of this image which
is basically what is the shape is
nothing but the overall size right so
this is 427 pixel by 640 pixel and then
there are three layers which is this
three basically is for RGB which is red
green blue so color image will have that
right so that is the shape of this now
what we need to do is data let's take a
look at how data is looking. So let me
just create a new cell and show you what
is in data. Basically we have captured
this information.
So data is what? Let me just show you
here.
All right. So let's take a look at
China. What are the values in China? And
uh if you see here, this is how the data
is stored. This is nothing but the pixel
values. Okay? So this is like a matrix
and each one has about for for this 427x
640 pixels. All right. So this is how it
looks. Now the issue here is these
values are large. The numbers are large.
So we need to normalize them to between
0 and one. Right? So that's why we will
basically create one more variable which
is data which will contain the values
between 0 and one. And the way to do
that is divide by 255. So we divide
China by 255 and we get the new values
in data. So let's just run this uh piece
of code and this is the shape. So we now
have also yeah what we have done is we
changed using reshape we converted into
the three-dimensional into a
two-dimensional data set. And let us
also take a look at how
let me just insert
probably a cell here and take a look at
how data is looking. All right. So this
is how data is looking and now you see
this is the values are between 0 and
one. Right? So if you earlier noticed in
case of China the values were large
numbers. Now everything is between 0 and
one. This is one of the things we need
to do. All right. So after that the next
thing that we need to do is to visualize
this and uh we can take random set of
maybe 10,000 points and plot it and
check and see how this looks. So let us
just plot this and uh so this is how the
original the color the pixel
distribution is. These are two plots one
is red against green and another is red
against blue and this is the original
distribution of the color. So then what
we will do is we will use K means
clustering to create just 16 clusters
for the various colors and then apply
that to the image. Now what will happen
is since the data is large because there
are millions of colors using regular K
means may be a little time consuming. So
there is another version of K means
which is called mini batch K means. So
we will use that which is which
processes in the overall concept remains
the same but this basically processes it
in smaller batches. That's the only
thing. Okay. So the results will pretty
much be the same. So let's go ahead and
execute this piece of code and also
visualize this so that we can see that
there are the this is how the 16 colors
uh would look. So this is red against
green and this is red against blue.
there is uh quite a bit of similarity
between this original color schema and
the new one. Right? So it doesn't look
very very completely different or
anything like that. Now we apply this
the newly created colors to the image
and uh we can take a look uh how this is
uh looking. Now we can compare both the
images. So this is our original image
and this is our new image. So as you can
see there is not a lot of information
that has been lost. uh it pretty much
looks like the original image. Yes, we
can see that for example here there is a
little bit uh it appears a little
dullish compared to this one right
because uh we kind of took off some of
the finer details of the color but
overall the highle information has been
maintained. At the same time, the main
advantage is that now this can be this
is an image which can be rendered on a
device which may not be that very
sophisticated. Now let's take one more
example with a different image. In the
second example, we will take an image of
the summer palace in China and we repeat
the same process. This is a high
definition color image with millions of
colors and also uh three-dimensional.
Now we will reduce that to 16 colors
using K means clustering. And um we do
the same process like before. We reshape
it and then we cluster the colors to 16
and then we render the image once again.
And we will see that the color the
quality of the image is slightly
deteriorates. As you can see here, this
has much finer details in this which are
probably missing here. But then that's
the compromise because there are some
devices which may not be able to handle
this kind of a high density images. So
let's run this code in Python notebook.
All right. So let's apply the same
technique for another picture which is
uh even more intricate and has probably
much complicated color schema. So this
is the image. Now once again uh we can
take a look at the shape which is 427x
640x3
and this is the new data would look
somewhat like this compared to the
flower image. So we have some new values
here and we will also bring this as you
can see the numbers are much big. So we
will much bigger so we will now have to
scale them down to values between 0 and
one. And that is done by dividing by
255. So let's go ahead and uh do that
and reshape it. Okay. So we get a
two-dimensional matrix and uh we will
then as the next step we will go ahead
and visualize this how it looks the the
16 colors and this is basically how it
would look 16 million colors. And now we
can create the clusters out of this. The
16 K means clusters we will create. So
this is how the distribution of the
pixels would look with 16 colors. And
then we go ahead and uh apply this and
visualize how it is looking for with the
with the new just the 16 color. So once
again, as you can see, this looks much
richer in color, but at the same time,
and this probably doesn't have, as we
can see, it doesn't look as rich as this
one, but nevertheless, the information
is not lost, the shape and all that
stuff. And this can be also rendered on
a slightly a device which is probably
not that sophisticated. Okay, so that's
pretty much it. So we have seen two
examples of how color compression can be
done uh using K means clustering and we
have also seen in the previous examples
of how to implement C means the code to
roughly how to implement C means
clustering and we use some sample data
using blob to just execute the C means
cluster that takes place after data
collection and before statistical
analysis. So before you conduct any
statistical formulas and analysis on the
data and squeeze the data to extract
some valuable insights, the process
which you perform is called as initial
data analysis. Like taking the data from
the source, cleaning the data,
transforming the data into a readable
format and using that readable data to
build some basic charts what exactly is
happening with this particular company,
brand or anything. Let's say I give you
some data from the company. Then you get
some insights of it. How many number of
traffic you received? How many number of
orders you received? What's the sale
that you made in a specific month,
specific quarter or specific year? And
what was the profit? So basic
information which you convert from the
data and create a dashboard. Right? That
is called as initial data analysis. So a
step beyond initial data analysis is
known as the exploratory data analysis.
This is where you perform some
statistics and probability and predict
the future. Right? So let's dive deep
and learn what exactly is exploratory
data analysis. So a simple definition
for exploratory data analysis is as
follows. Exploratory data analysis is a
key step in data analysis process that
helps you identify patterns, outliners
and relationships between variables
before making assumptions. It is not
like you just create a dashboard out of
the initial data analysis and you can
predict the future. No, it's not like
that. You might have to last. You might
have to go through some permutations and
combinations. You might have to check
the seasons. You might have to check the
possibilities, right? During some
particular seasons in the year, let's
say it's Christmas, then you can expect
some good sales. Let's say it's some
festival, it's some special occasion,
you can expect some good sales on the
product, right? and maybe a part of the
year, maybe a part of 10 years, a
decade, right? In a certain period of
time, there might be some reason due to
which the sales of a certain product
were high. So to make sure that your
assumptions to make sure that your
projection of the sales is 100% accurate
or at least 90 to 95% accurate then you
might have to go through the exploratory
data analysis where you make use of
statistics in your data analysis. Now
there are certain steps that you might
have to follow while going through
exploratory data analysis. So following
are the steps. So a first few steps
might be slightly relevant to initial
data analysis like connecting data,
cleaning it, transforming it and loading
it. After that you will import certain
libraries from Python and after that you
read the data what exactly you have in
your data. The number of columns, the
number of rows and if there are any null
values, if there are any uh entries
which are invalid, you might have to
check that, read that and you might have
to check for duplicate entries. It is
possible that one entry might have been
entered by two different people, right?
There might be a duplication. So you
might have to eliminate those
duplications. You might have to check
for some missing values. You might have
to calculate the total number of missing
values from the data set and try to
eliminate them from the calculation
during your exploratory data analysis.
Followed by that you have to do some
model engineering. Followed by that you
might have to do some feature
engineering creating features and then
you will get started with exploratory
data analysis and the at the end you
will generate a projection or a
prediction or give your assumption that
this might happen in the future and you
might have to take action to avoid it or
you might have to take action to
improvise it right so this is how the
steps in exploratory data analysis take
part now let's proceed and start with
our demo on Python's exploratory data
analysis and in this session we will be
using the use case that we discussed
before which happens to be the students
performance data set. So in this
particular data set we will be having
some columns based on physical activity
the distance from home parental
education the subjects the marks they
have scored in the previous exam. The
marks that they have scored in the
previous exam and if they have any
disabilities if they are having any
resources that they require to write the
exams right. So a few parameters the
important parameters that we will be
discussing in this session and followed
by that we will project the future that
how they will you know improve in their
exams and if there is a problem and if
there is a solution to it then we can
implement that solution and help
students to gain better marks in their
exam. So that's the overall use case for
this demonstration. Now let's get
started with our Jupyter notebook. Now
we are on Jupiter notebook. Now let's
get started. So I would like to have a
title for my um notebook. So I'll write
an HTML code for that.
So HTML code uh
and I have uh three apostrophes
and here I would like to write something
in H1. So I want my title to be in H1.
So
style will be
background color
name will be dark blue.
So we let's uh proceed with the simply
dance background which will be t usually
and the color
of text will be orange
and I also want to have the font let's
let me give the font size as 30.
And in the next line, I'd like to have
border
radius. I can give the border radius as
20 pixels
and padding
to be 16 pixels.
Text alignment, I'd like to keep it
center.
There you go. Let's code the Let's close
the H1. And now let's code the uh
background color or border color. So
B style.
So what we can do is we can basically
have this uh code here and what we will
do is we will reuse this particular code
segment because we will be having
multiple HTML codes in this particular
uh workbook which will explain the
results of the analysis that we are
doing. So the way we just run the code
and after that we will be getting some
visualizations and I will be writing
some textual content in XT in an HTML
page so that it will be easier for the
people to understand what's what exactly
is happening here right so the color
will be light blue and now comes the
text we will close this and here we will
write the text as
Python
explorate
data analysis
and we will
break here
the student performance
and here we will close the H1
and lastly we will display this in HTML
code. So basically we missed this
library. So we will be importing from
ipython display import html to display
this particular code. So without this
particular library we cannot display any
HTML codes in our notebook. So we will
quickly run that and we have a title
over here. Now let's proceed with the
next part. Now we will uh use some
libraries like CAD boost and light bgm.
So for that we might have to install the
these libraries. So we will be using pip
install here
pip install cat boost
and control enter to run this particular
code segment or you can also use run and
it's installed and after that we will
also install
light bgm
g sorry not bgm
so it's already installed
Now we will start importing the
libraries that we need. So we will be
needing numpy, panda, seabon, mattplot
lib and we will also import another
special library which is for ignoring
warnings. So we will import warnings and
after that from we will be importing
that warnings from IPython display
import clear output and after that we
will uh tell the jupyter notebook to
import if there are any warnings. So
that code will be warning dot filter
warnings and in the uh brackets we will
write ignore. So uh the basic libraries
which we will be needing are as follows.
import
numpy
as np. Let's quickly copy this and
proceed. Enter. And now we will be
needing pandas as pd.
Enter. And now we will be needing
seabbone
SNS.
And we will also use mattplot lib
py plot
as plt.
And after that we will import warnings
from
I Python
dot display
import
clear
output
warnings
dot filter warnings
ignore.
Now we'll just quickly run this query.
Run. And now we will proceed with
feature engineering.
So you can use a hashtag to ignore that
particular line from execution for
Jupyter notebook.
And here we will be importing import
from skarn
dot impute import
simple computer
from skarn
dot model
selection
import kf fold. There you go. Now let's
quickly run this query.
There you go. Now we will proceed with
modeling and model evaluation. Once we
are done with this then we will directly
input the data into our notebook. So for
modeling the data we will again use a
hash code so that this particular line
will not execute
and import some libraries lit
GBM
as LGB
from light GBM library.
import
LGBM regressor
from CAD boost import
boost regressor. There you go. Let's
quickly run this command.
There you go. Now, lastly, we have one
more task before importing the data that
is model evaluation. Once that is done,
we can proceed with importing the data.
There you go. Let's quickly run it. And
now so far so good. We have uh done the
basic library imports and feature
engineering is done, modeling is done
and model evaluation is also done. Now
we can begin with importing the data. So
we let's also add um the HTML code where
we have done this importing. So here I
will try to add another segment and I
will import the HTML code here which
displays a similar HTML format which
explains what exactly is happening here.
Just a moment I have the code ready.
I'll just paste it here. There you go.
Let's quickly run this so that we have a
HTML code here page here which explains
what exactly is happening here. So we're
importing libraries and also student
data. Now let's import the student data.
For that let's write the query. So we
are importing the data as data frame and
after that we will write pandas read CSV
right and here we will add the location
of the file. Right? So the file is
located in my downloads section. So
let's quickly copy that location and
paste it here. So this is the location
of my file. Let's quickly run it. So
there might be some error. It's okay. We
can resolve it. So in such scenarios
don't have to worry either you can add
an R but even if that doesn't work you
might have to change uh in this case it
worked but in case if it doesn't work
what you can do is uh you can eliminate
uh the r and you can just change this
from uh forward slash to backlash. This
will also help. So this could be worth
it. So this can also work. So these are
the situations where you can use this.
Now let's proceed with some more uh
interesting facts. Let's try to
understand what's going on with our uh
data set. Right? So what you can do is
now the data is stored in df variable as
a data frame. So what you can do is read
this particular data df.shape shape so
that you can understand what's the uh
what's happening with this data. Right?
So it can tell you that there are uh 6
sorry 6,67
rows and 20 columns. Now let's add
another query part and here you can uh
try to see the head right what head
means basically head means uh the column
headers. So what you can do is just
write df dot head
and run. Now you have the column headers
and couple of sample columns
sorry couple of sample rows. Now let's
proceed with uh checking the duplicates
and identifying the total number of
duplicate entries in this particular
data set. So you can write down df dot
duplicated
dot
sum and you will get the number of
duplicates present in this particular
file. So we have zero duplicate entries.
Now let's see if there are any null uh
elements in this particular data frame.
So df dot is null
dot sum.
So these are all functions. Control
enter. There you go. So in the column
teacher quality there are 78 null
entries and in parental educational
level there are 90 null entries and
distance from home there are u 67 null
entries. Now what we can do is uh from
the okay so from df shape we can add a
new cell here and we can write a HTML
code so that the viewer can understand
that we are trying to understand our
data. So let's write a quick code for
that. So let's not waste much time in
just writing the HTML code. So I've got
that HTML code written in a notepad
already. So I'll just paste it over here
and let's quickly run it so that we have
a HTML visibility here. So this was
supposed to be the result. So we will
add it here. So what we will do is
quickly edit this particular content.
We'll just quickly copy this code and
paste it over here.
Change the content from importing
libraries to
reading
student data and we will cut this code
from here and we will paste it in here
so that it get give us some information.
Let's also run this particular code
segment
so that we have trading student data.
There you go. Now the next part of this
session will be about creating a target
variable. So overall target of this
particular data analysis is about exam
score. Right? So we can name our target
variable as exam score. And let's
understand the distribution of this
particular exam score with uh the
variables we have. Now let's write down
plot dot figure.
Figure size should be around 15 comma 9
equals to let's add a bracket here 15
comma 9 or let's keep it as six 9 would
be a little bigger. Now enter now we
will use seabbond library here. SNS dot
count plot
x is equals to data frame cleaned
target. Okay. Uh before cleaned target
we might have to run a few more. Okay.
We did not perform data cleaning so far,
right? So let's proceed with data
cleaning so far. So we found some empty
entries, right? Null entries and we also
found some So here we have 299 rows
which have missing values. So we will
have to remove that. For that uh we
might have to create a new column which
has to be named as not um assigned right
df data frame not assigned which is
equals to df dot drop na. So we will be
dropping the null values here. It's a
function. And here let's print the
values. Print df dot df
na dot shape. So how many number of rows
and columns we have right and after that
let's also try to eliminate the null
values as well dot is null so we don't
have basically we don't have null values
but we have u some illegal entries maybe
some there you go now let's quickly run
this query so there you go so the new
data is about 678
20
so there is Um, okay. We did uh some
mistake here. So, we supposed to add it
as null. N U N L N N N N N N N N N N N N
N N N N N N L N N N N N N N N N N N N N
N N N N N N N N N N N N N N N N N N N N
N N N N N N N U L. Now quickly run this.
So we should not get any errors this
time.
There you go. No errors. So far so good.
Now let's describe the new data set. df
na dot describe. So these are the new
columns and rows that we have. And we
have uh mean, standard variation,
standard deviation, minimum, maximum. So
the scores are split into 25%, 50% and
75%. Which could be based on hours
studied which could be based on
attendance, sleep hours, previous
course, due training sessions,
everything. So uh minimum sleep hours,
maximum sleep hours, 25% of that, 50% of
that, 75% of that. So that is supposed
to be the u describe. Now what we will
do is we have a target variable which is
exam score. Right? Now we will do some
changes to it. We already know in an
exam there will be a threshold value. It
can be 25 marks per exam. It can be 50
marks per exam and it can be 100 marks
per exam. Right? In our situation let's
consider the threshold value is 100
marks. Right? If there is a situation
where marks is entered in a wrong way
right it can if they add if they wanted
to add 11 but by mistake if they added
uh another one right triple one it's not
a right entry right so we will try to
eliminate those kind of uh data
so we will create a new uh data frame
here which is dataf frame cleaned is
equals to dataf frame
not null so We have eliminated the null
values. BF NA. Now we will add our
target variable which is exam score
should be. So let's come out of this and
here we will add it as should be less
than or equal to 100 but not more than
100. Let's use square brackets.
Here we also the format is square
brackets. ing action now enter and
we will describe this particular
data set instead of the FNA we will copy
paste this here now let's run this
okay exam score is not identified let's
quickly check the error and resolve it
yeah so we missed out to add colons here
it's okay not a problem so this was
supposed to be how it is now let's run
this and we will have the answer over
here. So we have the output. Now let's
check the uh head of this particular
clean data set. So we can make use of
the same code here and paste it right
here and instead of describe let's write
head so that we have the header of uh
this particular data set. So we have our
study and everything normal and we will
categorize the data right. So we will
make use of three columns our study
attendance and previous scores and uh
after that we will also make use of
other columns in this particular data
set which happens to be the parental
involvement access to resources sleep
hours ting sessions etc. And now our
target will be the exam score that we
created over here. Right? This exam
score will be our target. And using this
particular exam score target, we will
categorize the data. Okay? We will
categorize the data in terms of uh let's
say uh first class uh second class and
uh pass or something like that. Right?
So if if a student is uh scoring below
64 and uh that is a separate category.
If the score student is scoring equals
to or above 65, that is a different
category. And if the student is scoring
beyond 70, that's a different category.
And uh before we proceed with that,
let's try to add a HTML code before this
particular data set so that we have u an
understanding of what exactly happened
here. So I let's uh I'll just quickly
copy paste this particular code here. So
we will run it and now next we will
describe check the data type column
separate and everything and we will you
know create data type category variables
right now so far so good. Now we will
create the categories.
So num call
equals to
hours studied
comma attendance.
So let's quickly add the data
previous exam scores.
Just a minute. Let's quickly add the
columns. Let me take a while. There you
go. I've added the columns. So we are
considering three different columns. uh
our study attendance and previous scores
for num call and cat call. We are
considering the other columns apart from
the first three and our target value is
exam score. Let's quickly run this.
There you go. And now we will try to
build some visualizations and before
that let's try to uh add some data uh
from in HTML. Let's try to create a
Okay, what we can do is simply copy this
particular HTML file here. We can take
this
and add it here so that we will
understand what exactly is happening
next. And in place of reading student
data, we will write data visualization
for student data.
And we will keep the colors same dark
blue background and u the color for data
visualization will be light blue and
student data will be orange. Let's
quickly run and there we have it. Now
our target variable is exam score. Right
now we will compare this particular
target variable with three other
parameters. So our parameters will be
the following uh as we discussed uh
creating the segregation in data set
right. So first will be distribution of
target variable with other parameters.
So we will create another HTML file for
that right here. Just a minute while I
paste the code for um HTML quickly run
this. There you go. So distribution of
target variable exam score against some
parameters. Now we will write the plot
for that
plot dot figure. So we are going to
consider the size equals to 15 6
big size
is equals to 15 6 the same one that we
considered before. And we will be using
Cbond SNS dot count plot
open bracket. This is a function x is
equals to df
clean. Okay. Uh what we can do is just
quickly take the column name so that we
don't create any mistakes here and we
will paste it over here instead of df.
There you go. Or target.
So our target is exam score,
and we will use the pellet as green
and the plot title will be distribution
of target variable exam score. We can
copy this. Okay, just a minute before
that plt do.
Should be let's use double quotes now.
Copy this and paste it here.
I think semicolons went off. Okay, not a
problem.
It's right here. Let's add a dot as a
full stop. And the next line, if you
need, you can add the full stop. If not
you can ignore plt dot grid true
which equals to major
comma
access is equals to y
comma line style equals to so I want
lines to be hyphen hyphen in this way I
want the lines and comma line width um
let's say 0.5 or 0.7. Let's go with 0.7
line width equals to 0.9 mm. There you
go. Now let's quickly run this query.
There you go. Done. And we have the
first visualization. So here you can see
there are some students which are
scoring 58 59 and you can see maximum
number of students are already scoring
good marks which is under 65 and uh
sorry which is under 70 and above 65 and
there is a good number of students uh
which are also scoring uh above 65 as
well right so sorry 70 70 as well. So we
have now less than or equal to 70 71.
And highest scorer in some situations
there is also 100. If you can see there
is slight growth here. There are a few
students toppers maybe which have
already scored 100 as well. Now we have
the list here. Now we what we need to do
is we need to segregate that is part one
which is less than or equal to 64 which
falls under 65 and another category
which falls in between 65 to 70 and
above 70. So we need to categorize these
three uh datas and segregate them as
bottom 65, top which is above 70 and mid
between 70 to 65. Right now before that
if you want to add an HTML document uh
sorry segment here which explains about
the distribution of target you can also
do that. It's already added here. Now
let's continue. But in case if you want
to uh add uh some data which explains
that we're trying to segregate, you can
also do that. I would like to do that.
Let's quickly uh add that HTML code
here. So what this particular code will
do is it will tell the percentage of
students which are scoring less than 65.
Number of stu uh percentage of students
uh scoring in between 65 and 70 and the
percent of students which are scoring
beyond 70. Right? Let's run this. And
here we have the result. Bottom 21.81%
scores under 64 while 24% scores over 70
and 50% are in between 65 to 69. Right
now let's uh continue with the
segregation part of the data. So for
segregation we will create three
different variables A, B, C. So first A
is equals to length of DF claimed. So
let's copy the column name sorry data
frame name length of DF cleaned inside
the square brackets we'll again add DF
cleaned of target variable which is exam
score
let's also add uh single quotes here
who are scoring in between or um less
than let's start with less than or equal
to 64 we'll not consider is 65 we'll
consider 64 divided by len of df cleaned
target variable exam score single quotes
into 100
which will give us the percentage now
similarly let's just copy and paste this
three more times for B and C. So here
instead of minus we're supposed to add
equals to and another one. So here
equals to and instead of A I will write
B and the last one is C. And instead of
64 we will add 70 here for top and here
we will make some changes. It should be
greater than or equal to
65. So this is the third category A B C
and then we will proceed with printing
the files. So print
the bottom
for the first one which is f of
a
is to do 2f
and we will add the percentage symbol
over here r under
64.
Now we can copy paste the same here and
we can change the variables.
So here we will be adding under over 70.
Lastly in between the ones in between
65 and 70.
There you go. Here we will change the
values from A to B and here A to C.
There you go. And we can quickly run
this query. So it's not 70. It was
supposed to be 69.
There you go. So we forgot to mention
this particular one. Now let's run this.
There you go. So we have 21%
of people who are scoring under 64, 24%
over 70 and 53% are in between average.
So I think the school is focusing on
improving this particular percentage,
reducing this particular percentage and
increasing this particular percentage
and try to eliminate if possible this
particular one which are under 64. So
that is the overall moto I guess. Now so
far so good. Let's now try to remove
infinite values from HTML, right? So
before that, let's add uh this
particular HTML code here. So which
explains what we are trying to do. So we
will first implement the code that
prevents warning about infinite values
during data visualization. And now let's
add the code which will try to eliminate
the uh infinite values. Let's copy this
particular data frame cleaned uh data
frame name here. Now df cleaned
dotreplace
square brackets np dot info
comma
minus np
dot info out of these square brackets
dot np
na
non na values we're trying to eliminate
na values in place of those values you
can write true and after that we will
try to eliminate the null. So if it is
null
dot sum give me the total number of null
values after this. Right? So let's try
to execute that. There you go. So all
the null entries have been removed here.
Now let's see the distribution of
numerical values here. So before that
let's add the HTML code for that. So
let's quickly run this. So distribution
of numerical values. So we will be
considering these three parameters. So
if you go back here you can see our
studied attendance and previous scores.
So we will be considering these three
values or these three columns and check
the distribution of these variables
against the exam score. So uh is it
making any um you know kind of variation
if the u attendance is high? If if the
attendance is high is the exam score
high and uh apart from that we have if
our studies is high is the uh mark score
is high and if the previous scores are
high is there a chance to get better
scores in this particular exam. So what
we are trying to do is we are trying to
see if there is any direct involvement
of number of study hours and number of
days attended and number of uh or the
number of marks they received in the
previous course and we'll try to build a
visualization on that front.
So let's go and build that. So we'll try
the try to write the code here.
Figure axis
equals tot
dot
subplots. So we will be having three
different plots here since we're
considering three different u
categories.
And the fixed size should be equal to
12A 4. There you go. Now
access
is equals to access dot variable
per idx
comma call
in enumerate
and we will try to import the seaborn
library here. We will try to create
histo plots here. Histograms here
plot.
So line style we will be selecting this
one
and the comma
line width will be 0.7.
There you go. Next will be access
dot set
title. So for this we will be uh setting
the title as distribution of columns. So
the columns will be the three uh ones
attendance, hours studied and uh what
was the third one that we considered
previous course. Right? So instead of
mentioning them specifically, what we
can do is we can just write columns
here. C L and close.
There you go. And lastly,
plt.tight Right.
And show the plot. There you go. Let's
quickly run this query. Run. And now we
will be having the visualizations here.
So um you can directly see the
involvement of these three parameters
here. If uh they are trying to help if
the number of hours are increased then
you can see if there is a better
improvement in scores. If the attendance
is increased, if there is a betterment
in scores or if the previous uh scores
are helping then you can find it out how
it is. There you go. Now we can write a
result here in the form of HTML page.
And if we run this, it will give you the
result. The breaks or gaps in the hour
study variable may be due to the
respondents answering appropriately. The
variables attendance and previous scores
which exhibit a uniform distribution
have a normal impact on exam scores
variable which is our target variable.
Now let's proceed with another part of
this session which will be about the
relationship between the numerical
values and the target variables. Now we
will copy paste the same code and make
some minute changes to it. So the only
change that we did to it is we're trying
to uh build a scatter plot. So we will
be getting a scatter plot here. But
before that let's try to add another
column here and try to add an HTML code
which explains why we are doing it.
There you go a scatter plot. So
basically these two are one and the
same. Here we use some column graphs. So
here we did the same using the scatter
plot which will help for a better
understanding. Now we will try to build
some correlations.
So basically a list is called
correlation is created containing the
names of the columns for which you want
to calculate the correlation. So here in
our situation it is the df c r which is
a data frame and it is created by
selecting only these columns for df
cleaned data frame effectively created a
new data frame containing only the
specified columns. Now the second one
which is the co r which is a calculate
correlation. So this method computes the
correlation matrix for selected columns
which is n df c r. The one indicates
perfect positive correlation minus one
indicates the perfect negative
correlation and zero indicates no
correlation. And lastly the setup of
plot. This line sets up the figure size
of the plot. In this particular
situation we are choosing five and four.
Right now let's quickly try to execute
this query and see the answer. And we
will also add the HTML code for this so
that we have a better understanding for
this. So we will be adding that HTML
uh box here which will explain what
exactly happened here. So this is our
plot and this is the correlation. Now
let's try to add that HTML code right
here. The result of this particular data
visualization will be maintained here.
So the hours studying and attendance
shows a positive correlation with the
target variable which is exams hour.
However, previous course appears to have
no or little relationship with the
target variable. Right now let's
continue with our next uh part of this
session. So now we will try to identify
the relationship between studies
hours and attendance and extracurricular
scores. Right? So we have other u
columns to consider which is
extracurricular activities. So there is
a belief that extracurricular activities
will also help students to study better.
So we will find if there is a relation
between the target variable and this
extracurricular activities attendance
and study hours. Now let's quickly add
the code here. Now let's quickly execute
the code. Now we have the visualization
which explains the relationship between
the number of hours studied
extracurricular activities etc. So here
it is and now let's add an HTML code
which explains about this result
in this particular code was supposed to
be added here.
So this uh is the resultant column here.
Now here it explains about the
influences that it performs. So the
extracal activities, parental income and
extra things that influence the scores
and there you go.
Now let us also consider other columns
right the other parameters like
resources are available or not parental
education and other things which also
might have influenced the exam scores of
students. So for that let's add an HTML
code so that we have an HTML page here
which explains what is the next
procedure that we are following. Right
now let's add the query here. So here we
are considering the other parameters
like family income, peer influence,
motivation level, gender, parental
involvement, parental educational level
and extracurricular activities. And we
are considering them against the target
variable which is exam score. And now
let's execute this query. There you go.
Now we have generated a graph which
explains about this particular
parameters against the target variable.
And now let's add the HTML page here
which explains about these results.
Let's quickly run it. And there you go.
So when certain factors affect Q1 and Q2
but not Q2, it can be understood that
individual has overcome challenges
through personal effort. Right? So if
government policies and corporate social
contributors are focused on addressing
these aspect, it seems that we could
create dynamic country with greater
social mobility and open opportunities
for all. Right? So if extracurricular
activities can outweigh the influence of
other variables in academic performance
then we should foster that kind of
environment right. So this is how u you
can get extract some statistical
analysis on this particular data set.
Now let's quickly rename this uh python
eda
students
performance
and you can quickly rename and save it.
Welcome to math refresher probability
and statistics.
In this lesson, we are going to explain
the concepts of statistics and
probability.
Describe conditional probability. Define
the chain rule of probability. Discuss
the measure of variance. Identify the
types of gshian distribution.
Basic of statistics and probability.
Probability and statistics. Data science
relies heavily on estimates and
predictions. A significant portion of
data science is made up of evaluations
and forecast.
Statistical methods are used to make
estimates for further analysis.
Probability theory is helpful for making
predictions. Statistical methods are
highly dependent on probability theory
and all probability and statistics are
dependent on data. Data is information
acquired for reference or research via
observations, facts, and measurements.
Data is a set of facts structured in the
form that computers can interpret such
as numbers, words, estimations, and
views. Importance of data. Data aids in
seeing more about the information by
identifying possible connections between
two features. Data assists in the
detection of distortion by uncovering
hidden patterns based on prior
information patterns. Data may be
utilized to anticipate the future or
predict the current state of affairs.
Also, data aids in determining whether
two pieces of information have any
instance in common or not. Types of
data. Data might be quantitative. That
is data that can be measured or counted
in numbers. Or it may be qualitative
which is data which is generally divided
into groups or in simpler words which
cannot be counted or measured in
numbers. Let's consider an example. A
customer information data of a bank may
contain quantitative and qualitative
data. Consider this snapshot where we
have customer ID, surname, geography,
gender, age, balance, has C or card is
active member. Amongst these variables
we can see surname is mostly qualitative
as it cannot be counted and measured in
numbers. Geography and gender are also
qualitative as they cannot be counted in
numbers and are mostly groups. has C or
card that is has credit card and is
active member although are containing
numerical in form but these are
categorical that means these have been
divided into groups of one and zero that
represent yes and no as an answer hence
these two variables are also qualitative
customer ID is again although a
numerical data however the significance
or intuition behind Customer ID is
categorical.
Hence, it may be kept in the qualitative
data also. However, age and balance
these are numerical information which
have been measured or counted and
numerical operations can be performed on
them. Hence, these are under
quantitative data categories.
Introduction to descriptive statistics.
Descriptive statistics. A descriptive
measurement is summary measure that
quantitatively portrays the most
important features of a set of data
allowing for a better comprehension of
the information. Data can be measured as
different levels. The levels of
measurement describe the nature of
information stored in the data assigned
to the variables. Qualitative data can
be measured as nominal or ordinal.
Quantitative data can be measured in
terms of interval and ratio type.
Nominal data. The data is categorized
using names, labels or qualities. For
example, brand name, zip code, and
gender. Ordinal data can be arranged in
order or ranked, and can be compared.
Examples include grades, star reviews,
position, and race, and date. Interval
data is the data that is ordered and has
meaningful differences between the data
points. Example, temperature in Celsius
and year of birth. Ratio data is similar
to the interval level with the added
property of inherent zero. Mathematical
calculations can be performed on both
interval as well as ratio data. For
example, height, age, and weight.
Population versus sample. Before
analyzing the data, it's important to
figure out if it's from a population or
a sample. Population is a collection of
all available items as well as each unit
in our study. Sample is a subset of the
population that contains only a few
units of the population. Population data
is used for study when the data pool is
very small and can give all the required
information. Samples are collected
randomly and represent the entire
population in the best possible way.
Measures of central tendency. The
central tendency is a single value that
aids in the description of the data by
determining its center position.
Measures of central tendency are
sometimes known as summary statistics or
measures of central location. The most
popular measurements of central tendency
are mean, median, and mode. The normal
distribution is a bell-shaped
symmetrical distribution in which mean,
median, and mode all are equal. The
curve over here shows the bell-shaped
curve or the normal distribution of
variable X. The point over here that is
X1 is the point which represents the
mean, median and mode of this
distribution. Mean mean is calculated by
dividing these sum of all data values by
the total number of data values. It gets
affected when there are unusual or
extreme values. It is sensitive to the
outliers. Mean can be calculated as
summation over all the values of X in a
collection divided by the size of the
collection.
For example, we have a collection where
we have values as 7 3 4 1 6 and 7.
We find out the sum of these values
which is 28 and there are total of six
values. So 28 / 6 gives us a mean value
of 4.66.
Median,
it is the middle value in the set of the
data that has been sorted in ascending
order.
It is a better alternative to mean since
it is less impacted by outliers and
skewess.
It is closer to the actual central
value.
Median is calculated differently for
different sizes of data.
Differentiated as if the total number of
values is odd or if the total number of
values is even. If the size of the data
is odd. For example, in this case we
have five elements.
After sorting whatever middle value we
get
that means n + 1 by 2 term in this case
5 + 1 / 2
that is the third term which is four is
the median value.
In case when the total number of values
is even like here there are six values.
The average or the mean of the two
central values is considered as the
median. In this case the median is the
mean of 6 and four which is five. Mode.
Mode represents the most common value in
the data set. It is not at all affected
by extreme observations.
It is the best measure of central
tendency for highly skewed or non-normal
distribution.
Mode for categorical data is determined
by estimating the frequencies for each
categories
and then the category with the highest
frequency is considered to be mode.
Like in this case 7 has the highest
frequency. Hence seven becomes the mode
value. However, in case of continuous
data or quantitative data, the
calculation of mode is slightly
different. The first step in calculation
of mode is dividing the data into
classes which are equal with then
getting the frequency of data points
lying in within that range of classes
and finally selecting the class with the
highest frequency.
Using the range of that class and the
frequencies, we can get the final mode
value.
Using the formula L+
minus F_sub_1 multiplied to H divided by
FM minus F_sub_1 plus FM minus F_sub_2.
Here L is the lower limit or the lower
observation of the mode class.
H is the size of the mode class.
FM is the frequency of the mode class.
F_sub_1 is the frequency of the class
proceeding to mode. And F_sub_2 is the
frequency of the class succeeding to
mode. This gives us the final mode
value.
Mean versus expectation.
Now let's talk about mean versus
expectation.
So in general we use the expected value
or expectation when we want to calculate
the mean of a probability distribution
that represents the average value we
expect to occur before collecting any
data. And mean on the other hand mean is
basically used when we want to calculate
the average value of a given sample.
This represents the average value of raw
data that we may have already collected.
We can understand this by using a simple
example.
Now to calculate the expected value of
this probability distribution, we can
use a specific formula from the previous
discussion.
This is going to be the expected value
where X is going to be the data value
and this PX is the probability of value.
For example, we could calculate the
expected value for this probability
distribution to be as shown.
So here it will be 1.45 goals.
So this represents the expected number
of goals that the team will score in any
given game.
And then if you talk about calculating
mean, so we typically calculate the mean
after we have actually collected raw
data.
For example, suppose we record the
number of goals that a soccer team will
score in 15 different games.
Now to calculate the mean number of
goals scored per game,
we can use the following formula
where sum of x is basically the sum of
all the goals divided by n and the
number of records or we can say the
sample size.
It is as shown on the screen.
So this represents the mean number of
goals scored per game by the team.
Measures of asymmetry.
The difference between the three
distinct curves can be studied in this
image.
The central curve is the normal or no
skeus curve. Here mean, median and mode
all lie on the same point. This normal
curve is symmetrical about its mean,
median and mode.
That means the left hand side of the
curve is a mirror image of the right
hand side of the curve.
However, in case of negatively skewed
data, the tail is elongated on the left
hand side
and the mean is smaller than the mode
and the median values or is on the left
hand side of the mode.
Hence indicating that the outliers are
in the negative direction.
On the other hand, in case of positively
skewed, the data is concentrated on the
left hand side of the curve.
While the tail is elongated or longer on
the right hand side of the curve,
the mean is greater than the mode and
median
or is on the right hand side of the mode
and median indicating that the outliers
are in the positive direction.
Let's consider an example.
The graph here shows the global income
distribution for the year 2003 2013 and
a projection for 2035.
If we see the global income distribution
statistics for 2003 it is highly right
skewed.
We can observe in the previous graph
that in 2003
the mean of $3,451
was higher than the median of $1090.
The global income is definitely not
evenly distributed. The majority of
people make less than $2,000 each year.
while only a small percentage of the
population earns more than $14,000.
Measures of variability.
Measures of variability.
Dispersion. The measure of central
tendencies provide a single value that
addresses the full worth. However, the
central tendency cannot depict the
viewpoint entirely. The metric of
dispersion helps us focus on the
inconsistency in the data spread.
Measures of dispersion describe the
spread of the data.
The range, intercortile range, standard
deviation and variance are examples of
dispersion measures.
Range.
The range of distribution is the
difference between the largest and the
smallest amount of data.
The range, for example, does not include
all of a series positive aspects.
It concentrates on the most shocking
aspects and ignores that aren't
considered critical. For example, for a
set 13, 33, 45, 67, 70.
The range is 57. That is the maximum of
this which is 70 minus the minimum over
here which is 13.
Variance.
Variance is the average of all squared
deviations.
It is defined as the sum of squared
distance between each point and the mean
or the dispersion around the mean.
The standard deviation is used as
variance suffers from a unit difference.
Variance can be computed as sigma square
summation over x - mu^ 2
divided by n
where mu is the mean of the data, x is
the individual data point
and n is the size of the data.
This representation is for a population
data.
for a sample data variance can be
computed as X minus
Xar whole square summation
over it divided by n minus one.
Here Xbar is the mean of these sample
data and n is the sample size.
The units of values and variance are not
equal.
So another variability measure is used.
Standard deviation.
Standard deviation is a statistical term
used to measure the amount of
variability or dispersion around a mean.
The standard deviation is calculated as
the square root of variance. It depicts
the concentration of the data around the
mean of the data set.
Standard deviation as indicated
previously can be computed as square
root of variance
for a population data. Standard
deviation sigma can be computed as
square root of summation over x i minus
mu^ square / n
where mu is the mean of the data x i are
the data points and n is the size. Let's
consider an example.
Let's find out the mean, variance, and
standard deviation for this data. The
data values are 3, 5, 6, 9, and 10. To
find out the mean, we first find the sum
of all these data values
that is 33 and divide it by the count,
which is five.
We get the mean of 6.6. To compute the
variance, we start by computing the
deviation.
That is X minus the mean of X. Here 3 is
one of the values of the data and 6.6 is
the mean.
So 3 - 6.6 squared and we do that
to find out sum of all the deviations
divided by the count
which is five.
We end up getting an overall variance of
6.64.
Standard deviation as we know is
measured at square root of variance that
is square of 6.64
which amounts to 2.576.
Measures of relationship.
Measures of relationship coariance.
Covariance is the measure of joint
variability of two variables.
It measures the direction of the
relationship between the variables. It
determines if one variable will cause
the other to alter in the same way.
Coariance between variable X and Y can
be computed as summation over the
product of X I - XR
and Y I - Y bar the whole divided by N
minus one.
Here Xar and Y bar are the mean of X and
Y respectively. The value of covariance
can range from minus infinity to a plus
infinity.
Correlation. Correlation is normalized
coariance.
It measures the strength of association
between two variables. The most common
measure for correlation is the Pearson
correlation coefficient.
Correlation between two variables
X and Y can be measured with respect to
coariance as coariance between X
and Y divided by the standard deviation
of X and standard deviation of Y.
The value of correlation ranges from a
negative 1 to positive 1.
Types of correlation.
Correlation can be either a positive
correlation,
zero correlation or a negative
correlation.
The first picture over here represents a
perfect positive correlation
wherein a straight line with a positive
slope
is representing the relationship between
the two variables.
Zero correlation means that the line
representing the relationship between
the two variables is horizontal to the
xaxis.
Perfect negative correlation can be
represented by a straight line with a
negative slope.
Correlation equals to 1 implies a
positive relationship. That is when one
variable increases the other variable
also increases. A correlation value of
negative one implies a negative
relationship. That is when one variable
increases the other decreases.
The correlation coefficient of zero
shows that the variables are completely
independent of each other.
Let's consider an example.
Here we have two variables height and
weight.
To compute the correlation between
height and weight,
we use the correlation formula as
covariance of X
and Y divided by standard deviation of X
and standard deviation of Y.
Here height is the X variable and weight
is the Y variable.
First to compute coariance we compute
the x - xar and y - y bar values and
then the product of them.
We then compute x - xrยฒ
and y - y bar square values to compute
the standard deviations of height and
weight respectively. Correlation as we
know has been defined as covariance of X
and I and Y divided by standard
deviations of X and Y.
This can also be represented as
summation over x - xr multiplied to y -
y bar
divided by square root of summation over
sum of squared deviations that is x - xr
square multiplied to square root of
summation over y - y bar square that is
sum of square deviations for y.
Now let's find out values to put into
this formula.
First we find out the overall sum of
height to get the mean of height which
is 5.14.
Similarly we get the sum of weight to
get the mean of weight as 50. We now get
the summation over x - xr multiplied to
y - y bar to get the numerator for the
formula. Then we compute x - xr square
summation
and y - y bar square that is sum of
squared deviation of x and y
respectively.
Now we put in the values in this final
correlation formula to get a correlation
value of 0.889.
This indicates that height and weight
have a positive relationship.
It is evident that as height grows,
weight also increases.
In this module, we will be talking about
expectation and variance.
So the expected value or we can say mean
of a given variable that we can denote
by X is a discrete random variable where
it is a weighted average of the possible
values that X can take and each value is
going to be according to the probability
of that specific event occurring.
So usually the expected value of X is
denoted by a simple formula where we can
define the expectation based on the X
parameter
which is going to be the sum of each
possible outcome multiplied by the
probability of the outcome occurring.
So in more concrete terms, the
expectation is what we would expect the
outcome of an experiment to be on
average.
We can take an example for the coin. If
a coin is being tossed 10 times, then
one is most likely to get five heads and
five tails.
Same logic can be discussed if we talk
about another example of rolling a
dieice. So there are six possible
outcomes when you roll a dieice. 1 2 3 4
5 6. And each of these has a probability
of 1x 6 of occurring. So we can say that
the expectation is going to be 1
multiplied by the probability of that
happening which is going to be 1x 6 + 2x
6 + 3x 6 + 4x 6 + 5x 6 + 6x 6 and that
is going to give us 3.5 as an output.
The expected value is 3.5.
So if you think about it, 3.5 is halfway
between the possible values that I can
take and this is what we should have
expected.
Next we talk about the concept of
variance. So variance of a random
variable allows us to know something
about the spread of the possible values
of the variable.
So for a discrete random variable X, the
variances of X is going to be denoted by
using a simple formula that is going to
be var=
E X - M the whole square where M is
basically the expected value of the
expectation of X. So this is more like a
standard deviation of X which can also
be represented by using this formula. So
the variance does not behave in the same
way as expectation when we multiply and
add constants to random variables.
So now there are two different type of
variance that we can have a fair
understanding on. First of all we have
low variance and then we have high
variance.
So low variance simply means that there
is a small variation in the production
of the target function with changes in
the trading data set and at the same
time high variance as we can see here
high variance shows a large variation in
prediction of the target function with
changes in the trading data set. So a
model that shows high variance learns a
lot and perform well with the training
data set and it does not generalize well
with the unseen data set and that's why
as a result such a model gives good
results with training data set but shows
high error rates on the test data set
and since the high variance a model
learns too much from the data set it
leads to an overfitting of the model. So
model with high variance will be having
couple of issues like it may lead to
overfitting or it may also lead to
increase in model complexities.
Next we have skewess.
So skewess in simple terms is basically
a measure of asymmetry of a
distribution. So distribution is
asymmetrical when its left and right
sides are not the mirror images.
Right now this is a mirrored image and a
distribution can have right positive or
we can say negative or it can have zero
skewess.
So right skewed in this scenario is
basically the distribution is longer on
the right side of its peak
and a left skew distribution is going to
be we can say where it is longer on the
left side.
So we can see we have this one as a part
of right side. It is more elongated
towards the right side and this one is
more elongated towards the left side. So
we can think of skewess in terms of
tails. A tail is long tampering and the
end of a distribution. So it simply
indicates that they are observations at
one end of the distribution but that
they are relatively infrequent. So a
right skew distribution has a long tail
on the right side as you can see here.
So the number supports observed. Let's
say we have a data on a per year basis.
So again we can have a more skewess
towards the right side where data is
being dropping as we continue to
increase the number of years. For
example we may have a high sales towards
the beginning of year suppose in 2022
but again as we proceed to 2023 second
half we are seeing the dip in
performance. So that is rightly skewed
and same way let's suppose if we started
with the sales figure it was really less
in suppose 2002
but again as we proceeded to 2023 now
our sales have been gradually
increasing. So it's more like skew
towards the left section as a part of
negative skew. Next we have curtosis.
So curtosis is basically a measure of
the tailness of a distribution.
So taeness is how often the outliers
occur and act as curtis is the tailness
of the distribution related to a normal
distribution. So a distribution with
medium curttosis is called as meocurtic.
A distribution with low curtosis like
this one. This is called as the
platicurtic and then distribution with
high curtosis like this one. This is
called as the leptoccuric.
So tails here they are tapering ends on
either side of a distribution like this.
So they represent the probability or the
frequency of values that are extremely
high or extremely low to the mean.
In other words, tails here represents
how often the outliers occur.
So there are three type of curtis. We
have platocurtic which is negative,
leptocortic which is a positive towards
the upper end and then we have messertic
which is a normal distribution. So
messertic is the medium tail. So normal
distributions they have a curtosis of
three. So any distribution with a curtis
of a prox value of three is going to be
messertic. And curtosis is described in
terms of excess curttosis which is
curtosis minus3. And since normal
distribution they have a curtosis of
three axis curtises makes comparing a
distribution curtosis to a normal
distribution even easier. Introduction
to probability.
Probability theory. Probability is a
measure of the likelihood that an event
will occur.
Let's consider an example of coin toss
where the chances of getting heads on a
coin are 1 by two or 50%.
The probability of each given event is
between zero and one both inclusive.
Sum of an events cumulative probability
cannot be greater than one.
Hence the probability of an event x lies
between zero and one. This means that
the integral of probability of
distribution over x equals to 1.
Conditional probability. Conditional
probability of any event A is defined as
the probability of occurrence of A given
that event B has previously occurred.
Condition probability of event A given B
can be estimated as probability of A
intersection B that is probability of
both A and B happening together
divided by the probability of B.
It is also written as that probability
of A intersection B equals to
probability of A given B multiplied to
probability of B.
Let's consider an example.
In a coin, we are doing a two coin flip.
Coin one gets heads, tails, heads, and
tails in subsequent flips.
while coin 2 gets tails, heads, heads,
and tails in the subsequent flips. Now,
the probability that coin one will get a
head is 2 out of four. While the
probability that coin two will get heads
is again two out of four.
The probability that both coin one and
coin two will have a heads is just one
out of the four flips.
Hence the probability that coin one will
get heads given that coin 2 is already
heads can be computed as probability of
coin one edge intersection coin 2 edge
that is 1x4 divided by probability of
coin 2 edge
that's a given that is 2x 4 which is
going to be 0.5 or 50% based
base theorem Base theorem calculates the
conditional probability of an event
based on its prior probabilities.
Basically base theorem incorporates the
prior probability distribution to
predict the posterior probabilities base
theorem for conditional probability
can be expressed as probability of A
given B equals probability of B given A
divided by probability of B multiplied
to probability of A.
Base theorem allows updating the
probability values by using new
information or evidence. Here
probability of A is known as prior
probability. That is the probability of
event that before any new data is
collected. Probability of A given B is
known as the posterior probability. It
is the revised probability of an event
occurring after taking into
consideration the new information
probability of B given A is known as the
likelihood and probability of B is
probability of observing an evidence B
model. An example consider an example
for calculating the likelihood of having
diabetes based on frequency of fast food
consumption. Here is the observed data.
Let's say the fast food audience is 20%.
Diabetes prevalence is 10% and 5% is
fast food and diabetes.
The chances of diabetes given fast food
that is the conditional probability of D
given B can be calculated as probability
of diabetes and fast food together
divided by probability of fast food.
That means 5% divided by 20%. that
equals 25%.
Define an analysis can state eating fast
food increases the chance of having
diabetes by 25%.
The multiplication rule of probability
if events A and B are statistically
independent and probability of A
intersection B can be given as
probability of A given B multiplied to
probability of B. However, probability
of A intersection B is also given as
probability of A multiplied to
probability of B. Here probability of A
given B equals to probability of A when
we assume that probability of B is non
zero. Similarly, probability of B equals
probability of B given A assuming
probability of A is non zero. Chain rule
of probability joint probability
distributions over many random variables
can be reduced into conditional
distributions over a single variable. It
can be expressed as probability of X1 X2
so on until XN equals probability of X1
intersection probability of X I given
probability of X1 till X I minus one.
For example, the joint probability of A,
B and C can be given as probability of A
given B. C multiplied to probability of
B given C multiply to probability of C.
Logistic sigmoid.
The logistics function is a type of
sigmoid function that aims to predict
the class to which a particular sample
belongs. Its outcome is discrete binary
value. a probability between zero and
one. The logistic sigmoid is a useful
function that follows the yes curve. It
saturates when the input is very large
or very small. Logistic sigmoid is
expressed as sigma of x= 1 upon 1 + e to
the power minus x.
The logistic sigmoid can be expressed as
sigmoid function of x is given as 1 upon
1 + e ^ minus x where e is the ooler's
number.
Gshian distribution.
The gossian distribution is a type of
distribution in which data tends to
cluster around a central value with
little or no bias to the left or right.
It is often referred to as normal
distribution.
In absence of prior information, the
normal distribution is frequently a fair
assumption in machine learning
equation.
The formula for calculating Gaussian
distribution is described as the normal
distribution of X.
That is the function of x given mean as
mu and variance is sigma square can be
calculated as 1 upon sigma square
roo of 2 pi e to the power -/ x -
mood / sigma square
where mu is the mean or peak value which
also is the expected value of x.
Sigma is the standard deviation. Sigma
square is the variance.
A standard normal distribution has a
mean of zero and a standard deviation of
one.
Goshan distribution can be univariate
which describes the distribution of a
single variable X.
It can also be multivariat where it can
just use to describe the distribution of
several variables.
It is represented in 3D of ND formats.
Law of large numbers.
Now let's talk about law of large
numbers. The law of large numbers states
that an observed sample average from a
large sample will be close to the true
population average and that it will get
closer in the larger sample. So the law
of large number does not guarantee that
a given sample spatially a small sample
will reflect the true population
characteristics or that a sample does
not reflect the true population will be
balanced by a subsequent sample. This is
for the law of large numbers to express
the relationship between scale and
growth rate.
So there are multiple examples through
which we can understand
and it is widely used in statistical
analysis in working with the central
limit theorem in terms of the business
growth. So there are multiple real time
setup in which these are going to be
used. So if you talk about tossing a
coin so tossing a coin in a number of
times will give us two different type of
outcomes.
the result will spread evenly between
head and tails and the expected average
value is going to be half.
That means 50 times tails and 30 times
heads. But again, if you toss a coin
1,000 times, then the result can be in
different manners because out of 1,000,
let's say 850 times it has been head and
only 150 times it has been tails and so
on. So that's why the possibility of one
event occurring is going to be changed
in large sample sets as compared to a
small sample sets as in let's say 10
times. So the number of heads and tails
unbalanced for lower number of trials.
So we can see it is unbalanced.
But again as soon as we toss more number
of coins more leans towards the balance
value or we can see the observed
averages.
Next we have p value.
So p value is basically a number
calculated from the statistical test
that describes how likely we are to have
found a particular set of observations
if the null hypothesis were true. So p
values are used in hypothesis testing to
help decide whether to reject the null
hypothesis.
And the smaller the p value, the more
likely we are to reject the null
hypothesis.
So we have a term called as null
hypothesis. So all statistical tests
they have null hypothesis. So for most
tests the null hypothesis is that there
is no relationship between our variables
of in first or that there is no
difference among groups. For example in
a two-taile t test the non-hypothesis is
that the difference between two groups
is going to be zero.
So p value is going to tell us how
likely it is that our data could have
occurred under the null hypothesis.
It is done by calculating the likelihood
of a test statistic
which is the number calculated by a
statistical test using our data. So p
value tell us how often we would expect
to see a test statistic as extreme or
more extreme
than one calculated by a statistical
test. if the null hypothesis of the test
was true.
So there are multiple limitations as
well. So first one is the results can be
significant but again they are they may
not be practical as we have compared it
can be based on multiple hypothesis for
a game for the healthcare test. If the
test is going to be positive or not it
may show even values of the effect of a
variable but not the magnitude in real
life. What exactly is going to be the
application of a drug test being failed
in pharma company? Therefore, it is
recommended to use confidence and levels
in addition to the p values to quantify
or we can say to give a solid figure to
the reserve which we are going to get.
The p values they are interpreted as
supporting or we can say refuting the
alternative hypothesis.
So p value can only tell you whether or
not the null hypothesis is supported. It
cannot tell us whether our alternative
hypothesis is true or why. So the risk
of rejecting the null hypothesis is
often higher than the p value. So
especially when we are looking at a
single study or when using small sample
sizes. So this is because the smaller
frame of reference, the greater are the
chance that as we stumble across a
statistically significant pattern
completely by accident.
Key takeaways.
Key takeaways. Probability and
statistics structure the premise of the
data. The data helps in anticipating the
future or gauging in view of the past
patterns of information.
The central tendency is a single value
that helps to describe the data by
identifying these central positions. The
mean, median and mode are the measures
of central tendencies.
The distribution where the data tends to
be around a central value with a lack of
bias or minimal bias towards the left or
right is called as gshian distribution.
So now let's dive into the definition of
the probability distribution function.
What is probability distribution
function? A function which defines the
relationship between a random variable
and its probability such that you can
find the probability of the variable
using the function is called a
probability density function.
In simple words, probability density is
the relationship between an observation
and the probability. Some outcomes of a
random variable will have low
probability density and other outcomes
will have a very high probability
density. Basically, the probability of a
variable X happening or occurring will
vary and it can sometimes take on a
lower value or it can take on a way
higher value.
The overall shape of the probability
density is referred to as probability
distribution. And the calculation of
probabilities for specific outcomes of a
random variable is performed by a
probability density function or PDF for
short. Now consider a variable X with a
PDF of f ofx.
This is what your probability density
function will look like. There might be
a point where the probability of X
occurring is very high. Hence your
probability distribution function or f
ofx will also be very high. At other
points the distribution or the
probability of X happening or occurring
is going to be very low. Hence your f
ofx is also going to have a very small
value. Basically given the random sample
of a variable we might want to know
things like the shape of the probability
distribution. This here is something
called a normal distribution where a
probability distribution function takes
on a bell shape.
However, this is not the probability
density function that might always
occur. There are different probability
distribution functions and all of their
graphs look very different from each
other. Knowing the probability
distribution for a random variable can
help you calculate movements of the
distribution like the mean and variance.
But it can also be useful for other more
general considerations like determining
whether an observation is unlikely or
very unlikely and might be an outlier or
an anomaly like consider this graph
itself. In this graph, these points over
here which have very less probability
distribution
are outliers which means that the chance
of them occurring is very low. And
basically this is not something that
you're going to see in your regular
scenario for your variable X. Now let's
consider two points A and B which are
values that a variable X can take. P of
A and P of B just represent the
probability of A and the probability of
B which can be found out by drawing a
straight line and coinciding it with our
graphs. The area under the graph over
here which is going to give you your
probability of this region occurring can
be written as probability of A less than
equal to X which is a probability that
we're searching for here less than equal
to probability of B. What does this mean
exactly? This means that this area is
always going to be greater than or equal
to the probability of A but less than or
equal to the probability of B. This
gives us the narrow region
of the probability which is present over
here. And doing this we can find the
probability of occurrence for any value
of X. Suppose you want to find the
probability of B happening. For a
probability distribution function, the
probability of B happening is not simply
this point here, but the entire area of
the graph which is taking place before
this point itself. So if you want to
find the probability between these
regions, you're going to have to find
the entire area and not simply the
probability at one point.
Now, so far we've been talking about
different types of variables which is
discrete random variables and continuous
variables. What exactly do these mean? A
variable which can only take a value
within a certain range is called a
discrete random variable. The value is
usually within a certain distance of
another finite value. An example of this
would be the sum of two dices. Basically
values which are well defined are called
discrete values or and a variable which
has well- definfined values will be
called a discrete random variable.
This variable can only take values which
fall within a certain set of values.
Let's say you roll a dice. The dice can
only give you specific outcomes which
range from 1 to six. This is what you
would call a discrete output.
On the other hand, a continuous random
variable can take on infinite different
values within a range of values. For
example, the height of a student. The
height of a student is not fixed. Even
if the height is 1.7 m, in reality, the
height can be 1.77 or 1.765
or 1.789.
The exact height is very hard to
determine because it's not easy for us
to find the precise value of the height
of a student. So basically the height
can take on an infinite different range
of values. When we're trying to define
the values that a continuous random
variable can take, we usually say it in
the form of a range of values which
means that the value can fall in that
range and can take on any value in that
range. It's not like a discrete random
variable where you can define definitive
values.
Now let's understand a probability
density function with the help of a
graph. Consider the graph below which
shows the rainfall distribution in an
year in a city. The x-axis has the
rainfall in inches or the amount of rain
that we're getting and the y-axis has
the probability density function of
getting that amount of rain.
The probability of some amount of
rainfall is obtained by finding the area
of the curve to the left of it. So let's
say we have a 3.
If you want to find the probability of 3
in of rainfall occurring, we would have
to find the area of the curve
which falls to the left of three. When
we draw a line from three which
intercepts the graph and further extend
it onto the yaxis, we get a value of
0.5.
Simply put, this means that the
probability of 3 in of rainfall
occurring is going to be lesser than or
equal to 0.5. The exact probability can
be found out by finding the area of the
curve
which falls to the left of three.
How do we find the probability
distribution function?
The first step is to summarize your
density with the help of a histogram.
The first step in a density estimation
is to create a histogram of the
observations
in the random sample.
Now what is a histogram? A histogram is
a plot which involves first grouping the
observation into bins and counting the
number of events that fall in each bin.
The counts or frequency of observation
in each bins are then plotted as a bar
graph with the bins on the x-axis and
the frequency on the y-axis. The choice
of the number of bins is important as it
controls the coarseness of the
distribution and in turn how well the
density of the observation is plotted.
It is a good idea to experiment with
different bin sizes for a given data
sample to get multiple perspectives or
views on the same data.
At the same time, the number of bins is
important as it determines how many bars
the histogram will have and their
widths. This will change not only the
shape of the graph but also how the
graph is read. This will also determine
how our density is plotted. Now let's
see how we can summarize our density
with histograms using Python. First
let's import all of our necessary
modules which we're going to require.
We're going to require Mattplot lib to
plot graphs. We're going to need the
normal random function so that we can
get a normal distribution. We're going
to import mean and standard deviation
from numpy to use on our graphs and also
going to normalize our uh data. So we're
going to import the nom function from
sci.
We finished importing all of our
necessary modules. Now let's generate a
sample
which has a size of thousand and it's
going to be a normal distribution. And
we're going to also plot this with the
help of a histogram in bins of 10.
So as you can see here you get a normal
distribution which is nothing but a
almost bell-shaped curve and we have 10
bins here which are centered at zero and
which extend from minus3 to 3.
How will our graph look if we change the
number of bins though?
Let's run it and see. So you still have
a normal distribution but it's not as
well defined because of how less the
number of bins are. you lose a majority
of the data which will contribute to
your normal distribution. It doesn't
look like a proper normal distribution
but it looks more like a discrete data
at this point. Now let's take a look at
the next step of finding a probability
distribution function.
The next step is called parametric
density estimation. What exactly is
parametric density estimation? The
probability density function is of many
types. The shape of your histogram will
help you determine what type of a
function it is. We can also calculate
the parameters associated with the
function to get our density. Now
different probability distribution
functions will have
different graphs which will have
different shapes and which will also
have different parameters like mean,
standard deviation etc associated with
them. Using these parameters, we can
find important points of our data.
Hence, it's very important for us to
recognize what type of a distribution it
is. Common distributions will occur
again and again in different and
sometimes unexpected domains.
Getting familiar with common probability
distributions will help you identify a
distribution from a histogram. And once
identified, you can attempt to estimate
the density of the random variable with
a chosen probability distribution. This
can be achieved by estimating the
parameters of the distribution from a
random sample of data. Now, an example
of this would be a normal distribution
which has two main parameters, the mean
and standard deviation. Given these two
parameters, we will now know the
probability distribution function. These
parameters can be estimated from data by
calculating the sample mean and sample
standard deviation. This entire process
is known as parametric density
estimation and it includes identifying
your probability distribution function
and getting the parameters which are
associated with it.
Now once we have estimated the density,
we can check if it's a good fit.
This can be done in three different
ways. One is plotting the density of the
function and comparing the shape to the
histogram. The next is sampling the
density function and comparing the
generated sample to the real sample. And
the last one is using a statistical test
to confirm if the data fits the
distribution. Now over here as you can
see all we've done is taken our data and
plotted the density function on top of
our histogram and we've compared the
shape. So the distribution so the
density function that we're actually
considering here is a normal
distribution and from this graph we can
see that it's almost an exact fit to our
histogram. Now let's see how we can
perform parametric density estimation
using Python. To begin with,
let's generate a random sample of
thousand observations from a normal
distribution with a mean of 50 which is
determined by the LOC parameter and a
standard deviation of five which is
determined by the scale parameter.
Now just to show you what the
distribution looks like, we're going to
plot it in the form of a histogram. So
this is what the histogram looks like.
But this is just to give you a basic
idea of our data and what it looks like
once plotted. But let's assume that we
don't know the probability distribution
and and we don't know what it looks like
as a histogram and we don't know that
that it's normal. So now if we just
assume that it's normal, we can
calculate the parameters of the
distribution specifically the mean and
the standard deviation.
We would not expect the mean and
standard deviation to be 50 and five.
Exactly given the small sample size and
the noise in the sampling data.
So because of this noise and the small
sample size, we have a mean of almost 50
and a standard deviation of a little
more than five.
Now let's define the distribution as
normal. So now using this we've defined
a normal distribution. We've used the
norm method of the sci-fi uh library and
uh we're doing this with the mean and
the standard deviation that we've
obtained from our samples. So up until
now we're just assuming that it's a
normal distribution and because of that
the parameters that we've calculated is
the mean and standard deviation
and using the calculated mean and
standard deviation we've gotten a normal
distribution.
And up until now again keep in mind we
do not know for certain that it is a
normal distribution. So far all we have
is this data.
So the next thing that we're going to do
is fit the distribution with these
parameters
and then sample the probabilities for a
distribution for a range of values in
our domain which in this case is 30 and
70. So all we're doing is we're
calculating probabilities for a range of
outcomes. And in this case we've taken
30 and 70 as our domain.
So these are the probability
distribution values for the normal
distribution that we've defined over
here. And this is going to uh this is
basically going to give you the outline
of your normal distribution.
These are the points at which your
normal distribution will be plotted. Uh
now what we're basically going to do is
we're going to plot our histograms using
the samples that we've already generated
along with the values and probabilities
of the normal function that we defined
over here.
So as you can see it's an all it's
almost a complete fit. The normal
distribution that we have here is made
using the mean and the standard
deviation of our actual samples.
The reason we took mean and standard
deviation was because we assumed it was
a normal distribution and the parameters
associated with the normal distribution
are mean and standard deviation. Using
the mean and standard deviation, we got
the normal distribution. We calculated
probabilities for this normal
distribution using a random domain of 30
and 70
and we plotted the probabilities and the
values on top of our histogram to see if
the normal distribution was a fit to our
histogram. If it was not a fit, you
would have to go and do the same
procedure with other common probability
density functions
until you found a function which was a
proper fit to your histogram. Now let's
move on to the final step which is used
in the calculation of a PDF. This final
step is called nonparametric density
estimation and it's only used when the
shape of a histogram doesn't match a
common probability density function or
it cannot be made to fit one. In this
case, we will calculate the density
using all samples in our data using
certain algorithms.
This is only done when a data sample
does not resemble a common probability
distribution or it cannot be easily made
to fit the distribution. And this is
often the case when the data has two
peaks. This is also called a biodal
distribution or it has many peaks which
is also called a multimodal
distribution. In this case, the
parametric density estimation will not
be feasible and alternative methods can
be used that do not use a common
distribution. Instead, you will use an
algorithm which is used to approximate
the probability distribution of the data
without a predefined distribution which
is also referred to as a non-parametric
method because we're not using any
predefined parameters. The distribution
will still have parameters but these are
not controllable in the same way as a
simple probability distribution. For
example, a non-parametric method might
estimate the density using all
observations in a random sample in
effect making all observations in the
sample parameters.
Now consider this graph which has two
peaks. You this is not a normal
distribution or any other sort of
distribution that we are familiar with.
So for this we're not going to use a
parametric estimation method but we're
just going to calculate the parameters
for every single sample point in this.
Perhaps the most common nonparametric
approach for estimating the probability
density function of a continuous random
variable is called kernel smoothing or
kernel density estimation or KDE for
short. Kernal density estimation is a
nonparametric method for using a data
set to estimate probabilities for new
points.
It uses a mathematical function and
smoothing probabilities. So the so the
sum of the resultant probabilities is
always one. Now in this case a kernel is
a mathematical function that returns a
probability for a given value of a
random variable. The kernel effectively
smooths or interpolates the
probabilities across a range of outcomes
for a random variable such that the sum
of probabilities always equals one. A
requirement of well- behaved
probabilities. You also have a parameter
called the smoothing parameter which
controls the scope or the window of
observations from the data samples that
contributes to estimating the
probability for a given sample. As such
the kernel density estimation is s is
sometimes referred to as your parsen
rosenbalt window. Now at the end you
also have a basis function which is a
function which is chosen to control the
contribution of samples in the data set
towards estimating the probability of a
new point. This is only done to make
sure that you're not learning from a lot
of noise and that you're not using a lot
of the outliers. Again let's see how we
can perform non-parametric density
estimation with the help of Python.
So first we'll start by importing all
the necessary modules along with the
kernel density estimation which can be
imported from skarn.
Now let's create a biodial distribution
by combining two different samples.
Sample one and sample two. Sample one
has 300 examples with a mean of 20 and a
standard deviation of five. While sample
two has 700 examples with a mean of 40
and a standard deviation of five.
We're then going to use it stack to com
to merge both of them together to get a
final sample.
The means that we've chosen which is 20
and 40 are chosen close together to
ensure that the distributions overlap in
the combined sample.
So this is what our distribution is.
Let's just plot it so you get a basic
idea of what our graph looks like.
So this is what our graph looks like.
Now we already know that none of the
various different uh uh probability
distribution functions fit these graphs.
So now we're going to perform
nonparametric estimations.
To perform nonparametric estimations,
we're going to use the scikitlearn
machine learning library which provides
the kernel density class that implements
kernel density sorry that implements
kernel density estimation. First the
class is constructed with the desired
bandwidth or window size of two
and your basis function
which in this case is a gshian function.
It's a good idea to at least test
different configurations to your data.
And in this case we're only going to try
a bandwidth of two and a gshian kernel.
Uh but usually there are multiple
different kernels that you can uh you
know like uh that you can play around
with and you can also tweak your
bandwidth to exactly fit the
distribution that you have.
Now let's run this.
Uh so now we've gotten our kernel
density estimation. Uh now we can
evaluate how well the density estimates
matches our data by calculating
probabilities
for a range of observations and
comparing shapes to the histogram just
like we did for the parametric case
before. So again we're going to just
calculate different probabilities using
the kernel density function
and we're just going to plot it on top
of a histogram to see how well this the
kernel density function is estimating
for our data.
So these are the probabilities that
we've gotten finally with the con uh
with the kernel density estimation. And
now we're going to plot it on top of our
histogram.
So as you can see it's almost a complete
fit. It's just left out some of these
outlier values which again are ranging
very high. But overall we have a pretty
good fit.
Uh the only problem is it's not very
smooth and you can uh again try tweaking
the bandwidth to different values. Uh so
let's just in this case try tweaking it
to three and see how well it runs. Okay.
So now we've got a new probabilities and
let's run it on top of our bandwidth. Uh
so again we using a bandwidth of three.
You can see that we're fitting our data
even better and we're again
ignoring a lot of the outliers which are
out there. So this is going to give us a
better estimation. The first question
that is probably in your mind is what's
in it for you? What can you expect from
this video?
First we will explain the concept of
regression a machine learning algorithm
to you.
Next we will take a look at the R squar
error which can be used to calculate the
error in regression models.
Next
we will teach you how to calculate the R
squar error and finally we will
implement the R squared error with the
help of Python.
So what is regression?
Regression is nothing but a machine
learning algorithm that helps us
determine the relationship between two
or more variables. It uses input or
independent variables to find the value
of the output or dependent variables.
Regression is a prediction algorithm
which means that given some variables we
can predict the value of an output
variable.
The predicted value is not going to be
from a set of values but it's going to
be a unique value in itself.
Now let's understand what exactly
regression is with the help of a few
independent input variables. In this
case, the variables that we'll be
looking at is rainwater, fertilizer, and
seeds. When we pass these independent
input variables through regression
model, we're going to get a predicted
output.
The output predicted is that a crop will
germinate when all three of these
components are put together in certain
quantities. So with the help of
regression given raw input data we can
find out the dependent output variable
that we'll get
when all of these input variables are
more or less combined. Regression is
nothing but a statistical method which
is used to determine the strength and
character of the relationship between a
dependent variable and a series of other
variables.
Now using regression if you have a set
of data points we can use a regression
model to fit a line which passes through
most of these data points and use it to
predict the outcome for new data points.
The line is fit using equation of a
straight line or a polomial equation.
Now in this graph consider that we have
our dependent variable or the output y
and our independent variable or the
output x. for a certain value of our
independent variable. We are going to
get a certain value of our dependent
variable or our output. Using the data
given here, we can see how our output
varies when our input varies. To predict
the value of our outcome Y, we're going
to need to find a relationship between
all of these data points. To do this,
we're going to plot a straight line
through it. Because, as you can see, all
the data points lie more or less along a
given straight line. Now using the
straight line for any value of x we can
find the approximate value of y. So
suppose you want to find the value of y
at a point x which is given here say
then all you have to do is extend a line
from this point onto a predicted model
which is this line here and then we can
see where this point on the line
coincides with the y-axis and get the
approximate value of the outcome.
The equation of a straight line is given
as y = b + b1 x + e. Where b is a
constant given by the y intercept of our
line or basically where the line
intersects on the y-axis.
B1 is the slope of our line and x is the
point for which we want to find the
output. E in this case is nothing but an
error correcting term.
So this is basically how regression
works and this is how prediction takes
place in a regression model. Next we
will explain the concept of the R squar
error to you. So what is R squar error?
R squared error is nothing but an error
measurement term which calculates how
well a regression model fits the data.
It determines the amount of variance in
a model caused by the input variables.
Now, R squar is a statistical measure
that represents the portion of variance
for a dependent variable that's
explained by an independent variable or
variables in a regression model. R 2 is
used to explain to what extent the
variance of one variable affects the
variance of a second variable.
So if the R square of a model is 0.5
then approximately half of the observed
variation can be explained by the
model's inputs.
In other words, an R squar of 60%
reveals that 60% of our data fits our
regression model exactly.
Now in this case in this graph if we
have a variance of 60% it means that 60%
of our data points fall exactly on our
regression line. However it is not
always the case that a high R squar is
good for a regression model. The quality
of the statistical measure depends on
many factors such as the nature of the
variables employed in the model, the
unit of measure of variables and the
applied data transformation.
Thus, sometimes a high R squar can
indicate the problems with the
regression model. A low R squar figure
is generally a bad sign for predictive
models. Now, all this time we've been
talking about variance in our model.
What exactly is variance? Variance is
nothing but a statistical term which
determines how spread out our data is
and tells us how many outliers are
present in it. Basically, it's a measure
of how far a set of numbers is spread
out from their average value.
Using variance, you can basically figure
out where your data is centered
and how spread it is from the mean and
also you can find out how many outliers
it has.
Now how can you calculate the R squar
error?
Let's start by considering the data that
we have been given as shown below. Now
we can find the relationship between the
input and the output variables by
plotting a straight line or a regression
model that passes through most of the
data. To get the perfect fit for a
model, we don't necessarily need to have
the line passing through as many data
points as possible.
A true measure of a good model is that
we reduce the error which is present in
our model. Now how do you find this
error? The error present in our model is
given by nothing but the distance
between our predicted line and the data
points which do not fall on the line.
This is what we have to minimize.
The distance between our data points and
our line can be calculated by
subtracting
our data point from the point at which
it coincides on our regression line. We
square this just to get rid of any
negative coefficients that may occur due
to finding the difference between the
two points. Now to find the variance in
our data, we're going to find the mean
and subtract the data points from the
mean. This will basically tell us how
spread out our data point is from the
average value. We can then square these
differences and add up the result to get
our total variance. This is also known
as the sum of squares total. And using
this we can find the total variance in
our data. The mean is nothing but the
average of our data. And using the mean
we can find the center of our data.
So this is exactly where our data is
centered. This is the average value that
occurs in our data. Now, the variance is
nothing but the distance of all of our
data points from the mean. If we do
this, we're basically going to find out
how spread apart our data is from the
average value or how far all of our data
points lie from each other.
When we subtract the position of our
data points from our mean and square it
and add all of that up, we get something
known as the sum of squares total.
Now the R squ error is the total
variance in our input data. It can be
obtained by dividing the SSR by the SST
and subtracting the results from one.
So the R squar error totally becomes 1
minus the sum of our squared errors
divided by the sum of the squared
difference between our data points and
the mean. Now this value gives us the
variance and this is why we can say that
R squ is used to find the portion of
variation in our data.
Now how can you implement R squ error
with Python?
To calculate R squared error with
Python, we're going to look at the data
which depicts the weather conditions
which were present during World War II.
And using the variables which are
present in our data, we're going to
create a model which predicts the daily
weather forecast during World War II.
And then we're going to use R squared
error to find the accuracy of our model.
So this is our R square uh model. So
we're going to start off by importing
all of our necessary modules. We're
going to use the model numpy to perform
numerical calculations on our database
and arrays and we're going to use
seaborn and mattplot lib to plot our
data.
So now we've managed to import all of
our data sets.
Let's also load our data by reading in
the CSV file that it is stored as in the
form of a data frame.
So over here as you can see we've read
in the CSV file into uh a variable
called weather and then after that we're
changing weather into a panda's data
frame called climate. This is what our
data frame finally looks like.
So as you can see in our data frame we
have five rows because we're only
looking at the top five rows and we have
31 columns. So these are values which
are not required and which are basically
going to increase our error value. So
let's drop them and get rid of them.
Now let's also drop any
empty values which may occur in the
remaining columns of our data set and
see what the final data set looks like.
So this is our final data set. We have
at the end we're only left with max
temperature, minimum temperature, and
mean temperature.
Now let's plot a count plot of a max
temperature.
A count plot is basically going to go
through the entire max temperature
column and figure out how many times
every temperature value occurs. So it's
going to figure out how many it's going
to count how many times 29.44444
has occurred and it's going to plot that
on this graph and it is going to do that
for every unique temperature value which
is present in our column. So finally
this is the value that we get. Uh so
over here as you can see the majority of
our temperature values are concentrated
within this range. This means that these
temperature values are the ones which
occur most frequently. The other ones
can be considered as outliers because
they rarely are seen in our data
and they can further skew the output
that we're going to get. Now let's do
the same with minimum temperature. Let's
plot a count plot for minimum
temperature.
So for the minimum temperature we can
see a very similar plot to the one that
we got for a maximum temperature. Most
of the values are concentrated around
this region but the outlier values here
are way fewer.
Now let's plot a regression plot between
our maximum and our minimum temperature.
So using the regression plot we can plot
a regression line for the two variables
in our x and y axis. So this is our
x-axis and this is our y-axis. This is
basically going to plot a straight line
which best fits the data
that we are getting here. So over here
as you can see this is a regression
line. This thin blue line is a
regression line which intersects our
minimum temperature at a value which is
between -30 and -40 and it passes
through our entire data. So uh for a
value of maximum temperature which is
zero using this we can predict the
minimum temperature that would have
occurred on the same day.
So for zero it'll be somewhere around -
10ยฐC. So if we saw maximum temperature
of 0ยฐ on that day, we would have seen a
minimum temperature of minus 10 on the
same day. Now let's plot a heat map to
see how these values are correlated with
each other. So the correlation is
basically
used to find which values affect each
other linearly
or which values
when changed will also affect the change
in other values. So over here as you can
see minimum temperature
and maximum temperature have a
correlation of 0.8. 88 which means if
minimum temperature changes then the
maximum temperature will also change to
0.88.
Now the best correlation is obviously
going to be between the mean
temperatures and the minimum and maximum
temperatures. The mean temperature is
nothing but the average temperature
value that we have. So this this is
basically going to lie in the middle of
all of our temperature values which is
why we going to have a better
correlation for mean temperature. But
minimum temperature and maximum
temperature are also pretty well
correlated with a correlation value of
0.88.
This means that if our minimum
temperature fluctuates, our maximum
temperature will also fluctuate
proportionately.
Now let's separate our input and output
values. We're going to predict the
maximum temperature given our minimum
temperature. Here x is our input
variable and y is our output variable.
So now we're basically just going to get
all the important values in our x and y
data sets. So after that this is what
our x and y data sets are going to look
like. Now let's split our data set into
training and testing sets.
The training set will be used to train
our regression model and the testing set
will be used to predict how well our
regression model is performing.
The training data is the data which will
be visible to our model or the data
which a model is allowed to have access
to. Testing data will be data which the
model has never seen before or which it
doesn't have access to and hence it will
be made to work on completely new data
to better test how well we've fitted to
our data set. Now we can split our data
set into training and testing sets by
using the train test split functionality
from our skarn.mmodel selection library.
Now finally from our scikitlearn library
let's import a linear regression model.
We're going to initialize a linear
regression model to a variable called
regressor and then we're going to fit a
linear regression model to our training
data set.
So we finally got our trained linear
regression model. Now let's use this
model to perform predictions on our
testing data set. So these are the
values that we've gotten after running
our linear regression model on our
testing data set. Let's see how well
we've performed.
We are going to import the R2 score from
our skarn metric.
Now the R2 score will directly perform R
squared error on our prediction and
testing data set and see how well our
testing data set matches a prediction
data set.
So now we've gotten an R squared error
of 0.9345
which basically means that 93% of our
output values are influenced by our
input values. This also means that a
model is 93% accurate
and that approximately
93% of our observed variation can be
explained by the model's inputs. Ever
wondered how to build an AI project that
actually gets noticed by Google, OpenAI,
or top startups, not just a chartboard
or recycled homework. Today I'm going to
walk you through 10 AI project ideas for
26 that are practical, futuristic and
portfolio ready. I'll tell you exactly
which models, framework and data sets to
you so that you can start coding
immediately. Now before we jump into hit
that like button, share and subscribe
because keeping up with future proof AI
projects is going to give you a massive
edge. Let's start with the AI shopping
buddy. This project acts like a personal
stylist and interior designer. Users
upload photos of the room, outfit, or
even face, and the AI suggest products
that match color, style, and
preferences. This isn't just about
throwing recommendations at someone.
It's about computer vision to understand
images, generative AI to create style
suggestions, and recommendation
algorithms to find the perfect products.
Personalized recommendation systems
drive massive engagement and conversions
which is why companies like Amazon,
Flipkart, Myntra or Urban Ladder would
be thrilled to hit someone who can build
this. Completing a project like this
demonstrates skills in deep learning,
computer vision, generative AI and full
stack deployment for web or mobile app.
While shopping and lifestyle AI is
exciting, the next project takes up to
our health and wellness. The smart
health analyzer predicts stress burnout
or sleep issues by analyzing voice,
facial microp expressions and variable
data. It uses multimodel AI that
integrates time series analysis for
variable data. NLP for voice and text
and computer vision for micro
expressions. Health tech startups in
India and globally like healthy, cure
fit, Fitbit and Apple Health are looking
for engineers who can make predictive
wellness tools. Building this project
demonstrates your ability to work with
multimodel AI, pre-process complex data
sets, train models, and visualize
result. Moving from personal health to
professional efficiency, the AI
productivity agent automates your daily
workflow. It reads emails, scans your
calendar, understands priorities, and
builds an optimized schedule. It uses
NLP to parse emails, API integration
with other tools like Gmail, Slack, and
Notion, and optimization algorithms to
prioritize task efficiently.
Productivity loss is a major issue for
companies which is why tech giants like
Google Workspace, Microsoft 365 and
startups in workflow automation would be
very interested in this project. It is a
great way to demonstrate automation, NLP
API integration and practical problem
solving skills. Taking automation to the
next level, the voice toaction system
allows users to speak commands and have
the AI perform multi-step action such as
booking flights, organizing files or
generating reports. It relies on
speechtoext models, intent
classification using NLP and task
automation pipelines. You can train
intent classification models using data
set into the snips NLU data set. This
project builds directly on productivity
AI and is exactly the kind of work that
Amazon openai or Apple would notice for
voiced driven automation solutions. Once
we have automated task, why not explore
creativity? Generative AI story maker
allows you to create full stories
including scripts, characters and
visuals based on just a few keywords.
Now it uses large language models for
text generation and image generation
models like stable diffusion deli3 for
visuals and you can also train
fine-tuned models on data sets like CMU
book summary corpus on writing prompts
text to speech library such as scope TTS
or GTTS can add narration media
companies like Netflix, Ubisoft and
Adobe are actively investing in
generative AI and a project like this
would definitely stand out. Building on
the idea of multiple AI capabilities
working together, the multi- aent AI
team project introduces collaboration
between AI agents. Multiple AI agents
are assigned specialized roles such as
researching, writing, criticing, and
summarizing. They communicate and
coordinate to complete complex task
using multi- aent reinforcement,
learning and communication protocols.
Enterprise AI and automation platforms
are investing heavily in this approach
and companies like Enthropic, OpenAI, AI
workflow startups are actively seeking
engineers who can build collaborative AI
systems from collaboration to
observation. The AI body language reader
analyzes micro expressions, tone of
voice and posture to provide feedback on
communication skills. It can be applied
in interviews, public speaking or remote
coaching. Computer vision tools such as
open pose or media pipe pose track
gestures and posture. While audio
processing libraries such as librosa
analyze tone models can be trained using
data sets like Raves for audio and CK
plus for facial expressions. HR tech
companies like High View, Pytrics and AI
coaching startups would highly value
this type of project as it help bridge
human behavior and AI analysis. Nucation
is also another area being transformed
by AI. The personalized tutor with
adaptive difficulty creates an AI tutor
that adjusts lessons in real time based
on student performance. Knowledge
tracing models like deep knowledge
tracing combined with transformers for
content generation allow the AI to adapt
to each learner. Data sets such as
assessments or edn nets can be used for
training. Now this AI can generate new
questions, explanations and motivational
feedback based on learning pace.
Companies like Baiju, Vidanto, Corsera
and Udemy are constantly looking for
talent that can build adaptive learning
platforms. Next, we move into research
augmentation with the autonomous
research agent. This AI can answer
research questions by reading academic
papers, extracting insights, summarizing
information, and citing sources
automatically. It uses the S2 or data
set for academic papers. Cybboard for
scientific text embeddings and hugging
face transformers or lang chain for
reasoning and summarization. Citation
extraction can be done with NLP passers
or rejects academic platforms. AI labs
and companies like Google research, open
AI, research gate and LCV would hire
engineers who can build this system.
This project connects perfectly with
education focused AI extending learning
into automated research capabilities.
Finally, we arrive at realworld robotics
control with the AI. The ultimate
demonstration of cuttingedge skill. This
project trains AI to control robotic
arms or humanoids based on goals rather
than just rigid instructions. It uses pi
bullet or vbots for simulation. Stable
baseline 3 or R lib for reinforcement
learning and open CV or media pipe for
vision input. Sim to real transfer
techniques bring simulations into realw
world scenarios. Robotics companies like
Boston Dynamics, Appronic, Amazon
Robotics and Agibot are seeking
engineers capable of endtoend AIdriven
robotic systems. After exploring
softwarebased AI projects, robotics is
the next step to show mastery of AI
applied in the physical world. These 10
AI projects are more than just ideas.
>> We will learn about some of the machine
learning and deep learning interview
questions.
So let's begin with our first question.
The first question is how to detect
outliers in data. So in data analytics
and machine learning, you often find
data points that lie at an abnormal
distance from other points in a random
sample from a population. Those are
called outliers. Now outliers in data
can significantly impact any prediction
analysis. There are majorly three
different methods to treat outliers.
First we have the univariate method. It
is one of the simplest methods for
detecting outliers. The univariate
method uses box plots. A box plot is a
graphical display for describing the
distributions of the data. Box plots use
the median and the lower and upper
quartiles.
This method looks for data points with
extreme values on one variable. Next, we
have the multivariate method. So, the
multivariate outliers can be found in an
n- dimensional space having n features.
We look for unusual combinations of all
the variables in this method. Finally,
we have Minowski error. This method
reduces the contribution of potential
outliers in the training process. The
Minowski error is a loss index that is
more insensitive to outliers than the
standard mean squared error. Now moving
on to the second question. What is a
confusion matrix? So a confusion matrix
is a table that is used to describe the
performance of a classification model on
a set of test data for which the true
values are already known. The target
variable has two values positive or
negative. The columns represent the
actual values of the target variable
which you can see here. The rows
represent the predicted values of the
target variable which you can see here.
Now there are four important terms that
are related to confusion matrix. First
we have true positive which is this one.
So in true positive the predicted value
matches the actual value. So the actual
value was positive and the model also
predicted a positive value. Then we have
true negative which is also represented
as tn. The true negative depicts the
predicted value matches the actual
value. Now the actual value was negative
and the model predicted a negative
value. Next we have false positive. Now
false positive is also known as a type
one error. In false positive the
predicted value was falsely predicted.
The actual value was negative but the
model predicted a positive value.
Finally we have false negative. A false
negative is also known as type two
error. So in false negative the
predicted value was falsely predicted.
The actual value was positive but the
model predicted a negative value. Now
moving to our third question which is
explain the ROC curve. Now the ROC curve
is one of the most important evaluation
metrics for checking the performance of
any classification model. ROC stands for
receiver operating characteristic.
Receiver operating characteristic or ROC
curve is a method to compare the
diagnostic tests. The ROC curve is
created by plotting the true positive
rate against the false positive rate at
various threshold settings. So here on
the y-axis you have the true positive
rate. On the x-axis we have the false
positive rate. The true positive rate
indicates the proportion of observations
that were correctly predicted to be
positive out of all positive
observations. Similarly, the false
positive rate is the proportion of
observations that are incorrectly
predicted to be positive out of all
negative observations.
You can take an example. Suppose in
medical testing, the true positive rate
is the rate in which people are
correctly identified to test positive
for the disease in question. Let's say
the corona virus testing. ROC does not
depend on any class distribution. This
makes it useful for evaluating
classifiers predicting rare events such
as diseases or disasters. Now moving to
the fourth question we have what are the
assumptions for linear regression. So
linear regression analysis is used for
modeling the relationship between a
single dependent variable Y and one or
more feature or predictor variables.
Some of the important assumptions for
linear regression are so first they
should have linearity. So linear
regression needs the relationship
between the independent and the
dependent variables to be linear. It is
also crucial to check for outliers since
linear regression is sensitive to
outlier effects. Next we have
homocyasticity.
Homoscadasticity
illustrates a situation in which the
error term that is the noise or random
disturbance in the relationship between
the features and the target variable is
the same across all levels of the
dependent variables. Third we have
independence. So observations should be
independent of each other. Finally, we
have no multi-olinearity.
So there should be little or no
multi-olinearity.
Independent variables should not be too
highly correlated. Now moving to our
fifth question in our list of interview
questions.
The question is what is regularization
in machine learning? Explain the L2
regularization.
So regularization is a machine learning
technique that is used to reduce the
errors by fitting the function
appropriately on the training set in
order to avoid overfitting of data. So
overfitting happens when a model learns
the detail and noise in the training
data to the extent that it negatively
impacts the performance of the model on
new data. So here you can see we have a
nice plot which shows how overfitting of
data can be visualized
and here we have a good fit line over
the same data points. So this is also
known as the regression line. Now L2
regularization is also known as ridge
regression. So ridge regression modifies
the overfitted model by adding the
squared magnitude of coefficient as a
penalty term to the loss function. So on
the right you can see a set of data
points plotted and we have our linear
regression line and here we are
calculating the cost function for the
ridge regression line. So our cost
function is actually loss plus lambda
into summation of w ^ 2 where loss is
actually the sum of squared errors or
squared residuals. Lambda stands for
penalty for the errors. W is called the
slope of the curve or line. Okay. Now
consider a case where there are two
points passing through the linear
regression line. Now if you calculate
the cost function, we get the value as
1.69. So here we have assumed that loss
is zero since the two points lie
directly on the line. We have taken
lambda to be 1 and w is 1.3. So if you
use this function or this formula, you
get the cost function as 1.69.
Now moving ahead, let's consider another
situation where we'll calculate the same
cost function for the ridge regression
line. There is some loss for both the
points as they are not on the same line.
So here you can see the sum of squared
residuals is 0.05 which actually is the
sum of 2 squared. I'm assuming this as 2
and 0.1 for this one.
So if you square both and add it the
value is 05 a lambda is again 1 and w we
have assumed to be 6. Now if you find
the cost function the value is 41. Let's
draw the linear regression line and the
ridge regression line with all the
points we find that the ridge regression
line as the best fit since its cost
function is less. Now coming to the
sixth question.
What are the different methods to split
a tree in a decision tree algorithm?
So there are three methods to split a
decision tree. First we have variance.
So reduction in variance is an algorithm
that is used for continuous target
variables. This algorithm uses the
standard formula variance to choose the
best split. So here you can see the
standard formula variance which is
summation of x that is all the
individual points minus xar which is the
mean squared divided by the total number
of observations. Now the split with
lower variance is selected as the
criteria to split the population.
Now the steps to calculate variance is
you need to calculate variance for each
node and then you need to calculate for
each split as the weighted average of
each node variance.
Moving ahead, the second method we have
is information gain. So information gain
is used for splitting the nodes when the
target variable is categorical.
It works on the concept of entropy. Now
the degree of disorganization in a
system is known as entropy. So here you
can see the formula for information gain
which is 1 minus entropy.
Finally we have genie impurity. So genie
impurity is the probability of
incorrectly classifying a randomly
chosen element in the data set if it
were randomly labeled according to the
class distribution in the data set. So
below you can see the formula for gen
impurity. So we have 1 minus summation
of pi whole square where n represents
the number of classes and p of i
represents the probability of randomly
picking an element of class i.
Now moving to the seventh question.
So the question is how do we find the
optimum cluster value in K means
clustering algorithm. Now there are two
methods to find the optimum cluster
value. So first we have the elbow method
which is one of the most wellknown for
finding the optimum number of clusters.
So in this method you need to calculate
the within cluster sum of squared errors
for different values of K and choose the
K for which within cluster sum of
squared errors first starts to diminish.
So in the below plot of squared errors
versus the number of clusters K you can
see at K is equal to 4 the squared error
starts to diminish. So hence our optimum
K value is four. Next we have the siloid
method.
So the celloid method measures how
similar a point is to its own cluster
compared to other clusters. The average
seloid method computes the average of
observations for different values of K.
The optimum number of clusters K is the
one that maximizes the average over a
range of possible values for K. The
seloid score reaches its global maximum
at the optimal K. So in our case the
average seloid reaches maximum at k is
equal to two which you can see here. So
our optimum cluster value will be two
here. Moving ahead the eighth question
in our list is how does the pooling
layer work in a convolutional neural
network. So the pooling layer performs a
downsampling operation in order to
reduce the dimensionality of the feature
map. So in the pooling operation, you
slide a two-dimensional filter over each
channel of feature map and summarize the
features lying within the region covered
by the filter.
It is a common practice to periodically
insert a pooling layer in between
successive convolutional layers in a
convolutional neural network
architecture.
So the pooling layer operates
independently on every depth slice in
the input and resizes it specially using
the max operation. So in the diagram
shown here you can see we have a
rectified feature map. We are using a 2
+2 filter and performing a max pooling
operation. So consider this as the
filter. If you perform the max operation
over the values
let's say 0 5 3 and 1. So considering
this one our pool feature map maximum
value will be five. Similarly for this
chunk of data it is going to be seven.
Next, if you slide the filter over this
square frame, you get eight. And
similarly here you get six. So this is
also known as a pulled feature map.
Moving ahead, the ninth question in our
list is how does LSTM network work? So
long short-term memory networks are a
type of recurrent neural networks that
are capable of learning order dependence
and sequence prediction problems. So
remembering information for long periods
of time is practically their default
behavior. Now, LSTMs also have this
chain-like structure which you can see
here.
But the repeating module has a different
structure. So, instead of having a
single neural network layer, there are
four interacting in a very special way.
Now, you can see these are called as
gates. These gates contain sigmoid
activations. A sigmoid activation is
similar to the tanh activation. Instead
of squishing values between minus1 and +
one, it squishes values between 0 and
one. An LSTDM has four gates.
Now these are called forget, remember,
learn and use or output. So if you see
this in the first step, we use the
forget gate that decides what
information should be thrown away or
kept. the information from the previous
hidden state and the information from
the current input is passed through the
sigmoid function. Values come out
between zero and one.
So if the value is closer to zero, it
means you need to forget that
information and if the value is closer
to one, it means you need to keep that
information.
Next we have the input gate. So the
input gate is used to update the cell
state. First, we pass the previous
hidden state and the current input into
a sigmoid function
that decides which values will be
updated by transforming the values to be
between 0 and 1. Zero means not
important and one means important. You
also pass the hidden state and current
input into the tan function to flatten
the values between minus1 and + one.
This helps to regulate the network.
Then you multiply the tan output with
the sigmoid output. The sigmoid output
will decide which information is
important to keep from the tanh output.
And finally in step three we have the
output gate. This output gate is used to
decide what the next hidden state should
be. First we pass the previous hidden
state and the current input into a
sigmoid function. Then we pass the newly
modified cell state into the tanage
function. We then multiply the tanage
output with the sigmoid output to decide
what information the hidden state should
carry. The output is the hidden state.
The new cell state and the new hidden
state is then carried over to the next
time step. Finally, talking about the
last question in our list of interview
questions, we have explained the concept
of gradient descent in deep learning.
Now gradient descent is an optimization
algorithm which is mainly used to find
the minimum of a function in machine
learning. Gradient descent is used to
update the parameters in a model.
Parameters can vary according to the
algorithms such as coefficients in
linear regression and weights in neural
networks.
You can see we have these maps and on
the y-axis we have the loss. On the
x-axis we have the weight and here we
are trying to find the local minimum or
the global minimum. Now this gradient
descent method is used to minimize the
cost function and update the parameters
of the learning model. The gradient
always points in the direction of the
steepest increase in the loss function.
The gradient descent algorithm takes a
step in the direction of the negative
gradient in order to reduce the loss as
quickly as possible. To determine the
next point along the loss function
curve, the gradient descent algorithm
adds some fraction of the gradient's
magnitude to the starting point. Now
this process is repeated to find the
global minimum.
>> Welcome to math refresher probability
and statistics.
In this lesson, we are going to explain
the concepts of statistics and
probability.
Describe conditional probability. Define
the chain rule of probability. Discuss
the measure of variance. Identify the
types of gshian distribution.
Basic of statistics and probability.
Probability and statistics. Data science
relies heavily on estimates and
predictions. A significant portion of
data science is made up of evaluations
and forecast.
Statistical methods are used to make
estimates for further analysis.
Probability theory is helpful for making
predictions. Statistical methods are
highly dependent on probability theory
and all probability and statistics are
dependent on data.
Data is information acquired for
reference or research via observations,
facts, and measurements. Data is a set
of facts structured in the form that
computers can interpret such as numbers,
words, estimations, and views.
Importance of data. Data aids in seeing
more about the information by
identifying possible connections between
two features. Data assists in the
detection of distortion by uncovering
hidden patterns based on prior
information patterns. Data may be
utilized to anticipate the future or
predict the current state of affairs.
Also, data aids in determining whether
two pieces of information have any
instance in common or not. Types of
data. Data might be quantitative. That
is data that can be measured or counted
in numbers or it may be qualitative
which is data which is generally divided
into groups or in simpler words which
cannot be counted or measured in
numbers. Let's consider an example a
customer information data of a bank may
contain quantitative and qualitative
data. Consider this snapshot where we
have customer ID, surname, geography,
gender, age, balance, has C or card is
active member. Amongst these variables
we can see surname is mostly qualitative
as it cannot be counted and measured in
numbers. Geography and gender are also
qualitative as they cannot be counted in
numbers and are mostly groups. has C or
card that is has credit card and is
active member although are containing
numerical in form but these are
categorical that means these have been
divided into groups of one and zero that
represent yes and no as an answer hence
these two variables are also qualitative
customer ID is again although a
numerical data however the significance
or intuition behind Customer ID is
categorical.
Hence, it may be kept in the qualitative
data also. However, age and balance
these are numerical information which
have been measured or counted and
numerical operations can be performed on
them. Hence, these are under
quantitative data categories.
Introduction to descriptive statistics.
Descriptive statistics. A descriptive
measurement is summary measure that
quantitatively portrays the most
important features of a set of data
allowing for a better comprehension of
the information. Data can be measured as
different levels. The levels of
measurement describe the nature of
information stored in the data assigned
to the variables. Qualitative data can
be measured as nominal or ordinal.
Quantitative data can be measured in
terms of interval and ratio type.
Nominal data. The data is categorized
using names, labels or qualities. For
example, brand name, zip code, and
gender. Ordinal data can be arranged in
order or ranked and can be compared.
Examples include grades, star reviews,
position, and race, and date. Interval
data is the data that is ordered and has
meaningful differences between the data
points. Example temperature in Celsius
and year of birth. Ratio data is similar
to the interval level with the added
property of inherent zero. Mathematical
calculations can be performed on both
interval as well as ratio data. For
example, height, age, and weight.
Population versus sample. Before
analyzing the data, it's important to
figure out if it's from a population or
a sample. Population is a collection of
all available items as well as each unit
in our study. Sample is a subset of the
population that contains only a few
units of the population. Population data
is used for study when the data pool is
very small and can give all the required
information. Samples are collected
randomly and represent the entire
population in the best possible way.
Measures of central tendency.
The central tendency is a single value
that aids in the description of the data
by determining its center position.
Measures of central tendency are
sometimes known as summary statistics or
measures of central location.
The most popular measurements of central
tendency are mean, median, and mode. The
normal distribution is a bell-shaped
symmetrical distribution in which mean,
median, and mode all are equal. The
curve over here shows the bell-shaped
curve or the normal distribution of
variable X. The point over here that is
X1 is the point which represents the
mean, median and mode of this
distribution. Mean mean is calculated by
dividing these sum of all data values by
the total number of data values. It gets
affected when there are unusual or
extreme values. It is sensitive to the
outliers. Mean can be calculated as
summation over all the values of X in a
collection divided by the size of the
collection.
For example, we have a collection where
we have values as 7 3 4 1 6 and 7.
We find out the sum of these values
which is 28 and there are total of six
values. So 28 / 6 gives us a mean value
of 4.66.
Median,
it is the middle value in the set of the
data that has been sorted in ascending
order.
It is a better alternative to mean since
it is less impacted by outliers and
skewess.
It is closer to the actual central
value.
Median is calculated differently for
different sizes of data.
Differentiated as if the total number of
values is odd or if the total number of
values is even. If the size of the data
is odd. For example, in this case we
have five elements.
After sorting whatever middle value we
get
that means n + 1 by 2 term in this case
5 + 1 / 2
that is the third term which is four is
the median value.
In case when the total number of values
is even like here there are six values.
The average or the mean of the two
central values is considered as the
median. In this case the median is the
mean of 6 and four which is five. Mode.
Mode represents the most common value in
the data set. It is not at all affected
by extreme observations.
It is the best measure of central
tendency for highly skewed or non-normal
distribution.
Mode for categorical data is determined
by estimating the frequencies for each
categories
and then the category with the highest
frequency is considered to be mode.
Like in this case 7 has the highest
frequency. Hence seven becomes the mode
value. However, in case of continuous
data or quantitative data, the
calculation of mode is slightly
different. The first step in calculation
of mode is dividing the data into
classes which are equal with then
getting the frequency of data points
lying in within that range of classes
and finally selecting the class with the
highest frequency.
Using the range of that class and the
frequencies, we can get the final mode
value.
Using the formula L+
minus F_sub_1 multiplied to H / FM minus
F_sub_1 plus FM minus F_sub_2.
Here L is the lower limit or the lower
observation of the mode class.
H is the size of the mode class.
FM is the frequency of the mode class.
F_sub_1 is the frequency of the class
proceeding to mode and F_sub_2 is the
frequency of the class succeeding to
mode. This gives us the final mode
value,
mean versus expectation.
Now let's talk about mean versus
expectation.
So in general we use the expected value
or expectation when we want to calculate
the mean of a probability distribution
that represents the average value we
expect to occur before collecting any
data. And mean on the other hand mean is
basically used when we want to calculate
the average value of a given sample.
This represents the average value of raw
data that we may have already collected.
We can understand this by using a simple
example.
Now to calculate the expected value of
this probability distribution, we can
use a specific formula from the previous
discussion.
This is going to be the expected value
where X is going to be the data value
and this PX is the probability of value.
For example, we could calculate the
expected value for this probability
distribution to be as shown.
So here it will be 1.45 goals.
So this represents the expected number
of goals that the team will score in any
given game.
And then if you talk about calculating
mean, so we typically calculate the mean
after we have actually collected raw
data.
For example, suppose we record the
number of goals that a soccer team will
score in 15 different games.
Now to calculate the mean number of
goals scored per game,
we can use the following formula
where sum of x is basically the sum of
all the goals divided by n and the
number of records or we can say the
sample size.
It is as shown on the screen.
So this represents the mean number of
goals scored per game by the team.
Measures of asymmetry.
The difference between the three
distinct curves can be studied in this
image.
The central curve is the normal or no
skewess curve. Here mean, median and
mode all lie on the same point. This
normal curve is symmetrical about its
mean, median and mode.
That means the left hand side of the
curve is a mirror image of the right
hand side of the curve.
However, in case of negatively skewed
data, the tail is elongated on the left
hand side
and the mean is smaller than the mode
and the median values or is on the left
hand side of the mode.
Hence indicating that the outliers are
in the negative direction.
On the other hand, in case of positively
skewed, the data is concentrated on the
left hand side of the curve.
While the tail is elongated or longer on
the right hand side of the curve,
the mean is greater than the mode and
median
or is on the right hand side of the mode
and median indicating that the outliers
are in the positive direction.
Let's consider an example.
The graph here shows the global income
distribution for the year 2003 2013 and
a projection for 2035.
If we see the global income distribution
statistics for 2003 it is highly right
skewed.
We can observe in the previous graph
that in 2003
the mean of $3,451
was higher than the median of $1090.
The global income is definitely not
evenly distributed. The majority of
people make less than $2,000 each year.
while only a small percentage of the
population earns more than $14,000.
Measures of variability.
Measures of variability.
Dispersion. The measure of central
tendencies provide a single value that
addresses the full worth. However, the
central tendency cannot depict the
viewpoint entirely. The metric of
dispersion helps us focus on the
inconsistency in the data spread.
Measures of dispersion describe the
spread of the data.
The range, intercortile range, standard
deviation and variance are examples of
dispersion measures.
Range.
The range of distribution is the
difference between the largest and the
smallest amount of data.
The range, for example, does not include
all of a series positive aspects.
It concentrates on the most shocking
aspects and ignores that aren't
considered critical. For example, for a
set 13, 33, 45, 67, 70.
The range is 57. That is the maximum of
this which is 70 minus the minimum over
here which is 13.
Variance.
Variance is the average of all squared
deviations.
It is defined as the sum of squared
distance between each point and the mean
or the dispersion around the mean.
The standard deviation is used as
variance suffers from a unit difference.
Variance can be computed as sigma square
summation over x - mu^ 2
divided by n
where mu is the mean of the data, x is
the individual data point
and n is the size of the data.
This representation is for a population
data.
For a sample data variance can be
computed as X minus
Xar whole square summation
over it divided by n minus one.
Here Xbar is the mean of these sample
data and n is the sample size.
The units of values and variance are not
equal.
So another variability measure is used.
Standard deviation.
Standard deviation is a statistical term
used to measure the amount of
variability or dispersion around a mean.
The standard deviation is calculated as
the square root of variance. It depicts
the concentration of the data around the
mean of the data set.
Standard deviation as indicated
previously can be computed as square
root of variance
for a population data. Standard
deviation sigma can be computed as
square root of summation over x i minus
mu^ square / n
where mu is the mean of the data x i are
the data points and n is the size. Let's
consider an example.
Let's find out the mean, variance, and
standard deviation for this data. The
data values are three, 5, 6, 9, and 10.
To find out the mean, we first find the
sum of all these data values
that is 33 and divide it by the count,
which is five.
We get the mean of 6.6. To compute the
variance, we start by computing the
deviation.
That is X minus the mean of X. Here
three is one of the values of the data
and 6.6 is the mean.
So 3 - 6.6 squared and we do that
to find out sum of all the deviations
divided by the count
which is five.
we end up getting an overall variance of
6.64.
Standard deviation as we know is
measured at square root of variance that
is square of 6.64
which amounts to 2.576.
Measures of relationship.
Measures of relationship. Coariance.
Coariance is the measure of joint
variability of two variables.
It measures the direction of the
relationship between the variables. It
determines if one variable will cause
the other to alter in the same way.
Coariance between variable X and Y can
be computed as summation over the
product of X I - XR
and Y I - Y bar the whole divided by N
minus one.
Here Xar and Y bar are the mean of X and
Y respectively. The value of covariance
can range from minus infinity to a plus
infinity.
Correlation.
Correlation is normalized coariance.
It measures the strength of association
between two variables. The most common
measure for correlation is the Pearson
correlation coefficient.
Correlation between two variables
X and Y can be measured with respect to
coariance as coariance between X
and Y divided by the standard deviation
of X and standard deviation of Y.
The value of correlation ranges from a
negative 1 to positive 1.
Types of correlation.
Correlation can be either a positive
correlation,
zero correlation or a negative
correlation.
The first picture over here represents a
perfect positive correlation
wherein a straight line with a positive
slope
is representing the relationship between
the two variables.
Zero correlation means that the line
representing the relationship between
the two variables is horizontal to the
xaxis.
Perfect negative correlation can be
represented by a straight line with a
negative slope.
Correlation equals to 1 implies a
positive relationship. That is when one
variable increases the other variable
also increases. A correlation value of
negative 1 implies a negative
relationship. That is when one variable
increases the other decreases.
The correlation coefficient of zero
shows that the variables are completely
independent of each other.
Let's consider an example.
Here we have two variables height and
weight.
To compute the correlation between
height and weight,
we use the correlation formula as
covariance of X
and Y divided by standard deviation of X
and standard deviation of Y.
Here height is the X variable and weight
is the Y variable.
First to compute coariance we compute
the x - xar and y - y bar values and
then the product of them.
We then compute x - xrยฒ
and y - y bar square values to compute
the standard deviations of height and
weight respectively. Correlation as we
know has been defined as covariance of x
and i and y divided by standard
deviations of x and y.
This can also be represented as
summation over x - xr multiplied to y -
y bar
divided by square root of summation over
sum of squared deviations
that is x - xr square multiplied to
square root of summation over y - yar
whole square that is sum of square
deviations for y.
Now let's find out values to put into
this formula.
First we find out the overall sum of
height to get the mean of height which
is 5.14.
Similarly we get the sum of weight to
get the mean of weight as 50. We now get
the summation over x - xr multiplied to
y - y bar to get the numerator for the
formula. Then we compute x - xr square
summation
and y - y bar square that is sum of
squared deviation of x and y
respectively.
Now we put in the values in this final
correlation formula to get a correlation
value of 0.889.
This indicates that height and weight
have a positive relationship.
It is evident that as height grows,
weight also increases.
In this module, we will be talking about
expectation and variance.
So the expected value or we can say mean
of a given variable that we can denote
by X is a discrete random variable where
it is a weighted average of the possible
values that X can take and each value is
going to be according to the probability
of that specific event occurring.
So usually the expected value of X is
denoted by a simple formula where we can
define the expectation based on the X
parameter.
which is going to be the sum of each
possible outcome multiplied by the
probability of the outcome occurring.
So in more concrete terms, the
expectation is what we would expect the
outcome of an experiment to be on
average.
We can take an example for the coin. If
a coin is being tossed 10 times, then
one is most likely to get five heads and
five tails.
Same logic can be discussed if we talk
about another example of rolling a die.
So there are six possible outcomes when
you roll a dieice 1 2 3 4 5 6. And each
of these has a probability of 1 by 6 of
occurring. So we can say that the
expectation is going to be 1 multiplied
by the probability of that happening
which is going to be 1x 6 + 2x 6 + 3x 6
+ 4x 6 + 5x 6 + 6x 6 and that is going
to give us 3.5 as an output. The
expected value is 3.5.
So if you think about it, 3.5 is halfway
between the possible values that I can
take and this is what we should have
expected.
Next we talk about the concept of
variance. So variance of a random
variable allows us to know something
about the spread of the possible values
of the variable. So for a discrete
random variable X the variances of X is
going to be denoted by using a simple
formula that is going to be var=
E X - M the whole square where M is
basically the expected value of the
expectation of X. So this is more like a
standard deviation of X which can also
be represented by using this formula. So
the variance does not behave in the same
way as expectation when we multiply and
add constants to random variables.
So now there are two different type of
variance that we can have a fair
understanding on. First of all we have
low variance and then we have high
variance.
So low variance simply means that there
is a small variation in the production
of the target function with changes in
the trading data set and at the same
time high variance as we can see here
high variance shows a large variation in
prediction of the target function with
changes in the trading data set. So a
model that shows high variance learns a
lot and perform well with the training
data set and it does not generalize well
with the unseen data set and that's why
as a result such a model gives good
results with training data set but shows
high error rates on the test data set
and since the high variance a model
learns too much from the data set it
leads to an overfitting of the model. So
model with high variance will be having
couple of issues like it may lead to
overfitting or it may also lead to
increase in model complexities.
Next we have skewess.
So skewess in simple terms is basically
a measure of asymmetry of a
distribution. So distribution is
asymmetrical when its left and right
sides are not the mirror images.
Right now this is a mirrored image and a
distribution can have right positive or
we can say negative or it can have zero
skewess.
So right skewed in this scenario is
basically the distribution is longer on
the right side of its peak
and a left skew distribution is going to
be we can say where it is longer on the
left side.
So we can see we have this one as a part
of right side. It is more elongated
towards the right side and this one is
more elongated towards the left side. So
we can think of skewess in terms of
tails. A tail is long tampering and the
end of a distribution. So it simply
indicates that they are observations at
one end of the distribution but that
they are relatively infrequent. So a
right skew distribution has a long tail
on the right side as you can see here.
So the number supports observed. Let's
say we have a data on a per year basis.
So again we can have a more skewess
towards the right side where data is
being dropping as we continue to
increase the number of years. For
example we may have a high sales towards
the beginning of year suppose in 2022
but again as we proceed to 2023 second
half we are seeing the dip in
performance. So that is rightly skewed
and same way let's suppose if we started
with the sales figure it was really less
in suppose 2002
but again as we proceeded to 2023 now
our sales have been gradually
increasing. So it's more like skew
towards the left section as a part of
negative skew. Next we have curtosis.
So curtosis is basically a measure of
the tailness of a distribution.
So taeness is how often the outliers
occur and act as curtis is the tailness
of the distribution related to a normal
distribution. So a distribution with
medium curttosis is called as messortic.
A distribution with low curtosis like
this one. This is called as the
platicurtic and then distribution with
high curtosis like this one. This is
called as the leptocortic.
So tails here they are tapering ends on
either side of a distribution like this.
So they represent the probability or the
frequency of values that are extremely
high or extremely low to the mean.
In other words, tails here represents
how often the outliers occur.
So there are three type of curtis. We
have platicurtic which is negative,
leptocortic which is a positive towards
the upper end and then we have messertic
which is a normal distribution. So
messertic is the medium tail. So normal
distributions they have a curtosis of
three. So any distribution with a
kurtosis of approx value of three is
going to be messertic. And curtosis is
described in terms of excess curttosis
which is curtosis minus3. And since
normal distribution they have a curtosis
of three axis curtises makes comparing a
distribution curtosis to a normal
distribution even easier. Introduction
to probability.
Probability theory. Probability is a
measure of the likelihood that an event
will occur.
Let's consider an example of coin toss
where the chances of getting heads on a
coin are 1 by two or 50%.
The probability of each given event is
between zero and one both inclusive. Sum
of an events cumulative probability
cannot be greater than one.
Hence the probability of an event X lies
between zero and one. This means that
the integral of probability of
distribution over x equals to 1.
Conditional probability. Conditional
probability of any event A is defined as
the probability of occurrence of A given
that event B has previously occurred.
Condition probability of event A given B
can be estimated as probability of A
intersection B that is probability of
both A and B happening together
divided by the probability of B.
It is also written as that probability
of A intersection B equals to
probability of A given B multiplied to
probability of B.
Let's consider an example.
In a coin, we are doing a two coin flip.
Coin one gets heads, tails, heads, and
tails in subsequent flips.
while coin two gets tails, heads, heads,
and tails in the subsequent flips. Now,
the probability that coin one will get a
head is 2 out of four. While the
probability that coin two will get heads
is again two out of four.
The probability that both coin one and
coin two will have a heads is just one
out of the four flips.
Hence the probability that coin one will
get heads given that coin 2 is already
heads can be computed as probability of
coin one edge intersection coin 2 edge
that is 1x4 divided by probability of
coin 2 edge
that's a given that is 2x 4 which is
going to be 0.5 or 50% based
base theorem Base theorem calculates the
conditional probability of an event
based on its prior probabilities.
Basically base theorem incorporates the
prior probability distribution to
predict the posterior probabilities.
Base theorem for conditional probability
can be expressed as probability of A
given B equals probability of B given A
divided by probability of B multiplied
to probability of A.
Base theorem allows updating the
probability values by using new
information or evidence. Here
probability of A is known as prior
probability. That is the probability of
event before any new data is collected.
Probability of A given B is known as the
posterior probability. It is the revised
probability of an event occurring after
taking into consideration the new
information probability of B given A is
known as the likelihood and probability
of B is probability of observing an
evidence B model. An example consider an
example for calculating the likelihood
of having diabetes based on frequency of
fast food consumption. Here is the
observed data. Let's say the fast food
audience is 20%. Diabetes prevalence is
10% and 5% is fast food and diabetes.
The chances of diabetes given fast food
that is the conditional probability of D
given B can be calculated as probability
of diabetes and fast food together
divided by probability of fast food.
That means 5% divided by 20%. that
equals 25%.
Define an analysis can state eating fast
food increases the chance of having
diabetes by 25%.
The multiplication rule of probability
if events A and B are statistically
independent and probability of A
intersection B can be given as
probability of A given B multiplied to
probability of B. However, probability
of A intersection B is also given as
probability of A multiplied to
probability of B. Here probability of A
given B equals to probability of A when
we assume that probability of B is non
zero. Similarly, probability of B equals
probability of B given A assuming
probability of A is non zero.
Chain rule of probability joint
probability distributions over many
random variables can be reduced into
conditional distributions over a single
variable. It can be expressed as
probability of X1 X2 so on until Xn
equals probability of X1 intersection
probability of X I given probability of
X1 till X I minus one.
For example, the joint probability of A,
B and C can be given as probability of A
given B. C multiplied to probability of
B given C multiply to probability of C.
Logistic sigmoid.
The logistics function is a type of
sigmoid function that aims to predict
the class to which a particular sample
belongs. Its outcome is discrete binary
value. a probability between zero and
one. The logistic sigmoid is a useful
function that follows the yes curve. It
saturates when the input is very large
or very small. Logistic sigmoid is
expressed as sigma of x= 1 upon 1 + e to
the power minus x.
The logistic sigmoid can be expressed as
sigmoid function of x is given as 1 upon
1 + e ^ minus x where e is the ooler's
number.
Gshian distribution.
The gossian distribution is a type of
distribution in which data tends to
cluster around a central value with
little or no bias to the left or right.
It is often referred to as normal
distribution.
In absence of prior information, the
normal distribution is frequently a fair
assumption in machine learning
equation.
The formula for calculating Gaussian
distribution is described as the normal
distribution of X.
That is the function of x given mean as
mu and variance is sigma square can be
calculated as 1 upon sigma square
roo of 2 pi e to the power -/ x -
mood / sigma square
where mu is the mean or peak value which
also is the expected value of x.
Sigma is the standard deviation. Sigma
square is the variance.
A standard normal distribution has a
mean of zero and a standard deviation of
one.
Gshian distribution can be univariate
which describes the distribution of a
single variable X.
It can also be multivariate where it can
just use to describe the distribution of
several variables.
It is represented in 3D of ND formats.
Law of large numbers.
Now let's talk about law of large
numbers. The law of large numbers states
that an observed sample average from a
large sample will be close to the true
population average and that it will get
closer in the larger sample. So the law
of large number does not guarantee that
a given sample spatially a small sample
will reflect the true population
characteristics or that a sample does
not reflect the true population will be
balanced by a subsequent sample. This is
for the law of large numbers to express
the relationship between scale and
growth rate.
So there are multiple examples through
which we can understand
and it is widely used in statistical
analysis in working with the central
limit theorem in terms of the business
growth. So there are multiple real time
setup in which these are going to be
used. So if you talk about tossing a
coin so tossing a coin in a number of
times will give us two different type of
outcomes.
the result will spread evenly between
head and tails and the expected average
value is going to be half.
That means 50 times tails and 30 times
heads. But again, if you toss a coin
1,000 times, then the result can be in
different manners because out of 1,000,
let's say 850 times it has been head and
only 150 times it has been tails and so
on. So that's why the possibility of one
event occurring is going to be changed
in large sample sets as compared to a
small sample sets as in let's say 10
times. So the number of heads and tails
unbalanced for lower number of trials.
So we can see it is unbalanced.
But again as soon as we toss more number
of coins more leans towards the balance
value or we can see the observed
averages.
Next we have p value.
So p value is basically a number
calculated from the statistical test
that describes how likely we are to have
found a particular set of observations
if the null hypothesis were true. So p
values are used in hypothesis testing to
help decide whether to reject the null
hypothesis. And the smaller the p value,
the more likely we are to reject the
null hypothesis.
So we have a term called as null
hypothesis. So all statistical tests
they have null hypothesis. So for most
tests the null hypothesis is that there
is no relationship between our variables
of in first or that there is no
difference among groups. For example in
a two-tail t test the non-hypothesis is
that the difference between two groups
is going to be zero.
So p value is going to tell us how
likely it is that our data could have
occurred under the null hypothesis.
It is done by calculating the likelihood
of a test statistic
which is the number calculated by a
statistical test using our data. So p
value tell us how often we would expect
to see a test statistic as extreme or
more extreme
than one calculated by a statistical
test. if the null hypothesis of the test
was true.
So there are multiple limitations as
well. So first one is the results can be
significant but again they are they may
not be practical as we have compared it
can be based on multiple hypothesis for
a game for the healthcare test. If the
test is going to be positive or not it
may show even values of the effect of a
variable but not the magnitude in real
life. What exactly is going to be the
application of a drug test being failed
in pharma company? Therefore, it is
recommended to use confidence and levels
in addition to the p values to quantify
or we can say to give a solid figure to
the reserve which we are going to get.
The p values they are interpreted as
supporting or we can say refuting the
alternative hypothesis.
So p value can only tell you whether or
not the null hypothesis is supported. It
cannot tell us whether our alternative
hypothesis is true or why. So the risk
of rejecting the null hypothesis is
often higher than the p value. So
especially when we are looking at a
single study or when using small sample
sizes. So this is because the smaller
frame of reference, the greater are the
chance that as we stumble across a
statistically significant pattern
completely by accident.
Key takeaways.
Key takeaways. Probability and
statistics structure the premise of the
data. The data helps in anticipating the
future or gauging in view of the past
patterns of information.
The central tendency is a single value
that helps to describe the data by
identifying these central positions. The
mean, median, and mode are the measures
of central tendencies.
The distribution where the data tends to
be around a central value with a lack of
bias or minimal bias towards the left or
right is called as gshian distribution.
>> My name is Richard Kersner with the
simply learn team. That's get certified,
get ahead. We're going to cover
mathematics for machine learning. So
today's agenda is going to cover data
and its types. Then we're going to dive
into linear algebra and its concepts,
calculus, statistics for machine
learning, probability for machine
learning, hands-on demos, and of course
throwing in there in the middle is going
to be your matrixes and a few other
things to go along with all this.
Data and its types. Data denotes the
individual pieces of factual information
collected from various sources. It is
stored, processed and later used for
analysis.
And so we see here uh just a huge
grouping of information, a lot of tech
stuff, money, dollar signs, numbers
uh and then you have your performing
analytics to drive insights and
hopefully you have a nice share your
shareholders gathered at the meeting and
you're able to explain it in something
they can understand. So we talk about
datas types of data we have in our types
of data we have a qualitative
categorical
you think nominal or ordinal and then
you have your quantitative or numerical
which is discrete or continuous
and let's look a little closer at those
data type vocabulary always people's
favorite is the vocabulary words okay
not mine uh but let's dive into this
what we mean by nominal nominal they are
used to label various just uh label our
variables without providing any
measurable value. Uh country, gender,
race, hair, color, etc. It's something
that you either mark true or false. This
is a label. It's on or off. Either they
have a red hat on or they do not. Uh so
a lot of times when you're thinking
nominal data labels, uh think of it as a
true false kind of setup. And we look at
ordinal. This is categorical data with a
set order or a scale to it. Uh and you
can think of salary range is a great
one. Uh movie ratings etc. You see here
the salary range if you have 10,000 to
20,000 number of employees earning that
rate is 150 20,000 to 30,000 100 and so
forth. Some of the terms you'll hear is
bucket. Uh this is where you have 10
different buckets and you want to
separate it into something that makes
sense into those 10 buckets. And so when
we start talking about ordinal, a lot of
times when you get down to the brass
bones, again, we're talking true false.
Uh so if you're a member of the 10 to
20k range, uh so forth, those would each
be either part of that group or you're
not. But now we're talking about buckets
and we want to count how many people are
in that bucket. Quantitative numerical
data uh falls into two classes, discrete
or continuous. And so data with a final
set of values which can be categorized
class strength questions answered
correctly and runs hit in cricket. A lot
of times when you see this you can think
integer uh and a very restricted integer
i.e. you can only have 100 questions um
on a test. So you can it's very
discreet. I only have a 100 different
values that it can attain. So think
usually you're talking about integers
but within a very small range. They
don't have an open end or anything like
that.
Uh so discrete is very solid, simple to
count, set number. Continuous on the
other hand uh continuous data can take
any numerical value within a range. So
water pressure, weight of a person etc.
Usually we start thinking about float
values where they can get phenomenally
small in their in what they're worth.
And there's a whole series of values
that falls right between discrete and
continuous. Um you can think of the
stock market. You have dollar amounts.
It's still discreet, but it starts to
get complicated enough when you have
like, you know, jump in the stock market
from $525.33
to $580.67.
There's a lot of point values in there.
It'd still be called discreet, but you
start looking at it as almost continuous
because it does have such a variance in
it. Now uh we talk about n we did we
went over nominal and ordinal uh almost
true false charts and we looked at
quantitative and numerical data which
we're starting to get into numbers.
Discrete you can usually a lot of times
discreet will be put into it could be
put into true false but usually it's
not. Uh so we want to address this stuff
and the first thing we want to look at
is the very basic which is your algebra.
So we're going to take a look at linear
algebra. You can remember back when your
uklidian geometry uh we have a line.
Well, let's go through this. We have
linear algebra is the domain of
mathematics concerning linear equations
and their representations in vector
spaces and through matrices. I told you
we're going to talk about matrices. Uh
so a linear equation is simply um uh 2x
+ 4 y - 3 z = 10. Very linear. 10 x +
12.4 4 y = z. And now you can actually
solve these two equations by combining
them. Uh, and that's we're talking about
a linear equation.
In the vectors, we have a + b= c. Now,
we're starting to look at a direction.
And these values usually think of an xyz
plot. Um, so each one is a direction.
And the actual distance of like a
triangle A is C. And then your matrix
can describe all kinds of things. Um, I
find matrixes uh confuse a lot of
people, not because they're particularly
difficult, but because of the magnitude
and the different things are used for.
And a matrix is a chart or a um, you
know, think of a spreadsheet, but you
have your rows and your columns. And
you'll see here we have a * b= c. Very
important to know your counts. Uh, so
depending on how the math is being done,
what you're using it for, making sure
you have the same rows and the number of
columns or a single number, there's all
kinds of things that play in that that
can make matrixes confusing. Uh, but
really it has a lot more to do with what
domain you're working in. Uh, are you
adding in multiple polomials where you
have like uh uh ax^2 plus b y plus, you
know, you start to see that can be very
confusing versus a very straightforward
matrix. And let's just go a little
deeper into these because these are such
primary this is what we're here to talk
about is these different math uh
mathematical computations that come up.
So we're looking at linear equations.
Let's dig deeper into that one. An
equation having a maximum order of one
is called a linear equation. Uh so it's
linear because when you look at this we
have uh ax plus b= c which is a one
variable. We have two variable ax plus b
y = c ax plus b y plus z c cz z= d and
so forth. But all of these are to the
power of one. You don't see x squar. You
don't see x cubed. So we're talking
about linear equations. That's what
we're talking about in their addition.
If you have already dived into say
neural networks, you should recognize
this ax plus b y plus cz um setup plus
the intercept. uh which is basically
your your neural network each node
adding up all the different inputs and
we can drill down into that most common
formula is your y = mx + c.
So you have your uh y equals the m which
is your slope, your x value plus c which
is your um y intercept. They kind of
labeled it wrong here
threw me for a loop but the the c would
be your y intercept. So when you set x
equal to zero, y equals c. And that's
that's your y intercept right there. Uh
and that's they they just had reversed
value of y. When x equals 0, it equals
the y intercept, which is c. And your
slope gradient line, which is your m. So
you get your y = 2x + 3. And there's
lots of easy ways to compute this. This
why this is why we always start with the
most basic one when we're solving one of
these problems. And then of course the
um one of the most important takeaways
is the slope gradient of the line. Uh so
the slope is very important that m
value. Uh in this case we went ahead and
solved this. If you have y = 2x + 3 you
can see how it has a nice line graph
here on the right.
So matrixes a matrix refers to a
rectangular representation of an array
of numbers arranged in columns and rows.
So we're talking m rows by n columns
here. A11 is denotes the element of the
first row in the first column. Similarly
a12 and it's really pronounced a11 in
this particular setup. So it's row one
column one. A12 is a of row one column 2
uh first row and second column and so
on.
And there's a lot of ways to denote
this. I've seen these as like a capital
letter A, smaller case A for the top row
or I mean you can see where they can go
all kinds of different directions as far
as the value. You just take a moment to
realize there's need to be some
designation as far as what row it's in
and what column it's in. And we have our
uh basic operations. We have addition.
So when you think about addition, you
have uh two matrices of 2x two and you
just add each individual number in that
matrix and then when you get to the
bottom you have uh in this case the
solution is 12, 10 + 2 is 12, 5 + 3 is 8
and so on. And the same thing with
subtraction.
Now again you're counting matrices you
want to check your um dimensions of the
matrix the shape you'll see shape come
up a lot in programming. So we're
talking about dimensions we're talking
about the shape. If the two shapes are
equal this is what happens when you add
them together or subtract them. And we
have multiplication. When you look at
the multiplication you end up with a
very slightly different setup going.
Now, if we look at our last one, we're
um uh we're like, why? This always gets
to me when we get to matrices. They
don't really say why you multiply
matrices. Um you know, my first thought
is 1 * 2, 4 * 3. But if you look at
this, we get 1 * 2 + 4 * 3, 1 * 3 + 4 *
5,
uh 6 * 2 + 3 * 3, 6 * 3 + 3 * 5. If
you're looking at these matrices, uh,
think of this more as an equation. And
so we have, uh, if you remember when we
back up here for our multiple line
equations, let's just go back up a
couple slides where we were looking at,
uh, two variable. So this is a two
variable equation. ax plus b y= c.
Um, and this is a way to make it very
quick to solve these variables. And
that's why you have the matrix, and
that's why you do
the multiplication the way they do. And
this is the dotproduct of uh 1 * 2 + 4 *
3
1 * 3 + 4 * 5
uh 6 * 2 + 3 * 3 6 * 3 + 3 * 5. And it
gives us a nice little 14, 23, 21, and
33 over here, which then can be used and
reduced down to a simple um formula as
far as solving the variables as you have
enough inputs. Uh and then in matrix
operations, when you're dealing with a
lot of matrices, uh now keep in mind
multiplying matrices is different than
finding the product of two matrices.
Okay? So we're talking about
multiplication, we're talking about
solving uh for equations. When you're
finding the product, you are just
finding one time two. Keep that in mind
because that does come up. I've had that
come up a number of times where I am
altering data and I get confused as to
what I'm doing with it. Uh transpose
flipping the matrix over it's diagonal.
Comes up all the time where you have you
still have 12, but instead of it being
uh 128, it's now 1214 821. You're just
flipping the columns and the rows. Uh
and then of course you can do an inverse
um changing the signs of the values
across this main diagonal. And you can
see here we have the inverse a to the
minus1 and ends up with uh instead of 12
8 14 12 it's now -22 -12 vectors uh
vector just means we have
a value and a direction and we have down
four numbers here on our vector.
uh in mathematics a one-dimensional
matrix is called a vector. Uh so if you
have your xplot and you have a single
value that values along the x- axis and
it's a single dimension. If you have two
dimensions you can think about putting
them on a graph. You might have x and
you might have y and each value denotes
a direction. And then of course the
actual distance is going to be the
hypothesis of that triangle. Uh and you
can do that with three dimensionals x y
and z. uh and you can do it all the way
to nth dimensions. So when they talk
about the k means uh for categorizing
and how close data is together, they
will compute that based on the
Pythagorean theorem. So you would take
uh the square of each value, add them
all together and find the square root
and that gives you a distance as far as
where that point is, where that vector
exists or an actual point value. And
then you can compare that point value to
another one and it makes a very easy
comparison versus comparing uh 50 or 60
different numbers. And that brings us up
to gene vectors and I gene values. Uh I
gene vectors the vectors that don't
change their span while transformation
and I gene values the scalar values that
are associated to the vectors.
Conceptually you can think of the vector
as your picture. you have a picture.
It's um uh two dimensions x and y. And
so when you do those two dimensions and
those two values or whatever that value
is um that is that point but the values
change when you skew it and so if we
take and we have a vector a and that's a
set value uh b is um your is your you
have a and b which is your hygiene
vector. Two is the i gene value. So,
we're altering all the values by two.
That means we're u maybe we're
stretching it out one direction, making
it tall if you're doing picture editing.
Um that that's one of the places this
comes in. But you can see when you're
transforming uh your different
information, how you transform it is
then your hygiene value. And you can see
here uh vector after line transition
uh we have 3 a is the hygiene vector.
Three is the hygiene value. So A doesn't
change. That's whatever we started with.
That's your original picture. And three
uh is skewing it one direction and maybe
uh B is being skewed another direction.
And so you have a nice tilted picture
because you've altered it by those by
the hygiene values.
So let's go ahead and pull up a demo on
linear algebra. And to do this, I'm
going to go through my trusted Anaconda
into my Jupiter notebook. and we'll
create a new uh notebook called linear
algebra. Since we are working in Python,
uh we're going to use our numpy. I
always import that as np or numpy array.
Probably the most popular um module for
doing matrixes and things in
given that this is part of a series. I'm
not going to go too much into numpy. Uh
we are going to go ahead and create two
different variables. A for a numpy array
10 15 and b 29.
We'll go ahead and run this. And you can
see there's our two arrays 105 29. And I
went ahead and added a space there in
between so it's easier to read. And
since it's the last line, we don't have
to put the print statement on it unless
you want. We can simp but we can simply
do a plus b. So when I run this, uh, we
have 10 15 29 and we get 30 24, which is
what you expect. 10 + 20 15 + 9. You
could almost look at this addition as
being um
just adding up the columns on here
coming down. And if we wanted to do it a
different way, we could also do a t plus
b dot t. Remember that t flips them. And
so if we do that, we now get them uh we
now have 304 going the other way. We
could also do something kind of fun.
There's a lot of different ways to do
this. Uh, as far as a plus b, I can also
do a plus b. T and you're going to see
that that will come out the same. The 30
24 whether I transpose a and b or
transpose them both at the end.
And likewise, we can very easily
subtract two vectors. I can go a minus
b. And we run that and we get - 106. Now
remember, this is the last line in this
particular section. That's why I don't
have to put the print around it. Um, and
just like we did before, we can
transpose either the individual or we
can transpose the main setup and then we
get a minus 106 going the other way.
Now, we didn't mention this in our
notes, but you can also do a scalar
multiplication.
Let me just put down scaler so you can
remember that. Uh what we're talking
about here is I have uh this array here
u and if I go a time u uh we'll take the
value two we'll multiply it by every
value in here. So 2 * 30 is 60 2 * 15
and just like we did before
um this happens a lot because when
you're doing matrices you do need to
flip them you get 6030 coming this way.
So in numpy uh we have what they call
dotproduct
and uh what this this in a
twodimensional vectors it is the
equivalent of two matrix multiplication
and remember we were talking about
matrix multiplication
uh where it is the well let's walk
through it
we'll go ahead and start by defining two
um numpy arrays we'll have uh 10 20 256
or our u and our E uh and then we're
going to go ahead and do if we take
the values uh and if you remember
correctly
an array like this would be 10 * 25 + 20
* 6. We'll go ahead and uh print that.
There we go.
And then we'll go ahead and do the uh np
dot of u comma
v.
And we'll find when we do this, we go
and run this uh we're going to get uh
370
370.
So this is a strain multiplication where
they use it to solve uh linear algebra
uh when you have multiple numbers going
across. And so this could be very
complicated. We could have a whole
string of different variables going in
here. But for this we get a nice uh
value for our dot multiplication
and we did um addition earlier which was
just your basic addition. Uh and of
course a matrix you can get very
complicated on these or in this case
we'll go ahead and do um let's create
two complex matrixes.
This one is a matrix of um you know 1210
46 431. We'll just print out A so you
can see what that looks like. Here's
print A.
We print A out. You can see that we have
a um 2x3
layer matrix for A. And we can also put
together always kind of fun when you're
playing with print values. Uh we could
do something like this. We could go in
here. There we go. Uh, we could print a.
We have it end with uh equals a run. And
this kind of gives it a nice look. Uh,
here's your matrix. That's all this is.
Comma, n means it just tags it on the
end. That's all all that is doing on
there. And then we can simply add in
what is a plus b. And you should already
guess because this is the same as what
we did before. There's no difference.
Uh, we do a simple vector addition. We
have 12 + 2 is 14, 10 + 8 is 18. And so
on. And just like we did the uh matrix
addition, we can also do a minus b and
do our matrix subtraction.
And we look at this uh we have what? 12
- 2 is 10. 10 - 8 um where are we?
Oh, there we go. 8 min
confusing what I'm looking at. I should
have reprinted out the original numbers.
Uh but we can see here 12 - 2 is of
course 10. 10 - 8 is 2. Uh 4 - 46 is -
42 and so forth. So same as a
subtraction as before, we just call it
matrix subtraction. It's identical.
Now if you remember up here, we had
scalar addition where we're adding just
one number to a matrix. You can also do
scalar multiplication. Uh and so simply
if you have a single value A and you
have B which is your array, we can also
do A * B. When we run that, uh, you can
see here we have 2 * 4 is 8. Uh, 5 * 4
is 20 and so forth. You're just
multiplying the four across each one of
these values. And this is an interesting
one that comes up. A little bit of a
brain teaser is matrix and vector
multiplication.
And so when we're looking at this,
uh, we are just do a regular arrays. It
doesn't necessarily have to be a numpy
array. We have a
which has our um array of arrays and b
which is a single array and so we can
from here
do the dot
a b and this is going to return two
values and the first value is that it's
you could say it's like uh um we're
doing the this array b array first with
a and then with a second one and so it
splits it up so you have a matrix of
vector multiplication and you mix and
match. When you get into really
complicated uh backend stuff, this
becomes more common because you're now
you got layers upon layers of data and
so you you'll end up with a matrix and a
set of uh vector matrices. Do you want
to multiply?
Now, keep in mind that if you're doing
data science, a lot of times you're not
looking at this. This is what's going on
behind the scenes. So if you're in um
the scikit looking at sklearn where
you're doing linear regression models,
this is some of the math that's hidden
behind the scenes that's going on. Other
times you might find yourself having to
do part of this and manipulate the data
around so it fits right and then you go
back in and you run it through the
scikit. And if we can do um up here
where we did a uh matrix and vector
multiplication, we can also do matrix to
matrix multiplication. And if we run
this where we have the two matrices, uh
you can see we have a very complicated
array that of course comes out on there
for our dot. And just to reiterate it,
we have our transpose a matrix which is
your T. And so if we create a matrix A
and then we do transpose it, you can see
how it flips it from 5 10 15 20 25 30 to
5 15 25 10 20 30 uh rows and columns.
And certainly with the math, uh, this
comes up a lot. Um, it also comes up a
lot with XY plotting. When you put it
into piplot, you have one format where
they're looking at pairs of numbers and
then they want all of X's and all Y's.
So, you know, the transpose is an
important tool both for your math and
for plotting and all kinds of things.
Another tool that we didn't discuss uh
is your identity matrix. Uh and this one
is more definition.
Uh the identity matrix. Um we have here
one where we just did uh two. So it
comes down as one 0 0 1 uh 1 0 0 1 0. It
creates a diagonal of one. And what that
is is when you're doing your identities,
you could be comparing all your
different features to the different
features and how they correlate. And of
course when you have uh feature one
compared to feature one to itself it is
always one uh where usually it's between
zero one depending on how well
correlates. So when we're talking about
identity matrix that's what we're
talking about right here is that you
create this preset matrix and then you
might adjust these numbers depending on
what you're working with and what the
domain is. And then another thing we can
do uh to kind of wrap this up. We'll hit
you with the most complicated uh um
piece of this puzzle here is an inverse
um a matrix. And let's just go ahead and
put the um it's a lengthy description.
Let's go and put the description. This
is straight out of the uh the website
for um numpy. Uh so given a square
matrix A, here's our square matrix A,
which is 2 1 0 0 1 0 1 2 1. Keep in mind
3x3, it's square. It's got to be equal.
It's going to return the matrix A
inverse satisfying dot A um A inverse.
So here's our matrix multiplication.
Um and then of course it equals the dot
uh yeah a inverse of a um with an
identity shape of uh a dotshaped zero.
This is just reshaping the identity.
That's a little complicated there. Uh so
we go and have our here's our array. Uh
we'll go ahead and run this. And you can
see what we end up with is we end up
with uh an array 0.5 minus 0.5 and so
forth with our 211 going down to 1 0 0 1
0 1 2 1. Um getting into a little deep
on the math understanding when you need
this is probably really is is what's
really important when you're doing data
science versus uh handwriting this out
and looking up the math and handwriting
all the pieces out. you do need to know
about the linear algorithm inverse of a.
Uh so if it comes up, you can easily
pull it up or at least remember where to
look it up. You took a look at the
algebra side of it. Let's go ahead and
take a look at the calculus side of uh
what's going on here with the machine
learning. So calculus, oh my goodness,
and differential equations, you got to
throw that in there because that's all
part of the bag of tricks, especially
when you're doing large neural networks,
but also comes up in many other areas.
The good news is most of it's already
done for you in the back end. Uh so when
it comes up, you really do need to
understand from the data science, not
data analytics. Data analytics means
you're digging deep into actually
solving these math equations. U and a
neural network is just a giant
differential equation. Uh so we talk
about calculus uh we're going to go
ahead and understand it by talking about
cars versus time and speed. uh so helps
to calculate the spontaneous rate of
change.
Uh so suppose we plot a graph of the
speed of a car with respect to time. So
as you can see here going down the
highway probably merged into the highway
from an on-ramp. So I had to accelerate
so my speed went way up uh stuck in
traffic merged into the traffic. Traffic
opens up and I accelerate again up to
the speed limit and u maybe it peters
off up there. So you can look at this as
as um the speed versus time. I'm getting
faster and faster because I'm
continually accelerating. And if I hit
the brakes, it go the other way. So the
rate of change of speed with respect of
time is nothing but acceleration. How
fast are we accelerating? The
acceleration is the area between the
start point of x and the end point of
delta x. Uh so we can calculate a simple
if you had x and delta x we could put a
line there and that slope of the line is
our acceleration.
Now that's pretty easy when you're doing
linear algebra but I don't want to know
it just for that line and those two
points. I want to know it across the
whole of what I'm working with. That's
where we get into calculus. So when we
talk about the distance between x and
delta x it has to be the smallest
possible near to zero in order to
approximate the acceleration.
Uh so the idea is that instead of I mean
if you ever did took a basic calculus
class they would draw bars down here and
you would divide this area up um let's
go back up a screen. you divide this
area of this time period up into maybe
10 sections and you'd use that and you
could calculate the acceleration between
each one of those 10 sections kind of
thing. Uh and then we just keep making
that space smaller and smaller until
delta x is almost uh infantismally
small. And so we get a function of a uh
equals a limit as h goes to zero of a
function of a plus h minus a function of
a over h. And that is you're computing
the slope of the line.
We're just computing that slope under
smaller and smaller and smaller samples.
Uh and that's what calculus is. Calculus
is the integral. You can see down here
we have our nice uh integral sign. Looks
like a giant s. And that's what that
means is that we've taken this down to
as small as we can for that sampling. Uh
so we're talking about calculus. Finding
the area under the slope is the main
process in the integration. Similar
small intervals are made of the smallest
possible length of x plus delta x where
delta x approaches almost an infantismly
small space. And then it helps to find
the overall acceleration by summing up
all the lengths together. Uh so we're
summing up all the accelerations from
the beginning to the end. And so here's
our integral. we sum of a of x * d ofx =
a + c. Uh that is our basic calculus
here. So when we talk about
multivvariant calculus, uh multivariate
calculus deals with functions that have
multiple variables and you can see here
we start getting into some very
complicated equations. Um uh change in w
over change of time equals change of w
over change of z. the differential of z
to dx differential of x to dt. It gets
pretty complicated. Uh and it really
translates into the multivariate
integration using double integrals. And
so you have the the sum of the sum of f
ofxy of d of a equals the sum from c to
d and a to b of f ofxy dx dy equals uh
the sum of a to b sum of c to d of fxy
dy dx.
understanding the very specifics of
everything going on in here and actually
doing the math is usually calculus one,
calculus 2, and differential equations.
Uh so you're talking about three
fulllength courses to dig into and solve
these math equations. What we want to
take from here is we're talking about
calculus. Uh we're talking about summing
of all these different slopes. And so
we're still solving a linear uh
expression. We're still solving y = mx +
b, but we're doing this for
infantismally small x's. And then we
want to sum them up. That's what this
integral sign means. The the sum of a of
x d of x= a plus c.
And when you see these very complicated
uh multivariate differentiation using
the chain rule uh when we come in here
and we have the change of w to the
change of t equals the change of w dz uh
and so forth. That's what's going on
here. That's what these means. We're
basically looking for the area under the
curve which really comes to how is the
change changing and speed's going up.
How is that changing? And then you end
up with a multiple layer. So if I have
three layers of neural networks, how is
the third layer changing based on the
second layer changing which is based on
the first layer changing? And you get
the picture here that now we have a very
complicated uh multivariate integration
um with integrals.
The good news is we can solve this uh
mathematically and that's what we do
when you do neural networks and reverse
propagation. Uh so the nice thing is
that you don't have to solve this on
paper unless you're a data analysis and
you're working on the back end of
integrating these formulas and building
the script to actually build them. So we
talk about applications of calculus. Uh
it provides us the tools to build an
accurate predictive model. Um so it's
really behind the scenes we want to
guess at what the change of the change
of the change is.
That's a little goofy. I I know I just
threw that out there. It's kind of a
meta term. But if you can guess how
things are going to change, then you can
guess what the new numbers are.
Multivariate calculus explains the
change in our target variable in
relation to the rate of change in the
input variables. So there's our multiple
variables going in there. If uh one
variable is changing, how does it affect
the other variable? And then in gradient
descent, calculus is used to find the
local and global maxima. And this is
really big. Uh we're actually going to
have a whole section here on gradient
descent because it is really I mean I
talked about neural networks and how you
can see how the different layers go in
there, but gradient descent is one of
the most key things for trying to guess
the best answer to something. So let's
take a look at the code behind gradient
descent. And uh before we open up the
code, let's just do real quick uh
gradient descent.
Let's say we have a curve like this. And
most common is that this is going to
represent your error. Oops.
Error. There we go. Error. Ah, hard to
read there. And I want to make the error
as low as possible. And so what I'm
looking at it is I want to find this
line here which is the minimum value. So
we're looking for the minimum and it
does that by uh sampling there and then
it based on this it guesses it might be
someplace here and it goes hey this is
still going down. It goes here and then
goes back over here and then goes a
little bit closer and it's just playing
a high low until it gets to that spot,
that bottom spot. And so we want to
minimize the error in uh on the flip
note, you could also want to be
maximizing something. You want to get
the best output of it. Uh that's simply
uh minus the value. Uh so if you're
looking for where the peak is, this is
the same as a negative for where the
valley is and looking for that valley.
Uh that's all that is and this is a way
of finding it. So the cool thing is um
all the heavy lifting's done. Um I
actually ended up putting together one
of these a while back as uh when I
didn't know about sidekick and I was
just starting. Boy, it's a long while
back and uh is playing high low. How do
you play high low? not get stuck in the
valleys, uh, figure out these curves and
things like that. Well, you do that and
the back end is all the calculus and
differential equations to calculate this
out. The good news is you don't have to
do those. Uh, so instead, we're going to
put together the code and let's go ahead
and see what we can do with that.
So, uh, guys in the back put together a
nice little piece of code here, which is
kind of fun.
uh some things we're going to note and
this is this is really important stuff
because when you start doing your data
science and digging into your machine
learning models uh you're going to find
these things are stumbling blocks. Uh
the first one is current x. Where do we
start at? Uh keep in mind your model
that you're working with is very
generic. So whatever you use to minimize
it the first question is where do we
start? Um, and we started at this cuz
the algorithm starts at x= 3. So, we
arbitrarily picked five. Learning rate
is uh how many bars to skip going one
way or the other. Uh, I'm in fact, I'm
going to separate that a little bit
because these two are really important.
Um, if we're dealing with something like
this where we're talking about um uh
well, here's our here's the function
we're going to use our um gradient of
our function um 2 * x + 5. Keep it
simple. So that's a function we're going
to work with. So if I'm dealing with
increments of a th00and 0.1 is going to
be a very long time. And if I'm dealing
with increments of 0.001,
uh 0.1 is going to skip over my answer.
So I won't get a very good answer. Um
and then we look at precision. This
tells us when to stop the algorithm. So
again, very specific to what you're
working on. uh if you're working with
money and you don't convert it into a
float value uh you might be dealing with
0.01 which is a penny that might be your
precision you're working with. Um and
then of course the previous step size
max iterations uh we want something to
cut out at a certain point. Usually
that's built into a lot of minimization
functions. And then here's our actual uh
formula we're going to be working with.
And then we come in, we go while
previous step size is greater than
precision and its is less than max its
say that 10 times fast. Um
we're just saying if it's uh if we're if
we're still greater than our precision
level, we still got to keep digging
deeper. Um and then we also don't want
to go past a thou or whatever this is, a
million or 10,000 uh running. That's
actually pretty high. um almost never do
max iterations more than like 100 or
200. Rare occasions you might go up to
four or 500 if it's depending on the
problem you're working with. Uh so we
have our previous equals our current.
That way we can track timewise.
Uh the current now equals the current
minus the rate times the formula of our
previous x. So now we've generated our
new version. Uh previous step size
equals the absolute current previous.
Uh, so we're looking for the change in x
itters equals iterations + one. That's
so we know to stop if we get too far.
And then we're just going to print the
local minimum occurs at x on here. And
if we go ahead and run this,
uh, you can see right here it gets down
to this point and it says, hey, um,
local minimum is minus 3.3222
for this particular series we created.
Uh, and this is created off of our
formula here. lambda x2 * x + 5. Now,
when I'm running this stuff, uh you'll
see this come up a lot
and uh with the sklearn kit and and one
of the nice reasons of breaking this
down the way we did is I could go over
those top pieces. Uh those top pieces
are everything when you start looking at
these minimization toolkits in built-in
code. And so from um we'll just do it's
actually docs.cipi.org
and we're looking at the scikit. There
we go. Um optimize minimize. You can
only minimize one value. You have the
function that's going in. This function
can be very complicated. Uh so we used a
very simple function up here. It could
be there's all kinds of things that
could be on there. And there's a number
of methods to solve this as far as how
they shrink down. Uh and your x knot.
There's your there's your start value.
So your function, your start value. Um
there's all kinds of things that come in
here that we can look at which we're not
going to. Um optimization automatically
creates constraints bounds. Some of this
it does automatically, but you really
the big thing I want to point out here
is you need to have a starting point.
You want to start with something that
you already know is mostly the answer.
Uh if you don't, then it's going to have
a heck of a time trying to calculate it
out.
Or you can write your own little script
that does this and and does a high low
guessing and tries to find the max
value. That brings us to statistics.
What this is kind of all about is
figuring things out. Lot of vocabulary
and statistics. Uh so statistics, well,
I guess it's all relative. It's
definitely not an ed class. Uh so a
bunch of stuff going on. Statistics.
Statistics concerns with the collection,
organization, analysis, interpretation,
and presentation of data. That is a
mouthful. Um so we have from end to end
we're
valid, what does it mean? How do we
organize it? Um how do we analyze it?
Then you got to take those analysis and
interpret it into something that uh
people can use. kind of reduce it to
understandable. Um, and nowadays you
have to be able to present it. If you
can't present it, then no one else is
going to understand what the heck you
did.
So, we look at the terminologies. Uh,
there is a lot of terminologies
depending on what domain you're working
in. So clearly if you're working in um a
domain that deals with
viruses and tea cells and and how does
you know where does that come from and
you're studying the different people
then you're going to have a population.
if you are working with um mechanical
gear um you know a little bit different
if you're looking for the wobbling
statistics uh to know when to replace a
rotor on a machine or something like
that uh that can be a big deal. You
know, we have these huge fans that turn
in our sewage processing systems. And so
those fans, they start to wobble and hum
and do different things that the sensors
pick up. At one point, do you replace
them? Instead of waiting for it to
break, in which case it cost a lot of
money. Instead of replacing a bushing,
you're replacing the whole fan unit. Uh
an interesting project that came up for
our city a while back. Uh so population,
all objects are measurements whose
properties are being observed. Uh so
that's your population all the objects.
It's easy to see it with people because
we have our population and large. Um but
in the case of the sewer fans we're
talking about how the fan units. That's
the population of fans that we're
working with.
You have a parameter a matrix uh that is
used to represent a population or
characteristic.
You have your sample a subset of the
population studied. You don't want to do
them all because then you don't have a
if you come up with a conclusion for
everyone, you don't have a way of
testing it. So you take a sample. Uh
sometimes you don't have a choice. You
can only take a sample of what's going
on. You can't u study the whole
population. And a variable, a metric of
interest for each person or object in a
population.
Types of sampling. We have probabilistic
approach. uh selecting samples from a
larger population using a method based
on the theory of probability
and we'll go into a little bit more
deeper on these. We have random
systematic stratified and then you have
nonprobabilistic approach selecting
samples based on the subjective judgment
of the researcher rather than random
selection. Uh it has to do with
convenience trying to reach a quota um
or snowball. Uh and they're very biased.
That's one of the reasons you'll see
this big stamp on it says biased. Uh so
you got to be very careful on that. So
probabilistic sampling uh when we talk
about a random sampling, we select
random size samples from each group or
category. So we it's as random as you
can get. Uh we talk about systematic
sampling. We're selecting randomsiz
samples from each group or category with
a fixed periodic interval. Uh so we kind
of split it up. This would be like a
time setup or different categories. And
you might ask your question, what is a
category or a group? Uh if you look at
I'm going to go back a window. Let's say
we're studying um economics of different
of an area. Um we know pretty much that
based on their culture, where they came
from, they might need to be separated.
And so uh and when I say separated, I
don't mean separated from their their uh
place where they live. I mean, as far as
the analysis, we want to look at the
different groups and make sure they're
all represented. So, if we had like an
80% uh of a group that is uh say
Hispanic and or Indian and also in that
same area, we have 20% 20% who are let's
call our expatriots. They left America
and they're nice and uh your Caucasian
group. We might want to sample a group
that is representative of both. Uh, so
we're talking about stratified sampling
and we're talking about groups. Those
are the groups we're talking about. And
it brings us to stratified sampling,
selecting approximately equalized
samples from each group or category. Uh,
this way we can actually separate the
categories and give us an insight into
the different cultures and how that
might affect them in that area. Uh so
you can see these are very very
different kind of depends on what you're
working with um as far as your data and
what you're studying. And so we can see
here just to go a little bit more we'd
have selecting 25 employees from a
company of 250 employees randomly. Don't
care anything about them. What groups
they're in, which office are in,
nothing. Um and we might be selecting
one employee from every 50 unique
employees in a company of 250 employees.
And then we have selecting one employee
from every branch in the company office.
So we have all the different branches.
There's our group or our categories by
the branch. And the category could
depend on what you're studying. So it
has a lot of variation on there. You see
this kind of grouping and categorizing
is also used to generate a lot of
misinformation.
Uh so if you only study one group and
you say this is what it is, then
everybody assumes that's what it is for
everybody. And so you got to be very
careful of that. and it's very unethical
thing to kind of do. So, types of
statistics. Uh we talk about statistics,
we're going to talk about descriptive
and inferential statistics. There are so
many different terms in statistics to
break it up. Uh so we so we're talking
about a particular setup. So we're
talking about descriptive and
inferential uh statistics. You the base
of the word describe is pretty solid.
you're describing the data. What does it
look like? With inferial statistics,
we're going to take that from the small
population to a large population. So, if
you're working with a drug company, uh
you might look at the data and say,
"These people were helped by this drug.
They did uh 80% better as far as their
health or 80% better survival rate than
the people um who did not have the drug.
So, we can infer that that drug will
work in the greater populace and will
help people. So that's where you get
your inferential. Uh so we are
predicting how it's going to affect the
greater population.
So descriptive statistics it is used to
describe the basic features of data and
form the basis of quantitative analysis
of data. So we have a measure of central
tendencies. We have your mean, median
and mode. And then we have a measure of
spread like your range, your
interquartile range, your variance and
your standard deviation. And we're going
to look at all these a little deeper
here in a second. Uh but one of them you
can think of is um how the data
difference differences you know what's
the max men range all that stuff is your
spread and anything that's just a single
number is usually your central uh
tendencies measure of central
tendencies. So we talk about the mean it
is the average of the set of values
considered. what is the average outcome
of whatever's going on? And then your
median separates the higher half and the
lower half of data.
Uh so where's the center point of all
your different data points? So your mean
might have some a couple really big
numbers that skew it uh so that the
average is much higher than if you took
those outliers out where the median
would by separating the high from the
low might give you a much lower number.
you might look at and say, "Oh, that's
that's odd. Why is the average so much
higher than the median?" Well, it's
because you have some outliers, or why
is it so much lower? And then the mode
is the most frequent appearing value.
Uh, this is really interesting. If
you're studying economics and how people
are doing, you might find that the most
common um income like in the US was at
one point 24,000 a year where the
average was closer to 80,000. And it's
like, wow, what a difference. Well,
there's some people have a lot of money
and so that skews that way up. So the
average person is not making that kind
of money. And then you look at the
median income and you're like, well, the
median income is a little bit closer to
the average. Uh so it does create a very
interesting way of looking at the data.
Again, these are all uh central
tendencies, single numbers you can look
at for the whole spread of the data.
And we look at the measure of central
tendencies. The mean is the average
marks of a students in a classroom. So
here we have the mean sum of the marks
of the students total number of students
and as we talked about the median uh if
we have 0 through 10 and we take half
the numbers and put them on one side of
the line half the numbers on the other
side of the line uh we end up with five
in the middle and then the mode what
mark was scored by most of the students
in a test in a simple case where most
people scored like an 82% and got
certain problems wrong easy to figure
out. uh not so easy when you have
different areas where like you have like
the um oh let's go back to economy a
little bit more difficult to calculate
if you have a large group that scores
that makes 30,000 and a slightly bigger
group that makes 26,000. So what do you
put down for the mode? Uh certainly
there's a number of ways to calculate
that and there's actually a different
variations depending on what you're
doing. So now we're looking at a measure
of spread uh range. What's the
difference between the highest and the
lowest value? First thing you want to
look at, you know, it's we had everybody
in the test scored between 60 and 100%,
somebody got 100% or maybe 60 to 90%. It
was so hard that a lot of people could
not get 100%. Um, and you have your
interquartile range. Quartortiles divide
a rankorder data set into four equal
parts.
very common thing to do as part of all
the basic packages whether you're
working in uh dataf frames with pandas
whether you're working in scala whether
you're working in R um you'll see this
come up where they have range your min
your max and then it'll have your
interquartile range how does it look
like in each quarter of data variance
measures how far each number in the set
is from the mean and therefore from
every other number in the set uh so you
have like a how much turbulence is going
on in this data. And then the standard
deviation, it is the measure of the
variance or the dispersion of a set of
values from the mean. And you'll usually
see uh if I'm doing a graph, I might
have the value graphed. Um and then
based on the the error, I might graph
graph the standard deviation and the
error on the graph as a background so
you can see how far off it is. Uh so
standard deviation is used a lot. So
measurement of spread uh marks of a
student out of a 100 uh we have here
from 50 to 63 or 50 to 90 uh so the
range maximum marks minimum marks we
have 90 to 45 and the spread of that is
45 90 - 45 and then we have the
interquartile range using the same marks
over there you can see here where the
median is and then there's the first
quarter the second quarter and the third
quarter based on splitting it apart by
those values
And to understand the variance and
standard deviation, we first need to
find out the mean. Uh so here's our our
you know calculating the average there.
We end up at approximately 66 for the
average. And then we look at that the
variance once we know the means we can
do equals the marks minus the mean
squared. Why is it squared? Uh because
one, you want to make sure it's you
don't have like if you if you're putting
all this stuff together, you end up with
an error as far as one's negative, one's
positive, one's a little higher, one's a
little lower. Uh so you always see the
squared value and over the total
observations. And so the standard
deviation equals the square root of the
variance, which is approximately 16. And
if you were looking at um a predictable
model, you would be looking at the
deviation based on the error. How much
error does it have? Uh that's again
really important to know if you're if
your prediction is predicting something,
what's a chance of it being way off or
just a little bit off.
Now that we've looked at the um tools as
far as some of the basics for doing your
statistics and what we're talking about,
let's go ahead and pull up a little demo
and show you what that looks like in
Python code. Uh so you can get some
little hands-on here. For that, let's go
back into our Jupyter notebook in
Python. Now, almost all of this you can
do in numpy. Last time we worked um in
numpy. This time we're going to go ahead
and use pandas. And if you remember from
pandas on here, uh this is basically a
data frame, rows, columns. Let's just go
ahead and do a print df. head
and run that.
And you can see we have uh the name
Jane, Michael, William, Rosie, Hannah,
and their salaries on here. And of
course, instead of having to do all
those hand calculations and add
everything together and divide by the
total, we can do something very simple
on this uh like use the command mean in
pandas. And so if I go ahead and do this
print df, pick our column salary because
we want to find the means of that
colery.
We want to find the means of that
column. Uh and we go and print this out.
And you can see that the uh average
income on here is 71,000.
Uh, and let's just go ahead and do this.
We'll go ahead and put in uh means.
And if we're going to do that, we also
might want to find the median.
And the median is uh very similar except
it actually is just median. Uh we're
used to means and average. It's kind of
interesting that those are they use the
two different words. Uh there can be in
some computations slight differences but
for the most part the means is the
average. Uh and then the median oops
let's put a
median here. DF salary that way it
displays a little better. We can see the
median is 54 um000. So the halfway mark
is significantly below the average. Why?
Because we have somebody in here who
makes 189,000.
Darn you Rosie for throwing off our
numbers. Uh but that's something you'd
want to notice. This is this is the
difference between these is huge and so
is what is the meaning behind that when
you're studying a populace and looking
at uh the different data coming in. And
of course we also want to find out hey
what's the most uh common income that
people make in this little tiny sample.
And so we'll go ahead and do the mode.
And you can see here with the mode uh
it's at 50,000.
So this is this is very telling that
most people are making 50,000. The
middle point is at 54,000. So half the
people are making more than that. What
that tells me is that if the most common
income is way is below the median, then
there's a few there's a SK, you know,
there's a a lot of high salaries going
up, but there's some really low salaries
in there. And so this trend which is
very common in statistic you when you're
analyzing the economy and different
people's income is pretty common and the
bigger difference between these is also
very important when we're studying
statistics. Uh and when you hear someone
just say hey the average income was you
might start asking questions at that
point. Why aren't you talking about the
median income? Why aren't you talking
about the mode the most common income?
What are you hiding? Uh and if you're
doing these analysis, you should be
looking at these saying, "Hey, why why
are this discrepancies? Why are these so
different?" And of course, with any uh
analysis, it's important to find out the
minimum
and the maximum. So, we'll go ahead.
It's just simply uh um min'll pull up
your minimum and then do max pulls up
the maximum. pretty straightforward on
as far as um translating it and knowing
what your you know what the your lowest
value and what your highest value is
here. Um which you'll use to generate
like a spread later on. And real quick
on no mode mode, uh note that it puts
mode zero. Like I said, there's a couple
different ways you can compute the mode.
Um although, you know, standard one's
pretty good. We can of course do the
range, which is your max minus your min.
So now we have a range of 149,000
between the upper end and the lower end.
And you might want to be looking up the
individual values on all of these. But
it turns out there is a describe
feature in pandas.
And so in pandas we can actually do df
salary describe. And if we do this you
can see we have that there's seven uh
setups. Here's our mean. Um, our
standard deviation, which we didn't
compute yet, which would just be a STD.
And you got to be a little careful
because when it computes it, it looks
for axes and things like that. Uh, we
have our minimum value, and here's our
cortiles,
uh, our maximum value, and then of
course the name salary. Uh, so these are
the these are the basic statistics. You
can pull them up and just describe. This
is a dictionary. So I could actually do
something like um in here I could
actually go uh count and run. And now it
just prints the count. Uh so because
this is a dictionary, you can pull any
one of these values out of here. It's
kind of a quick and dirty way to pull
all the different information and then
split it up and depending on what you
need. Now if I just walked in and gave
you this information um in a meeting, at
some point you would just kind of fall
asleep. That's what I would do anyway.
Um, so we want to go ahead and and see
about graphing it here. And we'll go
ahead and put it into a histogram and
plot that graph on it of the salaries.
And let's just go ahead and put that in
here. So we do our map plot inline.
Remember that's a Jupiter's notebook
thing. Uh, a lot of the new version of
the mapplot library does it
automatically, but just in case I always
put it in there. Uh, import mattplot
library piplot as plt. That's my
plotting.
And then we have our data frame. Uh I
don't I guess I really don't need to
respell the data frame. Maybe we could
just remind oursel what's in it. So
we'll go ahead and just uh print
DF. That way we still have it. And then
we have our salary. DF salary
salary.plot history title salary
distribution color gray. Uh plot AXV
line salary the mean value. So, we're
going to take the mean value um color
violet line style dash. This is just all
making it pretty. Uh what color dash
line width of two that kind of thing.
And the median. And let's go ahead and
run this just so you can see what we're
talking about.
And so up here we are taking on our
plot. Um so here's the data. Here's our
our data frame printed out so you can
see it with the salaries. We're looking
at the salary distribution and just look
at this the way they're the salary is
distributed. Um you have our in this
case we did let's see we had red for the
median we have violet
for our average or mean and you can just
see how it really here's our outlier.
Here's our person who makes a lot of
money. Here's the um average and here's
the median. Um, and so as you look at
this, you can say, "Wow." Um, based on
the average, it really doesn't tell you
much about what people are really taking
home. All it does is tell you how much
money is in this, you know, what the
average salary is. So, some of the
things you want to take away in addition
to this is that it's very easy to plot
um an AXV line. These are these up and
down lines for your markers. Um, and as
you display display the data, I mean,
you can add all kinds of things to this
and get really complicated. Keeping it
simple is pretty straightforward. I look
at this and I can see we have a major
outlier out here. We can definitely do a
histogram and stuff like that. Um, but
you know, picture's worth a thousand
words. What you really want to make sure
you take away is that we can do a basic
describe which pulls all this
information out and we can print any of
the individual information from the
describe uh because this is a
dictionary.
And so if we want to go ahead and look
up um the mean value, we can also do
describe mean. So if you're doing a lot
of statistics, uh being able to
doesn't have the print on there, so it's
only going to print um the last one,
which happens to be the mean. Uh you can
very easily reference any one of these.
And then you can also, if you're doing
something a little bit more complicated
and you don't need just the basics, you
can come through and pull any one of the
individual um
references from the from the pandas on
here. So now we've had a chance to
describe our data. Uh let's get into
inferential statistics. Inferial
statistics allows you to make
predictions or inferences from data. And
you can see here we have a nice little
picture movie ratings and um if we took
this group of people and said hey how
many people like the movie dislike it
can't say and then you ask just a random
person who comes out of the movie who
hasn't been in this study uh you can
infer that 55% chance of saying liked
35% chance of saying disliked or a 10 or
11% chance of can't say. So that that's
real basics of what we're talking about
is you're going to infer that the next
person is going to follow these
statistics.
Uh so let's look at point estimation. Uh
it is a process of finding an
approximate value for a population's
parameter like mean or average from
random samples of the population. Let's
take an example of testing vaccines for
COVID 19. Uh vaccines and flu bugs, all
that. It's a pretty big thing of how do
you test these out and make sure they're
going to work on the populace. A group
of people are chosen from the
population. Medical trials are
performed. Results are generalized for
the whole population. So here's a
protected here's our small group up here
where we've selected them. We run
medical trials on them and then the
results work for the population. You
nice diagram with the arrows going back
and forth and the very scary co virus in
the middle of one. And let's take a look
at the applications of inferial
statistics.
Very central is what they call
hypothesis testing uh and the confidence
interval which go with that. And then as
we get into
probability, we get into our binomial
theorem, our normal distribution and
central limit theorem. Hypothesis
testing. Hypothesis testing is used to
measure the plausibility of a hypothesis
assumption by using sample data. Now
when we talk about theorems, theory,
hypothesis,
uh keep in mind that if you are in a
philosophy class, theory is the same as
hypothesis where theorem is a scientific
uh statement that is something that has
been proven although it is always up for
debate because in science we always want
to make sure things are up to debate. So
a hypothesis is the same as a phil
philosophical class calling a theory
where theory in science is not the same.
Theory in science says this has been
well proven. Gravity is a theory. Uh so
if you want to debate the theory of
gravity try jumping up and down. If you
want to have a theory about why the
economy is collap collapsing in your
area that is a philosophical debate.
Very important. I've heard people mix
those up and it is a pet peeve of mine.
When we talk about hypothesis testing,
the steps involved in hypothesis testing
is first we formulate a hypothesis. We
figure out the right test to test our
hypothesis. We execute the test and we
make a decision. And so when you're
talking about hypothesis, you're usually
trying to disprove it. If you can't
disprove it and it works for all the
facts, then you might call that a
theorem at some point. So in a use case,
uh let's consider an example. We have
four students. were given a task to
clean a room every day. Sounds like
working with my kids. They decided to
distribute the job of cleaning the room
among themselves. They did so by making
four chits which has their names on it
and the name that gets picked up has to
do the cleaning for that day. Rob took
the opportunity to make chits and wrote
everyone's name on it. So here's our
four people, Nick, Rob, Imlia, Imlia,
and Summer.
Now Rick, Imlia and Summer are asking us
to decide whether Rob has done some
mischief in preparing the chits i.e
whether Rob has written his name on one
of the chit. For that we will find out
the probability of Rob getting the
cleaning job on first day, second day,
third day and so on till 12 days. The
probability of Rob getting the job
decreases every day. I.e. his turn never
comes up. Then definitely he has done
some mischief while making the chits. So
the probability of Rob not doing work on
day one is uh three out of four. There's
a 75 chance that he didn't do work. Uh
two days 34s * 34s equals.56.
3 days you have 3/4 34 34 which
equals42.
Uh when you get to day 12 it's 0032
which is less than 0.05.
Remember this 0.05 uh that comes up a
lot when we're talking about um certain
values when we're looking at statistics.
Rob is cheating as he wasn't chosen for
12 consecutive days. That's a very high
probability when on day 12 he still
hasn't gotten the job cleaning the room.
So we come up to our important important
terminologies.
We have null hypothesis.
a general statement that states that
there is no relationship between two
measured phenomenon or no assoc
association among the groups.
Alternative hypothesis contrary to the
null hypothesis it states whenever
something is happening a new theory is
preferred instead of an old one. And so
the two hypothesis go hand in hand. Uh
so your null this is always interesting
in in we're talking about data science
and the math behind it. It's about
proving that the things have no
correlation. Null hypothesis says these
two have zero relation to each other.
Where the alternative hypothesis says,
hey, we found a relation. This is what
it is. We have p value. The p value is
the probability of finding the observed
or more extreme results when the null
hypothesis of a study question is true.
And the t value, it is simply the
calculated difference represented in
units of standard error. The greater the
magnitude of t, the greater the evidence
against the null hypothesis. And you can
look at the t value as being specific to
the test you're doing where the p value
is derived from your t value and you're
looking for what they call the 5% or the
0.05
showing that it has a high correlation.
So digging in deeper, let's assume that
a new drug is developed with the goal of
lowering the blood pressure more than
the existing drug. And this is a good
one because uh the null value here isn't
that you don't have any drug. The null
value here is that it's better than the
existing drug. The new drug doesn't
lower the blood pressure more than the
existing drug. Now if we get that uh
that says our null hypothesis is
correct. There is no correlation and the
new drug is not doing its job. The
alternative hypothesis the new drug does
significantly lower the blood pressure
more than the existing drug. Uh, yay, we
got a new drug out there. And that's our
alternative hypothesis or the H1 or HA.
And we look at the p value results from
the evidence like medical trials showing
positive results which will reject the
null hypothesis. And again, they're
looking for um a 0.05 or 5%. And the t
value comparing all the positive test
results and finding means of different
samples in order to test hypothesis. So
this is specific to the test. how uh
what percentage of increase did they
have and this leads us to the confidence
intervals. Uh a confidence interval is a
range of values we are sure our true
values of observations lie in. Let's say
you asked a dog owner around you and
asked them how many cans of food do you
buy for your uh per year for your dog.
Through calculations you got to know
that the on an average around 95% of the
people bought around 200 to 300 cans of
food. Hence we can say that we have a
confidence interval of 230 where 95% of
our values lie in that spread data
spread. Uh and this the graph really
helps a lot. So you can start seeing
what you're looking at here where you
have the 95%. You have your peak in this
case it's a normal distribution. So you
have the nice bell curve equal on both
sides. It's not asymmetrical. And 95% of
all the values lie within a very small
range. And then you have your outliers
the 2.5% going each way.
So we touched upon hypothesis uh and
we're going to move into probability. Uh
so you have your hypothesis. Once you've
generated your hypothesis, we want to
know the probability of something
occurring. Probability is a measure of
the likelihood of an event to occur. Any
event can be predicted with total
certainty and can only be predicted as a
likelihood of its occurrence. So any
event cannot be predicted with total
certainty. It can only be predicted as a
likelihood of its occurrence. Uh score
prediction. how good you're going to do
in whatever sport you're in, weather
prediction, stock prediction, if you've
studied physics and chaos theory, even
the location of the chair you're sitting
on has a probability that it might move
3 ft over. Granted, that probability is
one in like uh I think we calculated as
under one in trillions upon trillions.
So, it's the better the probability, the
more likely it's going to happen. There
are some things that have such a low
probability that we don't see them. So
we talk about a random variable. Uh
random variable is a variable whose
possible values are numerical outcomes
of a random phenomena. So uh we have the
coin toss. How many heads will occur in
the series of 20 coin flips? Probably
you know the on average there are 10,
but you really can't know because it's
very random. How many times a red ball
is picked from a bag of balls if there's
equal number of of red balls and blue
balls and green balls in there. How many
times the sum of digits on two dice uh
result or five each? Um so you know
there's how often you're going to roll
two fives on your pair of dice. So in a
use case uh let's consider the example
of rolling two dice. We have a random
variable outcome equals y. You can take
values 2 3 4 5 6 7 8 9 10 11 12. So we
have a random variable and a combination
of dice and instead of looking at how
many times um both dice were roll five
let's go ahead and look at a total sum
of five and you have in as far as your
random variables you can have a one four
equals 5 4 1 2 3 32 so four of those
roles can be four if you look at all the
different options you have four of those
random rolls can be a five and if we
look at the total number
which happens to be 36 different
options. Uh you can see that we have
four out of 36 chance every time you
roll the dice that you're going to roll
a total of five. You're going to have an
outcome of five. And uh we'll look a
little deeper as to what that means. Uh
but you could think of that at what
point if someone never rolls a five or
they always roll a five, can you say,
"Hey, that person's probably cheating."
uh we'll look a little closer at the
math behind that but let's just consider
this as one of the cases is rolling two
dice and gambling. There's also a
binomial distribution. It is the
probability of getting success or
failure as an outcome in an experiment
or trial that is repeated multiple
times. And the key is is by meaning two
binomial. Uh so passing or failing an
exam, winning or losing a game and
getting either head or tails. So if you
ever see binomial distribution, it's
based on a um true false kind of setup.
You win or lose. Let's consider a uh use
case and let's consider the game of
football between two clubs Barcelona and
Dortmund. The teams will have to play a
total of four matches and we have to
find out the chances of Barcelona
winning the series. So we look at the
total games and we're looking at five
different games or matches. Let's say
that the winning chance for Barcelona is
75% or 75. That means at each game they
have a 75% chance that they're going to
win that game and losing chances are 25%
or 0.25. Clearly 75 plus 0.25 equals 1.
So that accounts for 100% of the game.
Probability for getting K wins in n
matches is calculated.
And we we're talking like so if you have
five games uh and you want to know if I
play um how many wins in those five
games should I get? What's a percentage
on those? And the probability for
getting k wins in n matches is
calculated by px= k= n k p the k q to
the n minus k. Here p is the probability
of success and q is the probability of
failure. And so we can do total games of
n equals 5 where k equals 012345.
P which is the chance of winning is 75.
Q the chance of losing equals 1 minus p
which equals 1 - 0075 which equals 0.25.
The probability that Barcelona will lose
all of the matches can then just plug in
the numbers and we end up with a
09765625.
So very small chance they're going to
lose all their matches.
And we can plug in uh the value for two
matches. Probability that Barcelona will
win at least two matches is 00878. And
of course we can go on to probability
that Barcelona will win three matches
the 26 and of course four matches and so
on. And it's always nice to take this
information um and let let's find the
cumulative discrete probabilities for
each of the outcomes where Barcelona has
won three or more matches x= 3 x= 4 x= 5
and we end up with the p =264 plus 395 +
237 which equals89.
In reality the probability of Barcelona
winning the series is much higher than
75. And it's always nice to uh put out a
nice graph so you can actually see the
number of wins to the probability and
how that pans out with our binomial
case. Continuing in our important
terminology, location, the location of
the center of the graph depends on the
mean value. And uh this is some very
important things. So much of the data we
look at and when you start looking at
probabilities almost always has a
normalized look like the graph in the
middle.
uh but you do have left skewed where the
data is skewed off to the left and you
have more stuff happening off to the
left and you have right skewed data and
so when this comes up and these
probabilities come up where they're
skewed it's really important to take a
closer look at that uh mostly you end up
with a normalized set of data but you
got to also be aware that sometimes it's
a skewed data and then the height height
of the slope inversely depends upon the
standard deviation
so you can see down here the standard
deviation is really large it kind of
squishes it out. And if the standard
deviation is small, then most of your
data is going to hit right there in the
middle. You're going to have a nice
peak. Um, and so being aware of this
that you might have a probability that
fits certain data, but it has a lot of
outliers. So you're if you have a really
high standard deviation, um, if you're
doing stock market analysis,
this means your predictions are probably
not going to make you much money. uh
where if you have a very small
deviation, you might be right on target
and set to become a millionaire. Which
leads us to the zcore. Zcore tells you
how far from the mean a data point is.
It is measured in terms of standard
deviations from the mean. Around 68% of
the results are found between one
standard deviation. Around 95% of the
results are found between two standard
deviations.
And you read the symbols. Of course,
they love to throw some Greek letters in
there. we have mu minus 2 sigma. Mu is
just a quick way. It's that kind of
funky u. It just means the mean. Uh and
then the sigma is the standard
deviation. And that's the o with a
little arrow off to the right or the
little waggly tail going up. The o with
a with a line on it. Uh so mu minus 2
sigma is your uh 95% of the results are
found between two standard deviations.
The central limit theorem. This goes
back to the skew. If you remember, we
were looking at the skew values on this
previous slide. Have left skewed,
normalized, and right skewed. When we're
talking about it being skewed or not
skewed, the distribution of the sample
means will be approximately normally
distributed, evenly distributed, not
skewed. If you take large random samples
from the population with the mean mu and
the standard deviation sigma with
replacement
and you can see here um uh of course we
have our uh mu minus 2 sigma and the
spread down here the mean the median and
the mode and so when you're talking
about very large populations
these numbers should come together and
you shouldn't have a skewed value. If
you do that's a flag that something's
wrong. That's why this is so important
to be aware of what's going on with your
data, where your samples are coming
from, and the math behind it. And if
you're going to do all this, we got to
jump into conditional probability. The
conditional probability of an event A is
a probability that the event will occur
given the knowledge that an event B has
already occurred. And you'll see this as
Baze theorem. B A Y S bay. Uh, and this
is read. I mean, you have these funky
looking little P brackets. A B. This is
the probability of A being true while B
is already true. And you have the
probability of B being true when A is
already true. So, P B of A probability
of A being true divided by the
probability of B being true. And we talk
about BA's theorem which occurred back
in the 1800s when he discovered this.
This is such an important formula and
it's really it's not if you actually do
the math you could just kind of do um um
XY equals J K and then you divide them
out and you're going to see the same
math but it works with probabilities
which makes it really nice. And so if
you have a s you might have uh eight or
nine different studies going on in
different areas different people have
done the studies they brought them
together. Um if we look at today's co
virus the virus spread uh certainly the
studies done in China versus the studies
the way they're done in the US that data
is different in each of those studies
but if you can find a place where it
overlaps where they're studying the same
thing together you can then compute the
changes that you need to make in one
study to make them equal and this is
also true if you have a study of uh um
one group and you want to find out more
about it. So this formula is very
powerful. Uh it really has to do with
the data collection part of the math and
data science and understanding where
your data is coming from and how you're
going to combine different studies in
different groups. And we'll go ahead and
go into a use case. Uh let's find out
the chance of a person getting lung
disease due to smoking. Uh and this is
kind of interesting the way they word
this. Um let's say that according to
medical report provided by the hospital
states that around 10% of all patients
they treated suffered lung lung disease.
Uh so we have kind of a generic medical
report. They further found out uh by a
survey that 15% of the patients that
visit them smoke. So we have 10% that
are lung disease and um 15% of the
patients smoke. And finally, 5% of the
people continued smoke even when they
had lung disease. Uh not the brightest
choice um but you know it is an
addiction so it can be really difficult
to kick. And so we can look at the
probability of a uh prior probability of
10% people having lung disease. And then
probability b probability that a patient
smokes is 15%.
Uh and the probability of B um if B then
A. The probability of a patient smokes
even though they have lung disease is
5%. And probability of A is B.
Probability that the patient will have
lung disease if they smoke. And then
when you put the formulas together, uh
you get a nice solution here. You get
the probability of A of B, probability
that the patient will have lung disease
if they smoke. And you can just plug the
numbers right in and we get a 3.33%
chance. Hence, there is a 3.33% chance
that a person who smokes will get a lung
disease. So, we're going to pull up a
little Python code, always my favorite,
roll up the sleeves. Keep in mind, we're
going to be doing this um kind of like
the backend way so that you can see
what's going on. And then later on we're
going to create um we'll get into
another demo which shows you some of the
tools that are already pre-built for
this. Let's start by creating a set. So
we're going to create a set with curly
braces. This means that our set has um
only unique values. So you have a list
uh you have your tupils which can never
change and then you have um in this case
the the set. So 47, you can't create a
47, 4. It'll delete the four out. So
it's only unique values. And if you use
dictionaries,
quick reminder, this should look
familiar because it is a dictionary uh
where you have a value and that value is
assigned to or that key is assigned to a
value. Uh so you could have a key value
set up as a dictionary. So it's like a
dictionary without the value. It's just
the keys and they all have to be unique.
And if we run this, we have a set of 47.
We can also take a list, a regular um
setup. And I'm going to go ahead and
just throw in another number in here,
four, and run it. Uh, and you can see
here if I take my list 1 2 3 4, and I
convert it to a set, and here it is. My
set from list equals set my list.
The result is 1 2 3 4. So, it just
deletes that last four right out of
there.
And with the sets, you can also go in
there and um print here is my set. My
set uh three is in the set. And then if
you do three in my set,
that's going to be a logic function. Uh
and one in my set, six is not in the
set, and so forth. If we run this,
we get three is in the set true one is
in the set false because 357 is another
one. Six is in the set uh six is not in
the set. So not in my set. You can also
use this with a list. We could have just
used 357 and it would have um the same
response on there is three and usually
you do if three is in but three in my
set is still works on a just a regular
list. And we'll go ahead and do a little
iteration. We're going to do kind of the
dice one. Remember um uh 1 2 3 4 5 6.
And so we're going to bring in an
iteration tool and import product as
product.
And uh I'll show you what that means in
just a second. So we have our two dice.
We have dice A and it's going to be a
set of values. Um they can only have one
value for each one. That's why they put
it in a set. And if you remember from
range, it is up to seven. So this is
going to be 1 2 3 4 5 6. It will not
include the seven. And the same thing
for our dice B.
And then we're going to do is we're
going to create a list which is the
product of A and B. So what's um a + b?
And if we go ahead and run this uh it'll
print that out. And you'll see um in
this case when they say product because
it's an iteration tool,
we're talking about creating a tupole of
the two. So we've now created a tupole
of all possible outcomes of the dice
where dice A is one to three one to six
and dice B is 1 to six. And you can see
one to one, one to two, one to three and
so forth. You remember we had a slide on
this earlier where we talked about um
the different all the different outcomes
of a dice. We can play around with this
a little bit. Uh we can do in dice
equals two divi dice faces 1 2 3 4 5 6.
Uh another way of doing what we did
before and then we can create an event
space where we have a set which is the
product of the dice faces repeat equals
end dice. And we'll go ahead and just
run this. And you can see here it just
again puts it through all the different
possible variables we can have. And then
if we wanted to take the same uh set on
here and print them all out like we had
before uh we can just go through for
outcome and event space. Outcome end
equals. So the event space is creating
a sequence and as you can see here when
we print it out it stacks them versus
going through and putting them in a nice
line.
and we'll go ahead and do something. Um,
let's go print. Since we have the end
printing with a comma, that just means
it's just going to it's not going to hit
the return going down to the next line.
Uh, and we'll go ahead and do the length
of our event space. Uh, that'll be an
important variable we're going to want
to know in a minute.
And of course, if I get carried away
with my typing of length, uh, we'll
print it twice and it'll give me an
error. Uh so we have 36 different
possible variations here
and we might want to calculate something
like um what about the multiple of
three? What if we want to have
uh the probability of the multiple of
three in our setup?
And so uh we can put together the code
for the outcome in event space of xy
equals outcome if x + y
remainder 3. So, we're going to divide
by three and look at the remainder and
it equals zero.
Then it's a favorable outcome and we're
going to pop that outcome on the end
there.
And we'll turn it into a set. So, the
favor outcome equals a set. Not
necessary uh because we know it's not
going to be repeating itself, but just
in case, we'll go ahead and do that.
And if we want to print out the outcome,
we can go ahead and see what that looks
like. And you can see here these are all
uh multiples of three. Uh 1 plus 2 is 3,
5 + 4 is 9, which divided by 3 is 3, and
so forth.
And just like we looked up the length uh
of the one before, let's go ahead and
print the length of our f outcome so we
can see what that looks like.
There we go.
And of course, I did forget to add the
print in the middle because we're
looping through and putting an end on
the on the setup on there. So, we're
going to put the print in there. And if
I run this, you can see um
we end up with 12. So, we have 36 total
options. Uh we have 12 that are multiple
that um add up to a multiple of three.
And we can easily conver compute the
probability of this uh by simply taking
the length of our favorable outcome over
the length of the event space.
And if we print it out, let me put that
in there. Probability
last line. So we just type it in. We end
up with a 3333 chance. And it's roughly
a third.
And we might want to make this look
nice. So let's go ahead and put in
another line there. The probability of
getting the sum which is a multiple of
three is
3333.
We can compute the same thing for five
dice.
And if we do this for five dice and go
ahead and run it, you can see we just
have a huge amount of choices. So it
just goes on and on down here. And we
can look at the uh length of the event
space.
And we have over 7,776
choices. That's a lot of choices.
And if we want to ask the question like
we did above, uh what is the sum where
the sum is a multiple of five but not a
multiple of three? We can go through all
of these different options. And then uh
you can see here uh d1 d2 d3 d4 d5
equals the outcome. And if uh you add
these all together and the
division by five does not have a
remainder of zero but the remainder is
also of a division by three is not equal
to zero. So the multiple of five is
equal to zero but the multiple of three
is not. We can just appin that on here
and then we can look at that uh
favorable outcome. We'll go ahead and
set that and we'll just take a look at
this. What's our length of our favorable
outcome?
It's always good to see what we're
working with. And so we have 94 out of
776.
And then of course we can just do a
simple division to get the probability
on here. What's the probability that
we're going to roll a multiple of five
when you add them together?
but not a multiple of three. And so
we're just going to divide those two
numbers. And you can see here we get
uh.16255
or 11.62%.
And so you can really have a nice visual
that this is not really complicated math
right here on probabilities. uh it's
just how many options do you have and
how many of those are you possibly going
to be able to um come up with with the
solution you're looking for. And this
leads us to a confusion matrix. A
confusion matrix is a table which is
used to describe the performance of a
classification model on a set of test
data for which the true values are
known. And so you'll see on the left we
have the predicted and the actual and we
have a negative uh false negative
positive true positive
um and then we have false positive and
true negative. And you can think of this
as your predicted model. What does that
mean? That means if you divided your
data and you use twothird of it to
create the model, you might then test it
against an actual case for the last
third to see how well it comes out. How
many times was it uh true positive
versus uh false positive? It gave a
false positive response. And you can
imagine in medical uh situations, this
is a pretty big deal. You don't want to
give a false positive. So you might
adjust your model accordingly so you
don't have a false positive. Say with a
co virus test, it'd be better to have a
false negative and then go back and get
retested than to have 30% false
positives where then the test is pretty
much invalid. So in a use case uh like
cancer prediction, let's consider an
example where a cancer prediction model
is put to the test for its accuracy and
precision. Actual result of a person's
medical report is compared with the
prediction made by the machine learning
model. And so you can see here here's
our actual predicted uh whether they
have cancer or not. You know cancer a
big one. You don't want to have a uh
false positive. I mean a false negative.
In other words, you don't want to have
it tell you that you don't have cancer
when you do. So that would be something
you'd really be looking for in this
particular domain. You don't want a
false negative. Uh and this is again,
you know, you've created a model, you
have hundreds of people or thousands of
pieces of data that come in. There's a
real famous case study where they have
the imagery and all the measurements
they take and there's about 36 different
measurements they take. And then if you
run the a basic model, you want to know
just how accurate it is. How many um
negative results do you have that are
either telling people they have cancer
that don't or telling people that don't
have cancer that they do? And then we
can take these numbers and we can feed
them into our accuracy, our precision,
and our recall. Uh so accuracy,
precision, and recall, accuracy metric
to measure how accurately the results
are predicted. And this is your um total
um true where you got the right results.
you add them together, the true
positive, the true negative over all the
results. So what percentage of them were
accurate versus what were wrong. We talk
about precision is a metric to measure
how many of the correctly predicted
cases are actually turned out to be
positive. Uh so we have a precision on
true positive. Again, if you're talking
about like uh COVID testing with the
viruses, uh you really want this to be a
a high number. you want this true um
that to be the center point where you
might have the opposite if you're
dealing with cancer where you want no
false negatives. Uh so this is your
metric on here. Precision is your test
positive uh true positive plus uh false
positive. And then your recall how many
of the actual positive cases we were
able to predict quickly with our model.
Uh so test positive is the test positive
plus the false negative on there. And
we'll want to go ahead and do a demo on
the naive bay classifier. Before I get
too far into uh naive baze classifier
because we're going to pull it from the
sklearn or the scikit. Um let's go ahead
kind of an interesting page here for
classifiers. When you go into the
sklearn kit, there's a lot of ways to do
classification. I'll just zoom up in
here so you can see some of the titles.
Uh there's everything from the nearest
neighbor linear
uh but we're going to be focusing on the
naive bays over here. And this is just
um a sample data set that they put
together. And you can see how some of
these have a very different output. The
naive bay remember is set up as probably
the most simplified uh calculator or um
set of predictions out there. And so
what we've been talking about with the
true false and stuff like that where
there's a uh
an belief that there is a independent
assumption between the features where
the features are very assumed to have
some kind of connection uh then we can
go ahead and use that for the
prediction. And so that's what we're
using as a naive bay classifier versus
many of the other classifiers that are
out there.
For this we're going to use uh the
social network ads. It's a little data
set on here and let me go and just open
that up the file. Uh here we go. It has
user ID, gender, age, estimated salary,
uh purchased. And so we have you can see
the user ID, male 19, uh estimated
salary 19,000 and purchased zero. Uh so
it's either going to make a purchase or
not. So look at that last one. 01. We
should be thinking of binomials. we
should be thinking of simple naive base
classifier kind of setup.
So if we close this out, we're going to
go ahead and import our numpy as np.
We're nice to have a a good visual of
our data. So we'll put in our mattplot
library. Here's our pandas, our data
frame.
Uh and then we're going to go ahead and
import the data set. And the data set's
going to be we're going to read it from
the social network ads.csv. Then we're
going to print the head just so you can
see it again uh even though I showed you
it in the file. And X equals the data
set I location uh two three values and Y
is going to be the four uh column 4. Let
me just run this so it's a little easier
to go over that. Um you can see right
here we're going to be looking at uh 012
is age and estimated salary. So 2 three
and that's what I location just means um
that we're looking at the number versus
a regular location. Uh regular location
you'd actually say age and estimated
salary.
And then column four is did they make a
purchase? They purchased something. Uh
so those are the three columns we're
going to be looking at when we do this.
And we've gone ahead and imported these
and imported the data. So now our data
set is all set with this information in
it.
And we'll need to go ahead and split the
data up. Uh so we need our from the
sklearn model selection we can import
train test split. Uh this does a nice
job. We can set the random state so it
randomly picks the data. And we're just
going to take uh 25% of it is going to
go into the test our x test and our y
test and the 75% will go to x train and
y train. That way once we create our
model, we can then have data to see just
how accurate or how well it has
performed with our um prediction.
The next step in pre-processing our data
is to go ahead and do feature scaling.
Now, a lot of this is start to look
familiar. If you've done a number of the
other modules and setup, you should
start noticing that we bring in our
data. We take a look at what we're
working with. uh we go ahead and split
it up into training and testing. Uh in
this case, we're going to go ahead and
scale it. Scale it means we're putting
it between a value of minus1 and one uh
or someplace in that middle ground
there. This way, if you have any huge
set, you don't have this huge um setup.
If we go back up to here where salary uh
salary is 20,000 versus age 35, well,
there's a good chance with a lot of the
back-end math that 20,000 will skew the
results and the estimated salary will
have a higher impact than the age
instead of balancing them out and
letting the calculations weigh them
properly.
And finally, we get to actually create
our naive bay model.
Um, and then we're going to go ahead and
import the Gazian naive bays.
And the Gazian is is uh the most basic
one. That's what we're looking at now.
It turns out though, if you go to the SK
um learn kit, uh they have a number of
different ones you can pull in there.
There's a um Bernoli. I I've never used
that one. Categorical
um compliment. And here's our Gazian. Uh
so there's a number of different options
you can look at. Gazian when you come to
the naive bays is the most commonly
used. Uh so we're talking about the
naive bays that's usually what people
are talking about when they when they're
pulling this in. And one of the nice
things about the gazian if you go to
their website um to sklearn the naive
bay gazian there's a lot of cool
features. One of them is you can do
partial fit on here. Um that means if
you have a huge amount of data, you
don't have to process it all at on you
once. You can batch it into the Gausian
uh NB model. And there's many other
different things you can do with it as
far as fitting the data and how you um
manipulate it. We're just doing the
basics. So we're going to go ahead and
create our classifier. We're going to
equal the Gausian NB.
And then we're going to do a fit. We're
going to fit our training data and our
training solution. So, X-Rain, Y train,
and we'll go ahead and run this. Uh,
it's going to tell us that it it ran the
code right there.
And now we have our trained classifier
model. So, the next step is we need to
go ahead and run a prediction. We're
going to do our Y predict equals the
classifier.predict
X test. So, here we fit the data and now
we're going to go ahead and predict.
And now we get to our confusion matrix.
Uh so from the sklearn matrix metrics
you can import your confusion matrix
just as saves you from doing all the
simple math. It does it all for you. And
then we'll go ahead and create our
confusion metrics with the y test and
the y predict. So we have our actual and
we have our predicted value.
And you can see from here this is the
chart we looked at. Here's predicted.
So, true positive, false positive, false
negative, true negative.
And if we go ahead and run this, there
we have it. 653725.
And in this particular uh prediction, we
had 65 uh or predicted the truth as far
as a a purchase. They're going to make a
purchase, and we guessed three wrong.
And then we had 25 we predicted would
not purchase, and seven of them did. So,
there's our our confusion matrix.
At this point, if you were uh with your
shareholders or a board meeting, um you
would start to hear some snoozing if
they were looking at the numbers and you
say, "Hey, here's my confusion mat uh
matrix." So, let's go ahead and
visualize the results.
We're going to pull from the map plot
library colors import listed color map.
Um, and this is actually my machine's
going to throw an error because this is
being um because of the way the setup
is. I have a newer version on here than
when they put together the demo. And we
need our um X set and our Y set, which
is our X train and Y train. And then
we'll create our X1, X2. And we'll put
that into a grid. Uh, and we set our X
set minimum stop and our X set max stop.
And if you come all the way over here,
we're going to step 0. 001. This is
going to give us a nice line, uh, is
what that's doing. And then we're going
to plot the contour, uh, plot the x
limit, plot the y limit, and put the
scatter plot in there. And let's go
ahead and run this. Uh, to be honest,
when I'm doing these graphs, there's so
many different ways to do that. There's
so many different ways to put this code
together to show you what we're doing.
it's uh a lot easier to pull up the
graph and then go back up and explain
it. So the first thing we want to note
here when we're looking at the data
is this is the training set.
And so we have those who didn't make a
purchase. We've drawn a nice area for
that that's defined by the naive bay
setup. And then we have those who did
make a purchase, the green. And you can
see that some of the green dots fall
into the red area and some of the red
dots fall into the green. So even our
training set isn't going to be 100%. Uh
we couldn't do that. And so we're
looking at our different data coming
down. Uh we can kind of arrange our x1
x2 so we have a nice plot going on. And
we're going to create the um contour.
That's that nice line that's drawn down
the middle on here with the red green.
Um that's what that's what this is doing
right here with the reshape and notice
that we had to uh do the t if you
remember from numpy um if you did the
numpy module um you end up with pairs
you know x uh x1 x2 x1 x2 next row and
so forth you have to flip it so it's all
one row you have all your x1's and all
your x2s. Um so this what we're kind of
looking for right here on this setup.
Uh, and then the scatter plot is of
course um your scattered data across
there. We're just going through all the
points that puts these nice little dots
onto our setup on here. And we have our
estimated salary and our H. And then of
course the dots are did they make a
purchase or not. And just a quick note,
this is kind of funny. You can see up
here where it says X set Y set equals uh
X train Y train, which seems kind of a
little weird to do. Um, this is because
this is probably originally a
definition. Uh, so it's its own module
that could be called over and over
again. And which is really a good way to
do it because the next thing we're going
to want to do is do the exact same
thing, but we're going to visualize the
test set results. Uh, that way we can
see what happened with our test group,
our 25%.
And you can see down here we have um the
test set. Uh, and it, if you look at the
two graphs next to each other, this one
obviously has um 75% of the data, so
it's going to show a lot more. This is
only 25% of the data. You can see that
there's a number that are kind of on the
edge as to whether they could guess by
age and income they're going to make a
purchase or not. U, but that said, it
still is pretty clear. It's pretty good
as far as how much the estimate is and
how good it does.
Now, graphs are really effective for
showing people what's going on, but you
also need to have the numbers. And so,
we're going to do from sklearn, we're
going to import metrics, and then we're
going to print our metrics
classification port from the Y test and
the Y predict.
And you can see here we have precision
uh precision of zeros is 90. There's our
recall
96. We have an F1 score and a support.
And we have our precision, the recall on
getting it right. Uh, and then we can do
our accuracy, the macro average, and the
weighted average. Uh, so you can see it
pulls in pretty good as far as um how
accurate it is. You could say it's going
to be about 90% is going to guess
correctly um that it that they're not
going to purchase. And we had an 89%
chance that they are going to purchase.
Um, and then the other numbers as you
get down have a little bit different
meaning, but it's pretty straightforward
on here. Here's our accuracy, and here's
our micro average, and the weighted
average, and everything else you might
need. And if you forgot the exact
definition of accuracy, it is the true
positive, true negative over all of the
different setups. Precision is your true
positive over all positives, true and
false. And recall is a true positive
over true positive plus false negative.
And we can just real quick flip back
there so you can see those numbers on
here. Uh here's our precision, here's
our recall, and here's our accuracy on
this.
>> Welcome to this exciting journey into
the world of statistics for data
science. Have you ever wondered how data
transforms from raw numbers into
powerful insights that drive decisions?
Well, statistics is the magic behind it
all. Today, we will uncover how
statistical methods help us summarize
data, model uncertaintity, test
hypothesis, and find relationships that
can predict the future. So, buckle up.
We are about to turn numbers into
knowledge. Without any further ado,
let's get started. Now, to start off,
here's a key question. What are
statistics in data science? Now,
statistics is the science of collecting,
analyzing and interpreting data. By
applying statistical methods, we can
uncover patterns in the data and make
informed decisions. Now, as we continue,
you notice how these essential concepts
will form the backbone of many data
science practices. Let's explore the key
functions of statistics in data science.
First, statistic helps summarize data
using measures like mean, median and
variance. Next, it models uncertaintity
with probability and distributions. So,
we can better understand risk and
variability in our data. It also tests
hypothesis such as when we use AB
testing to compare different outcomes.
Statistics finds relationship through
methods like regression and correlation
revealing how variables impact each
other. And finally, all these tools
enable datadriven decision-m turning raw
numbers into actionable insights. Now,
let's talk about why does statistics
matter in data science. Statistics form
the backbone of data science, providing
the mathematical framework needed to
make sense of data and draw reliable
conclusions. Without statistics, it
would be impossible to turn raw data
into meaningful insights or make
confident evidence-based decisions. Now
let's have a look at the main branches
of statistics. Descriptive statistics
and inferential statistic. So first
let's talk about the definition. Then we
have got methods and measures. Now in
the case of descriptive statistics, now
let's have a look at the main branches
of statistics. So basically there are
two core branches. Descriptive
statistics and inferial statistics.
Descriptive statistics summarizes and
describes data using measures like mean,
median, mode, range and standard
deviation. Its purpose is to organize
and present data typically with charts,
graph or summary tables. The scope of
descriptive statistics is limited to the
sample data itself. Now on the other
hand, inferential statistics makes
inferences about populations based on
samples. It uses methods like hypothesis
testing, confidence intervals and
regression. The purpose here is to draw
conclusion and make predictions with
common examples including AB testing and
survey analysis. The scope of inferial
statistics extends beyond the sample to
the larger population. Understanding
both these branches is essential for
analyzing and interpreting data in any
data science project. Now let's take a
closer look at the descriptive
statistics starting with measures of
central tendency. The first measure here
is mean which is the average of all the
values simply calculated as the sum of
all the data points divided by the total
count. Next is the median which
represents the middle value when the
data is sorted from lowest to highest.
This is especially useful when dealing
with skewed distributions. And finally,
the mode is the most frequently
occurring value in the data set, helping
us identify common patterns or repeated
outcomes. Let's continue our deep dive
into descriptive statistics by looking
at the measures of variability. First,
we've got the range. This is simply the
difference between the maximum and
minimum value in a data set showing us
the spread of our data which measures
the average of the squared differences
from the mean. This tells us how much
the values in our data set differ from
the average. Closely related is the
standard deviation which is the square
root of the variance. It gives us a more
intuitive sense of how much the values
typically deviate from the mean. And
finally, we've got the interquartile
range or we say IQR. This shows us the
range of the middle 50% of our data,
helping us understand how data is
distributed across the center and avoid
the effect of outliers. Understanding
these four measures allows us to
summarize not just the center of our
data, but how spread out and varied our
data set is. Now let's look at some
practical applications of descriptive
statistics. One major use is the data
exploration and summarization where we
quickly get an overview and basic
understanding of complex data sets
helping track performance and detect
problems early in fields like
manufacturing or operations. And
finally, they are central to business
reporting and dashboards where concise
summaries are essential for managers to
review trends and make datadriven
decisions. So in short, descriptive
statistics help transform raw data into
clear actionable information across many
business, scientific and operational
context. Let's explore the first type of
data in statistics, qualitative or
categorical data. This type of data
includes descriptive information that
cannot be measured numerically such as
categories or labels and yes or no
responses. These are all about qualities
or characteristics rather than
quantities. Qualitative data can be
further divided into two types. Nominal
where the categories have no specific
order and ordinal where the categories
do have an order or ranking. Recognizing
and classifying qualitative data is very
important as it affects how information
is analyzed and interpreted in
statistics. Now let's discuss the second
main type of data in statistics which is
quantitative or numerical data. This
type of data includes information that
can be measured and expressed with
numbers making it ideal for mathematical
analysis. Quantitative data is further
divided into two main categories. First,
there is discrete data. These are
countable values with specific fixed
points such as the number of students,
cars sold or website clicks. Second,
there is continuous data which includes
infinite possible values within a given
range. Examples of continuous data
include height, weight, temperature or
time. Understanding the distinction
between discrete and continuous data is
very important as it determines which
statistical methods and visualizations
will be most appropriate. Now let's dive
into the fundamentals of probability.
Probability measures the likelihood of
an event occurring and it's always
expressed as a value between 0 and 1.
Here are some key concepts. If the
probability or P equals to zero, that
means the event will never occur. If P
is equals to 1, the event will always
occur. And if P is equals to 0.5, the
event has an equal chance of occurring
or not occurring. It's truly a 50/50%
scenario. Now, understanding these basic
principle help us quantify uncertaintity
and make informed predictions about
future outcomes. Now let's look at the
different types of probability. First
there's classical probability. This is
based on equally likely outcomes such as
flipping a fair coin or rolling a
balanced die. Next empirical probability
which relies on observed frequency. It's
calculated from actual data such as the
proportion of rainy days over the past
month. Let's say for example the
probability of it's raining today given
that it's cloudy. Understanding these
three type help us choose the right
approach for different situations.
Whether we are predicting outcomes,
analyzing data or making decisions under
uncertaintity. Now probability has a
wide range of powerful applications in
data science. First of all, it is used
in predictive modeling and machine
learning where algorithms estimate
future outcomes based on existing data.
Probability also plays a key role in
risk management and decision making
helping businesses and researchers
evaluate the likelihood of different
scenarios and plan accordingly.
Conditional probability calculating the
chance of one event given that the
another has occurred. This is crucial in
fields like healthcare, fraud detection
and marketing analytics.
And finally, probability is foundational
in AB testing and experimental design,
allowing us to measure the effectiveness
of new strategies or products. These
applications show how probability
enables smarter evidence-driven progress
in modern data science. Let's look at
one of the most common probability
distributions, the normal distribution.
This distribution is famous for its
bell-shaped symmetric curve, which shows
that most values cluster around the mean
with fewer and fewer values appearing as
you move away from the center. Now many
natural phenomena like heights, test
scores and measurement errors tend to
follow this pattern making the normal
distribution a key concept in
statistics. It's characterized by two
main parameters. The mean which
determines the center of the curve and
the standard deviation which controls
its spread.
Recognizing the distribution help
analysts make predictions, calculate
probabilities, and apply statistical
techniques to real world data. Next,
let's explore the binomial distribution,
which is another common probability
distribution. The binomial distribution
is discrete and is used for situations
with binary outcomes like success or
failure. It is based on a fixed number
of trials where each trial has a
constant probability of success such as
flipping a coin a certain number of
times or tracking pass fail rates.
Examples of binomial experiment includes
coin flips and measuring how many
students pass or fail a test. This
distribution help us model and analyze
outcomes when only two possibilities
exist in each trial. Now let's focus on
the poison distribution. Another key
type of probability distribution. The
poion distribution is a discrete
distribution specifically used to model
rare events. It's particularly helpful
for modeling events that occur
independently over a fixed interval of
time or space. For example, it predicts
how many times an event like a customer
arriving on a website, receiving a visit
might happen in a certain period.
Typical examples include customer
arrivals at a store, defect rates, and
manufacturing or counts of website
visits over a set period. This
distribution is great tool for
understanding and predicting random
independent events that don't happen
very often, but are important to track.
It's a branch that allows us to make
inferences, predictions, or
generalizations about a larger
population using sample data. Here are
some key concepts. First is the
population versus sample. The population
represents the entire group we want to
know about while the sample is the
subset we actually collect data from.
Next concept is sampling distribution
which refers to the distribution of a
statistic across multiple samples from
the same population. Standard error
measures how much the sample statistic
is expected to vary due to random
sampling. And finally, margin of error
tells us how much we can expect our
estimates to differ from the true
population value. Together these concept
form the foundation for drawing reliable
insights from sample data in inferential
statistics.
Let's review some of the main techniques
used in inferial statistic. First, we've
got hypothesis testing. This method
allows us to test claims or ideas about
population parameters based on sample
data helping us determine if observed
results are statistically significant
and reliability of our estimate. Next
are confidence intervals. These provide
a range of likely values for a
population parameter giving us a sense
of possible variation reliability of our
estimate. And lastly, regression
analysis is used to model and analyze
the relationships between variables,
allowing us to make predictions, uncover
trends, and understand how changes in
one factor might affect another. Now,
let's walk through the hypothesis
testing process. The first step is to
formulate hypothesis. Start with null
hypothesis represented as Hnot, which
states that there is no effect or
difference. Then there's the alternative
hypothesis represented as H1 which
suggests that an effect or difference
does exist. Now after setting up the
hypothesis the next step is to choose a
significance level noted by alpha.
Common choices for significance levels
include 0.05 5% 01 1% 010 which is 10%.
The significance level is important
because it controls the likelihood of
making a type one error also known as
false positive. And now each step in the
process is critical for ensuring that
statistical results are both meaningful
and reliable. The next step is to
collect and analyze sample data. Ensure
the data is represented by choosing a
good sample. Then calculate the
appropriate and test statistic for your
hypothesis test. Once that's done, it's
time to make a decision. Compare the p
value to the chosen significance level
alpha. Now, if the p value is less than
or equals to alpha, you reject the null
hypothesis. If the p value is greater
than alpha, you fail to reject the null
hypothesis. And finally, interpret your
results. Draw your conclusions with
respect to the context and problem at
hand. Always keeping the bigger picture
in mind. Following these step helps
ensure your hypothesis test is robust,
clear and meaningful. Let's review the
common types of hypothesis test used in
statistic. The one sample test compares
the mean of a sample to a known value
and the one sample zed test is used when
the population standard deviation is
known. Next, we have two sample test.
The independent samples t test compares
to the means of two different groups
while paired samples t test compares
before and after measurements for the
same subject. For categorical data test,
the shear test checks for the
independence or goodness of fit. And
fiser's exact test is useful for small
sample sizes. And lastly, non-parametric
tests such as the man Whitney U test and
Wil Coxson signed rank test serve as
alternatives to the t test when data
doesn't meet certain parametric
assumptions. Choosing the right
hypothesis test depends on your data
type and the specific question you want
to answer. Let's explore the central
limit theorem or CLT of foundation for
inferial statistics. The central limit
theorem states that as the sample size
increases, the distribution of sample
means approaches a normal distribution
even if the original data is a normally
distributed. This sample works for any
population distribution making it
incredibly powerful. Now for good
results, the sample size should
typically be 30 or more. What's
interesting is the sample mean
distribution which will have the same
mean as the population. The standard
error which measures variability by the
sample mean equals sigma / the square
roo of n. And this gets smaller as
the sample size grows. And finally the
formula shown here which is z is equ= to
x -
sigma divided by sigma over the square
root of n. Let's standardize and compare
sample means. The CLT makes most
parametric statistics possible and is
the backbone for many statistical test.
Let's review the main types of
regression analysis which are used to
model and understand relationships
between variables. First up is linear
regression. This technique analyzes the
relationship between a continuous
variable and another variable resulting
in a straight line. It's commonly used
in predicting sales or prices. Next is
logistic regression. Unlike linear
regression, this method is used for
binary outcomes, helping estimate
probabilities such as whether an email
is spam or a patient has a disease.
Moving to multiple regression, this
allows us to account for the effect of
several variables at once, modeling more
complex relationships like determining
house prices. Lastly, polomial
regression which is used for nonlinear
relationships. The resulting curve
rather than a straight line lets us
capture growth trends and other patterns
that aren't linear. Understanding these
types of regression help analysts choose
the right model for the data and
business's problem at hand. Let's
clarify the important differences
between correlation and causation.
Correlation is statistical measure of
how two variables move together.
Correlation values range from minus 10
to + one. But remember correlation does
not imply causation. On the other hand,
causation means that one variable
actually causes changes in another.
Establishing causation requires
controlled experimentation and is much
stronger relationship than simple
correlation. Always be careful when
interpreting results. Just because two
variables move together doesn't mean one
cause the other. Let's look at some
common problems with interpreting
correlation and causation. First is the
third variable problem which happens
when a hidden variable affects both
variables in a question leading to a
misleading connection. There's also this
directionality problem where it's
unclear which variable is causing others
to change but both are actually caused
by a third variable which is the hot
weather. This demonstrate that
correlation does not mean one variable
causes the other. So always look for
hidden factors before assuming
causation. Let's talk about statistical
errors specifically type one and type
two errors. Type one error also called
as false positive occurs when we reject
the true null hypothesis. In other
words, we wrongly conclude there's an
effect when there's actually isn't. The
probability of making this error is
equal to the significance level alpha.
For example, concluding a drug works
when it actually doesn't. On the other
hand, type two error is false negative
means failing to reject a false null
hypothesis. That's when we miss a real
effect or difference. The probability of
type two error is beta. For example,
missing the real effect of a drug and
seeing it doesn't work when it actually
does. Understanding these errors is key
for designing good experiments and
interpreting statistical results
properly. Let's look at three common
sampling methods using statistics. The
first one is random sampling. Here every
individual in the population has a equal
chance of being selected where every nth
individual is chosen. Random sampling is
crucial for ensuring representative
sample. Next is stratified sampling. The
population is divided onto homogeneous
subgroups or strata and then a random
sample is drawn from each group. This
approach makes sure all the groups are
represented in the sample. And finally,
cluster sampling divides the population
into clusters, often based on geography.
Entire clusters are then randomly
selected. It's cost effective and useful
for large spread out populations.
Choosing the right method ensures the
data truly represents the whole
population and strengthens the study's
conclusions. So guys, let's wrap up with
some real world applications of
statistics. In business analytics,
statistics are used for AB testing to
optimize websites, customer segmentation
and targeting, sales forecasting and
demand planning and quality control and
also process improvement. These
techniques help businesses make smarter
datadriven decisions every day.
Statistics play a vital role in
healthcare and medicine as well. They
are key for analyzing clinical trial
results, conducting epidemological
studies, evaluating treatment
effectiveness, and identifying risk
factors. By using these approaches,
healthcare researchers and practitioners
can improve patient outcomes in public
health. From businesses to medicine,
statistics transform raw information
into actionable insights that create
real impact. Statistics has a huge
impact in technology, data science and
finance and power recommener systems
that personalize what users see. In
finance, statistic help with risk
assessment and management, optimizing
investment portfolios, determining
credit scores, and supporting market
research and analysis. Now, these
applications show how statistical
techniques help make smarter decisions
and solve complex challenges across high
impact industry. Are you one of the many
who dreams of becoming a data scientist?
Keep watching this video if you're
passionate about data science because we
will tell you how does it really work
under the hood. Emma is a data
scientist. Let's see how a day in her
life goes while she's working on a data
science project. Well, it is very
important to understand the business
problem first. In our meeting with the
clients, Emma asks relevant questions,
understands and defines objectives for
the problem that needs to be tackled.
She's a curious soul who asks a lot of
wise, one of the many traits of a good
data scientist. Now, she ges up for data
acquisition to gather and scrape data
from multiple sources like web servers,
logs, databases, APIs, and online
repositories. Oh, it seems like finding
the right data takes both time and
effort. After the data is gathered comes
data preparation. This step involves
data cleaning and data transformation.
Data cleaning is the most time consuming
process as it involves handling many
complex scenarios. Here Emma deals with
inconsistent data types, misspelled
attributes, missing values, duplicate
values and whatnot. Then in data
transformation she modifies the data
based on defined mapping rules. In a
project ETL tools like talent and
Informatica are used to perform complex
transformations that helps the team to
understand the data structure better.
Then understanding what you actually can
do with your data is very crucial. For
that Emma does exploratory data analysis
with the help of EDA. She defines and
refineses the selection of feature
variables that will be used in the model
development. But what if Emma skips this
step? She might end up choosing the
wrong variables which will produce an
inaccurate model. Thus, exploratory data
analysis becomes the most important
step. Now, she proceeds to the core
activity of a data science project which
is data modeling. She repetitively
applies diverse machine learning
techniques like KN&N, decision tree,
knives based to the data to identify the
model that best fits the business
requirements. She trains the models on
the training data set and tests them to
select the best performing model. Emma
prefers Python for modeling the data.
However, it can also be done using R and
SAS. Well, the trickiest part is not yet
over. Visualization and communication.
Emma meets the clients again to
communicate the business findings in a
simple and effective manner to convince
the stakeholders. She uses tools like
Tableau, PowerBI and ClickView that can
help her in creating powerful reports
and dashboards. And then finally, she
deploys and maintains the model. She
tests the selected model in a
pre-production environment before
deploying it in the production
environment which is the best practice.
Right? After successfully deploying it,
she uses reports and dashboards to get
realtime analytics. Further, she also
monitors and maintains the project's
performance. Well, that's how Emma
completes the data science project. We
have seen the daily routine of a data
scientist is a whole lot of fun, has a
lot of interesting aspects and comes
with its own share of challenges. Now,
let's see how data science is changing
the world. Data science techniques along
with genomic data provides a deeper
understanding of genetic issues and
reaction to particular drugs and
diseases. Logistic companies like DHL,
FedEx have discovered the best routes to
ship, the best suited time to deliver,
the best mode of transport to choose,
thus leading to cost efficiency. With
data science, it is possible to not only
predict employee attrition, but to also
understand the key variables that
influence employee turnover. Also, the
airline companies can now easily predict
flight delay and notify the passengers
beforehand to enhance their travel
experience. Well, if you're wondering,
there are various roles offered to a
data scientist like data analyst,
machine learning engineer, deep learning
engineer, data engineer, and of course,
data scientist. The median base salaries
of a data scientist can range from
$95,000 to $165,000.
So that was about the data science. Are
you ready to be a data scientist? If
yes, then start today. The world of
data. Picture this. You're shopping
online and suddenly you see a product
that feels like it was made just for
you. How did they know? It's not by
chance. It's data science. Data science
help businesses understand what you
like, predict what you'll need next, and
improve the way we shop and use
technology. And here's the best part.
Data science isn't just about watching
Netflix. It's one of the fastest growing
careers in the world right now. In fact,
the US Bureau of Labor Statistic says
that data science jobs are expected to
grow 36% by 2033, way faster than most
of the other jobs. Companies everywhere
are using data to make smarter
decisions. That means the demand for
data scientists is huge. And let's talk
about the salary. You're probably
wondering how much can I earn actually?
Well, for entry- level position, data
scientists in the US are earning around
$152,000
per year right now. And by 2025, some
can make as much as $230,000.
And in India, starting salaries range
from 50,000 rupees to 1 lakh per month.
And experienced professionals can earn
more than 5 lakh rupees per month.
That's impressive, right? But the best
part is as a data scientist, you won't
just stop here. The skills you develop
in this role like machine learning, data
visualization, and statistics are highly
transferable and crucial for moving into
AI roles. So whether it's becoming an AI
engineer or an AI specialist, the
foundation you build in data science
will help you level up and pursue
exciting hyping opportunities in the AI
field. Now if you're thinking this
sounds great but where do I even start?
Well that's exactly what the
professional certificate course in data
science from IIT Kpur and Simply Learn
is designed to do. It will get you
started and make sure you're ready for
this booming industry. In this 11 month
live online interactive program, you
will learn the skills you need to become
a data science professional. No more
theory, no more fluff. You get hands-on
projects, life classes and mentorship
from IT Kpur faculty plus real world
industry expert who will help you build
your skills. So this course comes with
exciting amazing features that makes
learning even more impactful. Eight
times high interaction in live online
classes with industry expert. Regular
live online classes conducted by
experienced professionals who bring real
world knowledge into every session.
You'll also get to master 13 plus key
skills including generative AI, prompt
engineering, charge, expendable AI,
conversional AI, NLP, and many more.
These are the skills that top companies
use every day. And you'll gain hands-on
experience with 14 plus industry tools
like Python, SQL, Tableau, Dali 2,
Midjourney, TensorFlow, and more. And by
the end of this course, you will be
ready to tackle real world challenges
using these powerful tools and
techniques. But we are not just talking
about textbook and theory. You'll work
on 25 plus real world projects giving
you hands-on experience with the tools
and skills you'll use in the industry.
For example, our first project would be
about sales analysis. You'll use Python
to analyze a clothing company fourth
quarter sales data across Australian
state, helping the company make informed
decisions. The next project would be
about employee performance analysis
where you learn how to build machine
learning models to understand the
factors influencing employees turnover.
Our third project would be about
e-commerce which will help Amazon
improve its recommendation engine to
offer better recommendation to
customers. Upon successful completion,
you'll receive a program certificate
directly issued by the ENI city academy
IT Kpool within 45 days of completing
your cohort. This prestigious
certificate will help you boost your
resume and show potential employees that
you have mastered the skills needed to
succeed. You'll also benefit from master
classes delivered by distinguished IT
Kur faculty who bring their deep
expertise in this course. Plus, you will
be exposed to trending tools like
chargeb2 geni and prompt engineering.
Now, if you're wondering who will be
teaching all of this, then the IT
carpool faculty is here for you. These
experts have been in this field for
years and have worked with the top
companies. You'll also get master
classes from them and they will guide
you through the learning process. You're
not just learning from a textbook. You
are learning from the people who have
been there and done that. And the best
part is once you have learned the
skills, simply learns career assistance
team will help you take to the next step
which will help you to build a killer
resume, show you how to stand out top
recruiters and even give you access to
mock interviews. Plus, you will also
gain access to exclusive networking
events and hackathons to connect with
industry professionals. Along with
Simply Learn's job assistant, you'll
also get access to ID Kpur's career
services helping you connect with top
recruiters and land interviews with
leading tech companies. So, upon
finishing this course, you'll be ready
for top data science roles like data
scientists, machine learning engineer,
AI specialist, and business analyst. The
good news is top companies like Amazon,
EY, Fidelity Investment, Johnson and
Johnson, Borafhone, Accenture, Infosys
and Nvidia is looking for professionals
just like you. And with salaries in the
US hitting around $230,000 plus and in
India reaching up to five lakh per
month, your career outlook looks great.
So what will you be actually learning in
this course? So here's a sneak peek of
the syllabus which you'll be learning in
this course which is the foundation in
Python, SQL and mathematics, core data
science like machine learning, data
visualization, NLP, special topics like
GNI, charge GBT, prompt engineering and
you'll also work on industry projects
which we have already mentioned before.
You'll also have the option to choose
electives like data storytelling with
PowerBI and business analytics with
Excel to tailor your learning experience
and focus on the areas that interest you
the most. So don't wait, hurry up and
enroll now and find the course link in
the description box below and in the pin
comments. Have you ever wondered how
your favorite online store seems to know
exactly what you are looking for? Every
time you browse, add to cart or wish
list an item, you are leaving clues
about your style, favorite colors,
brands, and even shopping times. Data
scientists jump in, analyze these
patterns, and create a super
personalized shopping experience.
Suddenly, the store is showing you just
the right pieces at just the right time.
Almost like it's reading your mind.
That's data science. Turning your clicks
into a shopping spree crafted just for
you. Hello everyone. Welcome back to
Simply Learn's YouTube channel. If
you're already a data science enthusiast
or just got curious about this exciting
field, you're in the right place. Today
in this video, I'm diving into 10
essential steps to help you become the
next in- demand data scientist and land
that dream job. No more waiting. Let's
dive right in and get you on the path to
your future in data science. So, let's
see the 10 essential steps to become the
next data scientist in demand. Step
number one is programming languages.
Starting with Python is a beginner is a
great move because it's simple,
versatile, and widely used in data
science. Python straightforward syntax
makes it beginner friendly, helping you
grasp programming basics quickly and
dive into data science libraries like
pandas, numpy and mattplotive with ease.
Adding R to your skill set is valuable
because it excels at statistical
analysis and data visualization two
essential parts of data science. You can
be comfortable with Python and R within
a month or two. So moving on to the next
step that is version control system.
Learning a version control system like
Git is essential because it allows you
to track, manage, and collaborate and
code effectively. With Git, you can save
different versions of your work, making
it easy to backtrack if something goes
wrong or to experiment without losing
progress. This is especially useful when
working with complex data science
projects where you might try out
different models of analysis techniques.
One or two weeks for practice along with
Python and R is good to get start. Now
moving on to the third step that is data
structures and algorithms. Learning data
structures and algorithms is crucial for
becoming a data scientist because they
provide the foundation for efficient
data handling and problem solving. Data
structures like arrays, stacks, cues and
trees help you store and organize data
in ways that make it easier and faster
to access, process and analyze.
Algorithms on the other hand give you
strategies to perform tasks like
searching, sorting and optimizing data
operations which are essential for
handling large data sets. While many
candidates struggle with the essay,
mastering it gives you an age helping
you stand out in the interviews and
shine as a skilled data scientist
capable of tackling the toughest data
problems. Spend about two months in
this, you will get in the shape for
sure. Now moving on to the step number
four that is SQL. Learning SQL is
essential for data scientists because it
enables you to access, manage, and
manipulate data directly within
databases where most real world data
resides. With SQL, you can create new
tables, alter existing ones, delete
unnecessary records, and run queries to
filter, sort, and aggregate data. These
abilities allow you to retrieve, clean,
and organize data effectively. Core
skills needed for any data science role.
It's easy and you don't have to spend
more than a month to have a deep
understanding of it. Now moving on to
the fifth step that is mathematics and
statistics. Mathematics and statistics
are essential for data science because
they form the backbone of data analysis,
model building and interpretation.
Topics like linear algebra, calculus,
probability and statistics gives data
scientists the tools to understand data
patterns, perform accurate analysis and
make datadriven decisions. Mastering
these areas enables you to build robust
models, validate results and tackle
complex problems confidently making you
a well-rounded and skilled data
scientist. Make sure you spend two
months to grasp this topics. Now moving
on to the step number six that is data
prep-processing and visualization.
Learning data prep-processing and
visualization is essential for a data
scientist because these skills make you
data accurate, insightful and easy to
understand. Python libraries like NumPy
and Panders are crucial for manipulating
and creating data, enabling you to
handle missing values, filter out noise,
and prepare data for analysis. Once the
data is ready, visualization lets you
uncover patterns and communicate results
effectively. Libraries like Mattplot tip
and Seaborn help create clear, impactful
visuals, allowing you to interpret
trends and convey insights in a way
that's easily understood by others.
Together with these tools make data
prep-processing and visualization
fundamentals for effective data science.
If you have a solid foundation on Python
and mathematics, you will get a good
understanding of data prep-processing
and visualization in a month or two. Now
moving on to the seventh step that is
machine learning fundamentals. Machine
learning fundamentals involve
understanding how algorithms enable
computers to learn from data and make
predictions on decisions without
explicit programming. The two main
categories are supervised learning and
unsupervised learning. In supervised
learning, models are trained on labelled
data to make predictions while in
unsupervised learning models find
patterns in unlabelled data. Popular
tools like TensorFlow, PyTorch help
build and train complex models
especially for deep learning. While
skyit learn is essential used for
simpler machine learning algorithms and
data prep-processing. These tools make
it easier to implement machine learning
fundamentals effectively and build
intelligent datadriven decisions.
Dedicate about three months to
understand the core of machine learning.
Now coming to the next step that is deep
learning. Deep learning is a subset of
machine learning that focuses on
algorithms inspired by the structures of
the human brain called neural networks.
Deep learning uses neural networks with
multiple layers often dozens or hundreds
to learn complex patterns from large
data sets. Specialized types like
convolutional neural networks that is
CNN's are great for image processing
while recurrent neural networks RNNs are
used for sequence data like text or time
series. Essential tools like TensorFlow,
PyTorch make building, training and
deploying deep learning models more
accessible, allowing you to create
powerful AI solutions across various
domains. I think it will take about 2
months to have a good hold on deep
learning concepts and how to implement
them. Now moving on to the ninth step
that is specializations. Once you have
grasped the deep learning, it's like
reaching a new level as a data
scientist. Just as doctors specialize in
areas in nephrology and cardiology, data
scientists often choose to specialize in
fields like natural language processing
or computer vision. Natural language
processing focuses on teaching machines
to understand and generate human
language enabling applications like
chatbot, sentiment analysis, and
language transition. It's about making
computers read, write, and even
interpret human emotions through text or
speech. Computer vision on the other
hand is all about enabling machines to
see and interpret images or videos. This
field powers innovations like facial
recognition, object detection and
autonomous driving. Now you don't need
to learn both. You can choose what
interests you the most. Now spend one to
two months diving deep into one of these
areas. Now moving on to the last but not
the least step that is big data. Big
data refers to extremely large volumes
of data generated rapidly from sources
like social media and sensors. For data
scientists, learning to handle big data
is crucial as it requires specialized
tools like Hadoop and Spark to analyze
and extract insights effectively. With
companies relying on datadriven
decisions, big data skills make you a
highly in- demand professional in the
field. Focus for about 2 months and you
will be able to spot trends and patterns
from data sets very easily. Once you're
ready, it's time to build a killer
resume packed with projects that
showcase your new skills. Start applying
to jobs on platforms like Noy and Indate
and supercharge your LinkedIn. Connect
with data scientists. See what skills
they are mastering and learn from their
journeys as well. Keep sharpening your
own skills and when the time comes, you
will be ready to crush those interviews
and land your dream data scientist role
in 2025.
>> Yeah. step by step we will go through
all of this and uh we'll make sure that
we learn everything and we bring
everything together towards the end
right without further ado let me just
straight away deep dive to business
right to learn data science
right
and with this data science there is also
something which is prefixed which is
applied data science
and suffix for this is with Python
right apply data science with Python
right so there are there are two key
concepts which are going to be a part of
this course the first one is the
knowledge about data science that what
data science is and then because we are
doing an applied course right we are
doing an applied course I will try to
tie up these concepts which we will
understand in data science with a tool
right which is Python for you right we
already know about uh 60 65% of Python
right which is the fundamental Python
and now we will be moving to the next
step to advanced Python
right and using Python right leveraging
Python we will be solving a lot of
problems of data science right using
this.
Okay. So, the first few sessions, right?
The first few sessions will be about
making you a breast with Python. What
Python is, right? What how and what
packages do we have? How do they work in
reality, right? And all those things.
And then we will be coupling it up with
data science concepts. And then finally
towards the end of the session in the
last few classes, we will be doing uh we
will be taking a real data set. And on
that data set we will be applying all
these concepts right to understand the
data better and we will be drawing
inferences from that to convert that
into information to take actionable
insights or using that actionable
insights taking a better decision.
Right? We'll do all of that in in the
actual way. Okay. So now guys if you
understand this then the next point of
contention is data sets
right one of the most famous keywords on
the planet right now right one of the
most famous keywords in the planet right
now do you think that these two things
okay let me put it different way what do
you think that can be the possible
explanation about this term data science
you know data You know science what do
you think is going to follow in these
sessions? What is data science to you as
per these two words? Okay. So this is
people made up of two words right data
and science right. So what we are trying
to do is we are
trying to understand data
right? We are trying to understand data
right and then do something to it right
understanding its science understanding
the uh nature the behavior of this data
and converting it in something called as
information
right do we know difference between data
and information data is something which
is completely raw okay it is completely
raw it has no meaning
right it has no meaning isn't it for
example
I give you these stats of some player
like suppose Sid Dhoni I give you stats
of Mahindra Singh Dhoni right that what
what what were his scores uh what is his
name what is his age and you know all
those things now everything is there
right but we don't know what to do about
it right do you think the score of Dhoni
has any context people it has any
context no right but when I deep down
but I when I go and deep dive about it
right what is the first thing you find
out of scores what is the first thing
you find out of score scores you try to
find the average of score isn't it that
in last 10 innings
right in last 10 innings before this
also you need something which is called
as a problem statement isn't it now for
example the problem statement is select
Selectors want to understand selectors
wants to understand that whether
Mahindra Singh Dhoni should be picked
up. So there's a problem now right
selectors want to see that whether Dhoni
is fit for the next tournament or not.
So what we will do we will now try to
take the mean of the scores right for
last 10 innings. And if this score is
suppose X, we will try to compare this
with Y. What is Y? Y is a reference,
right? Y is a reference that we want to
compare it against. Now when you are
doing this comparisons, when you are
applying these techniques to this score,
this is now slowly becoming information,
right? And at the end of the day once
you have the strike rate once you have
the mean score of Dhoni once you have
his age once you have his fitness score
all those things will now help you to
take this particular decision because
now what you have is called as
information right because this has
context
right this has meaning
and this is usually processed
Right? This is usually processed. Right?
This is usually processed. Now what did
we do? Now what did we do here? If you
will go and read about data science,
data science says,
data science says
it is the art of collecting,
right? cleaning,
analyzing,
modeling,
improving,
right? And visualizing,
right? Visualizing
the day, right? If a person is adept in
doing all these things, this person
people is cumulatively called a data
scientist. Right?
That person is called a data scientist.
Right? So before going to the definition
of data scientist, now I will give you
some more examples, right? I'll give you
some more examples. Data science people
as I said is a combination of these
things, right? You have to collect the
data,
right? You have to collect the data,
right? Right. And this has a lot of
things. Data can be connected from two
types in two types. One is primary
and the second one is secondary.
Right? What is the primary way of
collecting data? From your IoT devices,
right? From sensors,
from your inbuilt machines,
right? Then from your surveys
which you float, right? Questionnaires,
right? All these things are primary
ways. What is the secondary way of
collecting data?
Purchasing data,
right? Using internet data
because you have not generated it. You
are just using someone else's data.
Right? Something like uh transfer
learning.
What is transfer learning?
Transfer learning is a technique where
suppose I am bank A and you are bank B
right so bank A has created some model
right trained on their data now you are
going to use the exact same model right
you're going to use the exact same model
right maybe you're not seeing the data
but you are just using the property of
data like mean median mode and a lot of
modeling things which will come we will
learn about them and You use this model
on your particular data right so in a
way you did not have enough data to
create the model yourself but you are
now using someone else's model to run
your data on it right so this is called
as transfer learning so this kind of
collection is basically secondary data
collection so you can collect the data
right then you can perform data analysis
right you can perform data analysis
right How will you perform this data
analysis? Using complex
algorithms,
right? Using complex algorithms, right?
Some statistics,
right? You can use artificial
intelligence,
right? Artificial intelligence. You can
use machine learning.
Right.
Right. You can use all these things for
data analysis. Then you can transform
transform
the patterns
into predictions.
Right? You can transfer these patterns
into predictions, right? Which can be
used for business
decision making,
right? For business decision making.
Then you can validate the results,
right? And present the results,
right? So this is like a complete life
cycle of a data scientist, right? So
before going further, let me give you
what combinations do you need to have to
become a data scientist. The first one
is
domain knowledge,
right? So what is domain knowledge?
First of all, I told you right there
will be a problem, right? You'll be
solving a problem in any project of data
science. You'll be trying to solve a
problem right and the problem will be
belonging to a particular domain even if
you're working for yourself right even
if you're an entrepreneur then also
you'll be solving a problem. So this
domain knowledge part includes things
like understanding
it's a very important diagram
understanding
client requirement
right understanding the client
requirement right
important criterians
right important
criteria knowledge
Right? For example, to give an example,
suppose we have created a machine
learning model. Okay? Understanding the
data, we have created a machine learning
model whose accuracy is 90%. Right? Is
90% a good accuracy?
Yeah, fairly decent accuracy. Yes.
Suppose you have to predict sales,
right? You are selling something.
Suppose you are selling clothes and you
want to predict what will be the sales
for the next week. When you use this
model, whatever the output model gives
you, what is going to be the accuracy of
your output using this model?
How much accuracy?
90%.
But my question is model is 90%. But my
question to you is that if 90% accuracy
is on sales data, a person like me will
be very very happy. Okay? Very very
happy. I'll be probably dancing, right?
But if you try to apply the same model
right for a medical diagnosis case, will
you be interested in getting operated in
such a hospital or an institution where
the accuracy is coming as 90%. Domain
knowledge, right? Domain knowledge. We
need to understand what are the exact
requirements. We need to understand what
are the exact expectations,
right? And we need to know how much do
we need to pivot right? So the first
thing in data science is these
accuracies and everything are subjective
right they are subjective. So for that
you need domain knowledge. Domain
knowledge part very important guys very
important these three things which I'm
going to tell. Second part is people
the game changer right the second part
is
computer science.
Now if I take you back in history okay
if I take you back in history in 1980s
or somewhere then do you okay how many
of you think that data science is a new
concept
how many of you think that data science
is a new concept I hope you all it's not
a new concept everyone knows that yeah
it has been happening for ages just like
you guys will be shocked if you already
don't know AI was coined in the year
1956
1956 6 at the University of Dharma. AI
was coined by Paul McCarthy, right? And
we saw the boom of AI in the year 2010,
right? Such a long journey. Same case
with data science because people back in
the day data science was called as data
mining. Everyone heard about it data
mining, knowledge databases. Yeah, we
need we used to mine the data. Now, what
were the problems? What were the hiccups
of data mining? The hiccups for data
mining was that we were doing everything
everything manually
right now if I give you 100 points can
you calculate the mean
or let's say if I give you two points to
multiply 2 * 3 how much time will you
take?
2 seconds.
Yep. How much time a computer will take?
2 seconds. If I give you to multiply 2
489
multiplied by 200, how much time will
you take to calculate this? Say 5
seconds. How much time computer will
take? 2 seconds. Now if I give you to
multiply 2 48 9 into 15 1 95 4386
how much time will you take to calculate
this manually? Maybe say 1 minute
80 seconds 1 minute. How much time a
computer will take? Still 2 seconds
right? still 2 seconds right and if I
give you to calculate this over 200
times you will take 200 minutes right
using parallel computing computer will
still take about 3 to 5 seconds right so
are you understanding the power of
computer do you understand this concept
in this relationship what was happening
back then what computer science did
people was it revolutionized the way
data mining was happening and That thing
now is called as that again that thing
now is called as data science in which
computer science is one of the most
important contributors. So this is just
one reason. Now manually right manually
if I give you say 1 million rows of data
right 1 million rows of data right so
how many pages will you pages will you
need to store this data suppose your
notebook is like this
right these boxes and here you are
storing the data 1 million so maybe you
can buy n number of notebooks but now do
you think it's as easy in storing
something in computer because back in
the day people we had memory issues,
isn't it? Memory constraints.
So there is something called as Murray's
law, right? Which says as the
advancement in microprocessors will
increase, the price of microprocessor
will decrease, right? So this is what is
happening right now. Back in 1980s, if I
show you right guys, right? Yeah. 2.5
kg. Exactly. Right. It was size of a
fridge hard disk but now it fits in your
palm right. So this was enabled the
storage techniques right the processing
techniques the infrastructure right
things like big data what kind of data
do you think we will be dealing with
people in data science you all know the
term very famous term the kind of data
big data right everyone knows about big
data what is big data yes a data which
is fast right it has velocity veracity
variety right so This kind of data needs
to be stored. This kind of data needs to
be processed. So which thing brought all
these things into data science? It was
given to us by computer science, right?
So computer science people included
things like database management,
right? Data validation, right? Data
infrastructure,
right? Data infrastructure,
right? Then we had languages, computer
languages
which is Python right now for us. Right?
Again do you think when you do this
thing manually right suppose you do this
thing manually how easy do you think it
will become using something like Python
or any other computer language to create
complex models. How easy it will be to
do that to create the complexity in
models right where you can capture the
nonlinear nature isn't it people isn't
it
for example for example let me tell you
this
2 4 6 8 10 dash what do you think is the
next number guys 12 if I tell you to
define this to me in okay let leave
leave
What do you think is going to be the
next number here?
What is the next number? 11.
Next number
25.
Perfect. Now guys, if I ask you to write
these numbers, right, the way you
predicted them, can you give me a
function f ofx is equal to what is the f
of x here?
It's 2x, right? It's 2x. f ofx is equal
to 2x. If I tell you to create a
function, it will be f of x is equal to
2x. What will be the function here,
guys?
f ofx
will be equal to
x + 1. Yeah. x + 1. Yeah. And here f ofx
will be equal to
xยฒ. Yeah. Now the last example, right?
Last example.
What is the next number here? I don't
want the number. I want the function. I
want this so that I can generalize.
Isn't it? How did you reach this figure?
How many of you think it's not possible
to determine this? How many
of you
think
it is not
possible
to determine this?
Yeah. How many of you think what if I
just change this question and ask you
how many of you think it is not possible
to determine this
manually?
Same response. But now I say how many of
you think it is not possible to
determine this
with computers?
Will your answer still remain no? Do you
think I cannot approximate this function
using computers?
We have something called as deep
neural networks
and they are called as universal
function approximators.
Right? So this is the problem people.
This is the problem. Right? I will show
this to you when the time comes. Right?
I will remember this example and I will
show this to you. But now what I'm
trying to tell you is the things which
seemed impossible manually was solved by
what? It was solved by computers. The
distribution of this is like this. Can
you figure it out yourself? No. Right?
We cannot. Isn't it? We cannot do that.
So this kind of approximation will be
given by what? It will be only given by
machines. Right? And this is people what
data science is all about. Right? It is
what computer science did inside data
science. Right? I hope this is clear.
I'm assuming a lot of you will be going
for interviews and everything after
these course. Right? So this will be a
very very important thing for you to
know. Right? Often it is asked why data
science is having computer science in
it. Right? The reason is this. Okay? So
this is the role of computer science
inside data science. Now people the
third thing the third circle which is
one of the most
parts is
maths
and stats right mathematics statistics
which was optimization
right optimization of your models right
design
of model. Right? Now guys, if you look
carefully, if you look carefully, in
order to approximate this, what you what
will you be playing with? You will be
playing with a lot of data. You'll be
playing with a lot of mathematical to
mathematical concepts and statistical
concepts, isn't it? How did you what do
you call this? This is math, right? This
is statistics and mathematics. finding
mean, median, mode, standard deviations,
probability, statistics, all these will
lead to this kind of result, isn't it?
So, this becomes the third wheel of this
particular uh of of this particular
diagram. And this point of intersection,
right? The sweet point of intersection
is basically data science.
Yeah. is particularly data science.
Right? So now people this point okay
this point
is basically representing data
engineering right data engineering right
data engineering is people the part of
data science which enables us to capture
the correct data right how the data will
flow how the data will be stored how the
data will be cleaned right all this is
done by home it is done by data
engineering Just because so am I right
there you'll go for interview right
after this and try to fetch yourself
jobs in this domain data science AI ML
if my understanding is correct is that
the aim yes so now guys there will be
three types of companies
or let's say to simplify let's say two
types one is small and the other one is
big right so in a small organization if
you become a part of small or
organization and you are the data
scientist there You can be involved in
all of these things, right? All of these
things possible, right? Your bosses and
your management will expect you to
construct all these flows, right? Know
computer science, you should know maths
and stats and you should have the domain
knowledge and you will be asked to do
all of this. But if you are going to
become a part of a big organization,
usually all these roles are fragmented.
All of these roles are fragmented,
right? There's a separate data
infrastructure team. There's a separate
data governance team. Now guys, when you
go on to collect the data, can you
collect any sensitive data about is it
possible ethically it's not right? And
legally also it's not right. My question
to you is who will look after this
compliance? Whose responsibility indeed
it is to look after this compliance?
Data scientist. So this is about
fragmentation. If you are part of a big
organization, this thing will be taken
by someone else, right? But if you are a
part of a small organization, you will
be know you'll be expected to do this
all by yourself. In if if you are part
of a big organization then do you think
you need to have this domain knowledge?
The answer is no. Why? Because there
will be separate set of people who are
called as what? Who are called as
business analyst. Have you heard about
this position people? Business analyst.
What is a business analyst role? It is
basically a technical translator, right?
who knows technical, who knows domain
and that person goes and talks to the
client, talks to the client in a layman
language, convert it into technical
requirement coupled with the domain
knowledge and give you the document.
This is basically a medical engineering
problem. So we need this this this this
and you need to fulfill this this this
this criteria. But again, if you're part
of a small organization, who needs to
take care of that? Who needs to make
sure that you know everything about a
domain? You yourself, right? You
yourself, right? Then people, this area
usually represents whom? This area
represents research
and analysis.
Why?
Because they we have people who have
domain knowledge and we have people who
are knowledge of math, stats. Have you
heard about a position called as actury
in the world? Acturial science. Acturies
are people who are basically dealing
with uh domains which are very very
heavily data intrinsic. Right? For
example, finance domain, right? Finance
is all about numbers. So there we go
actal science and we have math, stats,
optimization, model development, all of
those things happening there. We are
also at at this point of time people
belonging to which section research and
data analysis right and then people
there's the third intersection right
there's a third the third intersection
which is this part and this people is
called as machine learning
right machine learning why machine
learning if you can combine the power of
maths stats right and you Combine the
power of computer science, you will find
yourself to be in a position where you
can call yourself a machine learning
engineer. Right? How does a machine
learning engineer becomes a data
scientist? When they couple it up with
the domain expertise, right? So in this
course people in this course we will
teach you computer science. We will
teach you a little bit about math stats.
But what we cannot teach you is domain
knowledge. Yeah.
and making the base of what we are about
to do. Very very important to
understand. Right? If you understand
this then half the battle is won. Right?
So now to answer the question which was
posted earlier uh answer to the question
which was posted earlier. There are
different things in the world of data
science. Right? So you can pick and
choose anything or you can do everything
by yourself. If you guys are engineers
then I think you can be at the sweet
spot going forward in life if you choose
a domain for yourself. For example, you
choose to be a part of automobile
industry, you choose to be a part of
medical industry, you choose to be a
part of say retail industry, you choose
to be a part of aeronautics industry,
right? You choose to be a part of
finance industry, right? So whatever you
will choose, this thing will get
developed over time, right? This is the
most difficult out of these three. I try
to give you the example right my domain
was agricultural industry right the agri
products I have worked extensively in
agricultural industry right so again for
now in my current role this is something
which I don't have I have this expertise
I have this expertise so same will be
with you and you guys will develop this
knowledge over the time now I will try
to give you an example right elections
to make you understand how data science
can be used in one particular use case.
Okay, maybe we can extend that to a lot
of other use cases and examples, right?
Talking about election season, right?
Talking about the election season,
right? We will we will try to understand
how do we use data science because it is
very extensively used in this data
science. Right? Now, let me talk about
the first phase. Let's say this is
pre-election
phase,
right?
Right. This is pre-election phase. In
this pre-election phase, what do you
think will be the tasks with which an
agency like Election Commission of India
will be doing? The first task can be
that they will be
doing the voter
registration,
right? Voter registration and data
management, isn't it?
Yeah, it will start with that.
And what will be the things inside this?
The first thing will be data collection,
right? First thing will be data
collection. So, you will collect the
data from all the registered voters,
right? Maybe it can be their demographic
information, where they live, what is
their age, what is their gender, right?
What is their past polling behavior?
Have they turned out previously or not?
Right? All these details we can collect.
Then can we also do data cleaning?
Because I've told you, right? That there
can be a lot of redundancies. Some
person's name can appear twice, right?
Some people can be a mismatch. Suppose
for example, we have learned this in
Python. Raghav.
Raghav.
Raghav.
Radha. Right.
Right. All these are what people?
This is belonging to the same name.
Right. This is me. But for a computer,
for a computer, how many ragavves are
there? Different ones. All are
different. Right? So example like these,
right? Some people might have died. They
might not be existing anymore. Right? So
all this part will be taken care where
people in the data cleaning process.
Right? Removing the duplicates, updating
the new addresses, right? Correcting the
information about every voter, all those
things, right? And now lastly, we can
also include a flavor of data analytics,
right? What will data analytics include
in this? We can analyze the demographic
data to identify eligible voters. Isn't
it? Yeah. People till now we haven't
understood this why it is not automated.
How will you pick up all the how how
will you pick up all the details and
nuances? Suppose you're filling a form
right by mistake you have. Suppose you
are 25 years of age. Suppose you have
written 250. So does that mean I remove
this? I remove this entry of yours
because by mistake you have written your
age as 250
is age 250 possible in our current world
never right so I will have to tell the
machine that because this is a mistake
please convert it to 25 isn't it this is
called as imputation so this is mostly a
manual task not manual but you have to
understand the problem manually and then
code it on
Suppose someone has written their state
as
E D L H I right and country
as I D I N right so what is this state
referring to what is this country
referring to the humans are very smart
India right we can say this is India and
if this is India then what is this
pointing out to
telly just because you said and it's a
it's a leading question I want you to
explain this tell me is it possible to
do this automatically no right we have
to employ manual rules right we have to
tell because we who is more intelligent
humans or machines humans right machines
are just more optimized right so we know
through human intelligence that this is
pointing to Delhi and this is not EDLHI
so this is data cleaning Right. Lastly,
we have data analytics. So, do you think
people based on these data points, we
can understand that who are the eligible
voters
and maybe who are not registered yet,
maybe who have not voted in the past.
Can we do all those analytics
and reach out to those peoples and
persons? That's the first part. This
just the first part pre-election phase.
Now moving on to the second part right
moving on to the second part let's say
uh we say public
opinion
regarding the polling right the polling
which is about to happen the first thing
will be you want to collect information
about people right so can can you go out
and reach all 1.8 8 billion people in
this country that what is their likable
vote for which party is it possible
1.8 8 billion do you think it's possible
for 1 billion
do you think it's possible for 500
million do you think it's possible for
100 million no right so basically I'm
talking about what I have something
which is called as a population
right and if I have to study about this
population which is 1.8 8 billion people
which is impossible which you just said
what do I need to do should I stop my
process no right I will go and collect
something which is called a sample
right we always work in samples right
suppose someone says that a Coca-Cola
bottle does not contain 500 ml of liquid
which it claims right suppose someone
has put this allegation possible that
Coca-Cola bottles do not have 500 ml
liquid which they promise. Now there are
two ways to deal with this. Right? There
are two ways to deal with this. Either I
go and collect all the bottles of
Coca-Cola in the world. Possible
never right. So what will I do? I will
go and pick up handful of bottles.
Right? Handful of bottles. So what is
that handful of bottles? Those are
called as samples. One last thing.
Suppose someone says that because of an
industry
all the fishes of the lake are dying or
they are infected. Is it possible to go
and collect and check all the fishes in
the pond or a lake? No. Right. What will
we do? We will collect again handful of
fishes and we will test them. Right?
Again samples. Now how does the raw data
collected? Raw data as in I hope you
understand this. This is no more about
population
with this step being told. Now you're
dealing with samples. So now do you want
to ask me how is sample created? Yeah,
now I'm coming to that. Now guys, there
are a lot of second point is how to
sample right? How to sample
right? So we have sampling techniques
people. One is called as probabilistic
and one is called as nonrobabilistic.
Right? I will not go in detail right
now. I just want to tell you an overview
probabilistic is suppose uh you are
manufacturing t-shirts right you are
manufacturing t-shirts right and suppose
you created
100 lots
of
thousand t-shirts right so this is box
one box two box three box four up to up
to 100 right 100 boxes and in each box
how many t-shirts are there 1 th00and
right now suppose you are a Quality
inspector. You're a quality inspector.
Is it possible you for you to go over
all the 100 lots with all checking all
the thousand t-shirts one by one? No.
Right. What will you do? You will sample
again. You will sample. Now the most
common way of sampling these kind of
problems is probabilistic sampling. What
is probability? What is the probability
of getting heads or a tails when you
spin the when you flip the coin? equally
likely 1x2 and 1x2. What is the
probability of getting 1 2 3 4 5 6 on a
roll of a dice? 1x 6. Now what is the
probability of picking any t-shirt from
this first slot out of thousand
t-shirts?
1 by,000.
Yes. So do you think all the t-shirts
have equally probable equal probability
of being picked up without any bias? If
you decide to draw five t-shirts, right,
from each of this box, right, and
suppose say two are defective and three
are not defective, what will you do?
Will you accept the lot or reject the
lot? We have majority of t-shirts of
non-deective
or let's say we will reject the lot. We
will reject the lot. We will reject the
lot. Let's say we will reject the lot.
Okay. Though this is basically
subjective as per the company policies
but let's say we rejected. Now people my
question to you is what if this entire
batch had only two defected t-shirts but
now what will happen? The entire batch
will be rejected.
Yes. Let's say you sampled one t-shirt.
Let's let's change the use case. I say
you sampled only one t-shirt and that
t-shirt was defected. Now you will
reject the batch and that is equally
likely case. So this is called as
probabilistic sampling people and there
is no way you can go back. There is no
way you cannot say that hey sir please
uh allow this batch to pass because
there is a chance that rest of the
t-shirts are not defective. No it is not
the way that happens. It happens
randomly. So right this is called as
random sampling.
Right? random sampling. Now suppose you
are doing a cancer research, right?
You're doing a cancer research, right?
So for your cancer research people, what
kind of people will you need? People who
had had cancer in the past, isn't it? To
know more about their problem, to know
more about their medical condition. So
is it possible people that in this use
case you can go and pick up any person
from the population and ask them
questions? No. Right? That is not
possible. So now is the probability
equally likely or it has changed when
you pick the sample? It has changed. Now
there is a bias which is introduced that
you only want people who had cancer.
Right? So that kind of sampling people
is called as nonprobabilistic sampling.
Right? Non-robabilistic sampling. Clear?
Now sampling technique. Right? Now third
thing in this same scheme can be people
what? It can be the data collection
right data collection mode right that
how do you collect the data? You can
float a survey
on say internet.
You can go and stand outside a mall
or office,
isn't it? How will you how will you can
probably interview someone,
right? Interview someone, right? You can
have a group discussion.
Yeah. All these techniques.
Yes. No, maybe. Right. In the same part,
public opinion polling, right? Now,
guys, uh so this was a brief
introduction, right? And this can be
extended to any industry. Right? As of
now, you can have example in the medical
science,
right? You can have an example in
automobile,
right? You can have an example in
retail,
right? Right. Then you can have example
in manufacturing.
Right? You can have example in
education,
right? You can have example in sports,
right? IPL analysis, cricket analysis,
all these sports analysis, right? These
are the most famous domains, right? They
are the most famous domains. They are
not topics, they are domains, right? In
which data science is used extensively.
So, I'm just going to check. Guys, in
automobiles, there's a biggest example,
self-driving cars.
Yeah, autonomous driving. How do you
think that's possible? Data science
again
like Tesla. Absolutely. Tesla is level
three. We have five levels.
Level three is narrow AI. Level four is
AGI and level five is super AI. We are
going to move first right with the
technical aspect right with the
technical aspect in our data science
course right
right which is based on Python
right because that is our base language
which we have learned so far right so in
Python people we will start and cover
four of the packages
right now and as we move on to machine
learning and other uh deep learning and
everything you will explore more and
more packages. The first package we have
to cover will be numpy.
Right? I'll explain you in detail what
numpy is. Then we will cover pandas.
Then we will cover mattplot lip
and then finally we will cover something
called as cbond.
Right? We will cover something called as
cbond. So these four packages inherently
we have to cover in Python to make sure
we are able to reduce
the time
in coding right we are able to reduce
the time in coding and using these
packages immediately help us in getting
the desired results right I hope you all
remember the concept of modules
we have covered in Python do we all
remember functions and modules. You can
use Jupyter notebook. If your Jupyter
notebook is not installed, you can use
something which is called as Google
Collab, right? Go to Google, type
Collab,
right? Let's say you write Collab,
right? And then you will see this
option. Click on Google Collab and it
will allow you to code in Python, right?
So we are now going to discuss about the
numpy package in python. Okay, numpy
package in python.
So numpy is a
fundamental
package
for data science in Python. Right? It is
one of the most fundamental packages for
practicing data science in Python.
Right? Why is that? Why is so why is
numpy so fundamental? What's is so
special? NumPy package
gives us a new data type
for handling
data in Python
called as
N D arrays, right? ND arrays which
stands for
this stands for
N dimensional
arrays right n dimensional arrays right
this stands for n dimensional arrays
so till now
till now we have studied
about
list integer
tpples,
strings,
right? Out of which
out of which
the data types
such as list
pupils have been used to store data,
right? store data, right?
And range
used for generating
new data which is primarily sequential.
Right?
Now there is one now there is one
problem right? There is one problem and
there should be a question that there
should be a question.
Why do we need a new data type
to work with data science,
right? Why do we need this? Yep. So the
answer to this question people the
answer to this question uh lies in a
small explanation right? Yeah lies in a
small explanation which is that
Python
is a
high
level
language right? Python is a highlevel
language, right? And
a highle language
is usually
very
distant
from
hardware.
A highle language is close to hardware
or distant from the hardware. Did you
not attend the Python programming
essentials?
What is the type of programming language
which is closest to hardware? A
low-level language.
If this is my hardware,
yeah, this is my OS.
This is my application layer. Right? So,
hard level langu language is here and
low-level languages here. Right? Which
is closest. So binary languages,
assembly languages,
right? All these are closest to the
hardware because where is the processing
happening? Where is the processing
happening of the data? At the hardware,
isn't it? Processing of data
is happening
at hardware.
No worry. Actually it is happening at
hardware right? What is processing?
Processing is signals of zeros and ones
right? What are zeros and ones? These
are electric signals.
These are the electric signals right?
Which is basically on and off. Right?
And it is communicated to the hardware
through the help of resistors
and microprocessors.
Right? Isn't it right? Why a computer
only knows zeros and ones? Because zero
is off and one is on which is the
electric current
right electric current to activate or
deactivate certain things right true and
false gates right so it is happening at
hardware so now people if you understand
this part then try to logically connect
it to what I'm going to say when you are
studying data science
what kind of data you'll be dealing with
big data,
right? And as the name suggests, it will
have a lot of volume,
right? It will have a lot of volume
other than a lot of other things, right?
It will be very very big. And on this
large volume of data, you'll be doing
processing.
You'll be doing processing. What does
processing means?
What does processing means? Operations.
So where is this operation happening?
This is happening in hardware
and for hardware which is the closest
language to hardware a low-level
language.
But now people but now we have a
situation in front of us. What is the
situation that what are we trying to do
data science with? What are we trying to
do data science with?
Python, right?
And Python is what?
A high level language,
isn't it? Yeah. So, there's a
discrepancy. Yeah. A big one
because hard level, high level language,
these
are slow
in processing,
right? These are very slow in
processing, right? So for these kind of
languages to handle this kind of data
and these kind of operations yeah is
very difficult right so let's let's keep
let's keep this part aside if you
understand this now let's go to the
second point right
when you learned Python
on a scale of 1 to 10 how easy was it
the ease of use of Python
It's relatively a higher number, right?
Relatively a higher number. So now guys,
when I talk about data science, right?
When I talk about
data science, okay?
Right? When I talk about data science,
my thing is that this will be used by
masses,
right? Will be used by masses. A lot of
people managers, programmers, business
analysts, data analysts, possible right
who are from nontechnical background who
don't know coding they also can do data
science because data science is a
general thing isn't it? Understanding
the data it should not be limited by
your capability to understand the uh
technicalities of a very complex
language. So for these people which
language is suitable which is Python
right? It is easiest to understand. It's
a high level language almost like
English. Yeah. So, Python is a simple
language. So, in this part people,
Python fits the bill, right? Which is a
bigger thing. In the second part, when
we talk about the operations,
we talk about the operations. In this
part, there is a problem, right? Python
fails,
right? Python fails, right? because it's
a highle language. We said that okay
there is a language called as C right
which is a middle level language
right and C language people is used to
create OS operating systems. It is used
to create networks
networking applications.
It is used to create games,
right? All the things which are close to
hardware,
C is used, right? So C fits this bill,
right? C fits this bill.
Python
said that okay, if C fits the bill and
Python is written
in C, written on C, right? It's written
on top of C language. Now what happened
was Python said okay if my intrinsic
data types my intrinsic processing is
not suitable for data science but my
highlevel nature is let's do one thing
let's take C language and use its power
right use its power that it is very
close to hardware and let's create a new
data type right let's create a new data
type which is written on top of C and
which can integrate with Python
seamlessly. And people this new data
type was called as array
and this was given to you by something
called as num py package. Right? Nump py
package. It defined a new data type
which was array. And along with defining
the array, it gave various operations.
Right? It gave various operations
one could
perform
on arrays,
right? One could perform on arrays.
It's simple, right? program the the
power of processing lied with C right it
was lying with C. So we developed a new
data type using C on top of Python and
that new data type was called as array
and this array was defined in a new
module which was called as numpy module
which told you how to create the arrays
and then how to manipulate those arrays
for doing data science. We had options
like Java, we had options like C, right?
We had options like forotron to be used
for data science but we chose Python
because of its simplicity and the simple
syntaxes that people from
non-programming background could also
use Python to do data science.
Right now the limitation was that
because it is slow because of being high
level we needed something which could
make it fast and that was using an
external data type which is not internal
to Python and that was array and this
array is defined inside a new uh module
or a library called as num py which is
numerical python right numerical python.
Back to the programming right where we
have understood right about this
question right.
So
numpy
essentially is
built on top of
C
language which is
compatible
with Python.
It leverages the power of closeness of C
with hardware,
right? C with hardware
which eventually
makes the processing
faster in
Python. Right?
So this is what the first part is right
this is what a first part is right now
the second thing is people so this is
about the performance bit right these
are the performance bit the above
is about the performance
of
right now coming to the next part which
is the memory efficiency Y right memory
efficiency right so
arrays created
by nump py in
python
are less memory
exhaustive
than lists in Python right and I will
prove these points later on to you
through code, right?
In list, right? In list,
each item is an object, right? Each item
is an object, right? I hope you remember
this, guys. Each item is an object,
right?
And it holds, right? It holds
meta information
like
references
and types,
right? Etc., right? A lot of information
it holds, right? This makes
list consume
more memory, right? But but
arrays in
numpy
are contiguous
which means
that they do not
create objects
but rather
directly store
the data
in
continuous
memory
blocks
one after another. Right? Also the
arrays are homogeneous in nature. Right?
You can only store the homogeneous data
in array unlike lists. In list you could
store different different data types.
Right? But in arrays you cannot right.
you have to store the same kind of data
in the array homogeneous. So it's a
contiguous memory block which is meaning
that you can store data in continuity
right one after another in the memory
block. So the access is faster the
memory location and the memory
efficiency is very higher right as
compared to the native data type like
lists or tpples in Python. Arrays
are basically
vectorzed operations
right I'll talk to about talk about to
you with vectors what are vectors right
they are the vectorzed operations and
they are way more convenient
to deal with as compared to list
in Python right in list you have to go
through a lot of loops, right? We saw
that we have to go through the list
comprehension. But you will see in
Python pandas, sorry, in Python numpy,
the vectorzed operations are very very
simple, right? They are very very
simple, right? Again, for all this, I
will give you examples, but uh it will
take some time because you have to
understand what arrays are first, right?
Okay. Then
the most important
other packages
which are pandas,
mattplot lib,
cb bond,
sklearn,
cypy
all are written
on top
of
numpy
package.
That's the reason for this reason
it's called as
fundamental
package.
All the other packages which make your
life easier as a data scientist where
you don't have to worry about code.
All these packages
make our data science
journey smooth
because we have to worry
less about code and more about
logic.
Yes. So for these for the understanding
of these packages it's very important
that we understand numpy first and then
we move forward right
now
then let's get started. So the first
step right the first step which you have
to uh see right the first step which you
have to see while using numpy package is
basically from where will you import the
package right from where will you import
the package.
So to import the package num py we write
import nump py as np right where np
is an alias right it's an alias
right so please do this import nump py
as np
right if you don't get any result for
this then you can write pip install
nump py right pip install nump py and
just execute this right when you will
execute this it will give you this
message or it will give you it will
download this package for you right pip
stands for
python
index package
right
it is basically like play store,
app store,
right? Or Windows store
for Python,
right?
So, you're just going to these Play
Store,
Windows Store, App Store of your Python
and asking them to download this for
you, right? It is also called as the
package manager
right pip
if it is done you can also check the
version you can say np dot
version
and it will give you the version of
numpy
right
numpy is an opensource
package,
right? Yeah, that's about it. Yeah, it's
an open-source package, people,
right? Open source package. And if you
want to see the code, you can go to
GitHub.
Go to Google, write the uh code for
numpy. It will show you it on GitHub.
Right. done. So let's say I okay so I
say
we have when we so okay before this let
me come to a little bit of theory before
I do this with you. So now guys I said
that num py
has arrays
as
data type
right and this is written on C which
runs directly
on hardware
also this numpy array
is basically basically a vector,
right? It's basically a vector. Now,
what is a vector, people? What is a
vector? A vector is a quantity which has
sign
plus magnitude,
right? It has a sign and magnitude. If I
say this is a cartition space, this is
I, this is J. And I say this this is 3 I
and 4 J right 3 I cap 4 Jcap. So this is
a vector right and this is the direction
guys. If you have studied elementary
maths you would know this.
Yes people this is a vector. If I draw
another like this
then this is another vector. So I will
call this say
uh 2 I and 5 J. Yeah, this is another
vector and this is the theta right. This
is the direction.
This is the direction. And what is the
magnitude?
3 I 3ยฒ + 4ยฒ which is 9 + 16 which is 25
under root which is 5. So the magnitude
of vector is five and direction is equal
to theta. This is a vector quantity.
Now what is this? What is this? This is
a scalar.
This is a scalar. Yeah. Only magnitude
isn't it? This is scalar only magnitude.
And then I have 2a 3. Right? This is
what? This is a vector.
It has two dimensions
or one dimension. Only one dimension,
right?
This is one dimension vector.
Yep. One dimension vector. Now if I say
this 2 3 4 5, what is this called? This
is called a matrix,
right? Right? This is called a matrix
which is what collection of vectors
and this collection of m vectors matrix
is called as two-dimensional. Right? It
is called as two-dimensional.
Now if you have this
so these are stacked behind each other.
This is one. This is two. So this is
three right? So we have three layers
in matrix.
So how many dimensions will be this
people?
One dimension, two dimension and three
dimension. This is a threedimension
matrix,
right? Or threedimension vector
or three-dimension array.
I'll repeat once again. What is a single
value? A single value is called as a
scalar. Right? It only has magnitude.
Now when you have multiple values,
right? This is called as a vector. It
has a direction. It has a magnitude. And
this is single dimension. Now multiple
vectors right like this or maybe you can
say like this, right? Are you
understanding why I'm calling it one
dimension? It can be either this
dimension or it can be this dimension.
In any dimension you stack two vectors,
you will get yourself a matrix. Right?
You will get yourself a matrix which is
now two dimensions. It has rows and it
has columns.
Right?
Now if you stack multiple such matrix
one after another, right? This becomes a
threedimensional matrix. And this can go
up to how many dimensions people? How
many dimensions this can go up to? It
can go up to this
can
go up to n dimensions,
right? N dimensions,
right? So now do we get it? Why do we
call it ND
arrays?
Yeah, n dimension arrays,
right? That why are we calling something
what?
Right. N dimension arrays. We are only
capable of viewing three dimensions
people. It can go up to 100 dimensions,
500 dimensions, 1,000 dimensions, any
dimensions.
Scalas are least important.
vectors are more important. So now we
have
multiple
dimensions in
arrays
namely 0D,
1D,
2D, 3D and so on till the N D, right? NV
arrays. Right now let's create our first
array. Right? Let's create our first
array. And this will be a zero
dimension
array, right? Zero dimension array. How
will you create this? You will say a r0
is equal to np dot array and you will
mention what a scalar value is. What is
a scalar value? It is a simple value. I
say two, right? np dot array equal to
two. And when you will now check or
print the type of ar r0, it will tell
you class num py nd array. Right? It is
a zero dimension array. And if you want
to check
the dimension,
you just have to write a r0 dot end div,
right? And it shows you that there is
zero dimensions present. So what is this
in short? This is a scalar. I have
entered a single value people 2 200 500
whatever you want to enter. And the
syntax is np dot array np dot array.
You're instructing nump py package to
fetch the function array on method array
and convert this into that particular
data type right array zero
and when you check the type type is nd
array but what is the dimension of this
nd array this is zero which is nothing
but a scalar right we have created this
this is what we have created
Right
now guys,
if I created this list, right?
Say Lis is equal to
Yeah. So this was this was a list,
right? This was a list and this was this
is what people what is this that we have
just studied according to that? What
dimension is this list?
So what I'm trying to do is I will
create a
one deal
array with a or let's say from a list
right let's let me show you how do we do
that okay
so I say
hurry read the error name ar r not
defined. Why? Because you have created
array from a a r0. Come on hurry.
Right.
I will create a onedimensional array.
How will I do that people? I will say a
ar r1 is equal to np dot array. And can
I pass 1D list inside this? Or can I say
I can pass lis inside this? When I do
this people now see what will happen. A
R R1 will be equal to this right and if
you say let me say print a ar a ar a ar
a ar a ar a ar a ar a ar a ar a ar a ar
r r r r r r r r r r r r r r r r r r r r1
it will be like this okay this is your
a ar r r1 dot nim
you will see that it gives you one right
this is a onedimension array right on
dimension array
Yes. So for example
when I say scalar right when I say
scalar I say
35s right? when I say vector
1D I say
uh
so these are suppose my marks
right now I say 35
40
50 right so now what are these my marks
in three subjects
yeah marks in three subjects this is my
say Hindi this is English and this is
maths right or let's say science because
not everyone has Hindi science English
and maths right guys now if I have to
create a matrix
of two dimension what will I write so
that means this is one student this is
one student people isn't it guys yes no
maybe so now in matrix we will have what
we will have multiple students
Yes. No. Maybe in a matrix people we
will have multiple students. Suppose
this was S1. Now you will have S_sub_1,
S_UB_2, S3, S4 like this.
And each student will have their own
individual list of marks. So can I say
that I'm making a nested list?
Can I say that people? I'm making a
nested list. So now let's make it okay
from a
nested list. Okay, a nested list.
So I'll say ar r2 is equal to np array
list. I will have to create a list
first. Lis2 is equal to. So this is my
first bracket. What is this bracket
representing? this bigger bracket. Now I
will put another bracket inside this and
I will write 1 1 22 33 3. I'll put a
comma again. Write a comma. Then I will
say
4455 666 comma 778899
right I'll do this right now I will say
list to
two
when you will do this you will see that
an array like this has been created
Right?
like this AR R2
right this has been created right when
we check the dimension it is two
dimension array right
yeah marks of three different students
in three different subjects
and this is the same technique you can
create a three-dimension array how will
you create a three-dimension array
people
if I go here how will you create a
threedimension array
Now suppose I have data in this and this
is my master list. Okay, this is my
master list. In this I have data
and this can be represented like this.
This is my
first matrix isn't it? And this is the
vector inside this
V_sub_1, V_sub_2, V3.
Then this can be called as M1. And now
to create a three-dimension setup, how
many M1s do you need? You need multiple
M1s, isn't it? You need M1, M2, M3,
multiple matrix like this people.
like this matrix 1, matrix 2, matrix 3.
So what will you do? You will have you
will have what people?
You will have another yellow, right?
And you will have inside this yellow
multiple purples.
Isn't it
right? Again like this.
Yeah. Like this you will have it people.
So can I say people can I say that as I
am increasing the dimensions as I am
increasing
the dimensions
I am putting 1D sorry 0D right and okay
again in this V_sub_1 in this V_sub1 do
you think you will have multiple scalers
people can I say that
can I say multiple scalers create a
vector multiple vectors create a matrix
And multiple matrix create one
three-dimensional matrix.
Can I say that? Let me talk to you about
an image. Right?
Image, right? What is an image made up
of people?
What is an image made up of?
H
pixels.
Yes or no? No,
not frames. Frames is basically videos.
Pixels are creating an image, right? So,
we have how many pixels here? 1 2 1 2 3
4 5 6 7 8 9 10 11 12 13 14 15 16. Right?
Suppose
this is one vector, right? This is one
vector and this is one scalar
right? Scalar 1, scalar 2, scalar 3,
scalar 4. And this will make vector
v_sub1.
This is v_sub_2, v_ub3, v4. And together
together can I call this m_sub_1 and
call this blue?
So for a colored image, how many
channels are there people? How many
channels are there? What do we call it?
We call it the image as
RGB
RGB image
that means red,
green,
blue.
So this is blue part. So similarly you
will have a red part in front of it.
Then you will have a green part and then
finally you will have a blue part. Yeah.
Are you understanding guys? Why do we
require three-dimensional arrays?
Yes. Suppose now you want to make a
change at this this pixel. So you will
go to the third layer which is the blue
layer. Then you will go to the third
column. You will go to the third column
and third row. And this is how you will
reach this pixel. Everyone? Yes. No.
Maybe.
Yes. So this is the reason why we need
to create a 3D array.
Right. To ingest information like this,
right? To ingest information like this.
So we can create a 3D array also. Right.
Right. And how did I tell you? How many
brackets will I have? Squared brackets.
I will have three squared brackets. So
now this is just one student.
Right. Now I'll put a comma here.
Right? Right, I'll put a comma here and
I will start.
Right, I'll do this. I'll say this list
three
a r3
list three ar r r3
a r3
r
and now you will see that there's a
threedimensional array
right guys
right this is a threedimensional
array
Yep. Moving on. There are multiple ways
create arrays, right? The first one we
have done.
So we have done
from lists.
From list we have done.
Then second will be
from uh we can create a zero array
right we can create
on's array
right then we can create custom array
right I will show this all to you
right I'll show this all to you so let's
start with the on's array sorry zero
those array
right what do you have to do you have to
write so let's say zero
ar r r0 okay a zero dimension zero array
so you say np dot zeros
right and you create
a two
right
a two right and if I say
uh 0
dot end
right you will see it is a onedimension
array
right it is a onedimension array
right so two is by default taken as so
let me just show this to you
it will look like this right it is taken
as horizontal what is the dimension of
this guys a vector what is the dimension
of this vector
no no it's 1. It is basically 2 + 1,
right? 2a 1.
Sorry, 1 comma 2. My bad.
1 comma 2. Isn't it? Now, what if you
had to create a 2 + 1? What? What if you
had to create a 2 + 1? Right? So, let me
just show that to you.
So, I say wait
like this. Okay.
Now when I do this, I say a ar r r once
and I say here
2, 1, right? I say 2a 1. Now you will
see people what will happen to this.
Now
because you have created right specified
two dimensions, right? What will this be
converted to now?
H what will this be converted to? This
will be converted to
a
two-dimension vector. By default, it was
this, right? Which was this vector,
right? By default, it was this vector,
right? What is the dimension of this
vector? 1 + how many values you put
here, right? 1 + 2 like this. So we call
this only one dimension. We call this
only one dimension. But now when I will
run this one, you will see yes 2 + 1.
And now you will see this is
two dimension, right? You see this is
two dimension. Now clear people just the
orientation has changed. But now you see
the brackets there are two brackets now
because what have you now instructed?
You have now instructed Python and
rather numpy to create the vector as 2 +
1. So you have said give me this zero
and give me this zero here. So the
moment you do this you are now
specifying the rows and columns.
Now this 21 right let's say this is 21.
Now let me create a another two cross
two dimension matrix for you. Let me
call this 10
comma 10. Right? So how many rows and
columns will it have people?
How many rows and columns will it have?
I will say 21.
It will have 10 rows and 10 columns like
this. You saw this? Yeah. 10 rows and 10
columns. Right?
By default, it is 1 +2 like this. It is
a one-dimensional vector and this is a
two-dimensional vector. What about a
three-dimensional vector? 0 dot 0
uh a ar r3
is equal to np dot zeros, right? np
do.zer,
right?
H should I just write 3a 3a 3? And I
should check for this.
Yeah, you will have a 3 +3 vector 3 + 3
matrix with three matrix stacked behind
each other. Right? If I say four, this
will be four. So how do we read this?
number of layers,
number of rows and number of columns.
So, can you help me with a syntax? Can
you help me with a syntax which can
create me a 3D matrix of five layers,
three rows and three columns? What will
I write?
Five layers, three rows and three
columns. What will I write?
5 33. Yeah, you'll get five layers. 1 2
3 4 5 right like this. Suppose suppose
you have to represent
an image
with
with say
RGB channel
and
uh say 256 and 256 pixels. How will you
create this? So I'll say img right image
is equal to np dot
zeros and I will say inside this
three channel 256 cross 256
right and when you will just run this
image it will be like this right it will
be like this this is one channel this is
two channel and this is three channel
And uh we understood that numpy is a
fundamental package for data science in
Python. Right? This package gives us a
new data type to work with which is
called as n- dimensional arrays. Right?
Now what are arrays? What are arrays?
Arrays are nothing but vectors right
which are stored in a contiguous block
of memory which means they are stored
continuously one after another and they
do not get converted into the object
unlike the list and they are way faster
they are more memory efficient than list
I will prove this fact to you today with
the help of example through the help of
code right so why numpy arrays because
numpy is a package which is built on top
of C language which is compatible with
python and C being a middle level
language interacts directly with the
hardware. So whatever operation you run
in a fact that is getting directly
executed on the hardware itself. Right?
That is the reason why the uh
performance is way better when we try to
use the n dimensional arrays. Along with
this these syntaxes the type of syntaxes
we used to type uh in list I will show
that to you also today with the help of
example are way simpler when you try to
do them with nd arrays right so arrays
are basically vectorized operations
right and they're also very convenient
to deal with as compared to lists and
other native data types in python right
and the most important part is that in
our data science journey whatever other
packages packages we will use. Right?
Again, what are packages? They have
predefined things stored for you so that
you can leverage them and focus less on
code and more on logic. Right? You need
to be aware about the logic more than
the knowledge of the code. Right? So, we
have packages for that which contains
methods inside them which you can use
and uh without any further calculations
you can work with them directly. Right?
So that is how we started with numpy
right and the syntax to import numpy was
import numpy as np where np was an alias
right. Uh you could do pip install numpy
if someone did not have access to numpy
if numpy was not coming by default. You
can use pip install numpy which is
python index package right and uh this
is like play store app store window for
python. All the packages are stored in
pip and you can call pip you can ask pip
to download that package for you so that
you can use that right it's basically
the package manager so numpy is an open
source and if you want to see the code
you can go to github and check the code
out for yourself right in numpy we have
n dimensional arrays now the question
arises people that why do we need numpy
right so my answer to this particular
question is that when you will deal in
data science right when you will become
a data scientist you will be dealing
with data Right now my question is how
will you ingest how will you make the
machine ingest the data right there has
to be a way right for you to input the
data to the machine right to make
manipulations to the data to read the
data so all this is started as the base
package of numpy arrays right other than
that it becomes very difficult and
cumbersome for us to deal with that and
this provides us a lot of ease and
flexibility to deal with such massive
amounts of data which you will along
with me as we will move forward in this
particular course. Yeah, perfect. Now
people, let me just uh pull up the PBTs.
This is what we're discussing people. We
start with something which is called as
scalar, right? We start with something
which is called as scalar which is a
quantity which only has magnitude.
Right? In numpy language, this is also
called as 0D, right? It is called as
zero dimension. Then people we have
vector right and vector has a constant
dimension. It has only one dimension
right you can interpret it as a row or
you can interpret it as a column it
doesn't really matter because this is
only one single dimension right so
usually it will be written as five comma
blank right there will be nothing
written in front of it so this in numpy
terminology and nomenclature is called
as one dimension right when you move on
then you get combine couple of vectors
you get a shape and now that is called
as a matrix and In numpy terminology it
is called as two-dimension right and
when you try to stack multiple
two-dimension matrices one before with
one after each other or one before each
other then they become something called
as three-dimensional and now in my
capability I don't know what four
dimension looks like but there is a high
possibility that you have n dimensional
data right it has it has n we are
dealing with n dimension data right
suppose with this I also add time right
at t equal to 1 at t=2 that will serve
as the fourth dimension for this data
but how do how does it look like I don't
really know that right because humans
are only capable of visualizing 3D three
dimensions at max right so you can go to
n dimensions and hence the name n
dimensional arrays right nd arrays post
this right post this we moved on to
create certain things and I tried to
explain you the data right so the data
will look to you like this right You
might have a scalar quantity which is
marks right one marks right now if I go
on to vectors in one day it can be marks
of one student
right marks of one student in science
English
and maths right 35
40 and 50 out of say 50 right three
subjects so this will be characterized
as a vector right what will be the
dimension written for For this it will
be 3 comma nothing. This will be zero
right shape will be zero. For this
matrix suppose we have 1 2 3 four
students and each student will have
three marks.
Right? Each student will have three
marks. So what will be the shape of
this? We have four rows and three
columns. Right? So this will be the
shape right of this 2D matrix right and
now if you stack images right one behind
each other then it will be like this
right image I gave you an example so
this has 4 + 4 pixels so the shape will
be 3 + 4 + 4 right this will be the
shape for this particular 3D matrix
right guys so this is how you input the
data just to tell you a little bit more
since generative AI is very popular
these days. Right? So what if I tell you
the fact that the Chad GPT
which you use or you might have used
right has
never seen
a single
word
in its lifetime.
Right? All it sees
is numbers, right? Only numbers. How do
we see numbers? Suppose I say my
name
is
Raghav.
Raghav
is a
nice
name. Right? So these are two data,
right? These are two data points. Now we
all know that computers do not
understand these right computers do not
understand these right there is nothing
no understanding for computers to know
what text is right it only knows 0 and
one yes dhika right it only knows zeros
and ones so see how we will convert this
so there is something called as
vocabulary
right so vocabulary are nothing but the
unique words
right how many unique words do I have in
this my name is Raga four. This is not
unique. This is repeating. This is
repeating. Five, six. And this is
repeating. So I have six words. So now
guys, I will do something called as word
to
right where I will convert these words
into vectors. How will I convert them?
Look at this. So I will have suppose
this is S_sub_1,
this is S_sub_1 and this is S_sub_2,
right? So I will represent
S1 as
right. I will have a vector
of size six. How? I will say 1 0 0 0.
Right? 1 0 0. How many elements does it
have? Six elements. Right? Name will be
0 1 0 0 0.
is will be 0 0 0 1 0 0 0
and ra will be 0 0 0 1 0 0 right this is
s1 my name is raghub now when it comes
to s_ub_2 right when it comes to s_ub_2
how will I enter this s2
0 1 0 0 ragh what is is here
00 0 1 0 0 0
what is
0 0 0 1 0 or nice is 0 0 0 1 and name
will be 0 1 0 0 0 0 right now this will
be the vector representation of these
two sentences just to tell you a fact
GBD3 right GPD3 model right GPD3 or
GPD3.5
they have vocabul vabulary
of 30,000 words, right? 30,000 words.
And each word, right? Each word
has a
dimension
of
12,500
numbers. Right? What do I mean? Suppose
I say Raghub.
So it will be one word and it will be
represented by five 11,500
different numbers
right and like raghub there will be
30,000 words in this GPD model right
30,000 words in this GPD model right and
this total number of parameters which
get trained in the neural network which
we will learn later on neural networks
they are almost close to 1
75
billion
parameters right 1.75 billion parameters
so why I'm telling you all this because
to showcase to you that what is the
importance of vectors right in the
entire machine learning and data science
yeah people so this is the reason people
now my question is how will you create
these vectors how will you read these
vectors Right? The answer is through
numpy package because it is the base
package. Clear people? Yeah. I hope
today's class will be uh interesting for
you because you will know the context.
Why are we doing it? Yeah. So I'll try
to show that to you how we convert
things to vectors.
Right? Okay. Let me go here now. Right.
Let me go here.
So people uh we started using numpy. So
I started with the zero dimension
arrays. Right? Zero dimension arrays. So
zero dimension is nothing but a scalar.
So I created a ar r0 which was np dot
array and I entered a single word single
uh element inside this which is nothing
but a scalar and then I checked the type
of uh ar0 also right and then I check
the dimension also. So the answer was
two class was numpy nd array and the
dimension was zero right exactly what I
had mentioned in my pb
right it will be having a zero dimension
like this right same thing has been
proven
right because it's a scalar now coming
on to one dimension right coming on to
one dimension I create a list which is
nothing but a one-dimension data type
right now I create an array which is a
ar r1 from array from this particular
list lis and then I check the type of
print ar1 check the type of ar1 and the
dimension right so when I execute
right then you will see that it was this
numpy array and the dimension was one
right exactly like this so if I show you
something else say print a ar r r1 one
dot shape
you will see it's 4 comma empty right
and I show it to you here
it will be empty right empty and this
will be comma 1 right so this means that
it has only four elements right if I
increase these elements to say 55 5 66
77
then it will become 7, blank, right?
Which means it has seven elements as a
vector. Now we create something with a
nested list right which is like this. So
with a nest I want one bracket which is
running outside right then inside this I
have one two and three lists inside one
list. Right? So this is a nested list.
This is marks of first student, second
student and the third student. Right? So
I do this and you see it is this right?
And I will show you the shape also
right. It will be 3 + 3 rows and three
columns. Three rows and three columns.
Right? Now similarly we can also create
a
the 3D matrix
right with
two levels right level one and level two
and 3 + 3 so the shape will be what
people can someone guess the shape
what will be the shape of this I've
shown you
Right? If this is 3 4 then what will be
this?
Three rows and three columns. Right? So
when you will execute you will get 2 3 3
right 2 3 3
right
right now people there are multiple ways
right there are multiple ways to create
arrays and we should know them because
all of these comes very very handy. Not
right now. I don't have enough context
to give you right now. But later on you
will see with me or with some other
trainer that how these will be used in
deep learning specifically, right? They
are the key of deep learning algorithms,
right? Where we initialize some weights,
we initialize some biases and those
initializations are nothing but
multi-dimensional numpy arrays, right?
Numpy arrays.
Okay.
Like for example, suppose I have this
I have to multiply this with some random
numbers, right? So how will you generate
these random numbers? You will generate
them through numpy. And you can generate
them in a specific kind of uh shape,
right? Which is 2 + 3. And then you can
multiply them. You can multiply the
matrices and you can get your output for
yourself. Right? So this is the way they
are used.
So we saw the first thing from list we
have already covered. Then now we are
moving on to creating zero arrays right.
So I create a zero array of one
dimension right of one dimension which
is 0 0. They are represented in floats
right. They are represented in floats
0.0 zero. Right? Now you can create a
two-dimensional zero array. Right? You
can create a two-dimensional zero array
which is you have to mention just the
shape inside 2 + 1. So it will have two
rows and one columns, right? Two rows
and one columns. The difference here is
the difference here is that these are
one dimension and these are two
dimensions. Right? You have explicitly
mentioned the rows and columns. So you
can expand this to 10 + 10 also.
Right? You can expand in 10 + 10 or 10 +
6 whatever you feel like yourself.
Right? It will have 10 rows and six
columns. Right?
Now you can also create
the 3D arrays 3D zero arrays.
Right?
Which is 5a 3a 3. What does five means?
What does five means? First element
represents the number of layers. So you
have five layers, right? It's five layer
deep. Then you have three rows and three
columns, right? So it will look
something like this.
1
2
1 2 3 4 5 right like this something like
this right it will look like this tab 1
2 3 1 2 3 1 2 3 right 5 33 okay yes the
number of matrices hurry what I
represented
right layers
rows
columns right layer rows and columns.
Clear?
So now when you execute this, you will
get an arrangement like this. Okay. I
try to show you this thing with another
example, right? Which is I created an
image of an RGB image of 256 cross 256
pixels, right? Which have all zeros
inside them. And this is how it was
created, right? Three layers RGB 256
256. So this is how the image will look
like
right.
This is how it will look like.
Now you can also create
arrays with ones. Right? Exactly the
same way you created it with zeros. I'll
give it to you. I'll give you 5 minutes
time to create them. I'll show you one.
So I say
a1 is equal to np dot
once
and inside I pass
two
I check a1. So this is a array like this
right? It is an array like this. Now you
create create two dimension
and
three dimension
arrays of one. It is basically
as a float. Hurry. It's represented as a
float. Right. It's represented as a
float. Okay.
Right.
Perfect. Right. Also guys uh with this
right also with this you can create the
custom arrays. Right. You can create the
custom arrays. Right. How do we create
custom arrays people? How do we create
the custom arrays?
You have created now zeros. You have
created now ones. Now what is left that
you create the custom arrays. Uh forget
about this. I will come to this later
on.
Right? Let's create
custom arrays. Right? So the syntax
remains the same. Right? I'll say cus
arr is equal to np.
Right? This is the syntax np.
Right? And you will say 6 + 6 and
suppose you want an array of all fours.
Right? This is the dimension 6 + 6. And
this value after comma is basically the
value which you want. You execute this
and you copy this paste this and you
will get the arrays of fours for
yourself. Right? If you want of 10, you
will get of 10. If you want 10.3,
you will get 10.3. Right? anything which
you want. If you want case, you will get
case, right? All the examples. So, let
me just show that to you.
Yep. Like this. Now guys, how did we
create
or how did we use
range in Python?
Can you use
range to generate
numbers between 20 to 50,
right? 20 to 50. Can you give me the
syntax quickly? How did you do that in
range?
How do we do that? We said R is equal to
range
20 to 51. Right? And then I said
I in R
print I,
right? And this is how I got the
numbers, right? So similar
to range in Python,
we have
a range in num py. Right? How do we use
a range? I say
uh a range
ar r is equal to np dot arange. Right?
And then same syntax I will say 20 to
51. Right? 20 to 51. And that's it. And
when I will check my AR range error, you
will see I have generated myself numbers
between 20 to 50 and a range in numpy.
Right? A range in numpy. Yes, if you
want a interval so you can use this say
a range one and after comma you pass the
third argument. Suppose it's three. So
now it will jump three times, right? 20
23 26 29 32 35 like this up till 50.
Now guys there is something which is
called as lind space
right. What is lindspace?
It stands for
linear spacing
which means
between two given numbers.
This function will fit the required
number of
numbers. Right? For example, suppose for
example,
we need to create an interval
from 0 to 1. People, there are infinite
numbers I can have between 0 to 1. Isn't
it?
Infinite numbers I can have between 0 to
1. 0.0000000000001
0 0 1 0 1 01 right I can go in the
infinite manner right now for example
you need to create numbers between 0 to
10 right and you want to create and want
to have
10 numbers in it right so how will you
do this it's not float it's about the
number theory right between 0 and one
you have infinite finite numbers, right?
So you say lindspace is equal to np dot
lindspace, right? np.tlind space. You
mention from 0 to 10, you want to have
10 numbers, right? And when you will
create lindspace,
you will see that these are the numbers
are there which have been created,
right? These are the numbers which have
been created,
right? Nine numbers. Now I say 100
numbers. I want evenly spaced 100
numbers, right? Evenly spaced 100
numbers. How are they even? You can
simply subtract one number from another
and the difference for all the numbers
will be exactly the same. 0.01 0 1 01.
Subtract any two numbers. It will be
0.01 0 1 01. Right? Where do we need
this? We need this to plot the axises.
Right? When you will plot graphs, you
will need access between this interval.
You need five values that works like
this. Okay? Suppose you want from 0 to
10 five different values, right? You
will have five different values like
this, right? Between the gap of 0.5,
right? If I say 1 to 10, you will have
values like this,
right? Like this. Clear? 100 values like
this. Yeah. Clear guys. How do we use
lin space? Suppose you want from 0 to
10. Interval from 0 to 10 and 10 will be
included. Zero will not be included.
Right? We'll start from one. So I go to
one it will be from
one. Why is 0 not included then? Yeah. 0
is included. Right? 0 is also included
and 10 is also included. 100 numbers
between them. Right?
Clear? This is what lin space is. Now
guys, now
suppose you
lohan l space is basically used if you
want to create n numbers between the
range of numbers right between 0 to 1
right suppose between 0 to 1 you are
trying to plot a graph okay and your
values are 0.2 0.3 0.6 six right and you
want to draw a graph so you will have to
mark the axis right the x axis and the
y- axis so you can use lindspace there
and what will it do it will take the
range it will take the interval in
between you want to add the equal space
numbers and then the third argument here
will be that how many numbers do you
want between them so this syntax tells
you that
from
0 to 1
give me 100 numbers. How are these
numbers? Equally
spaced
numbers. Equally spaced numbers, right?
So when you will execute this, you will
see that all there are 100 numbers which
have been generated, right? 100 numbers.
And all the numbers are equidistant from
each other because difference of every
single number from the next number is
0.01
01.
Yep, that's what it does. Right
now guys, now suppose we want to
generate
random numbers, right? We want to
generate random numbers, right? Now
we're interested in generating random
numbers. So we have something called as
random
dot random right. What will it do?
Random.random
will generate
random
float numbers
between 0 to 1. Random float numbers
between 0 to 1. How will this happen?
You will say
rand rand is equal to np do. random dot
random and inside you will mention what
is the dimension that you seek. Suppose
I want 6 + 6. So when you will check
this you will have all numbers for 6 + 6
dimension right 6 + 6 matrix right now
every time you rerun this the numbers
will change because all of these are
random numbers
right all of these are random numbers
now I say 100 multiplied by rand rand
you will see all of them all these
numbers will be multiplied by 00 right
all of them
in one shopping.
Now just like this we can also create
random integers right how will we create
random integers guys
I say rand intore
rand is equal to np dot random dot rand
right and here you will specify that
what is the range of numbers you want
from so I say between 20 to 25 I need
random numbers and then I want it from
in a 3 + 3 format. Right? And now when
you will check your random you will get
random numbers generated like this.
Okay? Random numbers generated like
this.
Right?
If you say 3 + 3 + 3 you will get a 3 +
3 + 3 matrix. Even if you will only say
3, you will get a 1D.
Right guys? You can change the
dimension. So this is the range from
which you want to choose the random
numbers and this is the dimension you
want this matrix or vector to be in.
Now guys, we'll move on to the next part
which is basically properties and again
there are a lot of operations you'll
have to see it yourself right
properties and
attributes
of numpy
arrays right property and attributes of
numpy arrays.
Okay. Now guys, the first one in this
scheme of things is shape of array,
right? Shape of array. What is shape of
array? It tells you
the
dimensions of the array
stored in a
tle. Right?
For example, I say a ar r r r r r r r r
r r r r r r r r r r r r0
right a r r r r r r r r r r r r r r r r
r r r r r 1 a ar a ar a ar a ar a ar a
ar a ar a ar a ar a ar a r r r r r r r r
r r r r r r r r r r r r r 2 a ar r r3
right and then I say
print this
dot shape
right like this and you will see that it
will give you the shape of each array
Right? 0D, 1D, 2D and 3D. Right people?
Shape of the array.
Please try it out. We have used it one
or two times. But this is what shape of
array actually means.
Second is people
end
right is end
right end of array right it tells you
the rank of the array whether it's one
dimensional two dimensional zero
dimensional threedimensional four
dimensional
so again I will do the same and
right I'll say end and you will see it
will give you 0 1 2 3 zero dimension
zero rank one rank two rank and three
rank and it can go all the way up to end
rank
before this I should have also
printed these arrays
Right. These are the arrays.
Yep.
These are the arrays which we have and
these are the subsequent things, right?
Rank copy array. Then guys, the third
thing is the size of
array, right? It tells you
the number of elements inside. How many
elements do we have inside this array in
a ar r1? How many elements do we have? 1
2 3 4 5 6 7. How many elements do we
have in this 2 + 2 ar2? 1 2 3 4 5 6 7 8
9 which is rows multiplied by columns 3
* 3 right and how many elements do we
have in this 3D which is 2 * 3 * 3 which
is 18 right so now when you copy this
right you can use this
and say
size right it will say 17 918
Yep.
Fourth is people.
The D type
of array tells you the data type of the
array. Right? And I've told you we place
only
homogeneous
data in array right what will happen if
we don't do this I will show that to you
also right so we do this right and we
say
retype
and you will see in 64 all of them are
integers right all of them are integers
that's the reason we are getting in 64
right suppose I create a new array a ar
r new right let me say head
right hetro
heterogenous
and I say it is like n dot array
let's say like this
right like this now people when you will
Check
ar r
dot d type you will see it will give you
float just because of one floating point
number inside this entire array it gets
converted to float right between all the
integers if you put one float then it
will be taking float directly right now
let me just copy this and let's say
heterogenous one and let me add another
value which is string and say rather
right and when you will execute this it
will give you U32 U32 here is
representing strings right it is all
called as objects right these are all
string values right so precedences
strings greatest then float and then
your uh integers right if you place the
heterogenous data inside the numpy array
right you only need to put homogeneous
data in the array
now fifth is the item size
of array. Right? What is item size?
It gives you
the bite
occupied
by each element of an array. Right?
Because we assume that elements will be
homogeneous. It will give you the bite
occupied by each element of the array.
Only one element. Okay? So how will it
happen? So let's say a ar r r0
or let's say ar r r1 dot item size
right and you will get eight right. So
why eight? Because
because each data point right each data
point is occupying
the result is 8 bytes. Let me put this
here.
Right. Eight bytes
because each element in ARR1 is
occupying
64 bits which are
equivalent to which are equivalent to 8
bytes. Right? one bite is equal to 8
bits. So 64 bits will be equal to 8
bytes. Right? That is how it is giving
you the result. Now if you are
interested in knowing the entire bytes
right entire bytes then you say n bytes
will give you the
total bytes
occupied
by the elements of the array. Right? You
say print
a ar r1 dot
n bytes right and write
bytes it will be 56 bytes
right why because how many elements do
we have in our ar r1 1 2 3 4 5 6 7 right
7 8 are 56 right 56 total bytes are
being occupied with by ar r1 one. Now
guys, the seventh one
is
as type right
in array.
This will help you change the data type
of the array. Right? Change the data
type of the array. Suppose I have a arr
type which is uh so I'll say print
d type right this is end 64 right and
now what I do is I say print
uh wait let me give you a structured way
print a r1 right let me say
array
R1
right D type of array.
Now
a ar r2 sorry a ar a ar a ar a ar a ar a
ar a ar a ar a ar a ar a r r r r r r r r
r r r r r r r r r r r r r1 is equal to a
ar r1
dot as type right dot as type and let's
say I want to convert this in np dot
right np dot
uh
int 32 right I want to create convert
this in uh ar r r int 32 right when I do
this and now when I will copy these same
things you will see for yourself. Right?
Now the D type was int 64 and now the DT
type is int 32. Right?
If I want I can do this conversion in
float also
float 64. Right? And then I will just
copy this
and I will paste it here.
Right? And now you will see now the
floating point has been activated.
Right? It has been now activated. We can
go till int. We can go till int 8.
Right?
I can go to 16.
Right? And I can do this.
I can go to int
8 also.
Yeah, like this in date also 64 32.
So first one has to be
Yeah.
32
whatever right like this okay you can
convert this
also people this is later on conversion
you can define the data type of the
array while creation time also. How
would you do that? Suppose you are
creating a ar r11 and you say np dot
array right and suppose you take it from
a list and then you just put a comma and
say d type. So what will be the default
data type here people? If I just do this
if I just execute this what will be the
default data type? Int 64 is the default
isn't it? But now suppose I want to
change it right here. I say data type is
equal to float 32 right sorry float 64
np dot
sorry my bad
float 64 right and I say a ar r r11 and
this will be float 64 if you want float
32 it will also become float 32 right
right here while you define
Instead of using as type, you can do it
right here. Right? These things will
come in very handy people because you
will have to save memory because when
your data becomes very very big, you
will be always in a crunch for memory
like this. You want integers, then
integers will be like this
random.randent
like this. Suppose you want to generate
0 to six, right? And suppose you want to
generate
say 100 numbers like this 0 to 6 the
scores
1 to six like this randomly
right suppose you want to generate 100
scores for five different batsmen
randomly it will be like this batsman
number one batsman number to bat number
three, fourth and fifth. Right.
Yep.
Understand the data guys. Now it's the
time to understand the data.
Yep. Now guys, we have methods in numpy
arrays, right? Methods in numpy arrays.
So what are these methods? Now the first
method we have to learn is called as
reshape right. Reshape.
Yeah. So reshape is you can use this to
create
you can use this to create
a new shape of the array. Very very
powerful guys. Very powerful. One of the
most powerful methods in numpy is uh the
reshape right and how do we use reshape
suppose
we have a 1D array of 20 elements right
now to reshape this
reshape it we need to find the factors
Right. Factors of 20. They are what? 1
20 4 5
2 10.
Right. The other factors.
The other factors. Now see what will I
do. Right. Now see what will I do. Let
me create a say random array. Right?
Random array. I say random
arr is equal to entprandom
dot rand right and let me say I want to
create it from 1 to 50 right and I want
it to be having 20 elements right so I
say random arrand
values inside this right random 20
values now see Now reshape
first
I will reshape in 1 + 20 right 1A 20 how
will I do that you just have to write
uh
a ar r r sorry sorry random dot ar r
random ar
dot reshape dot reshape and you just
pass in the dimension I say 1 20
right and when you will do this you will
see it is coming now in 1 20 format
right so let me just also write print
right 2D
1A 20
right and now what I'll do is I'll copy
this and I will paste this and say 20
comma 1 right you will see it will be
like this 20 comma 1 immediately with
reshape right let me copy this
let me say 2D I'm still at 2D let me say
2 10 right and you to see this is 2A 10.
Now I can reshape it in
10 2
right 10 2 right then I can reshape the
same thing
in
4A 5 and I can reshape this in 5A 4
right like this guys are you able to see
the power one dimension I'm able to
create two dimensions
And now I will take it a step further
and I will write it in three dimensions.
Right? How will I write it in three
dimension? Let me say this 1 comma
2a 10. Right? This is will also be three
dimensions. Let's sorry 2a 2a 5.
Let me say this. And now you will see I
can have this in three dimensions.
Right?
Yep. I can also say in three dimensions
like this. I want to have five layers
with two rows and two columns. Right? So
you will have it like this also. Right?
So this is how people we can reshape the
array. Very powerful. Very very
powerful.
Right? Very very powerful.
Right? And you can take this
and save print
random dot this and you can say print
dot shape.
Right? So this was the first one.
Yep. like this.
Also people if you want to visualize we
can also go this route.
We can have 10 comma 2 comma 1, right?
It will look like this, right? 10
layers. 10 layers you can have,
right? 10 layers you can have.
Great. Now, second method which we have
to learn is called as
transpose,
right? Transpose method. Right? What
does that do? It interchanges the
dimensions
like rows and columns, right? Yes.
Absolutely. Right. Suppose you have a
matrix, right? Which is
22, 33, 44, 55, 66, 77, right? This is
a. So now when you will a transpose it,
right? The dimension right now the shape
right now is 3A 2. Now this will become
2a 3. And how this will happen? Rows
will now become columns and columns will
now become rows. Right? So let's make
column the rows. Right? Sorry columns
the rows. It will be 22 44 66
33 55 77. Right? So people in transpose
no information is lost. It is just a
change in the view right which is
happening right?
22 44 66 33 55 77.
Why do we need transposition? Suppose we
have two matrix.
This is 11th class mathematics. Right?
one has
a dimension of n cross m and the second
has dimension of a cross b. If you
want to multiply
these two matrix say
M_sub_1 and M_sub_2,
right? There needs to be a satisfaction
of condition. M should be equal to A.
Right? M should be equal to A. Right? M
should be equal to A. This should be
equal to this and the resultant vector
the resultant matrix which you will get
will be of n crossb dimension right. So
often times suppose this is n cross m
this is n cross m right m cross n and
you know that m is equal to a. So what
will you do? You will transpose this
matrix right? You will transpose this
matrix then it will become n cross m and
then m can be equivalent to a. Right?
For example, what I'm saying, we have
one matrix which is 2 + 3 and this
matrix is 2 + 5, right? So, can you
multiply these matrix people? Is 3 equal
to 2? The answer is no. Right? The
answer is no. So, what will you do? You
will just transpose this and this will
become 3 + 2 and this is 2 + 5. And now
you can multiply this and the resultant
will become 3 + 5 matrix. Right? So for
operations like these we need
transposition. Right? So how do we
transpose it?
How do we transpose it? So let's say
again
uh a ar r2 right this is a ar r2 and I
want to transpose it. So I say a r2 t is
equal to uh np.transpose transpose
sorry a r2 dot
transpose
right
and now when you will see ar r2 ts you
will see rows and columns have
interchanged right rows and columns have
interchanged with each other
or let me give you one more example
Uh if this is not clear, let me pick up
this again
right now. I say
dot reshape
into say
2 + 10. Right? So this is 2 + 10. And
now when you want to transpose this so
I'll say this t is equal to this dot
transpose
t or a this and we can check this now
and it will be this. Sorry guys. So this
is going to be it will be like this
right transposed.
And now if you want to see this,
this was the original shape, right? Rows
and columns have now been interchanged,
transposed with each other.
Now guys, the third method,
the third method which is there with us
is called as flatten, right? It is
called as flatten.
Right? What does flatten do? It reduces
the dimension to one dimension. Right?
Any dimension you have, it reduces it to
one dimension. For example, I have this,
right? And now when I say
this
dot_f,
this will become this dot platin.
And when you will check this up, you
will see that it has now become one
dimension. No matter how many dimensions
you have, it will become one dimension.
Right? Let me take this again to show
you one more example.
Right?
And here I say dot reshape into
uh
uh 5 + 2 + 2 right I do this my random
ar r is this right it has five layers
two rows and two columns right so now I
say this
flatten is equal to this dot flatten
and If you will check it now again, you
will see it has now flattened it out.
Yep.
Now why do we need this? We need this
for a lot of statistical operations. We
need this to feed the data into the
algorithms. Right? As we will move
forward, you will understand the use of
flattening.
Right? Now guys, moving on and uh as
discussed, let me now show you
the power of
numpy,
right? Numpy
over
lists and other data types, right? I
will not take a lot of examples. Just a
second, guys.
Yeah. Okay.
Now guys, I told you that
less
take up
much more memory
as compared
to numpy arrays. Right? And I'm going to
prove this to you. Right? Now let me use
let me create a random sequence of
random numbers using range in Python
right and say I create range of 10,000
numbers right range of 10,000 numbers so
what will this give me this will give me
numbers from 0 to 99999 right continuous
numbers right so this is range I will
use
a range in
numpy Y to create
a similar
series of numbers
right so let's say array is equal to np
dot arange
right same thing same done by both right
I've shown you above also now let me
import sis
library right sis package and I will use
something called as get size of right
get size of. What does this do? Get size
of it's a method
which calculates
the bytes
occupied
by a single
element in
vanilla Python. What is vanilla Python?
It is the traditional Python,
right? vanilla Python.
So let me just show that to you. I'll
use this and I will say print. Now guys,
if I get
size of any random number from this
range, right? Any random number. Say I
get size of five, right? And I then
multiply that byes with the length of
random, right? With the length of random
this rand, right?
Right? With the length of random, do you
think I will get the bytes for the
entire
data structure? What am I saying is
suppose
uh I used range
five. So what will this give me? 0 1 2 3
4. Right? This will be the output. So
now I say get
size of say I say two. Right? So suppose
2 is x and then I multiply this with the
length of this series which is five. So
do you think I will get 5x which will
represent the number of bytes occupied
by the entire data type. Anything
randomly any random number this can be
three right? Why not hurry?
Why not?
All of these are integers. So integers
all of 64 bits assuming. So if you
calculate the side of size of this and
if you multiply with the total number of
numbers you will get the total size
isn't it?
Huh? Index is in
no no it's not about that it's about the
element right homogeneous elements
inside this.
I am saying when you use range five what
is going to be the output? 0 1 2 3 4
right now all these are elements
elements of range.
Right? All of them are elements of
range. Right? Now I'm saying if I fetch
the size of one element and multiply it
with the length of the entire range,
will I get the bytes occupied by the
entire range? For example, if I do this,
right? If I do this,
this is 28, right? 28 bytes people. 28
bytes
bytes are occupied
by one element of range right
one element of
range right now if I just I'm saying I'm
just saying if I multiply to find how
many numbers range has
how many numbers
range has
equal to 10,000
right so total
memory occupied
will be will be how much it will be 28
ult*lied by 10,000 which will be equal
to 28,000
yeah and how will you find this you will
say Print
this multiplied by length of RAM,
right? 28,000 bytes. Clear? Now, yes.
Now, this is for the range. Now, let me
use another thing. So, how will you
calculate the length of this array? What
what property and attribute will you use
people?
N bytes, right? N bytes will give you
total bytes occupied by the elements of
the array. Right? We will use n bytes
here. So I come back down and I say
using n bytes for arrays. Right? And you
will see what the result comes. Print
uh array dot n bytes. Right? And I say
bytes. Are you ready to see the result?
Do you see what has happened?
How many bytes this was taking? It was
taking 28,000 bytes. How many bytes this
is taking? This is taking 80,000 bytes.
This was taking 2 lakh 80,000. This is
taking 80,000. Two lakh extra bytes of
memory is taken by range.
Ran is range.
Yep. And if I just go to million
numbers,
right? Million numbers in both.
See the difference it becomes,
right? This is now 3 three 28 million
bytes it is taking and it is taking 8
million bytes. 20 million extra bytes
are occupied right now people do you
believe me? Yeah, that numpy wy are way
more efficient in memory management as
compared to the traditional data types
of Python. Yes. Okay, that's the first
part. Now second is people performance,
right? Performance. So what I'm going to
do is what I'm going to do is I am going
to
import
time, right? It's a it's a module in
Python, right? Suppose I say x is equal
to range
this much right. Okay. And then I have y
is equal to range say
this
to
this. Right? Both of them will have
equal amount of numbers. Same numbers
both of them will have. Right?
This will have say
uh 1 2 3 1 2 3 10 million values. 10
million values. This will also have 10
million values,
right? Both of them will have 10 million
values. Now what I'm trying to do is I
want to add them up right by bit by bit.
I want to add them up right. I want to
add first element of this to first
element of this. second of this to
second of this, third of this to third
of this like this. Okay, I want to do
this. Now what I'll do is I will run a
counter, right? I will run a counter
which is the start time,
right? And this is given by time dot
time, right? Which will give you the
this will give you the
current time, right? After this I will
run the operation. I will say C is equal
to X + Y
for X Y
in zip
X Y right in zip X Y right add X + Y bit
by bit element by element for X and Y in
zip zip is a function right which allows
you to do this operation sequentially
right sequentially right add the
elements of X and Y
element by element right element by
element
right element by element right and then
I'm going to print right so start time
will start and now I will say time dot
time which is now the end time minus
start time so this will give me the Time
taken for execution isn't it guys
will give me
delta of time which is equal to time
taken for operation
seconds
right these many seconds will be taken
right so let me run this and it takes
around say
4.3 seconds right to do this right 4.3
seconds now Guys, see what happens. You
had to write this complex syntax in the
traditional Python. Now let me show this
on arrays. Right? What will happen on
arrays? I will say a is equal to np dot
a range.
Right? And inside a range I will pass
the same values what I have taken above.
Right? And I will say b is equal to np
dot
a range and I will pass the same values
inside
right exactly the same now what I'll do
is I will say same thing
right just I will change the execution
of C will now simply become people A + B
what is simple this or this
this or this
two right do you see the power if not I
will show this to you again right later
on and let me run this and you see the
difference now let me just increase a
couple of zeros right a couple of zeros
two zeros I'm increasing in both the use
cases
it is going on and on right let's see
See how much time it will take
to add say 2 million 1 billion numbers.
1 billion numbers I have asked my system
to add and I want to see how much time
it takes.
Running running running.
Yep. Colonel has died. Kernel has died.
People,
I'll have to restart.
Right. I will have to import
numpy
as np. So let me just remove one zero
from both.
It is taking 4 seconds. Removing one
zero from here also. Right?
And when I do this it takes 1 second. Do
you see guys what is the difference in
performance also right for both of
these? Yeah.
And if you didn't understand this, let
me give you an example.
Range five. This is 5 to 10, right?
And this is basically
adding elements,
right? 5 7 9 11 13 Right. So this will
be what will the output of this? This
will be 0 1 2 3 4 and this will be
output what 5
6 7 8 9 right so 0 + 5 5 6 + 1 7 7 + 2 9
8 + 3 11 9 + 4 13 and the same thing if
I do here
then what will happen I say 5 I say 5
and 10 right and I say C is equal to a +
b and I say c. Same thing you get here.
Right? We can move on. Right? The next
bit guys which we have to understand the
next bit which we have to understand is
called as the indexing in numpy arrays.
Right? Indexing in numpy arrays. Right?
How do we index the elements? Right? How
do we index the elements?
indexing in
nump py
arrays right indexing in numpy arrays
so again you know indexing from basic
python so let's start with 1d for 1
arrays right I will use ar r r1
yeah this is a ar r1 now right this is a
ar r1 Okay.
Now people what I want to do is what I
want to do is I want to fetch right I
want to fetch right you can slice and
dice let's say dice
33 right so again as per our normal
indexing of list what is 33
what is the index of 33 people
two so you will say the same thing print
right A R R1 squared bracket 2 and you
will get 33 for yourself. Right? If you
wish to slice same things, right? 33 to
say 66. What is the index?
33 is 2. 2. Which one? 3 4 5 and 6.
Right? We will write 3 to six. Not five.
Hurry. Five is not included. Remember?
We print
a ar r r1
2 is to 6 and you will get 33 44 55 66.
Right? Simple indexing. Please try it
out. Please try it out. And if you want
you can have this code also. You can
write this code. You will always have
clarity that why do we use it.
Moving on people. Moving on. Let's see
indexing.
in a 2D array. Right? And before I
explain this to you, let me take you
here. Right? So a 2D array will be what?
Right? This is a 2D array. It has three
rows and three columns. Right? Rows
columns. So now for rows indexing will
start from zero. So if you have to pitch
this particular row, right? So what will
be the index? It will be row 0. If you
have to fetch this particular row, the
index will be one. And if you have to
fetch this particular row, this the
index will be two. Similarly, for
column, if you have to fetch this
particular column, right, you will have
column is equal to zero. This particular
column, column equal to 1. This
particular column, column equal to two.
Right? indexing will start from minus
one again right 012
so let me just show that to you right
let's say I call ar r r2 this is my a r2
now
indexing
first
row right how will I do that I will say
print a ar r r2 and I will write how how
will I write this
h I will Write row 0, right? Row 0. So
what will this give me? This will give
me this, right? The syntax is
row space column. Right? So now if you
just pass one, it will give you only
rows, right?
row one, row two, row three. Right?
Similarly,
if you want to create it for columns,
right? What will you say?
Sorry,
uh
columns.
Uh
uh it was
zero
column 1 column 2
column 3. Right?
Right.
So what is my column 1? 76 89 98 76 89
98 right and if you want column 2
uh sorry
if you want column two this is the
column two and this is column 3 right
guys 90 999 99 independently
now if I ask you to fetch me a
particular element
element, right? Say I want you to fetch
me 78, right? How will you fetch 78? You
will say print a ar r2. What is the row
for 78? Which row does it belong to?
012. To which row 78 belongs? One row.
Which column it belongs to? 012.
1. Hurry. Check again whether it belongs
to zero column, first column or second
column.
Right? And you will get uh sorry
01. So this is two. Right? 78. Right?
Second row first column. Right? Second
row first column.
How to define it as a
right? This is your matrix. Right? Now,
what are the index positions for this?
This row is zero. Row, first row, second
row. This is your zeroth column, first
column, second column. Yeah.
Yes. No. Maybe. Are we understanding
this? This much is clear. The indexing
of rows and columns. Now, now if you
have two, so the syntax is
syntax is say this is matrix A. So you
will say A in this A matrix you will
write row,
column. Right? Suppose I want to access
only zeroth row. So there will be no
column. So what will you get? Row number
zero.
Right? Row number zero. And you will
just put it like this. Or you can put a
comma and put colon. Colon means what?
Take everything. So I want all three
columns together. Zero row and all three
columns. So this will be your
show. Let me show that to you.
Right? See this zero row and all the
columns. So what will be your answer? 76
88 90. Right? Then if you want to access
the second row 89 90 99 89 90 99 one and
like this. Clear? Now similarly for
columns what will happen? You will take
all the rows
comma which column do you want? If you
say two what will be the result for
this? What numbers will you get for
this? Colon, 2, you will get 33 66 99.
Now coming to the element, right?
Suppose now you want to fetch 55s,
right? So what will you write? You will
write which row does it belong? Follow
the syntax. It belongs to the first row.
Which column does it belong to?
First column. So what will you get? 55.
Come to come come to this example. Now
you want 78 in this particular array or
matrix. Right? Where is 78? Which row
does it belong to?
Is it first row? Check carefully.
Zero row. First row. Second row.
Yeah. So I put two here. Now which
column does it belong to? First column.
Second column. Sorry. zero column, first
column, second column belongs to the
first column. So I put one here. So when
you put this syntax, you will get 78.
Clear? Now the third thing is slicing
through the array. Right? Now for this I
want I want 90 99 78 99. That means
what? I want this, this, this, and this.
Let me take you to the PPT first. Now,
what I'm asking you to fetch me? I'm
asking you to fetch me these four
numbers. Right? These four numbers. So,
here what will you write? Which rows are
included in this people? Which rows are
included in this?
Row one to all.
Right. One to all.
Yeah. Not two. One to all.
Right? And you leave everything like
this. If you have the last row, you
leave it empty after the colon, comma.
Which columns do you want for this? One
and two. Right? And when these will
intersect, when these will intersect,
you will get this area. You will get
this shaded area, green shaded area. So
I will say I need from column one to all
the columns. Right? So what will this
fetch you? What will this fetch you?
This will fetch you 44
55 66
and 77 88 99. Right? What will this
fetch you? This will fetch you 22 55 88
33 66 99. What are the commonalities
between both of these
access
both of these slices? What are the
commonalities? It is only this much
right
55 66 88 99 right so do you think you
will get your result yeah let's check it
here right I say print a ar r2 right and
this I say
I need row zero sorry row one to empty
and then column one empty and you will
get 90 99 78 90 indexing.
Yeah.
Now guys, let's check it for
three dimensions, right? 3D.
I say AR R3, right? This is my AR r3,
right? So, let me just put an example
here. Suppose my arrays are 1 1 22 2 3 4
4 5 6 7 88 91. Right? This is my first.
Then behind this I have another matrix
which is uh
111 222 333
444 555 666
777 888 9999.
Right. And in my
third one, I have 11 1 1 222 33 33 3 4
44 44 44 44 44 44 44 44 44 44 44 44 44 4
55 555 66 66 66 66 66 6 7 8
9
is that is it qualifying for a 3D array?
Can you access anything on this
particular array if given a chance
separately?
like this one. Now people, if you want
to move between layers, right? If you
want to move between layers, do you
think I have told you one particular
access which is what? Which is the
layers.
So what number will be given to this
layer? Layer number zero,
layer number one
and layer number two. So now if you have
to access this 555
which layer you will go to first you
will go to first layer. Then which row
will you go to? You will go to first row
and first column. So if you pass this
syntax what will you get? You will get 5
five5.
Right? Problem solved. The only
bottleneck was the layers part and you
have additional parameter or argument
for this layer.
So if I come back to my example and if
you have to access this 555 right how
will you do that? I say print a ar r2 a
r3 right and this I say and in this I
say what I have to access this 555 which
which layer number is this? This is
layer number zero. Right? This is layer
number zero. And this is layer number
one. So I say layer number one. Which
row is this
in this particular layer? It is layer
number. Sorry, it is row number one and
column number one. And when you do this,
you will get 5x5.
Right? If you want 999, what you will
do? You will change this to 2, 2, right?
If you want 777 or let's say if you want
98 what will you do for 98
layer number zero row number two column
number zero then you will get this 98
clear guys? Yep. Layer, row, column. In
the same way, you will slice it. Right.
You will slice it.
Yep.
Right. I'll write the syntax
so that you don't get confused. It is
array square bracket. Layer row column.
Right? Layer
row
column.
Right. This is the syntax.
Perfect.
Right. Great. We can also perform some
operations, right? Plus, minus,
multiplication, division between two
arrays very very easily, right? It
should not pose any problem to us,
right? We can do all the operations
which we want to, right? Between two
arrays, right? For example, right? You
had you have
right operations on arrays.
So let's say uh
let's see as a list right list. So we
have list is equal to 1 2 3 4 5 right
now suppose you want to square
each element of the list. Right? What
will you do? You will say s sq l is
equal to x to the power of 2 for x in
l right and then you will say print xq l
and you will get the squared of list
Right?
Original list squared list. Now
in array what will happen? Suppose I say
a 1 is equal to np dot array and l I
create the same array out of this list.
I say print
original
array
right I say A1 right this is my original
array same as this now I want to square
it
square the elements of arrays very very
simple nothing you require you just say
you just say
sq
a1 is equal to sq a1 1 is equal to a1 to
the power of 2, right? A1 to the power
of 2, right? And then you print
the
right you get the squared r. Suppose
you want to find the mean of
the mean of numbers
using list. Right? So what will you do?
You will say
mean is equal to sum of
l right list divided by length of list
right and you will get the means
right which is three for this one right
this original list. Now let me show you
in arrays
right. How will you do this? You will
say mean is equal to np dot mean and you
will just pass a1 right and when you
will check mean you will get 3.2
Right? Nothing like this direct right
direct like this right you can you have
I've already showed you add I've already
showed you uh square and then I believe
you can understand that what all
operations are possible using the arrays
right leveraging the power of arrays
right also guys in the arrays right in
the arrays what you can do is you can
perform
you can perform some string operations
right very powerful string operations
right so for example let's say I have a
I I have a array right I have an array
of names of people right so I say names
is equal to np dot array right and I say
Radha
uh then I say D
right and then I say
Maduk right these three names I have
right so you can check the names they
will be like in the array right and the
data type will be U6 which is a
representation of strings right now guys
suppose you want to capitalize you wish
to capitalize the names right of people
what will you do you will say print me
np docare right npcare dot capitalize
right capitalize and inside this you
will pass names
and you will see all the names have been
capitalized
right you see this R has been
capitalized D has been capitalized. M
has been capitalized.
Right?
You can convert them into upper if you
want. Print
np.care dot upper
names and you will have all of them in
caps lock. You can say print np.care
dot lower.
You will have them in lower.
Right? You can put the title. Right?
Suppose I say uh
title e title is equal to np dot array
and I say
rahov
go
right then I say
dhapati
right and I say mad
warm
right I say these three things now if I
say print
npcare
dot
title right and I say
titles you will see that all the words
will be in capitalized mode rael d and s
of dh sinapati m and V of MaduMa are now
capitalized. I can also
replace something if I wish to suppose I
want to replace Madhu with say suri
right I will say print
right np do.care care dot replace
right and you will say where you want to
replace I say title
and in this I want to replace
madu
with
suri
right and you will say it will be suri_1
suri one Right. Madu has been replaced
with
Right.
Right. If you want to calculate
the characters of strings,
right, you can do that. Print np.car
dot str length of
titles and you will get 11 characters
are there in here.
15 are there here and 11 are here in
this particular thing.
Right? You can do much more powerful
things also. Let me show you one complex
function. Right? Suppose I have f name
is equal to np dot array.
Right? And we have Ra
Madu
right and we have L name
arapati
worma. Right, we have these two things.
Now I can create a new array full name
by simply right by simply saying np.car
car dot add right and I say uh f name
right
comma
l
name right
Yes.
Yep.
Full name.
Yeah. Ra. I just was trying to add a
space
in between.
Anyway,
right. We can do that,
right? You can also people search in
arrays, right? Very powerful. Again,
search in arrays using where,
right?
Right. You can say suppose a 2 is equal
to
np dot array right and I'll say 1 1 22
33 3 4 4 5 66 right and now you have to
say a is equal to np dot where right and
in this you say a2 greater than 20
right and when you will check A
uh it's giving me the index. Why is it
giving me the index
or is it giving me the index?
Does it always return index?
One more thing is you can find a you can
find an
a letter
right through a letter. You can find an
element
through
a letter. Guys, these are all some
tricks which you should know because you
will be dealing with data and you need
to pull data, right? You need to pull
data a lot, right? Based on conditions
and based on things. Suppose you want to
find out the names which have G in them
right or R A in them. So how will you do
this? I will say print right and I will
say uh np do.care care dot find right
and I will say find this inside full
names right and find me ra a right ra a
two p
it uh right it returns true because it
has found it here right so I don't want
to tell you indexing through this but
anyway you should know this just just
assume this that I'm telling you to
write this okay because this is much
easier when we will go to pandas right
just uh write it as a syntax okay
greater than equal to zero I hope this
is clear
>> so let's start with the data science
interview questions and answers and The
number one problem we would be facing is
real world problem solving. And the
question one is handling missing data in
predictive modeling. So imagine you have
given a data set where 30% of the data
for key predictive variable is missing.
This variable is crucial for a
predictive model. How would you handle
this situation to ensure the integrity
and performance of your model? And
please describe your approach step by
step. So starting with the answer you
can start with handling missing data set
is a common challenge in data science
and it's important to address it
carefully to maintain the accuracy of
your model and here's how you could
approach this situation. The number one
point could be identify the missing
data. So first you need to understand
where the missing values are in your
data set. You can do this by using a
simple code in Python with libraries
like mandas. For example, you can use
the data dot isnull dot sum function
that will show you the count of missing
values in each column. Then you can
analyze the pattern. Determine if
there's a pattern to the missing data.
Is it random or is it missing for a
reason? This can affect your approach.
If the data is missing at random, the
methods you use might be different than
if the data is missing systematically.
So choosing a method for imputation.
Let's see the next method that is
choosing a method for imputation. So if
the missing data is numeric, you might
replace missing values with the mean or
median of that column. This is simple
and effective but can be used primarily
when the data is missing completely at
random. Then comes model based
imputation. Sometimes you can use other
variables in the data to predict missing
values using a regression model. This
can be more accurate but is also more
complex. Then we'll use the k nearest
neighbors can algorithm. But before that
we have a code snippet here that could
be used for the implementation of
imputation. You could use Python or R.
And now moving on we'll see the K
nearest neighbors algorithm. So this
method predicts the missing values based
on how closely related the data points
are to each other. So after imputation
it's crucial to check how your changes
have affected the overall data set and
model performance. Sometimes filling in
too many missing values can introduce
bias. And then we have visualization. To
help understand before and after the
imputation, you could visualize the
distribution of the variable using
histograms or box plots. This helps in
seeing how the imputation has changed
the statistical properties of the data.
And by following these steps, you can
handle missing data thoughtfully and
maintain the integrity of your
predictive model. Now moving to the
question number two that is based on
evaluating model overfitting. So the
question is you have developed a
predictive model but you suspect it
might be overfitting the training data.
How would you test and address the
issue? Please explain your steps and the
techniques you would use. So you could
start the answer by explaining what is
overfitting. So overfitting is a common
problem where model performs well on
training data but poorly on unseen data
indicating it's too closely fitted to
the training data specific details and
noise. So now we'll see a step-by-step
guide on how to address this. The number
one step is cross validation. So one
effective way to test for overfitting is
by using cross validation technique.
Cross validation involves splitting your
training data into multiple smaller sets
that is false and then training a model
on some of these set and validating it
on the others. So this helps you
understand if the model's good
performance is consistent across
different subsets of data. For example,
in Python you can use the cross value
score function from skarn.mmodel
selection. So this is the code and this
is the code snippet of Python that you
can use for the cross validation and
here we are importing from skarn that is
the module and we're importing
cross_well
score and here we have used the cross
val score function and then we have
printed the average cross validation
score and the next step we will do is
running cross validation model. So this
is your predictive model that you have
already built using scikit learn and
here's the x train these are the x input
features of your training data and y
train these are the output labels of
training data. So we are running gross
validation model here this is your
predictive model that you have already
built using scikitlearn. So x train here
that means these are the input features
of your training data and y train here
means these are the output labels of
training data and cv equal to 5. This
parameter tests the function to split
the data into five parts that is false.
And the model is trained on four of
these parts and the remaining part is
used for testing. So this process
rotates until each part has been used
for testing once and the printing
results that is score dot mean. So this
calculates the average of the scores
obtained from each gross validation for
this average score gives you an idea of
how well your model is likely to perform
on unseen data. A consistent score
across different polls suggests your
model is generalizing well rather than
overfitting to the training data. So now
moving to the next point that is
training versus validation error. So
plot the training and validation errors
as a function of training epochs or
complexity of the model. A model that
overfits will show a low error on
training data and a high error on
validation data as it trains further.
Then we have pruning the model. If you
confirm that the model is overfitting,
consider simplifying it. This might mean
reducing the number of parameters by
selecting fewer features using
regularization techniques like lasso or
ridge or choosing a less complex model.
After this step, we will move to
regularization technique step. So these
techniques add a penalty to the loss
function used to train the model which
can discourage complex models that
overfeit. Then we have common methods
that include L1 that is lasso and L2
ridge regularization. And here's how you
can add L2 regularization in Python. So
this is the code snippet here. And what
we have done here is we are creating the
ridge model and we have applied alpha
equal to 1.0. So this parameter controls
the strength of the regularization. A
higher alpha value increases the
regularization effect which helps reduce
model complexity and combat overfitting.
The alpha value can be tuned to find the
optimal balance between bias and
variance. And now coming for the fitting
the model. So model do fit and in that
we have X train and Y train that trains
the ridge model on the training data. It
adjusts the weight of the feature in X
train to predict the Y train while also
considering the regularization term.
This helps prevent the model from
fitting too closely to the noisy aspects
of the training data. And then we are
re-evaluating the model. After making
adjustments, it's important to
re-evaluate the model again using the
same cross validation technique to see
if the issue of overfitting has
improved. And then we have
visualization. To help illustrate or
ffitting, you could create a plot
showing the training and validation
errors or the number of epochs or model
complexity. So by using these
techniques, you can identify if your
model is all fitting and take steps to
correct it ensuring it performs well not
only on the training data but also on
new unseen data. So now moving to the
next question that is question number
three and it is based on realtime data
stream processing and the question is
you are tasked with building a model to
predict stock prices in real time. The
data comes in every second and you need
to update your predictions accordingly.
Describe how you would set up your
system to handle this type of data
effectively and what tools and
techniques would you use and why. So you
could start answering this question with
handling real-time data. So handling
real-time data especially for something
as volatile and fastpaced as stock
prices requires a robust system that can
process and analyze data quickly and
accurately. So here's how you could
approach this. We will set up such a
system and we'll have some steps. So
starting with the steps. So the first
step is choosing the right tools. The
right tool would be Apache Kafka. So
this is a popular tool for handling
real-time data that streams because it
allows you to publish and subscribe to
streams of records that is data and it
can handle high throughput with low
latency. Kafka acts as a buffer and
manages the flow of data ensuring that
your system doesn't get overwhelmed and
you can also use Apache Spark especially
Spark streaming is excellent for
processing the data. It can process data
in real time and perform complex
operations like windowing, grouping data
into chunks of a specified time period
and aggregating that is summarizing
data. So you can modify it and perform
the predicting of stock prices. And then
the step is data processing pipeline.
And the first step comes here is
injection. Data first enters the system
typically through Kafka which collects
data sent from the stock market and then
we do the processing. So spark streaming
takes over here. Here you can apply
transformations and run your predictive
models on the data. For example, you
might calculate moving averages or other
indicators that feed into your stock
price prediction model. And then comes
the output. Finally, the predictions are
outputed. This could be to a dashboard
for traders, an automated trading system
or even stored for further analysis. And
then we develop the model. Now comes the
model development. You would likely use
a machine learning model that can update
quickly and incorporate new data as it
arrives. models such as aim for time
series forecasting or more complex
machine learning models like rect neural
networks RNNs can be suitable. The model
should be retrained or fine-tuned
periodically with new data to ensure it
stays accurate. Now we'll come to
scalability and reliability. So ensure
your system can scale as data volume
increases. This might mean adding more
servers or optimizing your data
processing code. Implement monitoring to
catch any issues early like delays in
data processing or model performance
drops. And now we'll see the step that
is visualization and monitoring.
Consider setting up a real-time
dashboard that shows key metrics like
prediction accuracy and processing time.
This helps in quickly spotting when
something goes wrong. By setting up your
system with these tools and strategies,
you can effectively handle the challenge
of predicting stock prices in real time.
So now we'll move to the next question
that is question number four and this
will based on scalable data analytics.
So we have covered two questions that
were a bit code based questions and now
we'll see other questions that would be
based on scalable data analytics or they
might be on different areas and with the
13th question we'll start again with the
coding ones. So moving with the question
four that is based on scalable data
analytics and the question is given a
scenario where your organization
suddenly needs to scale its data
analysis capabilities due to an influx
of data that would be 10 times the
normal volume. How would you handle this
situation to ensure your data analytics
processes remain efficient and accurate?
What technologies would you consider and
what steps would you take? So you can
start answering this question with
handling a sudden increase in data
volume requires a strategic approach to
scaling your analytics infrastructure
without compromising on efficiency or
accuracy. So we'll see some steps from
that you could effectively manage this
scenario that you would start answering
the interviewer that we can start by
evaluating the current infrastructure's
ability to handle increased loads. This
includes assessing your databases,
servers and analytical tools to identify
potential bottlenecks or limitations.
Then you could move to next step that
would be choosing scalable technologies
to manage the increased data volume.
Consider leveraging cloud-based
solutions such as Amazon web services,
Google cloud platform or Microsoft
Azure. These platforms offer scalable
resources which can be adjusted
accordingly to the data load ensuring
you only pay for what you use. integrate
big data technologies like Apache Hadoop
for distributed storage and Apache Spark
for fast data processing. These tools
are designed to handle massive volumes
of data efficiently and can scale up to
meet standard increased demands. Now we
move to the next step that would be
optimizing data processing. So implement
data partitioning and indexing
strategies to improve the efficiency of
data queries. This will help in managing
large data sets by breaking them into
smaller manageable chunks and speeding
up search operations and use real-time
data processing frameworks like Apache
Kafka or Apache Flink which can handle
high throughput and provide timely
insights from large data streams. And
the next step would be automation and
monitoring. Automate routine data
processing task to reduce the manual
effort and speed up the analysis. This
can be done through scripting or using
workflow automation tools. Set up
comprehensive monitoring systems to
track the performance of your data
processes. Tools like Prometheus for
system monitoring and Graphana for
analytics and monitoring dashboards are
useful here. They help ensure that the
system is running smoothly and alert you
to potential issues before they become
critical. And the next step will be
regular evaluation and scaling.
Continuously evaluate the performance of
analytics infrastructure. As your data
grows, keep adjusting and scaling your
resources to maintain optimal
performance. Plan for periodic reviews
of your technology stack and
infrastructure to ensure they remain
aligned with your data needs and
organizational goals. By following these
steps, you can ensure that your data
analytics processes are prepared to
handle sudden surges in data volume
effectively maintaining the integrity
and speed of insights. So this was all
for the question four. Now moving to the
question five and this is based on
integrating machine learning models into
production and the question is you have
developed a machine learning model that
performs well in testing environment.
Now you need to integrate it into your
production environment where it will be
used in realtime applications. What
steps would you take to ensure the
successful deployment and operations of
the model in production? So we'll start
answering this by successfully deploying
a machine learning model into production
involves several critical steps to
ensure it performs as well in real time
operations as it does in testing. So you
would have a clear pathway to make the
interviewer understand. We will start
with the pathway with the first step
that would be model validation. So
before moving anything into production
revalidate your model's performance
using a separate validation data set.
This helps confirm that the model
generalizes well to new unseen data. The
next step will be preparing the
production environment. Ensure that the
production environment is ready to
handle the model. This includes setting
up the necessary hardware and software
ensuring that it can handle the expected
load and that all dependencies are
correctly installed and configured. Then
the next step comes that is model
wrapping. Wrap your model in an API that
is application programming interface
making it accessible to other parts of
your software infrastructure. Frameworks
like flask for Python can be used to
create a simple web server that listens
for data inputs and provides model
outputs. Then comes the next step that
is deployment strategies. Consider using
containerization tools like doer which
can help encapsulate your model and its
environment ensuring that it works
uniformly across different development
and production settings. And then we'll
use deployment strategies like blue
green deployment or canary releases to
minimize downtime and reduce the risk of
introducing a faulty model into
production. And then comes the next step
that is monitoring and logging.
Implement logging and monitoring to
track the model's performance and health
in real time. Tools like prompts for
monitoring and ELK elastic search log
statch kibbana for logging help in
quickly identifying and diagnosing
issues in production. And then comes the
next step that is performance tuning.
Monitor the model's performance over
time. If the model's performance
degrades or if new data shows different
patterns, you may need to retrain or
fine-tune the model to maintain
accuracy. And after this step, there's a
step for feedback loop. Set a feedback
loop where predictions and outcomes can
be compared. This feedback is crucial
for continuously improving the model and
catching any drift in data or changes in
external conditions that affect the
model. And after this comes a last step
that is legal and compliance checks.
Ensure all the data used by the model in
production complies with privacy laws
and regulations. This is crucial for
maintaining trust and legality
especially when handling sensitive
information. So by carefully planning
and executing these steps you can
smoothly transition your machine
learning model from a testing
environment to a fully functional
component of a production system. So
this was all about the question number
five. Now moving to the question number
six that would be based on datadriven
decision making. And the question is
your company wants to shift towards more
datadriven decision making. You have
been tasked with developing a strategy
to implement this. What steps would you
take to ensure that the data at all
levels of the organization is utilized
effectively to make informed decisions
and what challenges might you face and
how would you address them? So you can
start answering this by implementing a
datadriven decision-m strategy that will
require a comprehensive approach to
ensure that reliable data is accessible
and effectively used across all levels
of the organization. And now we can
develop and deploy this strategy. And
similarly you could tell this strategy
to the interviewer. So the number one
step will be assessing current data
infrastructure. Start by evaluating the
existing data infrastructure to
understand what data is available, how
it is stored and how it is currently
used. This assessment will help identify
gaps in data collection, storage and
access that need to be addressed. Now we
move to the next step that is developing
a data governance framework. Implement a
data governance framework that defines
who can access data, how it can be used
and who is responsible for maintaining
its quality. This framework ensures data
integrity and security which are
critical for making reliable decisions.
Now we'll move to the next step that is
training and empowerment. So train
employees at all levels on the
importance of datadriven decision making
and provide them with the tools and
knowledge necessary to analyze and
interpret data. This might include
training sessions, workshops and ongoing
support to ensure everyone can use data
effectively. Now move to the next step
that is implementing analytical tools.
So deploy user-friendly analytical tools
that can integrate seamlessly into the
daily workflows of employees. Tools like
Tableau, Microsoft PowerBI or even
advanced Excel techniques can provide
powerful data analysis capabilities
without requiring extensive technical
knowledge. After this we'll move to the
step that would be creating a
centralized data platform. Developer
centralized data platform where all
organizational data can be accessed and
analyzed. This platform should be
scalable and secure providing a single
source of truth for the organization.
And then we have the promoting a
datadriven culture. So foster culture
that values datadriven decision-m
encourage experimentation and learning
from datadriven initiatives. celebrate
successes and learn from failures to
continually improve the use of
datadriven in decision making and there
would be some challenges and solutions
for that. So one major challenge we know
here is resistance to change as some
employees may prefer traditional
decision-m methods. So address this by
demonstrating the tangible benefits of
datadriven decisions through pilot
projects and success stories. So data
silos can also hinder effective data use
promote cross department collaboration
and integrate disparate data sources to
overcome this challenge. After that you
can monitor and do continuous
improvement. So by systematically
implementing these steps you can
transform your organization into one
that leverages data at all levels to
make informed and effective decisions.
And after answering in these steps you
could make the interviewer have a truth
and a faith in you that you could make
these models. Now move to the next
question that is question number seven
and that is based on handling large data
set and the question is your project
involves analyzing extremely large data
sets potentially exceeding terabytes in
size. What strategies would you use to
manage and analyze such large data sets
effectively? Describe the tools and
techniques you might employ and you
could start this with answering that
working with large data sets especially
those in terabyte range presents unique
challenges in terms of storage
processing and analysis. So we'll have a
structured approach to handle these
challenges effectively. We'll start with
the data storage that would be use
distributed file systems. Consider using
a distributed file systems like Hadoop
distributed file system HDFS or Amazon
S3. These systems are designed to store
vast amounts of data across many servers
offering high availability and port
tolerance. And then comes the next step
that is data processing. Leverage big
data processing frameworks. Tools like
Apache Spark are ideal for processing
large data sets because they handle
distributed computing effectively. Spark
can perform data processing task much
faster than traditional disk based
processing due to its in-memory
computing capabilities. And next we
could start with efficient data
sampling. So there are many sampling
techniques that we can use. So when the
data set is too large to handle even
with powerful tools consider using data
sampling techniques to reduce the size
to a manageable level without losing
significant insights. Ensure that the
sample represents the whole data set
accurately. And then comes optimization
of data queries. Indexing and
partitioning. Optimize your data queries
by implementing indexing and
partitioning. This can drastically
reduce the time it takes to perform
queries by limiting the amounts of data
scan. And then we can do scalable
analytics. And then we'll move to the
next step that is scalable analytics.
And in that we could start with the
parallel computing. Use parallel
computing capabilities of frameworks
like spark or dask to analyze data
across multiple nodes. This helps in
scaling up your analytics operations to
handle large data sets effectively. And
now we'll move to the cloud-based
analytical tools. So consider using
cloud services like Google BigQuery or
AWS Red Shift which are designed to
handle massive data sets and complex
analytics with ease. And after this step
we'll move to data cleaning and
pre-processing. Here we will automate
pre-processing task. We'll use automated
tools to clean and pre-process data.
This includes handling missing values,
normalizing data and removing duplicates
which can be particularly challenging
with large data set. And after this
step, we'll move to the step that will
visualize large data set. So we'll use
specialized tools. That tools could be
Tableau or PowerBI that can handle large
data set by aggregating data and using
efficient backend technologies. For more
detailed exploration, tools like plotly
or bouquet can be used as they offer
capabilities to interactively visualize
large volumes of data. And after that,
there would be step for regular
maintenance and updates. That could be
continuously monitoring the data
quality. As new data comes in, you can
continuously monitor its quality. And
after this step, you could integrate all
these strategies and tools into your
workflow. And you can effectively manage
and extract valuable insights from
extremely large data sets thereby
supporting robust datadriven decision
making. And you could answer the whole
strategy to the interviewer. Now moving
to the question number eight that is
based on optimizing machine learning
models and the question is during model
development you have noticed that your
machine learning model is
underperforming. What steps would you
take to diagnose the problem and
optimize the model's performance? What
techniques and tools would you use? So
you can answer this by starting with the
optimizing and optimizing a machine
learning model that is underperforming
involves several steps to diagnose and
improve its accuracy and efficiency. And
here we will have structured approach to
tackle this issue and you could start
this with the number one step that is
diagnosing the problem. Evaluate model
metrics. Start by thoroughly evaluating
the performance metrics of your model.
For classification task, for
classification task, look at accuracy,
precision, recall and the F1 score. For
regression task, consider R squ, mean
squared error that is MSE and mean
absolute error that is MA. And then you
can move to the next step that is use
plots like ROC curves for classification
models and residual plots for regression
to visually assess with the model is
going wrong. After that, we'll move to
the next step that is data quality and
quantity check. Inspect the data that is
sometimes the quality and quantity of
data can be the root cause of poor model
performance. Ensure the data is clean,
well pre-processed and sufficient. Look
for issues like missing values, outliers
or imbalanced classes. And after this
we'll move to the feature engineering
step that would be experiment with
creating new features or transforming
existing ones to provide better
predictive power. And then we have the
next step that is model tuning and
configuration. After feature
engineering, we'll move to the next step
that is model tuning and configuration.
So, hyperparameter tuning. Use
techniques like grid search or random
search to find the optimal settings for
your model's parameters. Tools like
scikit learns, grid search CV or
randomized search CV can automate this
process. And there's a cross validation
that would implement cross validation to
ensure that the model's performance is
consistent across different subsets of
the data set. And then we have the next
step that is trying different models. So
experiment with algorithms here. If
initial models are underperforming, try
different algorithms that might be
better suited for the problem. For
instance, if you started with linear
regression and it's not performing well,
consider more complex models like random
forest or gradient boosting machines.
And after this, we have nseml methods
that we can use techniques like bagging,
boosting or stacking to combine the
predictions of multiple models to
improve overall performance. After this
step, we have feature selection that
includes reduce dimensionality. Use
techniques like principal component
analysis that is PCA to reduce the
number of features which might help in
improving model performance by removing
noise and redundancy. And then we have
select important features. So use model
based technique to identify and keep
only the most important features that
impact the outcome. And then comes the
last step that is regular updates and
retraining. So here you can monitor and
update that could be continuously
monitoring the model's performance over
time as new data becomes available
update and retrain the model to adapt to
any changes in underlying patterns and
after that you could have a consultation
and collaboration work with the other
teams and by methodically addressing
each of these areas you can diagnose why
your machine learning model is
underperforming and can take steps to
optimize its accuracy and efficiency. So
this was all about question number
eight. So let's start with the question
number nine and this is based on
handling unstructured data. So the
question is you are given a large amount
of unstructured data including text,
images and videos. What strategies would
you use to manage and analyze this type
of data effectively? Describe the tools
and techniques you might employ. So you
can start answering this question by
describing that dealing with
unstructured data can be challenging due
to its lack of predefined format or
structure. However, with the right
strategies and tools, you can
effectively manage and analyze it to
extract valuable insights and there will
be a approach how you can do that. So,
we will discuss the approach here and
starting with the steps. So, the number
one step will be data categorization and
organization. So, the number one step in
this step will be sorting and tagging.
We will begin by categorizing the data
into types that will be text, images or
videos. Use tagging to add metadata
which helps in organizing the data and
makes it easier to access and analyze
later. Then and after that particularly
for text data we'll use natural language
processing NLP. We will employ NLP
techniques to extract useful information
from text. Tools like NLTK, spacy or
even more advanced models like BERT can
help you perform tasks such as sentiment
analysis, entity recognition and topic
modeling. After that we will do text
indexing. We can use elastic search or
Apache sle to index large volumes of
text. These tools provide powerful
search capabilities and can handle
complex queries efficiently. And after
that we'll move to image data. And to
structure image data we'll use image
processing. We'll use libraries like
OpenCV for basic image processing tasks
such as filtering and transformations.
For more advanced image analysis,
consider deep learning models using
frameworks like TensorFlow or PyTorch.
And then we'll feature extraction. Apply
techniques to extract features from
images such as edges, textures or key
points which can be used for further
analysis or machine learning. And then
we'll come to video data. And here we'll
do video processing. and we'll use the
tools like fmpg that can be used for
basic video processing tasks such as
format conversion or extracting frames
for analyzing video content look at
machine learning models that can
classify or recognize activities in the
video and after this we'll move to
temporal analysis for videos temporal
components are important techniques like
sequence modeling or recurrent neural
networks RNNs can be useful to analyze
sequences of frames for activities or
events and then we'll move to data
storage and management. Here we'll use
the given volume and complexity of
unstructured data and use big data
platforms like Hadoop or cloud services
like AWS S3 for storage. These platforms
can scale up to handle large data sizes
and provide the necessary infrastructure
to store and retrieve unstructured data
efficiently. And then we have
visualization and reporting custom
dashboards that we'll create here. We
will develop custom dashboards using
tools like Tableau or PowerBI which can
integrate different data types and
provide a unified view of the analyzed
data. And after that we will do data
summarization. Tools that provide
summarization capabilities can help in
considering large volumes of
unstructured data into more manageable
and interpretable forms. And after that
we'll leverage these strategies and
tools and can effectively manage,
analyze and derive insights from
unstructured data which can be crucial
for making informed decisions in various
applications. And this is the path that
you can explore and explain to the
interviewer if this question has been
asked. Now moving to the question number
10 and that will be based on scaling AI
solutions in enterprise and the question
is your company wants to scale its AI
operations from a few initial pilot
projects to enterprisewide
implementation. What are the key
considerations and steps you would take
to ensure the successful scaling of AI
solutions across the organization and
what challenges might you face and how
would you address them? So you can start
answering this question with the scaling
AI solutions. You could answer him that
scaling AI solutions across an
enterprise requires careful planning and
strategic implementation to ensure
success and alignment with business
objectives and there should be a
strategic approach to implement this. So
starting with the approach and the
number one step will be that will be
strategic alignment. So identify
business objectives. Start by
identifying the business objectives that
the AI solutions are intended to
support. This ensures that the AI
initiatives are aligned with the company
strategic goals and can demonstrate
clear business value. And then comes the
stakeholder engagement. So engage
stakeholders from various departments
early in the process to gather input and
build support. This helps in
understanding diverse needs and ensures
broader acceptance of the AI solutions.
And after that comes the infrastructure
and technology. So there's an option
that is assess and upgrade
infrastructure. Evaluate whether your
current IT infrastructure can support
the expanded use of AI. You might need
to upgrade hardware, invest in cloud
solutions or adopt technologies that
facilitate AI processing and data
handling. And after that we have
standardization of tools. Standardize
the tools and platforms used for AI
development to ensure compatibility and
ease of maintenance across the
organizations. And after that we'll move
to data management. So robust data
governance that is to implement a strong
data governance framework to manage
enterprise data effectively. This
includes policies for data quality,
security and compliance especially
important when scaling AI solutions that
rely on vast amounts of data. And after
that we will come to data accessibility.
So ensure that data is accessible across
the organization but also secure against
unauthorized access. This involves
setting up secure data leaks or
warehouses that centralize data while
allowing controlled access. And then we
come to the next step that is talent and
training. So build AI competency that is
develop in-house AI expertise through
training programs and hiring. So this
build the necessary skills within the
organization to develop, manage and
scale AI solutions and after that you
can also perform cross functional AI
teams that could be forming cross
functional teams that include data
scientists, IT professionals and domain
experts. So this fosters collaboration
and ensure that AI solutions are
developed with a comprehensive
understanding. And after forming these
collaborative teams, we move to scalable
deployment models. So pilot test and
phase roll out. Before a full-scale
rollout, conduct pilot test to go the AI
solution effectiveness and integration
capabilities based on feedback, adjust
and then gradually deploy the solutions
across the organization. And then we
have modular and flexible design. So
design AI systems to be modular and
scalable allowing for adjustments and
expansions as needs and then we'll
monitor and do the continuous
improvement. So there will be
performance metrics that would establish
metrics to regularly assess the
performance of AI systems. We will
monitor these systems to ensure they met
expected outcomes and adapt as
necessary. And after that we have next
step that is addressing challenges. So
there could be cultural resistance that
there could be employees that would be
resisting to the changes but we have to
address this through continuous
education and by showcasing successful
AI use cases within the organizations
and by carefully considering these
aspects and methodically implementing
steps you can successfully scale AI
solutions across your enterprise driving
significant business value and
innovation. And that's all for question
number 10. Now we'll move to question
number 11 and that is based on ethical
considerations in data science. So the
question is in your data science
projects how do you ensure that ethical
considerations are addressed? Describe
the steps you take to identify and
mitigate ethical risk in your projects.
What frameworks or guidelines do you
follow? So you could start answering
this question with ethical
considerations that they're crucial in
data science to ensure that the
solutions and analyzes do not
advertently cause harm or bias. Here's
how you can ensure that. So there are
some steps and we will discuss those
steps. Starting with the number one that
is educate on ethical standards. So stay
informed about the ethical standards in
data science such as fairness,
accountability, transparency and
privacy. Organizations like the data
science association and the ACM have
codes of ethics that we refer to as
guidelines. And then we have ethical
risk assessment. Identify potential
ethical issues. That would be at the
beginning of each project. Conduct a
thorough assessment to identify any
potential ethical risk such as biases in
data or impact on vulnerable groups.
This involve reviewing the source of
data, the methodologies used for data
collection and the intended use of the
data analytics results. And then we have
stakeholder analysis. Engage with
stakeholders to understand the diverse
perspectives and potential impact of the
project. This helps in identifying
ethical issues that may not be apparent
from a purely technical standpoint. And
then we'll move to mitigation
strategies. Implementing bias mitigation
techniques. We will use statistical and
machine learning techniques to detect
and mitigate biases in data. This might
involve techniques like resampling,
reeing or using algorithms designed to
be fair. And then we have privacy
preserving methods. Employ methods such
as data anonymization, encryption or
differential privacy to protect
individual privacy when analyzing
sensitive data. Then we have other
methods that is transparency and
explanability. There we have model
explanability and after that coming to
documentation and reporting. So we have
to maintain thorough documentation of
data sources, model decisions and
methodologies. And then we have
continuous monitoring and feedback.
There you have to monitor outcomes and
the feedback mechanisms should be
applied. And then we have the panels
that is collaboration and advisory
panels. Then we have ethical review
boards. So for complex projects setting
up or consulting with an ethical review
board can provide oversight and diverse
perspectives on the ethical implications
of project methodologies. So by
proactively addressing ethical
considerations through these steps you
can ensure that your data science
projects uphold high ethical standards
and positively contribute to society
while minimizing harm. So this was all
about question 11. Now moving to
question number 12 that is based on time
series forecasting for business
decisions. So the question number 12 is
you are tasked with forecasting monthly
sales for a retail company using time
series data from the past 5 years. What
steps would you take to prepare and
analyze this data to make accurate
forecast? What specific tools or
techniques would you use and why? So we
can start answering this by time series
forecasting and we could address them
that it's a powerful tool for predicting
future events based on past data
especially in business context like
retail sales. So we will have a
structured approach here and we'll start
with data collection and cleaning.
First, you will gather data and ensure
that you have collected all relevant
data including monthly sale figures from
the past five years and also considering
including external factors that might
affect sales such as economic
indicators, holidays and promotional
activities. And then we'll proceed to
clean data. We will check for and handle
any inconsistencies or missing values.
And then we have data visualization.
Here we will plot the data. We'll use
plotting libraries like Matt Lib or
Seabbone in Python to visualize the
data. This will help in identifying
patterns, trends and seasonality. And
then we have decomposition of data. So
there's a seasonal decomposition and
we'll use statistical techniques to
decompose the data into trend seasonally
and residuals. So this can be
accomplished with tools like the
seasonal decompose function from the
stats models library in Python. And
we'll understand these components
separately and can improve the accuracy
of our forecast. And then the next step
is model selection and forecasting. So
there are two models that is a ara and s
IMA models. So we have to choose
appropriate forecasting models based on
data's characteristics. For instance,
AMA that is auto reggressive integrated
moving average. It is effective for
non-season data while SMA that is
seasonal AMA that is suitable for data
with seasonal patterns. And after
choosing the model we'll move to cross
validation. We will implement time
series specific cross validation
techniques like timebased splitting to
evaluate model performance and this will
ensure your model generalizes well on
unseen data and then we have model
fitting and diagnostics. We will fit the
model that is by using the cinemax class
from stat models that will fit your
model to the data and then we will
carefully select parameters based on AIC
that is a cake information criterion
that scores or thorough grid research
technique and then we can do the
diagnostics and forecast and validation
and after forecast validation we'll move
to iterative improvement. So there's a
feedback loop that should be mandatory
and there should be a regular update for
the model with new sales data and
refining the model as needed. So this
continuous improvement cycle helps adapt
to changing patterns in sales data. And
by following these steps and using these
tools, you can create robust forecast
that help the retail company plan better
and make informed decisions. So this was
all about question number 12. Now move
to question number 13 that is based on
customer segmentation using machine
learning. So the question is you are
given a data set containing demographic
and purchasing behavior data for a group
of customers. Your task is to segment
these customers into distinct groups
based on similarities in the purchasing
behavior and demographics. So what steps
would you take to perform this
segmentation and can you provide a
sample Python code snippet to illustrate
the initial stages of data handling and
model application. So we can start this
by explaining customer segmentation that
it's a powerful approach to tailor
marketing strategies and improve
customer service by identifying distinct
groups based on their behavior and
characteristics. And here also we have a
detailed approach for this task. So
we'll start with number one step that
would be data exploration and
pre-processing. So there will be initial
exploration that is beginning by
examining the data set to understand the
features available such as age, income,
purchase frequency etc. Then we'll look
for missing values or anomalies and
decide how to handle them. That could be
using imputation and then we'll move to
feature engineering. We'll create new
features that might be useful for
segmentation such as customer lifetime
value or average transaction amount.
We'll also use normalization that is
normalize the data to ensure that one
feature doesn't disproportionately
influence the model due to its scale.
We'll use standard scaling or minmax
scaling as appropriate. So then we'll
come to the next step that is choosing
the segmentation technique. And here we
have k means clustering. So this is a
popular method for customer
segmentation. Here we will decide on the
number of clusters by using techniques
like the elbow method or analysis to
determine the optimal cluster count. And
then we have model implementation and in
that we will use data preparation and
we'll prepare the data by selecting the
relevant features and applying any final
transformations and then we have model
fitting. We fit the C means clustering
model to the data and evaluate and
interpret analyzing clusters and after
analyzing clusters we'll move to the
next step that is strategic insights. We
will provide actionable insights based
on cluster characteristics such as
targeted marketing strategies for each
segment. And then we have iterative
refinement that is feedback
incorporation and we'll use business
feedback to refine the segmentation. If
additional data becomes available
incorporated to enhance the model and
now we'll see the sample Python code. So
for this first we'll import the
libraries and modules. As you can see on
the screen we have imported pandas
random forest classifier train test
split standard scaler classification
report and after that we will load the
data and and for that we have used the
pandas to read the data that is read ssv
and after that we are processing the
data that is data prep-processing we are
handling missing values and using the
forward fill or fil to fill missing
values in the data set and then we are
featuring scaling that is normalizing
the selected features that is feature
one, feature two and feature three using
standard scaler and then we'll move to
the next step that is data splitting.
We'll split the data set into training
and testing sets. So that test size
equal to 0.2 parameters specifies that
20% of the data will be used for testing
and then we'll train the model. We'll
initialize and train a random forest
classifier with 100 trees and a random
state for reproductibility and then
we'll evaluate the model. will make
predictions on the test set that is X
test using the train model and print a
classification report showing precision
recall F1 score and support for each
class. So this code demonstrates the
process of loading, pre-processing,
training and evaluating a machine
learning model that is random forest
classifier for predicting equipment
failures in a manufacturing plant. The
use of techniques such as data
prep-processing and splitting along with
the random forest classifier highlights
a standard flow for building predictive
maintenance models. So this was all
about the question number 13. So now
moving to the question number 14 that is
based on predictive customer churn and
the question is you are tasked with
developing a model to predict which
customers are likely to churn from a
subscription service. So what steps
would you take to build this model and
can you provide a sample Python code to
illustrate the data preparation and
model training process? So we'll start
answering this question about depicting
what is predicting customer churn. So
predicting customer churn is crucial for
businesses to implement detention
strategies proactively and we'll have a
detailed approach for building a
predictive model for this purpose.
Starting with data collection and
exploration and in this we will collect
data and after that we'll perform the
exploratory data analysis that is EDA.
We'll perform an initial analysis to
understand patterns and trends and then
we have feature engineering. We will
create new features and derive new
feature that might influence churn such
as change in usage pattern or service
upgrades. And then we'll handle the
missing values if we found any. And then
we'll encode categoral variables. We'll
use techniques like one hot encoding or
label encoding for categorial variables.
And then we have scale features to
normalize or standardize numerical
features to ensure they contribute
equally to the model's performance. And
then we'll select the model that is
we'll choose the appropriate model and
start with for the knowing handling
binary classification task that could be
with logistic regression, random forest
or gradient boosting machines. And after
selecting the model, we'll train the
model and evaluate it. So fit your model
on the training data and after that
evaluate the model using appropriate
metrics like accuracy, precision,
recall, F1 score and ROC to go its
performance. And then we'll optimize the
model using hyperparameter tuning. We'll
optimize the model parameter using grid
search or random search to improve
performance. And then we have feature
importance that is analyze and rank
features by their importance in
predicting churn to refine the model
further. And then and then the last step
is deployment and monitoring. We'll
deploy the model once validated deploy
the model into a production environment
where it can predict real-time churn. So
after deploying the model regularly
monitor the model to ensure it remains
effective over time as new data comes
in. So now we'll see the sample Python
code for this example. So starting with
the importing of libraries we will
import pandas numpy scikitlearn skarn
tensorflow and the tensorflow kas and
callbacks and after importing the
modules we'll start with data loading
we'll load the data set from a CSV file
named equipment data dot csv and that
with the pandas data frame and after
that we'll do the data prep-processing
we'll handle missing values and for that
we'll use forward fill to fill missing
values in the data set and then we have
feature scaling that will normalize the
selected features that is feature one,
feature two, feature three using
standard scaler and after that we'll use
the data splitting. We'll split the data
set into training and testing sets and
the test size will be equal to 0.2 and
this parameter specifies that 20% of the
data will be used for testing and after
that we'll start with building the
model. First we'll see sequential model
that initializes a sequential model
technique. And then we have dense layers
that adds two dense layers with 64 units
and value activation function. Then we
have dropout layers that adds two
dropout layers with a dropout rate of
0.5 to reduce overfitting. After that
we'll do the model compilation. We'll
compile the model using the atom
optimizer and binary cross entropy loss
function for binary classification. And
there will be an early stopping that
will define an early stopping call back
to stop training when the validation
loss metric has stopped improving after
three blocks. And after training the
model, we will evaluate the model. And
evaluating the model on the test data
and print the loss and accuracy metrics.
So this code demonstrates the process of
loading, pre-processing, building,
compiling, training and evaluating a
deep learning model using TensorFlow and
KAS for predicting equipment failures in
a manufacturing plant. So the use of
techniques such as data prep-processing,
dropout regularization and early
stopping helps in building a robust deep
learning model for predictive
maintenance. So that's all with question
number 14. Now we'll start with question
number 15 that is based on deep learning
and NLP. And your question is you are
tasked with developing a sentiment
analysis model using deep learning to
understand customer opinions from
reviews. So what steps would you take to
build this model and can you provide a
sample Python code snippet to illustrate
how you would pre-process data and train
a simple deep learning model? So we'll
start answering this with sentiment
analysis that sentiment analysis using
deep learning allows businesses to coach
customer sentiment from text data like
reviews or comments effectively. And
we'll have a detailed approach for
building a sentiment analysis model.
We'll start with data collection and
cleaning. We will collect the data,
gather a substantial data set of text
reviews and their associated sentiments
typically labeled as positive, negative,
or neural. And then we'll clean the
data, pre-process the data by removing
noise such as HTML tags, special
characters, and so words. And we'll
normalize the text by converting it to
lower case. And then we have text
prep-processing. We'll convert text into
tokens, words, or phrases. And then we
have vectorization that transforms
tokens into numerical format using
techniques like word embeddings or TF
that is term frequency in document
frequency and then we'll use the padding
and then we have the option of model
selection. will choose a model
architecture based on a basic approach
and use a RNN or more advanced
architecture like LSTM that is long
short-term memory or GRU that is gated
recurrent units which are effective for
sequence data like text and then we have
model training we'll compile the model
define the model architecture and
compile it with a loss function suited
for classification like categoral cross
entropy and an optimizer like Adam and
then we'll train the model. We'll fit
the model on our pre-processed data.
We'll evaluate and optimize it.
Evaluating model performance. Here use
the metrics such as accuracy, precision,
recall, and F1 score to assess the
model. And then we have hyperparameter
tuning. We'll optimize the model by
adjusting parameters like learning rate,
number of layers and units per layer.
And then coming to deployment. We'll
deploy the model and integrate the model
into the existing review processing
pipeline. So it can automatically
classify new reviews. So let's see the
sample Python code and we'll have a
basic approach for that. Here we'll
import numpy tensorflow sequential
embedding LSTM dense stroke out. So
embedding converts positive integers
that is indexes into dense vectors of
fixed size and LSTM that is long
short-term memory layer that is used for
learning dependencies in sequence data.
And then we have dense that is a
regularly densed connected NN layer. And
then we will import pad sequences. And
after that we have the data set and the
sample text data representing customer
reviews that will store in variable
text. And then we have labels that has
binary labels indicating sentiment one
for positive zero for negative. And now
we'll start with the pre-processing of
data. Here we have declared that
tokenizer. We will initialize a
tokenizer that will help only the top
thousand most frequent words. And then
we have fit_on
text that is update the internal
vocabulary based on the list of text. It
essentially creates a dictionary of word
to index pairs. And then we have text to
sequences that will transform each text
in text to a sequence of integers. And
then we have pad sequences that will
ensure all sequences have the same
length by padding shorter sequences with
zeros up to the maximum length. And then
we'll start building the model. Here we
have sequential model that will set up a
linear stack of layers. And then we have
embedding layer that will map each word
index to an embedding vector of size 64.
So the input length is set to 10 that is
the length of the input sequences. Then
we'll start with LSTM layers. So two
LSTM layers are added. The first one
returns sequences to allow the next LSTM
layer to process these sequences. And
after that we have the dropout layer
that applies dropout with a rate of 0.5
of the first LSTM layer to reduce
overfitting. And after that we'll come
to dense layer that has output of a
single scalar that represents the
predicted setment and using sigmoid
activation to output a probability. And
now we'll start with model compilation
and training. So we'll configure the
model for training and we'll use binary
cross entropy as the loss function that
is suitable for binary classification
and the atom optimizer and tracks
constantly accuracy as a metric and then
we have the fit that trains the model
for a specified number of epochs that is
iterations or the entire data set and
then we'll predict the model that is
after training the model can predict the
sentiment of the reviews in the data
set. This is useful for checking how the
model performs on the training data
itself. So this breakdown explains each
step of the coding process detailing how
the data is prepared and how the model
is configured and then we'll compile it
and use for training and prediction. So
it's detailed explanation should help in
understanding how to implement a simple
LSTM model for sentiment analysis in
TensorFlow. Now moving to the question
number 16. So let's start with question
number 16 that is based on anomly
detection in transaction data. So the
question is you are tasked with
identifying unusual transactions in a
company's financial data that might
suggest fraudulent activity. So what
steps would you take to develop an
anomaly detection model and can you
provide a sample Python code snippet to
illustrate how you would pre-process the
data and apply an anomaly detection
technique. So we'll start answering this
with anomaly detection technique that is
anomaly detection is essential for
preventing fraud by identifying
transactions that deviate significantly
from typical patterns. And now we'll see
the structured approach to building an
anomaly detection model for transaction
data. We'll start with data collection
and cleaning and we'll collect all the
compiling transaction data which should
include details like transaction amount,
time, user ID and transaction type. Then
we'll move to feature engineering and
develop features that capture the
essence of transaction such as time of
day and the day of the week. And then we
have data normalization. We'll use
scaling techniques such as minmax
scaling or standardization to ensure
that the model is perfectly normalized.
And then we have choosing the anomaly
detection technique. So here we have to
choose the technique which is effective
for highdimensional data sets and works
for isolating anomalies instead of
profiling normal data points. After
choosing the anomaly technique will
train anomaly identification. We'll fit
the chosen model to the data and the
anomalies that would have been chosen
will be those transactions that the
model identifies and after this we come
to the last step that is review and
action. Here we have manual review that
is transactions flagged as potential
anomalies should be reviewed manually to
confirm fraudent activity and then we
have continuous improvement that is we
can regularly update the model with the
new data and feedback from the review
process to improve accuracy. And now
moving to the prediction that is after
training the model we can predict the
sentiment of the reviews in the data set
and this is useful for checking how the
model performs on the training data
itself. Now we'll see the Python code to
see how you can set up this model for
anomaly detection. We'll start by
importing the libraries and module and
after that we'll load and prepare data.
That is we'll load transaction data from
a CSV file into the pandas data frame.
And after that we'll convert the
transaction time column to date time
format which allows the extraction of
additional time based features. And
after that we'll perform feature
engineering that will extract the hour
of the day from the transaction time
column. This feature can be important as
transactions occurring at unusual hours
may be indicative of fraud. And then
we'll move to the normalization of data.
This will apply standard scaling to the
amount n of the day feature. This
normalization process involves
subtracting the mean and dividing by the
standard deviation for each feature
ensuring that the feature contribute
equally to the analysis and improving
the performance of many machine learning
algorithms. And after that we'll start
with anomaly detection with isolation
forest. That's a technique. We'll
initialize an isolation forest model
with 100 trees that is n estimators
equal to 100. Setting the proportions of
outliers that is contamination to 1% of
the data. So this parameter is crucial
as it influences the threshold of
marking an observation as an anomaly.
Then we'll fit the model to the scaled
amount n of the data and predict the
anomaly status for each transaction. And
then we'll start with filter and display
anomalies. We'll filter out transactions
identified as anomalies that is anomaly
equal equal to minus one. We'll display
these transactions which can be reviewed
manually to determine if they represent
actual fraud net activity. So this code
snippet provides a systematic approach
to detecting anomalies in transaction
data leveraging the isolation forest
algorithms ability to handle complex and
highdimensional data set effectively. So
the pre-processing steps ensured that
the data is appropriately formatted and
normalized for optional model
performance. So this was all about
question number 16. Now moving to
question number 17 and that is based on
integrating machine learning models into
web applications. And your question is,
you have developed a machine learning
model to predict real estate prices
based on various features like location,
size, and amenities. How would you
integrate this model into a web
application to allow users to get
real-time price predictions? Can you
provide a sample Python code snippet to
illustrate how you would prepare the
model for integration and handle user
request? So, starting with the approach
that is integrating a machine learning
model into a web application. This will
involve several steps to ensure the
model is accessible and perform well in
a live environment. So here's how you
can approach this task. We could divide
into steps and we'll start with number
one step that is model preparation.
We'll finalize and save the model. So
once your model is trained and
validated, save it using a format that
can be easily loaded into a web
application. So Python's pickle module
or TensorFlow's save model format are
commonly used for this purpose. Then we
can use web application backend setup.
For this, select a suitable web
framework. So, Flask is popularly known
for its simplicity and effectiveness in
integrating Python based machine
learning models. And after that, we'll
develop the API. After developing the
API within your Flask app that you can
receive user inputs for model features,
load the model, make prediction, and
return the result. And after this, we'll
develop the UI. We'll design a
user-friendly interface. We'll create a
simple and intuitive UI that lets users
input the feature like location, size
and submit them for prediction. And
after that we'll move to the deployment
phase. We'll use a cloud platform like
Heroku, AWS or Google Cloud to deploy
your Flask application. And then we have
the maintenance and updates. We'll
monitor and update regularly for the
performance and use the model as needed
based on user feedback. So now moving to
the Python code and see how this model
can be created. So here we'll start
importing the libraries and module and
we are using flask ple and jsonify and
we will start with app initialization.
We'll initialize a new flask web
application that would be a special
variable which gives python files a
unique name to differentiate between
them when they are important into other
scripts. And after that we'll load the
model. So loading a pretend machine
learning model from the file system. So
this model is assumed to be saved in the
same directory as this script. So the
model is loaded in RB mode which stands
for read binary. And after that we'll
move to API route and prediction
function. So we will define an API
endpoint at predict that listens for
post request. This is the URL that the
front end of the web application will
call to send data to the back end. And
after that we'll start with predicting
the function. And here we have extract
features that retrieves data sent into
the JSON format from the post request
that is request get_json and the force
we have set it as true here and
forcefully formats the request data into
JSON ensuring compatibility and then
we'll extract the relevant features that
is location size and amenities from the
JSON object and store them in a list as
expected by the model and after
preparing the features we'll make the
prediction we'll use the loaded model to
make a prediction based bas on the
provided feature and then we have the
return prediction method. Here we will
convert the prediction result into JSON
format using JSON and send it back to
the client and this will ensure that the
response can be easily handled by the
client application. So this was all
about the question number 17. Now moving
to the question number 18 that is based
on analyzing. And now we move to the
question number 18 that is based on
analyzing geospatial data. And your
question is you are tasked with
analyzing geospatial data to help a city
improve its public transportation
system. The data includes GPS
coordinates of bus stops, ridership
numbers and traffic patterns. What steps
would you take to analyze this data? And
can you provide a sample Python code
snippet to illustrate how you might
visualize bus stop location and
ridership? So you can start answering
this question that juice better data
analysis can provide critical insights
into how effectively a public
transportation system serves its city
and guide improvements and there's an
detailed approach for this and we can
start with data preparation and in this
we'll do data collection and data
cleaning and after this step we'll move
to the next step that is explorative
data analysis and in this we'll have
statistical summary we'll generate
descriptive statistics and then we have
correlation analysis
And after moving that we have geospatial
visualization that is mapping bus stop.
We'll plot the locations of bus stop on
a map to visually assess their
distribution across the city. And after
that we have heat maps that will create
ridership data to identify hot sports
and areas with potential service gaps.
And after geospatial visualization we'll
move with spatial analysis. We have
proximity analysis that will analyze the
proximity of bus stop to key areas like
commercial centers or residential areas.
And now moving to the fifth step that is
optimization and recommendation. So
we'll have a route optimization that
will suggest modifications to route
based on traffic patterns and ridership
demand and the policy recommendations
that will provide actionable
recommendations for improving bus
frequencies. Now move to the sample
Python code where we can define this
model and use it accordingly. And here
we will start importing the libraries
and modules. And here we'll start with
importing geopandas and m lib dotpipo.
And after importing we'll start with
data loading. So we will declare a
variable bus stops and load the bus
stops data from a shape file. So shape
files are popular geospatial vector data
formats for geographic information
system software and then we have the
wrership that will load wrership data
from a CSV file which includes columns
for longitude latitude and ridership
levels and after that we'll create geo
data frame that will convert the
wrership data frame into a geo data
frame and this step involves creating a
geometry column from the longitude and
latitude columns and then we have the
plotting one here we will plot the
graphs that would figures and axis and
create a figure for the single subplot
with a specified size that is 10 + 10
in. And then we have city map.plot. It
is assumed that there is a base map of
the city loaded as a geo data frame
named city map. This is plotted first
with a light gray color to serve as a
background for the other layers. So this
was all about question number 18. Now
moving to question number 19 that is
based on predictive maintenance using
machine learning and the question is you
are tasked with developing a predictive
maintenance system for a manufacturing
plant that relies heavily on automated
machinery. So the data available
includes machine operational parameters,
maintenance history and failure
incidents. What steps would you take to
develop a predictive model and can you
provide a sample Python code? So you can
start with predictive maintenance that
is essential in manufacturing as it
helps prevent equipment failures
reducing downtime and maintenance cost.
And here you would have a detail
approach or predictive model for this
starting with data collection and
integration. Then you can do EDA that is
exploratory data analysis and then we
can perform feature engineering and then
move to data prep-processing task and
then the selection model and training
and after that we have model evaluation
and deployment technique that we can do
for the model using appropriate metrics
such as precision, recall and F1 score.
So this was all about question number
19. So now move to question number 20
that is based on personalization using
machine learning and your question is
you are tasked with developing a machine
learning model to personalize content
recommendations for users on a media
streaming platform. The data available
includes user demographic retails
viewing history and ratings. So what
steps would you take to build a model
for personalized recommendations and can
you provide a sample Python code for
that? So you can start answering this
with creating a personalized
recommendation systems. This would be
essential for engaging users by
providing content that is relevant to
their interest. And there will be a
systematic approach or personalized
content recommendation. We'll start with
data collection and integration. And
after that, we'll perform EDA that is
explorative data analysis. And then we
have feature engineering. In this we'll
interact features and the temporal
features. We'll include time based
features to capture trends and
seasonality in viewing behavior. And
then we'll select the model that is by
collaborative filtering and hybrid
models. And then we'll train the model
and validation and implement and monitor
them. And after that we'll deploy the
model. So let's start with beginner
level questions. And number one is what
is machine learning? So machine learning
is a subset of artificial intelligence
that involves the use of algorithms and
statistical models to enable computers
to perform task without explicit
instructions. that is by relying on
patterns and interference. And now
moving to number second question that is
what are the different types of machine
learning. So the three main types of
machine learning are number one is
supervised learning and then comes
unsupervised learning and then there is
reinforcement learning. Now moving to
next question that is third that is what
is supervised learning. So supervised
learning involves training a model on a
label data set which means each training
example is paired with an output label.
The model learns to predict the output
from the input data. Now moving to the
fourth question that is what is
unsupervised learning. So unsupervised
involve training a model on data that
does not have labeled responses. The
model tries to learn the patterns and
the structure from the input data. So
guys, these are the beginner level
questions and now we'll move to the
fifth question that is what is
reinforcement learning. So reinforcement
learning is a type of machine learning
where an agent learns to make decisions
by performing actions and receiving
rewards or penalties. The goal is to
maximize the cumulative reward. So now
moving to the sixth question that is
what is a model in machine learning. So
a model in machine learning is a
mathematical representation of a real
world process. It is trained on data to
recognize patterns and make predictions
or decisions based on new data. So now
moving to seventh question that is what
is overfitting? So overfitting occurs
when a machine learning model performs
well on the training data but poorly on
new unseen data. It indicates that the
model has learned the noise and details
in the training data instead of the
actual patterns. So now coming to
question number eight that is what is
underfitting? So underfitting occurs
when a machine learning model is too
simple to capture the underlying
patterns in the data. It performs poorly
on both the training data and new data.
Now move to the next question that is
ninth question and the question is what
is a confusion matrix? So confusion
matrix is a table used to evaluate the
performance of a classification model.
It summarizes the number of correct and
incorrect predictions made by the model
and that is categorized by each class.
Now moving to the 10th question that is
what is cross validation? So cross
validation is a technique for assessing
how the results of a statistical
analysis will generalize to an
independent data set. It involves
partitioning the data into subsets.
Training the model on some subsets and
validating it on the remaining subsets.
This was all about that is the 10th
question or the overall 1 to 10
questions for beginner level. Now we'll
move to intermediate level and here
we'll cover 10 questions. So we'll start
with 11th question that is what is a ROC
curve. So ROC that is receiver operating
characteristic curve. It is a graphical
representation of a classifier's
performance across different thresholds.
It plots the true positive rate that is
TPR against a false positive rate that
is FPR. Now moving to 12th question that
is what is precision and recall. So
precision is the ratio of correctly
predicted positive observations to the
total predicted positives and recall is
the ratio of correctly predicted
positive observations to all actual
positives. So the formula is precision
equal to TP/TP
plus FP and the recall is TP/TP
+ F_sub_1. So now we'll move to the 13th
question that is what is the F1 score?
So the F1 score is the harmonic mean of
precision and recall. It provides a
balance between the two metrics and is
useful when you need to balance
precision and recall. F1 score is equal
to twice into precision into recall and
that is divided by precision plus
recall. Now we'll move to 14th question
and here we will cover regularization.
So the question is what is
regularization? So it is a technique
used to prevent overfitting by adding a
penalty to the model's complexity and
the common types of regularization
include L1 that is lasso and L2 ridge
regularization. Now we'll move to the
15th question that is what is the bias
variance tradeoff. So the bias variance
trade-off is a fundamental issue in
machine learning that involves balancing
the error introduced by the model's
assumptions and the error due to model
complexity. So a good model should have
low bias and low variance. Now we'll
move to the question number 16 that is
what is feature engineering. So feature
engineering is the process of creating
new features or modifying existing ones
to improve the performance of a machine
learning model. It involves techniques
like normalization and coding
categorical variables and creating
interaction terms. So now we'll move to
question number 17 and that is about
gradient descent. So the question is
what is gradient descent and your answer
is gradient descent is an optimization
algorithm used to minimize the cost
function in machine learning models and
it iteratively adjust the model
parameters in the direction of the
steepest descent of the coast function.
So with this we'll move to the 18th
question and that will cover with the
difference between bagging and boosting.
So the question is what is difference
between bagging and boosting and you
could answer this with starting with
bagging that is bootstrap aggregating
that involves training multiple models
on different subsets of the data and
averaging their predictions. Then comes
boosting that involves training models
sequentially with each new model
focusing on correcting the errors of the
previous ones. And then we have the
question number 19 that is what is a
decision tree? So a decision tree is a
nonparametric supervised learning
algorithm used for classification and
regression. It splits the data into
subsets based on the value of input
features resulting in a treel like
structure of decisions. Now we'll move
to question number 20 that is what is a
random forest. So random forest is an
ansemble learning method that combines
multiple decision trees to improve the
accuracy and robustness of the model. It
builds each tree using a random subset
of features and data points and then
averages their predictions. So these
were the questions that are for the
intermediate level and these are just
the basic questions or I will just say
the theoretical questions that can be
asked in an interview. So be prepared
for that. Now we'll move to the advanced
level interview questions and we'll
start with question number 21. And here
also we'll cover the 10 questions. So
number one question or that is 21
question and the question is what is a
support vector machine? So support
vector machine is a supervised learning
algorithm used for classification and
regression. It finds the optimal hyper
plane that maximizes the margin between
different classes in the feature space.
And then comes question number 22 that
is what is principal component analysis.
So principal component analysis is a
dimensionality reduction technique that
transforms highdimensional data into a
lower dimensional space by finding the
directions that is principal components
that maximize the variance in the data.
And then comes the question number 23
that is what is a neural network? So a
neural network is a series of algorithms
that attempt to recognize underlying
relationships in a set of data through a
process that mimics the way the human
brain operates. It consists of layer of
interconnected nodes or neurons. And
then comes the question number 24 that
is what is deep learning? So deep
learning is a subset of machine learning
that involves neural networks with many
layers that is deep neural networks and
it is particularly effective for task
like image and speech recognition. Now I
move to question number 25 that is what
is convolutional neural network that is
CNN. So we will start the answer by
answering the interviewer that a
convolutional neural network is a type
of deep learning model specifically
designed for processing structured grid
data like images. It uses convolutional
layers to extract special features or
the spatial features and patterns from
the input data. Now we move to the
question number 26 that is what is a
recurrent neural network or RNN. So a
recurrent neural network is a type of
neural network designed for sequential
data and it has connections that form
directed cycles allowing it to maintain
a memory of previous inputs and process
sequences of data. So this was all about
question number 26 and now we will cover
the question number 27 that is what is
the difference between batch gradient
descent and stoastic gradient descent.
So batch gradient descent computes the
gradient of the coast function using the
entire training data set while
stochastic gradient descent that is SGD
computes the gradient using only one
training example at a time. So SGD is
faster but noisier. Now we move to
question number 28 that is what is
dropout in neural networks. So dropout
is a regularization technique used in
neural networks to prevent overfitting
and it involves randomly setting a
fraction of the neurons to zero during
training forcing the network to learn
more robust features. And now we'll move
to question number 29 and that will be
about transfer learning. And your
question is what is transfer learning?
So we'll answer this to the interviewer
by starting that transfer learning is a
technique in machine learning where a
model developed for one task is reused
as the starting point for a model on a
second related task. It is particularly
useful when there is limited data
available for the second task. Now we'll
move to the last question and the 30th
question. So that is what is a
generative adversial network that is GN.
So you can start answering this. So
generative adversial network is a type
of deep learning model consisting of two
neural networks a generator and a
discriminator that are trained
simultaneously. The generator creates
fake data while the discriminator tries
to distinguish between real and fake
data leading to the generator producing
increasingly realistic data. And these
questions and answers are over and these
covers a wide range of topics in machine
learning and should help prepare for
interviews at various levels.
>> And with that we have reached the end of
the session on the AI and machine
learning engineer full course for
beginners. If you have any doubts or
questions about this video let us know
in the comment section below and a team
of experts will be happy to help you.
Until next time thank you and keep
learning. Stay tuned for more from
SimplyLearn.
Can't find what you're looking for?
Get subtitles in any language from opensubtitles.com, and translate them here.